Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

121 to 180 of 18,807

  1. FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks

    Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia +3

    cs.CVarXiv:1612.01925v12016
  2. An Empirical Study of World Model Quantization

    Zhongqian Fu, Tianyi Zhao, Kai Han +3

    cs.LGcs.CVarXiv:2602.02110v12026
  3. Understanding Convolutional Neural Networks with A Mathematical Model

    C. -C. Jay Kuo

    cs.CVarXiv:1609.04112v22016
  4. Unsupervised Monocular Depth Estimation with Left-Right Consistency

    Clément Godard, Oisin Mac Aodha, Gabriel J. Brostow

    cs.CVcs.LGstat.MLarXiv:1609.03677v32016
  5. It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces

    Leonhard F. Feiner, Manuel Nickel, Martin Menten +6

    cs.LGcs.CVarXiv:2608.24518v12026
  6. Continual Visual Learning under Evolving Semantic Concept Shift

    Ismail Lamaakal, Chaymae Yahyati, Yassine Maleh +2

    cs.CVcs.CLarXiv:2608.23903v12026
  7. Deep Image Homography Estimation

    Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich

    cs.CVarXiv:1606.03798v12016
  8. Improved Techniques for Training GANs

    Tim Salimans, Ian Goodfellow, Wojciech Zaremba +3

    cs.LGcs.CVcs.NEarXiv:1606.03498v12016
  9. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

    Marie-Lisa Eich, Kai Standvoss, Timo Milbich +30

    cs.CVcs.AIcs.LGarXiv:2608.23803v12026
  10. Segmentation from Natural Language Expressions

    Ronghang Hu, Marcus Rohrbach, Trevor Darrell

    cs.CVarXiv:1603.06180v12016
  11. Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

    Federico Stella, Fei Jiang, Zhongshi Jiang +4

    cs.CVcs.LGarXiv:2608.23410v12026
  12. Loopy: Seamless Video Loop Generation via Anchored Looping Shift of Positional Embedding

    Haotian Dong, Wenjing Wang, Chen Li +3

    cs.CVarXiv:2608.23090v12026
  13. From Generation to Simulation: How Far Are World Models from Being True Simulators?

    Tong Wang, Huan Deng, Mucheng Yang +3

    cs.AIcs.CVarXiv:2608.23070v12026
  14. Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

    Geo Ahn, Inwoong Lee, Taeoh Kim +3

    cs.CVcs.AIarXiv:2601.16211v32026
    Summaries:한국어
  15. Transition Matching Distillation for Fast Video Generation

    Weili Nie, Julius Berner, Nanye Ma +3

    cs.CVcs.AIcs.LGarXiv:2601.09881v22026
  16. Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

    Luka Ribar, Jeevan Bhoot, Douglas Orr

    cs.CVcs.LGarXiv:2608.21134v12026
  17. CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

    Bokai Zhao, Yiyang Zhang, Hanqing Chao +6

    cs.AIcs.CVarXiv:2608.21060v12026
  18. On learning optimized reaction diffusion processes for effective image restoration

    Yunjin Chen, Wei Yu, Thomas Pock

    cs.CVarXiv:1503.05768v22015
  19. Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds

    Changjin Kim, Seunghun Paik, Dongsoo Kim +1

    cs.CVarXiv:2608.20884v12026
  20. LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic Residuals

    Haozhen Yan, Ruoxin Chen, Jiahui Zhan +6

    cs.CVarXiv:2608.20882v12026
  21. 4DAnyone: Create Anyone in 4D from a Casual Monocular Video

    Yudong Jin, Tao Xie, Qihang Zhang +6

    cs.CVarXiv:2608.20335v12026
  22. Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

    Pardis Taghavi, Reza Langari, Gaurav Pandey

    cs.CVcs.AIcs.LGarXiv:2608.18484v12026
  23. Seedream 4.0: Toward Next-generation Multimodal Image Generation

    Team Seedream, :, Yunpeng Chen +48

    cs.CVarXiv:2509.20427v32025
  24. MapAnything: Universal Feed-Forward Metric 3D Reconstruction

    Nikhil Keetha, Norman Müller, Johannes Schönberger +14

    cs.CVcs.AIcs.LGarXiv:2509.13414v32025
  25. Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

    Yifei Wu, Yicheng Wu, Qiang Ma +5

    cs.CVcs.AIarXiv:2608.17255v12026
  26. Kwai Keye-VL 1.5 Technical Report

    Biao Yang, Bin Wen, Boyang Ding +58

    cs.CVarXiv:2509.01563v32025
  27. Factors of Transferability for a Generic ConvNet Representation

    Hossein Azizpour, Ali Sharif Razavian, Josephine Sullivan +2

    cs.CVarXiv:1406.5774v32014
  28. Nuclear Norm based Matrix Regression with Applications to Face Recognition with Occlusion and Illumination Changes

    Jian Yang, Jianjun Qian, Lei Luo +2

    cs.CVarXiv:1405.1207v12014
  29. MLLM-Guided Semantic Correction for Text-to-Video Generation

    Junhao Chen, Zheqi Lv, Keting Yin +6

    cs.CVcs.AIarXiv:2608.16513v12026
  30. Point Transformer

    Nico Engel, Vasileios Belagiannis, Klaus Dietmayer

    cs.CVarXiv:2011.00931v22020
    Summaries:한국어
  31. From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

    Le Zhuo, Liangbing Zhao, Sayak Paul +6

    cs.CVarXiv:2504.16080v12025
  32. LaSOT: A High-quality Large-scale Single Object Tracking Benchmark

    Heng Fan, Hexin Bai, Liting Lin +11

    cs.CVarXiv:2009.03465v32020
  33. Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

    Hongbo Jiang, Jie Li, Yunhang Shen +2

    cs.LGcs.CVarXiv:2608.14655v12026
  34. Accurate RGB-D Salient Object Detection via Collaborative Learning

    Wei Ji, Jingjing Li, Miao Zhang +2

    cs.CVarXiv:2007.11782v12020
  35. NSGANetV2: Evolutionary Multi-Objective Surrogate-Assisted Neural Architecture Search

    Zhichao Lu, Kalyanmoy Deb, Erik Goodman +2

    cs.CVcs.LGcs.NEarXiv:2007.10396v12020
  36. HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting

    Wei Zhang, Shengkai Yu, Shiqiang Gong +3

    cs.CVcs.AIarXiv:2608.14136v12026
  37. ContactPose: A Dataset of Grasps with Object Contact and Hand Pose

    Samarth Brahmbhatt, Chengcheng Tang, Christopher D. Twigg +2

    cs.CVarXiv:2007.09545v12020
  38. V-RAE: Rethinking Video Latent Spaces for Generation

    Minghui Guo, Shengqiong Wu, Hao Fei

    cs.CVarXiv:2608.13556v12026
  39. SegFix: Model-Agnostic Boundary Refinement for Segmentation

    Yuhui Yuan, Jingyi Xie, Xilin Chen +1

    cs.CVarXiv:2007.04269v42020
  40. Evidence-RL: Towards Evidence-intensive Visual Reasoning

    Haojie Huang, Xinlei Yu, Chengming Xu +6

    cs.CVcs.AIarXiv:2608.08021v12026
  41. Long Context Tuning for Video Generation

    Yuwei Guo, Ceyuan Yang, Ziyan Yang +5

    cs.CVarXiv:2503.10589v12025
  42. Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

    Xuechao Zou, Shun Zhang, Kai Li +6

    cs.CVcs.AIcs.MAarXiv:2608.06865v12026
  43. SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

    Katrin Renz, Long Chen, Elahe Arani +1

    cs.CVcs.ROarXiv:2503.09594v12025
  44. Training Generative Adversarial Networks with Limited Data

    Tero Karras, Miika Aittala, Janne Hellsten +3

    cs.CVcs.LGcs.NEarXiv:2006.06676v22020
  45. Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

    Zelong Sun, Jun Wang, Kaicheng Yang +3

    cs.CVarXiv:2608.06060v12026
  46. FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

    Zhuoran Zhang, Bowen Li, Jingcheng Ju +5

    cs.CVarXiv:2608.04530v12026
  47. Dataset Distillation with Neural Characteristic Function: A Minmax Perspective

    Shaobo Wang, Yicun Yang, Zhiyuan Liu +4

    cs.CVcs.AIcs.LGarXiv:2502.20653v12025
    Summaries:한국어
  48. Wavelet Integrated CNNs for Noise-Robust Image Classification

    Qiufu Li, Linlin Shen, Sheng Guo +1

    cs.CVarXiv:2005.03337v22020
  49. Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment

    Harrish Thasarathan, Julian Forsyth, Thomas Fel +2

    cs.CVcs.LGarXiv:2502.03714v22025
  50. Maximum Density Divergence for Domain Adaptation

    Li Jingjing, Chen Erpeng, Ding Zhengming +3

    cs.CVcs.LGcs.MMarXiv:2004.12615v12020
  51. Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

    Kairui Hu, Penghao Wu, Fanyi Pu +5

    cs.CVcs.CLarXiv:2501.13826v12025
  52. Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

    Krishna Kumar Singh, Dhruv Mahajan, Kristen Grauman +3

    cs.CVcs.LGarXiv:2001.03152v22020
  53. ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

    Ting Huang, Zhenyu Zhang, Wenyuan Huang +2

    cs.CVarXiv:2607.17599v12026
  54. Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

    Yucheng Zhou, Lingran Song, Jianbing Shen

    cs.CLcs.AIcs.CVarXiv:2501.01377v22025
  55. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou +17

    cs.CVcs.AIcs.IRarXiv:2412.07626v22024
  56. Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

    Wenzheng Zeng, Siyi Jiao, Chen Gao +2

    cs.CVcs.AIcs.MMarXiv:2607.02963v12026
  57. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    Seohyun Lee, Seoung Choi, Dohwan Ko +2

    cs.CVcs.AIarXiv:2607.00446v12026
  58. Guiding Monocular Depth Estimation Using Depth-Attention Volume

    Lam Huynh, Phong Nguyen-Ha, Jiri Matas +2

    cs.CVarXiv:2004.02760v22020
  59. PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

    Shaowei Liu, Zhongzheng Ren, Saurabh Gupta +1

    cs.CVcs.AIcs.LGarXiv:2409.18964v12024
  60. LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

    Yijia Xiao, Edward Sun, Tianyu Liu +1

    cs.AIcs.CLcs.CVarXiv:2407.04973v12024