Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

181 to 240 of 18,815

  1. ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

    Ting Huang, Zhenyu Zhang, Wenyuan Huang +2

    cs.CVarXiv:2607.17599v12026
  2. Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

    Yucheng Zhou, Lingran Song, Jianbing Shen

    cs.CLcs.AIcs.CVarXiv:2501.01377v22025
  3. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou +17

    cs.CVcs.AIcs.IRarXiv:2412.07626v22024
  4. Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

    Wenzheng Zeng, Siyi Jiao, Chen Gao +2

    cs.CVcs.AIcs.MMarXiv:2607.02963v12026
  5. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    Seohyun Lee, Seoung Choi, Dohwan Ko +2

    cs.CVcs.AIarXiv:2607.00446v12026
  6. Guiding Monocular Depth Estimation Using Depth-Attention Volume

    Lam Huynh, Phong Nguyen-Ha, Jiri Matas +2

    cs.CVarXiv:2004.02760v22020
  7. PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

    Shaowei Liu, Zhongzheng Ren, Saurabh Gupta +1

    cs.CVcs.AIcs.LGarXiv:2409.18964v12024
  8. LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

    Yijia Xiao, Edward Sun, Tianyu Liu +1

    cs.AIcs.CLcs.CVarXiv:2407.04973v12024
  9. MAIRA-2: Grounded Radiology Report Generation

    Shruthi Bannur, Kenza Bouzid, Daniel C. Castro +18

    cs.CLcs.CVarXiv:2406.04449v22024
  10. Solving Inverse Problems with Piecewise Linear Estimators: From Gaussian Mixture Models to Structured Sparsity

    Guoshen Yu, Guillermo Sapiro, Stéphane Mallat

    cs.CVarXiv:1006.3056v12010
  11. UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

    Jiehui Huang, Yuechen Zhang, Bin Xia +7

    cs.CVarXiv:2606.21661v12026
  12. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

    Zhe Chen, Weiyun Wang, Hao Tian +32

    cs.CVarXiv:2404.16821v22024
  13. DeepGCNs: Making GCNs Go as Deep as CNNs

    Guohao Li, Matthias Müller, Guocheng Qian +4

    cs.CVcs.LGeess.IVarXiv:1910.06849v32019
  14. World Action Models: A Survey

    Qiuhong Shen, Shihua Zhang, Yue Liao +5

    cs.ROcs.CVarXiv:2606.20781v12026
  15. Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

    Mu Hu, Wei Yin, Chi Zhang +7

    cs.CVarXiv:2404.15506v42024
  16. The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

    Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng +4

    cs.LGcs.CVarXiv:2606.19162v12026
  17. ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback

    Ming Li, Taojiannan Yang, Huafeng Kuang +4

    cs.CVcs.AIcs.LGarXiv:2404.07987v42024
  18. GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

    Yinghao Xu, Zifan Shi, Wang Yifan +5

    cs.CVarXiv:2403.14621v12024
  19. Anomaly Detection in Video Sequence with Appearance-Motion Correspondence

    Trong Nguyen Nguyen, Jean Meunier

    cs.CVcs.LGcs.NEarXiv:1908.06351v12019
  20. SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

    Dongyang Liu, Renrui Zhang, Longtian Qiu +16

    cs.CVcs.AIcs.CLarXiv:2402.05935v32024
  21. InstructIR: High-Quality Image Restoration Following Human Instructions

    Marcos V. Conde, Gregor Geigle, Radu Timofte

    cs.CVcs.LGeess.IVarXiv:2401.16468v52024
  22. Supervised Dictionary Learning

    Julien Mairal, Francis Bach, Jean Ponce +2

    cs.CVarXiv:0809.3083v12008
  23. Brain-Inspired Stochastic Joint Embedding Representation Learning

    Makoto Yamada, Kian Ming A. Chai, Ayoub Rhim +3

    cs.CVcs.AIcs.LGarXiv:2505.11129v22025
  24. Siamese Masked Autoencoders

    Agrim Gupta, Jiajun Wu, Jia Deng +1

    cs.CVcs.LGarXiv:2305.14344v12023
  25. Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

    Alexandre Eymaël, Renaud Vandeghen, Anthony Cioppa +3

    cs.CVarXiv:2403.17823v22024
  26. Visual Representation Learning with Stochastic Frame Prediction

    Huiwon Jang, Dongyoung Kim, Junsu Kim +3

    cs.CVcs.AIcs.LGarXiv:2406.07398v22024
  27. Anatomy-aware 3D Human Pose Estimation with Bone-based Pose Decomposition

    Tianlang Chen, Chen Fang, Xiaohui Shen +3

    cs.CVarXiv:2002.10322v52020
  28. Pushing Stochastic Gradient towards Second-Order Methods -- Backpropagation Learning with Transformations in Nonlinearities

    Tommi Vatanen, Tapani Raiko, Harri Valpola +1

    cs.LGcs.CVstat.MLarXiv:1301.3476v32013
  29. Harnessing Large Language Models for Training-free Video Anomaly Detection

    Luca Zanella, Willi Menapace, Massimiliano Mancini +2

    cs.CVarXiv:2404.01014v12024
  30. Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

    Xiaowei Mao, Bowen Sui, Weijie Zhang +7

    cs.CVcs.AIarXiv:2604.23724v42026
  31. Towards Automatic Concept-based Explanations

    Amirata Ghorbani, James Wexler, James Zou +1

    stat.MLcs.CVcs.LGarXiv:1902.03129v32019
  32. LaME: Learning to Think in Latent Space for Multimodal Embedding via Information Bottleneck

    Peixi Wu, Biao Yang, Feipeng Ma +7

    cs.CVarXiv:2606.13061v32026
  33. Pose-conditioned Spatio-Temporal Attention for Human Action Recognition

    Fabien Baradel, Christian Wolf, Julien Mille

    cs.CVarXiv:1703.10106v22017
  34. A Hybrid Model for Identity Obfuscation by Face Replacement

    Qianru Sun, Ayush Tewari, Weipeng Xu +3

    cs.CVcs.CRarXiv:1804.04779v22018
  35. Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

    Lanyun Zhu, Deyi Ji, Tianrun Chen +2

    cs.CVarXiv:2510.02745v22025
  36. ViStoryBench: Comprehensive Benchmark Suite for Story Visualization

    Cailin Zhuang, Ailin Huang, Yaoqi Hu +12

    cs.CVarXiv:2505.24862v52025
  37. Actor-Critic Sequence Training for Image Captioning

    Li Zhang, Flood Sung, Feng Liu +4

    cs.CVarXiv:1706.09601v22017
  38. MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

    Junjie Zhou, Zheng Liu, Ze Liu +6

    cs.CVcs.CLarXiv:2412.14475v12024
  39. MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

    Kai Zhang, Yi Luan, Hexiang Hu +5

    cs.CVcs.AIcs.CLarXiv:2403.19651v22024
  40. CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning

    Hao Yu, Zhuokai Zhao, Shen Yan +7

    cs.CVcs.AIcs.CLarXiv:2503.19900v12025
  41. MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

    Kejian Zhu, Zhuoran Jin, Hongbang Yuan +6

    cs.CVcs.CLarXiv:2506.04141v22025
  42. Active Learning for Convolutional Neural Networks: A Core-Set Approach

    Ozan Sener, Silvio Savarese

    stat.MLcs.CVcs.LGarXiv:1708.00489v42017
  43. RepMLPNet: Hierarchical Vision MLP with Re-parameterized Locality

    Xiaohan Ding, Honghao Chen, Xiangyu Zhang +2

    cs.CVcs.AIcs.LGarXiv:2112.11081v22021
  44. Continual Contrastive Learning for Image Classification

    Zhiwei Lin, Yongtao Wang, Hongxiang Lin

    cs.CVcs.AIarXiv:2107.01776v42021
  45. WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition

    Marius Bock, Hilde Kuehne, Kristof Van Laerhoven +1

    cs.CVcs.HCarXiv:2304.05088v42023
  46. R-Transformer: Recurrent Neural Network Enhanced Transformer

    Zhiwei Wang, Yao Ma, Zitao Liu +1

    cs.LGcs.CLcs.CVarXiv:1907.05572v12019
  47. TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning

    Christian Greisinger, Steffen Eger

    cs.AIcs.CLcs.CVarXiv:2603.03072v32026
  48. AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ

    Jonas Belouadi, Anne Lauscher, Steffen Eger

    cs.CLcs.CVarXiv:2310.00367v22023
  49. DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning

    Abhay Zala, Han Lin, Jaemin Cho +1

    cs.CVcs.AIcs.CLarXiv:2310.12128v22023
  50. Multi-task Learning of Hierarchical Vision-Language Representation

    Duy-Kien Nguyen, Takayuki Okatani

    cs.CVarXiv:1812.00500v12018
  51. DeepPermNet: Visual Permutation Learning

    Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian +1

    cs.CVarXiv:1704.02729v12017
  52. Attention Guided Anomaly Localization in Images

    Shashanka Venkataramanan, Kuan-Chuan Peng, Rajat Vikram Singh +1

    cs.CVeess.IVarXiv:1911.08616v42019
  53. Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

    Chen Li, Peng Zhang, Hanyu Zhou +5

    cs.CVarXiv:2608.26902v12026
  54. Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering

    Yang Liu, Guanbin Li, Liang Lin

    cs.CVcs.AIarXiv:2207.12647v82022
  55. OnePose: One-Shot Object Pose Estimation without CAD Models

    Jiaming Sun, Zihao Wang, Siyu Zhang +4

    cs.CVarXiv:2205.12257v12022
  56. Multimodal Whole Slide Foundation Model for Pathology

    Tong Ding, Sophia J. Wagner, Andrew H. Song +20

    eess.IVcs.AIcs.CVarXiv:2411.19666v12024
  57. ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias

    Yufei Xu, Qiming Zhang, Jing Zhang +1

    cs.CVarXiv:2106.03348v42021
  58. On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law

    Damien Teney, Kushal Kafle, Robik Shrestha +3

    cs.CVcs.LGarXiv:2005.09241v12020
  59. Smoothed Dilated Convolutions for Improved Dense Prediction

    Zhengyang Wang, Shuiwang Ji

    cs.CVcs.LGarXiv:1808.08931v22018
  60. Binary Generative Adversarial Networks for Image Retrieval

    Jingkuan Song

    cs.CVarXiv:1708.04150v12017