Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,941 to 6,000 of 18,839

  1. Improved Anomaly Detection in Crowded Scenes via Cell-based Analysis of Foreground Speed, Size and Texture

    Vikas Reddy, Conrad Sanderson, Brian C. Lovell

    cs.CVarXiv:1304.0886v12013
  2. ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

    Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer +7

    cs.LGcs.AIcs.CVarXiv:2608.26083v12026
  3. Domain-Size Pooling in Local Descriptors: DSP-SIFT

    Jingming Dong, Stefano Soatto

    cs.CVarXiv:1412.8556v32014
  4. Reconstructing Hand-Object Interactions in the Wild

    Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa +1

    cs.CVarXiv:2012.09856v22020
  5. PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

    Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro +4

    cs.CVcs.AIarXiv:2608.25970v12026
  6. WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

    Jack Hong, Shilin Yan, Jiayin Cai +3

    cs.CVcs.AIarXiv:2502.04326v32025
  7. Less Contouring, More Accuracy: Lesion-Guided ROI Deep Learning for Ovarian Ultrasound Classification

    Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi +3

    cs.CVarXiv:2608.25965v12026
  8. Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

    Yijun Yang, Shenghe Zheng, Wenbo Li +8

    cs.CVarXiv:2609.03729v12026
  9. RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition

    Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin +2

    cs.CVarXiv:2007.07542v22020
  10. Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

    Haoyu Wang, Songchun Zhang, Haoran Li +3

    cs.CVcs.GRarXiv:2609.03557v12026
  11. GRIT: Teaching MLLMs to Think with Images

    Yue Fan, Xuehai He, Diji Yang +6

    cs.CVcs.AIcs.CLarXiv:2505.15879v22025
  12. Multiple Expert Brainstorming for Domain Adaptive Person Re-identification

    Yunpeng Zhai, Qixiang Ye, Shijian Lu +3

    cs.CVarXiv:2007.01546v32020
  13. Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

    Ye Wang, Ziheng Wang, Boshen Xu +14

    cs.CVcs.AIcs.CLarXiv:2503.13377v32025
  14. Metrics reloaded: Recommendations for image analysis validation

    Lena Maier-Hein, Annika Reinke, Patrick Godau +71

    cs.CVarXiv:2206.01653v82022
    Summaries:한국어
  15. Swarm-SLAM : Sparse Decentralized Collaborative Simultaneous Localization and Mapping Framework for Multi-Robot Systems

    Pierre-Yves Lajoie, Giovanni Beltrame

    cs.ROcs.CVarXiv:2301.06230v32023
  16. Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

    Yuxiang Lai, Jike Zhong, Ming Li +4

    cs.CVarXiv:2503.13939v52025
  17. Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?

    Erjin Zhou, Zhimin Cao, Qi Yin

    cs.CVarXiv:1501.04690v12015
  18. Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation

    Sun Yu, Ye Yun, Liu Wu +3

    cs.CVarXiv:1908.07172v22019
  19. Compact 3D Scene Representation via Self-Organizing Gaussian Grids

    Wieland Morgenstern, Florian Barthel, Anna Hilsmann +1

    cs.CVarXiv:2312.13299v22023
  20. Deep Self-Taught Learning for Weakly Supervised Object Localization

    Zequn Jie, Yunchao Wei, Xiaojie Jin +2

    cs.CVarXiv:1704.05188v22017
  21. MMSearch-R1: Incentivizing LMMs to Search

    Jinming Wu, Zihao Deng, Wei Li +5

    cs.CVcs.CLarXiv:2506.20670v12025
  22. Denoising Hyperspectral Image with Non-i.i.d. Noise Structure

    Yang Chen, Xiangyong Cao, Qian Zhao +2

    cs.CVarXiv:1702.00098v12017
  23. Functional Adversarial Attacks

    Cassidy Laidlaw, Soheil Feizi

    cs.LGcs.CVarXiv:1906.00001v22019
  24. KiU-Net: Towards Accurate Segmentation of Biomedical Images using Over-complete Representations

    Jeya Maria Jose, Vishwanath Sindagi, Ilker Hacihaliloglu +1

    eess.IVcs.CVarXiv:2006.04878v22020
  25. Rethinking the Heatmap Regression for Bottom-up Human Pose Estimation

    Zhengxiong Luo, Zhicheng Wang, Yan Huang +2

    cs.CVarXiv:2012.15175v42020
  26. Quaternion Convolutional Neural Networks

    Xuanyu Zhu, Yi Xu, Hongteng Xu +1

    cs.CVarXiv:1903.00658v12019
  27. Unpaired Deep Image Deraining Using Dual Contrastive Learning

    Xiang Chen, Jinshan Pan, Kui Jiang +5

    cs.CVarXiv:2109.02973v42021
  28. Epona: Autoregressive Diffusion World Model for Autonomous Driving

    Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu +9

    cs.CVarXiv:2506.24113v12025
  29. Re-IQA: Unsupervised Learning for Image Quality Assessment in the Wild

    Avinab Saha, Sandeep Mishra, Alan C. Bovik

    cs.CVcs.LGcs.MMarXiv:2304.00451v22023
  30. Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details

    Zeqiang Lai, Yunfei Zhao, Haolin Liu +23

    cs.CVcs.AIarXiv:2506.16504v12025
  31. Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

    Zhaochen Su, Peng Xia, Hangyu Guo +12

    cs.CVarXiv:2506.23918v32025
  32. SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation

    Zhongang Cai, Wanqi Yin, Ailing Zeng +10

    cs.CVarXiv:2309.17448v32023
  33. InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

    Zigang Geng, Binxin Yang, Tiankai Hang +8

    cs.CVarXiv:2309.03895v12023
  34. EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning

    Chenxin Xu, Robby T. Tan, Yuhong Tan +4

    cs.CVcs.MAarXiv:2303.10876v22023
  35. DeepEyesV2: Toward Agentic Multimodal Model

    Jack Hong, Chenxiao Zhao, ChengLin Zhu +3

    cs.CVcs.AIarXiv:2511.05271v42025
  36. FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling

    Haonan Qiu, Menghan Xia, Yong Zhang +4

    cs.CVarXiv:2310.15169v32023
  37. Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

    Duo Zheng, Shijia Huang, Yanyang Li +1

    cs.CVcs.AIarXiv:2505.24625v32025
  38. Neural-Collapse-guided Task-Free Continual Anomaly Detection

    Xiaotong Kong, Chaoyang Song, Ziai Zhou +3

    cs.CVarXiv:2609.03406v12026
  39. Improved Mean Flows: On the Challenges of Fastforward Generative Models

    Zhengyang Geng, Yiyang Lu, Zongze Wu +3

    cs.CVcs.LGarXiv:2512.02012v22025
  40. GP-VTON: Towards General Purpose Virtual Try-on via Collaborative Local-Flow Global-Parsing Learning

    Zhenyu Xie, Zaiyu Huang, Xin Dong +5

    cs.CVarXiv:2303.13756v12023
  41. Sonata: Self-Supervised Learning of Reliable Point Representations

    Xiaoyang Wu, Daniel DeTone, Duncan Frost +7

    cs.CVarXiv:2503.16429v12025
  42. VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

    Hritik Bansal, Clark Peng, Yonatan Bitton +3

    cs.CVarXiv:2503.06800v12025
  43. Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

    Yingmao Miao, Pengfei Zhang, Chaoran Xu +5

    cs.CVarXiv:2609.03673v12026
  44. Osprey: Pixel Understanding with Visual Instruction Tuning

    Yuqian Yuan, Wentong Li, Jian Liu +5

    cs.CVarXiv:2312.10032v42023
  45. VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

    Rui Meng, Ziyan Jiang, Ye Liu +10

    cs.CVcs.CLarXiv:2507.04590v12025
  46. MambaIRv2: Attentive State Space Restoration

    Hang Guo, Yong Guo, Yaohua Zha +5

    eess.IVcs.CVcs.LGarXiv:2411.15269v22024
  47. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

    Xiyao Wang, Zhengyuan Yang, Chao Feng +6

    cs.CVarXiv:2504.07934v32025
  48. MonoPerfCap: Human Performance Capture from Monocular Video

    Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer +4

    cs.CVcs.GRarXiv:1708.02136v22017
  49. A multicenter benchmark and clinically structured metric for coronary CTA report generation

    Zhiyu Ye, Yue Sun, Limiao Zou +7

    cs.CVarXiv:2609.00909v12026
  50. CMRVision: A Foundation Model for Cardiac MR Image Analysis

    Athira J. Jacob, Puneet Sharma, Daniel Rueckert

    cs.CVarXiv:2609.01308v12026
  51. Pose Estimation for Non-Cooperative Spacecraft Rendezvous Using Convolutional Neural Networks

    Sumant Sharma, Connor Beierle, Simone D'Amico

    cs.CVarXiv:1809.07238v12018
  52. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

    Tilman Räuker, Anson Ho, Stephen Casper +1

    cs.LGcs.AIcs.CLarXiv:2207.13243v62022
  53. TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval

    Yuqi Liu, Pengfei Xiong, Luhui Xu +2

    cs.CVarXiv:2207.07852v12022
  54. Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

    NVIDIA, :, Alisson Azzolini +51

    cs.AIcs.CVcs.LGarXiv:2503.15558v32025
  55. Content-Aware Unsupervised Deep Homography Estimation

    Jirong Zhang, Chuan Wang, Shuaicheng Liu +5

    cs.CVarXiv:1909.05983v22019
  56. Multi-Adapter RGBT Tracking

    Chenglong Li, Andong Lu, Aihua Zheng +2

    cs.CVarXiv:1907.07485v12019
  57. Illiterate DALL-E Learns to Compose

    Gautam Singh, Fei Deng, Sungjin Ahn

    cs.CVcs.LGarXiv:2110.11405v32021
  58. Making Better Mistakes: Leveraging Class Hierarchies with Deep Networks

    Luca Bertinetto, Romain Mueller, Konstantinos Tertikas +2

    cs.CVcs.LGarXiv:1912.09393v22019
  59. PromptIR: Prompting for All-in-One Blind Image Restoration

    Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan +1

    cs.CVarXiv:2306.13090v12023
  60. WorldMem: Long-term Consistent World Simulation with Memory

    Zeqi Xiao, Yushi Lan, Yifan Zhou +4

    cs.CVarXiv:2504.12369v32025