Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,941 to 6,000 of 18,839
Improved Anomaly Detection in Crowded Scenes via Cell-based Analysis of Foreground Speed, Size and Texture
Vikas Reddy, Conrad Sanderson, Brian C. Lovell
cs.CVarXiv:1304.0886v12013ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer +7
cs.LGcs.AIcs.CVarXiv:2608.26083v12026Domain-Size Pooling in Local Descriptors: DSP-SIFT
Jingming Dong, Stefano Soatto
cs.CVarXiv:1412.8556v32014Reconstructing Hand-Object Interactions in the Wild
Zhe Cao, Ilija Radosavovic, Angjoo Kanazawa +1
cs.CVarXiv:2012.09856v22020PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology
Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro +4
cs.CVcs.AIarXiv:2608.25970v12026WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Jack Hong, Shilin Yan, Jiayin Cai +3
cs.CVcs.AIarXiv:2502.04326v32025Less Contouring, More Accuracy: Lesion-Guided ROI Deep Learning for Ovarian Ultrasound Classification
Mehran Ahmad, Ali Abbasian Ardakani, Afshin Mohammadi +3
cs.CVarXiv:2608.25965v12026Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Yijun Yang, Shenghe Zheng, Wenbo Li +8
cs.CVarXiv:2609.03729v12026RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin +2
cs.CVarXiv:2007.07542v22020Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation
Haoyu Wang, Songchun Zhang, Haoran Li +3
cs.CVcs.GRarXiv:2609.03557v12026GRIT: Teaching MLLMs to Think with Images
Yue Fan, Xuehai He, Diji Yang +6
cs.CVcs.AIcs.CLarXiv:2505.15879v22025Multiple Expert Brainstorming for Domain Adaptive Person Re-identification
Yunpeng Zhai, Qixiang Ye, Shijian Lu +3
cs.CVarXiv:2007.01546v32020Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Ye Wang, Ziheng Wang, Boshen Xu +14
cs.CVcs.AIcs.CLarXiv:2503.13377v32025Metrics reloaded: Recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Patrick Godau +71
cs.CVarXiv:2206.01653v82022Summaries:한국어Swarm-SLAM : Sparse Decentralized Collaborative Simultaneous Localization and Mapping Framework for Multi-Robot Systems
Pierre-Yves Lajoie, Giovanni Beltrame
cs.ROcs.CVarXiv:2301.06230v32023Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
Yuxiang Lai, Jike Zhong, Ming Li +4
cs.CVarXiv:2503.13939v52025Naive-Deep Face Recognition: Touching the Limit of LFW Benchmark or Not?
Erjin Zhou, Zhimin Cao, Qi Yin
cs.CVarXiv:1501.04690v12015Human Mesh Recovery from Monocular Images via a Skeleton-disentangled Representation
Sun Yu, Ye Yun, Liu Wu +3
cs.CVarXiv:1908.07172v22019Compact 3D Scene Representation via Self-Organizing Gaussian Grids
Wieland Morgenstern, Florian Barthel, Anna Hilsmann +1
cs.CVarXiv:2312.13299v22023Deep Self-Taught Learning for Weakly Supervised Object Localization
Zequn Jie, Yunchao Wei, Xiaojie Jin +2
cs.CVarXiv:1704.05188v22017MMSearch-R1: Incentivizing LMMs to Search
Jinming Wu, Zihao Deng, Wei Li +5
cs.CVcs.CLarXiv:2506.20670v12025Denoising Hyperspectral Image with Non-i.i.d. Noise Structure
Yang Chen, Xiangyong Cao, Qian Zhao +2
cs.CVarXiv:1702.00098v12017Functional Adversarial Attacks
Cassidy Laidlaw, Soheil Feizi
cs.LGcs.CVarXiv:1906.00001v22019KiU-Net: Towards Accurate Segmentation of Biomedical Images using Over-complete Representations
Jeya Maria Jose, Vishwanath Sindagi, Ilker Hacihaliloglu +1
eess.IVcs.CVarXiv:2006.04878v22020Rethinking the Heatmap Regression for Bottom-up Human Pose Estimation
Zhengxiong Luo, Zhicheng Wang, Yan Huang +2
cs.CVarXiv:2012.15175v42020Quaternion Convolutional Neural Networks
Xuanyu Zhu, Yi Xu, Hongteng Xu +1
cs.CVarXiv:1903.00658v12019Unpaired Deep Image Deraining Using Dual Contrastive Learning
Xiang Chen, Jinshan Pan, Kui Jiang +5
cs.CVarXiv:2109.02973v42021Epona: Autoregressive Diffusion World Model for Autonomous Driving
Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu +9
cs.CVarXiv:2506.24113v12025Re-IQA: Unsupervised Learning for Image Quality Assessment in the Wild
Avinab Saha, Sandeep Mishra, Alan C. Bovik
cs.CVcs.LGcs.MMarXiv:2304.00451v22023Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
Zeqiang Lai, Yunfei Zhao, Haolin Liu +23
cs.CVcs.AIarXiv:2506.16504v12025Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
Zhaochen Su, Peng Xia, Hangyu Guo +12
cs.CVarXiv:2506.23918v32025SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng +10
cs.CVarXiv:2309.17448v32023InstructDiffusion: A Generalist Modeling Interface for Vision Tasks
Zigang Geng, Binxin Yang, Tiankai Hang +8
cs.CVarXiv:2309.03895v12023EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning
Chenxin Xu, Robby T. Tan, Yuhong Tan +4
cs.CVcs.MAarXiv:2303.10876v22023DeepEyesV2: Toward Agentic Multimodal Model
Jack Hong, Chenxiao Zhao, ChengLin Zhu +3
cs.CVcs.AIarXiv:2511.05271v42025FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
Haonan Qiu, Menghan Xia, Yong Zhang +4
cs.CVarXiv:2310.15169v32023Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
Duo Zheng, Shijia Huang, Yanyang Li +1
cs.CVcs.AIarXiv:2505.24625v32025Neural-Collapse-guided Task-Free Continual Anomaly Detection
Xiaotong Kong, Chaoyang Song, Ziai Zhou +3
cs.CVarXiv:2609.03406v12026Improved Mean Flows: On the Challenges of Fastforward Generative Models
Zhengyang Geng, Yiyang Lu, Zongze Wu +3
cs.CVcs.LGarXiv:2512.02012v22025GP-VTON: Towards General Purpose Virtual Try-on via Collaborative Local-Flow Global-Parsing Learning
Zhenyu Xie, Zaiyu Huang, Xin Dong +5
cs.CVarXiv:2303.13756v12023Sonata: Self-Supervised Learning of Reliable Point Representations
Xiaoyang Wu, Daniel DeTone, Duncan Frost +7
cs.CVarXiv:2503.16429v12025VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
Hritik Bansal, Clark Peng, Yonatan Bitton +3
cs.CVarXiv:2503.06800v12025Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation
Yingmao Miao, Pengfei Zhang, Chaoran Xu +5
cs.CVarXiv:2609.03673v12026Osprey: Pixel Understanding with Visual Instruction Tuning
Yuqian Yuan, Wentong Li, Jian Liu +5
cs.CVarXiv:2312.10032v42023VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
Rui Meng, Ziyan Jiang, Ye Liu +10
cs.CVcs.CLarXiv:2507.04590v12025MambaIRv2: Attentive State Space Restoration
Hang Guo, Yong Guo, Yaohua Zha +5
eess.IVcs.CVcs.LGarXiv:2411.15269v22024SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Xiyao Wang, Zhengyuan Yang, Chao Feng +6
cs.CVarXiv:2504.07934v32025MonoPerfCap: Human Performance Capture from Monocular Video
Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer +4
cs.CVcs.GRarXiv:1708.02136v22017A multicenter benchmark and clinically structured metric for coronary CTA report generation
Zhiyu Ye, Yue Sun, Limiao Zou +7
cs.CVarXiv:2609.00909v12026CMRVision: A Foundation Model for Cardiac MR Image Analysis
Athira J. Jacob, Puneet Sharma, Daniel Rueckert
cs.CVarXiv:2609.01308v12026Pose Estimation for Non-Cooperative Spacecraft Rendezvous Using Convolutional Neural Networks
Sumant Sharma, Connor Beierle, Simone D'Amico
cs.CVarXiv:1809.07238v12018Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Tilman Räuker, Anson Ho, Stephen Casper +1
cs.LGcs.AIcs.CLarXiv:2207.13243v62022TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval
Yuqi Liu, Pengfei Xiong, Luhui Xu +2
cs.CVarXiv:2207.07852v12022Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
NVIDIA, :, Alisson Azzolini +51
cs.AIcs.CVcs.LGarXiv:2503.15558v32025Content-Aware Unsupervised Deep Homography Estimation
Jirong Zhang, Chuan Wang, Shuaicheng Liu +5
cs.CVarXiv:1909.05983v22019Multi-Adapter RGBT Tracking
Chenglong Li, Andong Lu, Aihua Zheng +2
cs.CVarXiv:1907.07485v12019Illiterate DALL-E Learns to Compose
Gautam Singh, Fei Deng, Sungjin Ahn
cs.CVcs.LGarXiv:2110.11405v32021Making Better Mistakes: Leveraging Class Hierarchies with Deep Networks
Luca Bertinetto, Romain Mueller, Konstantinos Tertikas +2
cs.CVcs.LGarXiv:1912.09393v22019PromptIR: Prompting for All-in-One Blind Image Restoration
Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan +1
cs.CVarXiv:2306.13090v12023WorldMem: Long-term Consistent World Simulation with Memory
Zeqi Xiao, Yushi Lan, Yifan Zhou +4
cs.CVarXiv:2504.12369v32025