Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
181 to 240 of 18,815
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Ting Huang, Zhenyu Zhang, Wenyuan Huang +2
cs.CVarXiv:2607.17599v12026Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
Yucheng Zhou, Lingran Song, Jianbing Shen
cs.CLcs.AIcs.CVarXiv:2501.01377v22025OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou +17
cs.CVcs.AIcs.IRarXiv:2412.07626v22024Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, Chen Gao +2
cs.CVcs.AIcs.MMarXiv:2607.02963v12026VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Seohyun Lee, Seoung Choi, Dohwan Ko +2
cs.CVcs.AIarXiv:2607.00446v12026Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen-Ha, Jiri Matas +2
cs.CVarXiv:2004.02760v22020PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta +1
cs.CVcs.AIcs.LGarXiv:2409.18964v12024LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Yijia Xiao, Edward Sun, Tianyu Liu +1
cs.AIcs.CLcs.CVarXiv:2407.04973v12024MAIRA-2: Grounded Radiology Report Generation
Shruthi Bannur, Kenza Bouzid, Daniel C. Castro +18
cs.CLcs.CVarXiv:2406.04449v22024Solving Inverse Problems with Piecewise Linear Estimators: From Gaussian Mixture Models to Structured Sparsity
Guoshen Yu, Guillermo Sapiro, Stéphane Mallat
cs.CVarXiv:1006.3056v12010UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating
Jiehui Huang, Yuechen Zhang, Bin Xia +7
cs.CVarXiv:2606.21661v12026How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Zhe Chen, Weiyun Wang, Hao Tian +32
cs.CVarXiv:2404.16821v22024DeepGCNs: Making GCNs Go as Deep as CNNs
Guohao Li, Matthias Müller, Guocheng Qian +4
cs.CVcs.LGeess.IVarXiv:1910.06849v32019World Action Models: A Survey
Qiuhong Shen, Shihua Zhang, Yue Liao +5
cs.ROcs.CVarXiv:2606.20781v12026Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
Mu Hu, Wei Yin, Chi Zhang +7
cs.CVarXiv:2404.15506v42024The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL
Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng +4
cs.LGcs.CVarXiv:2606.19162v12026ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
Ming Li, Taojiannan Yang, Huafeng Kuang +4
cs.CVcs.AIcs.LGarXiv:2404.07987v42024GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation
Yinghao Xu, Zifan Shi, Wang Yifan +5
cs.CVarXiv:2403.14621v12024Anomaly Detection in Video Sequence with Appearance-Motion Correspondence
Trong Nguyen Nguyen, Jean Meunier
cs.CVcs.LGcs.NEarXiv:1908.06351v12019SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Dongyang Liu, Renrui Zhang, Longtian Qiu +16
cs.CVcs.AIcs.CLarXiv:2402.05935v32024InstructIR: High-Quality Image Restoration Following Human Instructions
Marcos V. Conde, Gregor Geigle, Radu Timofte
cs.CVcs.LGeess.IVarXiv:2401.16468v52024Supervised Dictionary Learning
Julien Mairal, Francis Bach, Jean Ponce +2
cs.CVarXiv:0809.3083v12008Brain-Inspired Stochastic Joint Embedding Representation Learning
Makoto Yamada, Kian Ming A. Chai, Ayoub Rhim +3
cs.CVcs.AIcs.LGarXiv:2505.11129v22025Siamese Masked Autoencoders
Agrim Gupta, Jiajun Wu, Jia Deng +1
cs.CVcs.LGarXiv:2305.14344v12023Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
Alexandre Eymaël, Renaud Vandeghen, Anthony Cioppa +3
cs.CVarXiv:2403.17823v22024Visual Representation Learning with Stochastic Frame Prediction
Huiwon Jang, Dongyoung Kim, Junsu Kim +3
cs.CVcs.AIcs.LGarXiv:2406.07398v22024Anatomy-aware 3D Human Pose Estimation with Bone-based Pose Decomposition
Tianlang Chen, Chen Fang, Xiaohui Shen +3
cs.CVarXiv:2002.10322v52020Pushing Stochastic Gradient towards Second-Order Methods -- Backpropagation Learning with Transformations in Nonlinearities
Tommi Vatanen, Tapani Raiko, Harri Valpola +1
cs.LGcs.CVstat.MLarXiv:1301.3476v32013Harnessing Large Language Models for Training-free Video Anomaly Detection
Luca Zanella, Willi Menapace, Massimiliano Mancini +2
cs.CVarXiv:2404.01014v12024Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Xiaowei Mao, Bowen Sui, Weijie Zhang +7
cs.CVcs.AIarXiv:2604.23724v42026Towards Automatic Concept-based Explanations
Amirata Ghorbani, James Wexler, James Zou +1
stat.MLcs.CVcs.LGarXiv:1902.03129v32019LaME: Learning to Think in Latent Space for Multimodal Embedding via Information Bottleneck
Peixi Wu, Biao Yang, Feipeng Ma +7
cs.CVarXiv:2606.13061v32026Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
Fabien Baradel, Christian Wolf, Julien Mille
cs.CVarXiv:1703.10106v22017A Hybrid Model for Identity Obfuscation by Face Replacement
Qianru Sun, Ayush Tewari, Weipeng Xu +3
cs.CVcs.CRarXiv:1804.04779v22018Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
Lanyun Zhu, Deyi Ji, Tianrun Chen +2
cs.CVarXiv:2510.02745v22025ViStoryBench: Comprehensive Benchmark Suite for Story Visualization
Cailin Zhuang, Ailin Huang, Yaoqi Hu +12
cs.CVarXiv:2505.24862v52025Actor-Critic Sequence Training for Image Captioning
Li Zhang, Flood Sung, Feng Liu +4
cs.CVarXiv:1706.09601v22017MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Junjie Zhou, Zheng Liu, Ze Liu +6
cs.CVcs.CLarXiv:2412.14475v12024MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
Kai Zhang, Yi Luan, Hexiang Hu +5
cs.CVcs.AIcs.CLarXiv:2403.19651v22024CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
Hao Yu, Zhuokai Zhao, Shen Yan +7
cs.CVcs.AIcs.CLarXiv:2503.19900v12025MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
Kejian Zhu, Zhuoran Jin, Hongbang Yuan +6
cs.CVcs.CLarXiv:2506.04141v22025Active Learning for Convolutional Neural Networks: A Core-Set Approach
Ozan Sener, Silvio Savarese
stat.MLcs.CVcs.LGarXiv:1708.00489v42017RepMLPNet: Hierarchical Vision MLP with Re-parameterized Locality
Xiaohan Ding, Honghao Chen, Xiangyu Zhang +2
cs.CVcs.AIcs.LGarXiv:2112.11081v22021Continual Contrastive Learning for Image Classification
Zhiwei Lin, Yongtao Wang, Hongxiang Lin
cs.CVcs.AIarXiv:2107.01776v42021WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
Marius Bock, Hilde Kuehne, Kristof Van Laerhoven +1
cs.CVcs.HCarXiv:2304.05088v42023R-Transformer: Recurrent Neural Network Enhanced Transformer
Zhiwei Wang, Yao Ma, Zitao Liu +1
cs.LGcs.CLcs.CVarXiv:1907.05572v12019TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
Christian Greisinger, Steffen Eger
cs.AIcs.CLcs.CVarXiv:2603.03072v32026AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
Jonas Belouadi, Anne Lauscher, Steffen Eger
cs.CLcs.CVarXiv:2310.00367v22023DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
Abhay Zala, Han Lin, Jaemin Cho +1
cs.CVcs.AIcs.CLarXiv:2310.12128v22023Multi-task Learning of Hierarchical Vision-Language Representation
Duy-Kien Nguyen, Takayuki Okatani
cs.CVarXiv:1812.00500v12018DeepPermNet: Visual Permutation Learning
Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian +1
cs.CVarXiv:1704.02729v12017Attention Guided Anomaly Localization in Images
Shashanka Venkataramanan, Kuan-Chuan Peng, Rajat Vikram Singh +1
cs.CVeess.IVarXiv:1911.08616v42019Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation
Chen Li, Peng Zhang, Hanyu Zhou +5
cs.CVarXiv:2608.26902v12026Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering
Yang Liu, Guanbin Li, Liang Lin
cs.CVcs.AIarXiv:2207.12647v82022OnePose: One-Shot Object Pose Estimation without CAD Models
Jiaming Sun, Zihao Wang, Siyu Zhang +4
cs.CVarXiv:2205.12257v12022Multimodal Whole Slide Foundation Model for Pathology
Tong Ding, Sophia J. Wagner, Andrew H. Song +20
eess.IVcs.AIcs.CVarXiv:2411.19666v12024ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
Yufei Xu, Qiming Zhang, Jing Zhang +1
cs.CVarXiv:2106.03348v42021On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law
Damien Teney, Kushal Kafle, Robik Shrestha +3
cs.CVcs.LGarXiv:2005.09241v12020Smoothed Dilated Convolutions for Improved Dense Prediction
Zhengyang Wang, Shuiwang Ji
cs.CVcs.LGarXiv:1808.08931v22018Binary Generative Adversarial Networks for Image Retrieval
Jingkuan Song
cs.CVarXiv:1708.04150v12017