Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,421 to 3,480 of 18,781
Video Transformers: A Survey
Javier Selva, Anders S. Johansen, Sergio Escalera +3
cs.CVarXiv:2201.05991v32022Contextual Object Detection with Multimodal Large Language Models
Yuhang Zang, Wei Li, Jun Han +2
cs.CVcs.AIarXiv:2305.18279v22023A Review of Uncertainty Estimation and its Application in Medical Imaging
Ke Zou, Zhihao Chen, Xuedong Yuan +3
eess.IVcs.CVarXiv:2302.08119v32023Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges
Edgar Galván, Peter Mooney
cs.NEcs.CVcs.LGarXiv:2006.05415v12020Combining Optimal Control and Learning for Visual Navigation in Novel Environments
Somil Bansal, Varun Tolani, Saurabh Gupta +2
cs.ROcs.AIcs.CVarXiv:1903.02531v22019HRDNet: High-resolution Detection Network for Small Objects
Ziming Liu, Guangyu Gao, Lin Sun +1
cs.CVarXiv:2006.07607v12020pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
Zhanpeng Luo, Ce Zhang, Silong Yong +6
cs.CVarXiv:2603.00905v12026Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta +7
cs.ROcs.CVcs.LGarXiv:2409.16283v12024Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
Raphi Kang, Hongqiao Chen, Georgia Gkioxari +1
cs.CVarXiv:2601.12626v12026GenSmoke-GS: A Multi-Stage Method for Novel View Synthesis from Smoke-Degraded Images Using a Generative Model
Qida Cao, Xinyuan Hu, Changyue Shi +3
cs.CVarXiv:2604.03039v22026One-Class Convolutional Neural Network
Poojan Oza, Vishal M. Patel
cs.CVarXiv:1901.08688v12019DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
Chenxu Dang, Sining Ang, Yongkang Li +7
cs.CVarXiv:2602.14577v12026From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
Davide Testa, Hugh Mee Wong, Alessandro Lenci +2
cs.CLcs.CVarXiv:2609.05149v12026Domain Adaptive Relational Reasoning for 3D Multi-Organ Segmentation
Shuhao Fu, Yongyi Lu, Yan Wang +4
cs.CVarXiv:2005.09120v22020Retinal OCTA Phenotyping with LLM Reporting for Alzheimer's Disease
Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir
cs.CVcs.CLarXiv:2609.04689v12026Latent-Aligned Reasoning for Multimodal Recommendation
Jiarui Jin, Anyang Ji
cs.IRcs.CLcs.CVarXiv:2609.04645v12026To Create What You Tell: Generating Videos from Captions
Yingwei Pan, Zhaofan Qiu, Ting Yao +2
cs.CVarXiv:1804.08264v12018COMBOOD: A Semiparametric Approach for Detecting Out-of-distribution Data for Image Classification
Magesh Rajasekaran, Md Saiful Islam Sajol, Frej Berglind +2
cs.CVarXiv:2602.07042v12026ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
Xialin He, Sirui Xu, Xinyao Li +4
cs.ROcs.CVarXiv:2603.03279v12026Automatic Lung Cancer Prediction from Chest X-ray Images Using Deep Learning Approach
Worawate Ausawalaithong, Sanparith Marukatat, Arjaree Thirach +1
eess.IVcs.CVarXiv:1808.10858v12018IPOD: Intensive Point-based Object Detector for Point Cloud
Zetong Yang, Yanan Sun, Shu Liu +2
cs.CVarXiv:1812.05276v12018NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results
Shuhong Liu, Chenyu Bao, Ziteng Cui +103
cs.CVarXiv:2604.04135v22026A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking
Chang Liu, Yinpeng Dong, Wenzhao Xiang +7
cs.CVarXiv:2302.14301v12023Coupled Convolutional Neural Network with Adaptive Response Function Learning for Unsupervised Hyperspectral Super-Resolution
Ke Zheng, Lianru Gao, Wenzhi Liao +4
eess.IVcs.CVarXiv:2007.14007v12020FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
Peng Li, Zihan Zhuang, Yangfan Gao +16
cs.ROcs.CLcs.CVarXiv:2601.12799v12026Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
Erik Thiringer, Fredrik K. Gustafsson, Kajsa Ledesma Eriksson +1
eess.IVcs.CVcs.LGarXiv:2601.04163v12026CLIP-Guided Data Augmentation for Night-Time Image Dehazing
Xining Ge, Weijun Yuan, Gengjia Chang +2
cs.CVarXiv:2604.05500v22026Attend to You: Personalized Image Captioning with Context Sequence Memory Networks
Cesc Chunseong Park, Byeongchang Kim, Gunhee Kim
cs.CVcs.CLarXiv:1704.06485v22017WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
Hao Bai, Alexey Taymanov, Tong Zhang +2
cs.LGcs.CVarXiv:2601.02439v62026Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Raz Lapid, Ron Langberg, Moshe Sipper
cs.CLcs.CVcs.NEarXiv:2309.01446v42023Multi-Sensor Data Fusion for Cloud Removal in Global and All-Season Sentinel-2 Imagery
Patrick Ebel, Andrea Meraner, Michael Schmitt +1
eess.IVcs.CVarXiv:2009.07683v12020Reflection-aware Generative Novel View Synthesis
GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh
cs.CVcs.AIarXiv:2609.05382v12026Stable Low-rank Tensor Decomposition for Compression of Convolutional Neural Network
Anh-Huy Phan, Konstantin Sobolev, Konstantin Sozykin +6
cs.CVarXiv:2008.05441v12020Dual-Branch Remote Sensing Infrared Image Super-Resolution
Xining Ge, Gengjia Chang, Weijun Yuan +6
cs.CVarXiv:2604.10112v22026CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction
Lixin Yang, Xinyu Zhan, Kailin Li +3
cs.CVarXiv:2012.00924v42020ZeroSense:How Vision matters in Long Context Compression
Yonghan Gao, Zehong Chen, Lijian Xu +3
cs.CVarXiv:2603.11846v12026FastDraw: Addressing the Long Tail of Lane Detection by Adapting a Sequential Prediction Network
Jonah Philion
cs.CVarXiv:1905.04354v22019Omni2Sound: Towards Unified Video-Text-to-Audio Generation
Yusheng Dai, Zehua Chen, Yuxuan Jiang +4
cs.SDcs.CVcs.MMarXiv:2601.02731v32026Group Component Analysis for Multiblock Data: Common and Individual Feature Extraction
Guoxu Zhou, Andrzej Cichocki, Yu Zhang +1
cs.CVcs.LGarXiv:1212.3913v42012Robo3D: Towards Robust and Reliable 3D Perception against Corruptions
Lingdong Kong, Youquan Liu, Xin Li +6
cs.CVcs.ROarXiv:2303.17597v42023M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation
Vishal Chudasama, Purbayan Kar, Ashish Gudmalwar +3
cs.CVcs.SDeess.ASarXiv:2206.02187v12022Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions
Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith +6
cs.CVcs.AIcs.LGarXiv:2609.05334v12026From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Niu Lian, Yuting Wang, Hanshu Yao +5
cs.CVcs.AIcs.CLarXiv:2603.01455v32026Context-aware Human Motion Prediction
Enric Corona, Albert Pumarola, Guillem Alenyà +1
cs.CVarXiv:1904.03419v32019RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Zhenxuan Fan, Bo Zhang, Yutong Lin +9
cs.ROcs.AIcs.CVarXiv:2609.05324v12026What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
Vivek Chavan, Pengtao Xie, Yahuan Shi +3
cs.ROcs.AIcs.CVarXiv:2609.05376v12026Beyond Model Design: Data-Centric Training and Self-Ensemble for Gaussian Color Image Denoising
Gengjia Chang, Xining Ge, Weijun Yuan +4
cs.CVarXiv:2604.11468v22026Training-Free Model Ensemble for Single-Image Super-Resolution via Strong-Branch Compensation
Gengjia Chang, Xining Ge, Weijun Yuan +4
cs.CVarXiv:2604.11564v22026SiamMOT: Siamese Multi-Object Tracking
Bing Shuai, Andrew Berneshawi, Xinyu Li +2
cs.CVarXiv:2105.11595v12021OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir +8
cs.CVarXiv:2602.17665v42026Learning Deep Bilinear Transformation for Fine-grained Image Representation
Heliang Zheng, Jianlong Fu, Zheng-Jun Zha +1
cs.CVarXiv:1911.03621v12019Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Wenjing Wang, Huan Yang, Zixi Tuo +4
cs.CVarXiv:2305.10874v42023Invertible Denoising Network: A Light Solution for Real Noise Removal
Yang Liu, Zhenyue Qin, Saeed Anwar +4
eess.IVcs.CVarXiv:2104.10546v12021CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion
Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas +7
cs.CVarXiv:2210.10716v22022S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
Haodong Yan, Zhide Zhong, Jiaguan Zhu +10
cs.CVcs.ROarXiv:2603.16195v22026Variational Autoencoders Pursue PCA Directions (by Accident)
Michal Rolinek, Dominik Zietlow, Georg Martius
cs.LGcs.CVstat.MLarXiv:1812.06775v22018FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
Ruiteng Zhao, Wenshuo Wang, Yicheng Ma +4
cs.ROcs.CVarXiv:2602.02142v22026Learning to drive from a world on rails
Dian Chen, Vladlen Koltun, Philipp Krähenbühl
cs.ROcs.CVcs.LGarXiv:2105.00636v32021Graph Degree Linkage: Agglomerative Clustering on a Directed Graph
Wei Zhang, Xiaogang Wang, Deli Zhao +1
cs.CVcs.SIstat.MLarXiv:1208.5092v12012Zero-Shot Visual Recognition via Bidirectional Latent Embedding
Qian Wang, Ke Chen
cs.CVarXiv:1607.02104v42016