Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,421 to 3,480 of 18,781

  1. Video Transformers: A Survey

    Javier Selva, Anders S. Johansen, Sergio Escalera +3

    cs.CVarXiv:2201.05991v32022
  2. Contextual Object Detection with Multimodal Large Language Models

    Yuhang Zang, Wei Li, Jun Han +2

    cs.CVcs.AIarXiv:2305.18279v22023
  3. A Review of Uncertainty Estimation and its Application in Medical Imaging

    Ke Zou, Zhihao Chen, Xuedong Yuan +3

    eess.IVcs.CVarXiv:2302.08119v32023
  4. Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges

    Edgar Galván, Peter Mooney

    cs.NEcs.CVcs.LGarXiv:2006.05415v12020
  5. Combining Optimal Control and Learning for Visual Navigation in Novel Environments

    Somil Bansal, Varun Tolani, Saurabh Gupta +2

    cs.ROcs.AIcs.CVarXiv:1903.02531v22019
  6. HRDNet: High-resolution Detection Network for Small Objects

    Ziming Liu, Guangyu Gao, Lin Sun +1

    cs.CVarXiv:2006.07607v12020
  7. pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning

    Zhanpeng Luo, Ce Zhang, Silong Yong +6

    cs.CVarXiv:2603.00905v12026
  8. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

    Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta +7

    cs.ROcs.CVcs.LGarXiv:2409.16283v12024
  9. Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models

    Raphi Kang, Hongqiao Chen, Georgia Gkioxari +1

    cs.CVarXiv:2601.12626v12026
  10. GenSmoke-GS: A Multi-Stage Method for Novel View Synthesis from Smoke-Degraded Images Using a Generative Model

    Qida Cao, Xinyuan Hu, Changyue Shi +3

    cs.CVarXiv:2604.03039v22026
  11. One-Class Convolutional Neural Network

    Poojan Oza, Vishal M. Patel

    cs.CVarXiv:1901.08688v12019
  12. DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving

    Chenxu Dang, Sining Ang, Yongkang Li +7

    cs.CVarXiv:2602.14577v12026
  13. From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

    Davide Testa, Hugh Mee Wong, Alessandro Lenci +2

    cs.CLcs.CVarXiv:2609.05149v12026
  14. Domain Adaptive Relational Reasoning for 3D Multi-Organ Segmentation

    Shuhao Fu, Yongyi Lu, Yan Wang +4

    cs.CVarXiv:2005.09120v22020
  15. Retinal OCTA Phenotyping with LLM Reporting for Alzheimer's Disease

    Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir

    cs.CVcs.CLarXiv:2609.04689v12026
  16. Latent-Aligned Reasoning for Multimodal Recommendation

    Jiarui Jin, Anyang Ji

    cs.IRcs.CLcs.CVarXiv:2609.04645v12026
  17. To Create What You Tell: Generating Videos from Captions

    Yingwei Pan, Zhaofan Qiu, Ting Yao +2

    cs.CVarXiv:1804.08264v12018
  18. COMBOOD: A Semiparametric Approach for Detecting Out-of-distribution Data for Image Classification

    Magesh Rajasekaran, Md Saiful Islam Sajol, Frej Berglind +2

    cs.CVarXiv:2602.07042v12026
  19. ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

    Xialin He, Sirui Xu, Xinyao Li +4

    cs.ROcs.CVarXiv:2603.03279v12026
  20. Automatic Lung Cancer Prediction from Chest X-ray Images Using Deep Learning Approach

    Worawate Ausawalaithong, Sanparith Marukatat, Arjaree Thirach +1

    eess.IVcs.CVarXiv:1808.10858v12018
  21. IPOD: Intensive Point-based Object Detector for Point Cloud

    Zetong Yang, Yanan Sun, Shu Liu +2

    cs.CVarXiv:1812.05276v12018
  22. NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

    Shuhong Liu, Chenyu Bao, Ziteng Cui +103

    cs.CVarXiv:2604.04135v22026
  23. A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking

    Chang Liu, Yinpeng Dong, Wenzhao Xiang +7

    cs.CVarXiv:2302.14301v12023
  24. Coupled Convolutional Neural Network with Adaptive Response Function Learning for Unsupervised Hyperspectral Super-Resolution

    Ke Zheng, Lianru Gao, Wenzhi Liao +4

    eess.IVcs.CVarXiv:2007.14007v12020
  25. FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions

    Peng Li, Zihan Zhuang, Yangfan Gao +16

    cs.ROcs.CLcs.CVarXiv:2601.12799v12026
  26. Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models

    Erik Thiringer, Fredrik K. Gustafsson, Kajsa Ledesma Eriksson +1

    eess.IVcs.CVcs.LGarXiv:2601.04163v12026
  27. CLIP-Guided Data Augmentation for Night-Time Image Dehazing

    Xining Ge, Weijun Yuan, Gengjia Chang +2

    cs.CVarXiv:2604.05500v22026
  28. Attend to You: Personalized Image Captioning with Context Sequence Memory Networks

    Cesc Chunseong Park, Byeongchang Kim, Gunhee Kim

    cs.CVcs.CLarXiv:1704.06485v22017
  29. WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks

    Hao Bai, Alexey Taymanov, Tong Zhang +2

    cs.LGcs.CVarXiv:2601.02439v62026
  30. Open Sesame! Universal Black Box Jailbreaking of Large Language Models

    Raz Lapid, Ron Langberg, Moshe Sipper

    cs.CLcs.CVcs.NEarXiv:2309.01446v42023
  31. Multi-Sensor Data Fusion for Cloud Removal in Global and All-Season Sentinel-2 Imagery

    Patrick Ebel, Andrea Meraner, Michael Schmitt +1

    eess.IVcs.CVarXiv:2009.07683v12020
  32. Reflection-aware Generative Novel View Synthesis

    GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

    cs.CVcs.AIarXiv:2609.05382v12026
  33. Stable Low-rank Tensor Decomposition for Compression of Convolutional Neural Network

    Anh-Huy Phan, Konstantin Sobolev, Konstantin Sozykin +6

    cs.CVarXiv:2008.05441v12020
  34. Dual-Branch Remote Sensing Infrared Image Super-Resolution

    Xining Ge, Gengjia Chang, Weijun Yuan +6

    cs.CVarXiv:2604.10112v22026
  35. CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction

    Lixin Yang, Xinyu Zhan, Kailin Li +3

    cs.CVarXiv:2012.00924v42020
  36. ZeroSense:How Vision matters in Long Context Compression

    Yonghan Gao, Zehong Chen, Lijian Xu +3

    cs.CVarXiv:2603.11846v12026
  37. FastDraw: Addressing the Long Tail of Lane Detection by Adapting a Sequential Prediction Network

    Jonah Philion

    cs.CVarXiv:1905.04354v22019
  38. Omni2Sound: Towards Unified Video-Text-to-Audio Generation

    Yusheng Dai, Zehua Chen, Yuxuan Jiang +4

    cs.SDcs.CVcs.MMarXiv:2601.02731v32026
  39. Group Component Analysis for Multiblock Data: Common and Individual Feature Extraction

    Guoxu Zhou, Andrzej Cichocki, Yu Zhang +1

    cs.CVcs.LGarXiv:1212.3913v42012
  40. Robo3D: Towards Robust and Reliable 3D Perception against Corruptions

    Lingdong Kong, Youquan Liu, Xin Li +6

    cs.CVcs.ROarXiv:2303.17597v42023
  41. M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation

    Vishal Chudasama, Purbayan Kar, Ashish Gudmalwar +3

    cs.CVcs.SDeess.ASarXiv:2206.02187v12022
  42. Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

    Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith +6

    cs.CVcs.AIcs.LGarXiv:2609.05334v12026
  43. From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

    Niu Lian, Yuting Wang, Hanshu Yao +5

    cs.CVcs.AIcs.CLarXiv:2603.01455v32026
  44. Context-aware Human Motion Prediction

    Enric Corona, Albert Pumarola, Guillem Alenyà +1

    cs.CVarXiv:1904.03419v32019
  45. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

    Zhenxuan Fan, Bo Zhang, Yutong Lin +9

    cs.ROcs.AIcs.CVarXiv:2609.05324v12026
  46. What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

    Vivek Chavan, Pengtao Xie, Yahuan Shi +3

    cs.ROcs.AIcs.CVarXiv:2609.05376v12026
  47. Beyond Model Design: Data-Centric Training and Self-Ensemble for Gaussian Color Image Denoising

    Gengjia Chang, Xining Ge, Weijun Yuan +4

    cs.CVarXiv:2604.11468v22026
  48. Training-Free Model Ensemble for Single-Image Super-Resolution via Strong-Branch Compensation

    Gengjia Chang, Xining Ge, Weijun Yuan +4

    cs.CVarXiv:2604.11564v22026
  49. SiamMOT: Siamese Multi-Object Tracking

    Bing Shuai, Andrew Berneshawi, Xinyu Li +2

    cs.CVarXiv:2105.11595v12021
  50. OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

    Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir +8

    cs.CVarXiv:2602.17665v42026
  51. Learning Deep Bilinear Transformation for Fine-grained Image Representation

    Heliang Zheng, Jianlong Fu, Zheng-Jun Zha +1

    cs.CVarXiv:1911.03621v12019
  52. Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

    Wenjing Wang, Huan Yang, Zixi Tuo +4

    cs.CVarXiv:2305.10874v42023
  53. Invertible Denoising Network: A Light Solution for Real Noise Removal

    Yang Liu, Zhenyue Qin, Saeed Anwar +4

    eess.IVcs.CVarXiv:2104.10546v12021
  54. CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion

    Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas +7

    cs.CVarXiv:2210.10716v22022
  55. S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

    Haodong Yan, Zhide Zhong, Jiaguan Zhu +10

    cs.CVcs.ROarXiv:2603.16195v22026
  56. Variational Autoencoders Pursue PCA Directions (by Accident)

    Michal Rolinek, Dominik Zietlow, Georg Martius

    cs.LGcs.CVstat.MLarXiv:1812.06775v22018
  57. FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

    Ruiteng Zhao, Wenshuo Wang, Yicheng Ma +4

    cs.ROcs.CVarXiv:2602.02142v22026
  58. Learning to drive from a world on rails

    Dian Chen, Vladlen Koltun, Philipp Krähenbühl

    cs.ROcs.CVcs.LGarXiv:2105.00636v32021
  59. Graph Degree Linkage: Agglomerative Clustering on a Directed Graph

    Wei Zhang, Xiaogang Wang, Deli Zhao +1

    cs.CVcs.SIstat.MLarXiv:1208.5092v12012
  60. Zero-Shot Visual Recognition via Bidirectional Latent Embedding

    Qian Wang, Ke Chen

    cs.CVarXiv:1607.02104v42016