Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,281 to 5,340 of 18,866

  1. A Very Big Video Reasoning Suite

    Maijunxian Wang, Ruisi Wang, Juyi Lin +53

    cs.CVcs.AIcs.LGarXiv:2602.20159v22026
  2. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark

    Chang Xu, Jinwang Wang, Wen Yang +3

    cs.CVarXiv:2206.13996v12022
  3. PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

    Selim Kuzucu, Alessio Tonioni, Vasile Lup +3

    cs.CVcs.AIcs.CLarXiv:2605.30126v12026
  4. Cosmos World Foundation Model Platform for Physical AI

    NVIDIA, :, Niket Agarwal +76

    cs.CVcs.AIcs.LGarXiv:2501.03575v32025
  5. Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

    Wenli Xiao, Haotian Lin, Andy Peng +9

    cs.CVcs.ROarXiv:2511.00091v12025
  6. Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction

    Alfonso López-Ruiz, Diego Royo

    cs.DCcs.CVarXiv:2608.28183v12026
  7. Parallel Decoding Distillation for Fast Image and Video Generation

    Neta Shaul, Chao Liu, Arash Vahdat +1

    cs.CVcs.LGarXiv:2607.26004v12026
  8. RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

    Tianxing Chen, Yue Chen, Zixuan Li +41

    cs.ROcs.AIcs.CVarXiv:2607.04434v32026
  9. Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

    Chiara Tappermann, Steffen Renisch, Lars Ole Schwen +3

    cs.CVcs.AIarXiv:2608.16725v12026
  10. Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

    Alex Fang, Gabriel Ilharco, Mitchell Wortsman +4

    cs.CVcs.CLcs.LGarXiv:2205.01397v22022
  11. Deep learning for cardiac image segmentation: A review

    Chen Chen, Chen Qin, Huaqi Qiu +4

    eess.IVcs.CVcs.LGarXiv:1911.03723v12019
  12. EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

    Ryan Punamiya, Simar Kareer, Zeyi Liu +37

    cs.ROcs.CVarXiv:2604.07607v22026
  13. The Evolution of First Person Vision Methods: A Survey

    Alejandro Betancourt, Pietro Morerio, Carlo S. Regazzoni +1

    cs.CVarXiv:1409.1484v32014
  14. Predicting Citywide Crowd Flows in Irregular Regions Using Multi-View Graph Convolutional Networks

    Junkai Sun, Junbo Zhang, Qiaofei Li +3

    cs.CVcs.LGarXiv:1903.07789v22019
  15. Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

    Orcun Cetintas, Guillem Brasó, Tim Meinhardt +1

    cs.CVcs.AIarXiv:2609.00924v12026
  16. Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

    Teresa DiMeola, Charles Walter, Hong Xiao

    cs.CVcs.AIarXiv:2609.00628v12026
  17. VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents

    Ryota Tanaka, Taichi Iki, Taku Hasegawa +3

    cs.CLcs.AIcs.CVarXiv:2504.09795v12025
  18. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

    Chengzhi Liu, Zhongxing Xu, Qingyue Wei +5

    cs.CLcs.AIcs.CVarXiv:2505.21523v32025
  19. Automated and Interpretable Patient ECG Profiles for Disease Detection, Tracking, and Discovery

    Geoffrey H. Tison, Jeffrey Zhang, Francesca N. Delling +1

    cs.CVarXiv:1807.02569v12018
  20. Unpaired Image Captioning via Scene Graph Alignments

    Jiuxiang Gu, Shafiq Joty, Jianfei Cai +3

    cs.CVarXiv:1903.10658v42019
  21. From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network

    Yuxin Wang, Hongtao Xie, Shancheng Fang +3

    cs.CVarXiv:2108.09661v12021
  22. On the Adversarial Robustness of Vision Transformers

    Rulin Shao, Zhouxing Shi, Jinfeng Yi +2

    cs.CVcs.AIcs.LGarXiv:2103.15670v32021
  23. Towards Efficient Model Compression via Learned Global Ranking

    Ting-Wu Chin, Ruizhou Ding, Cha Zhang +1

    cs.CVcs.LGarXiv:1904.12368v22019
  24. Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

    Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot +2

    cs.CVcs.AIcs.LGarXiv:2504.02821v32025
  25. PointVLA: Injecting the 3D World into Vision-Language-Action Models

    Chengmeng Li, Junjie Wen, Yan Peng +3

    cs.ROcs.CVcs.LGarXiv:2503.07511v12025
  26. Stack-Captioning: Coarse-to-Fine Learning for Image Captioning

    Jiuxiang Gu, Jianfei Cai, Gang Wang +1

    cs.CVarXiv:1709.03376v32017
  27. Benchmarking and Error Diagnosis in Multi-Instance Pose Estimation

    Matteo Ruggero Ronchi, Pietro Perona

    cs.CVarXiv:1707.05388v22017
  28. High-Resolution Semantic Labeling with Convolutional Neural Networks

    Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat +1

    cs.CVarXiv:1611.01962v12016
  29. Assessing bikeability with street view imagery and computer vision

    Koichi Ito, Filip Biljecki

    cs.CVarXiv:2105.08499v32021
  30. StreamScout: Learning When to Look Deeper for Streaming Video Understanding

    Ce Zhang, Jing Bi, Jinxi He +9

    cs.CVarXiv:2609.00291v12026
  31. Correlation-Aware Deep Tracking

    Fei Xie, Chunyu Wang, Guangting Wang +3

    cs.CVarXiv:2203.01666v12022
  32. NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation

    Xiangyan Liu, Jinjie Ni, Zijian Wu +5

    cs.CVarXiv:2504.13055v42025
  33. ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

    Fanbin Lu, Zhisheng Zhong, Shu Liu +2

    cs.CVarXiv:2505.16282v12025
  34. STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

    Yun Li, Yiming Zhang, Tao Lin +4

    cs.CVarXiv:2503.23765v62025
  35. BACON: Band-limited Coordinate Networks for Multiscale Scene Representation

    David B. Lindell, Dave Van Veen, Jeong Joon Park +1

    cs.CVcs.GRcs.LGarXiv:2112.04645v22021
  36. Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

    Jungseob Lee, Seongtae Hong, Dongyub Jude Lee +4

    cs.AIcs.CLcs.CVarXiv:2609.00355v12026
  37. LAS-AT: Adversarial Training with Learnable Attack Strategy

    Xiaojun Jia, Yong Zhang, Baoyuan Wu +3

    cs.CVarXiv:2203.06616v12022
  38. Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable Prototypes

    Jon Donnelly, Alina Jade Barnett, Chaofan Chen

    cs.CVcs.AIcs.LGarXiv:2111.15000v32021
  39. FedGAN: Federated Generative Adversarial Networks for Distributed Data

    Mohammad Rasouli, Tao Sun, Ram Rajagopal

    cs.LGcs.CVcs.MAarXiv:2006.07228v22020
  40. HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition

    Hossein Rahmani, Arif Mahmood, Du Q. Huynh +1

    cs.CVarXiv:1408.3809v42014
  41. PQ-NET: A Generative Part Seq2Seq Network for 3D Shapes

    Rundi Wu, Yixin Zhuang, Kai Xu +2

    cs.CVcs.GRcs.LGarXiv:1911.10949v32019
  42. Vision-Language Models Do Not Understand Negation

    Kumail Alhamoud, Shaden Alshammari, Yonglong Tian +4

    cs.CVcs.CLarXiv:2501.09425v22025
  43. LongCat-Video Technical Report

    Meituan LongCat Team, Xunliang Cai, Qilong Huang +8

    cs.CVarXiv:2510.22200v22025
  44. Histogram of Oriented Principal Components for Cross-View Action Recognition

    Hossein Rahmani, Arif Mahmood, Du Huynh +1

    cs.CVarXiv:1409.6813v22014
  45. Dynamic Label Graph Matching for Unsupervised Video Re-Identification

    Mang Ye, Andy J Ma, Liang Zheng +2

    cs.CVarXiv:1709.09297v12017
  46. Edge-Assisted Multi-Robot Visual-Inertial SLAM with Efficient Communication

    Xin Liu, Shuhuan Wen, Jing Zhao +2

    cs.ROcs.CVcs.MAarXiv:2603.11085v12026
  47. Virtual Multi-view Fusion for 3D Semantic Segmentation

    Abhijit Kundu, Xiaoqi Yin, Alireza Fathi +4

    cs.CVeess.IVarXiv:2007.13138v12020
  48. DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

    Jingxiang Sun, Bo Zhang, Ruizhi Shao +4

    cs.CVcs.CGarXiv:2310.16818v22023
  49. RayZer: A Self-supervised Large View Synthesis Model

    Hanwen Jiang, Hao Tan, Peng Wang +8

    cs.CVarXiv:2505.00702v12025
  50. MultiGait: A Multi-Sensor Multi-Perspective Multi-Session Biometric Inference Benchmark and its Dataset

    Julian Todt, Felix Morsbach, Philip Dissert +1

    cs.CRcs.CVarXiv:2609.01036v12026
  51. Object-aware Aggregation with Bidirectional Temporal Graph for Video Captioning

    Junchao Zhang, Yuxin Peng

    cs.CVarXiv:1906.04375v12019
  52. On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios

    Zhe Shen, Liyuan Lou, Yifei Yu +4

    cs.CVarXiv:2609.00923v12026
  53. Learning Auxiliary Monocular Contexts Helps Monocular 3D Object Detection

    Xianpeng Liu, Nan Xue, Tianfu Wu

    cs.CVarXiv:2112.04628v12021
  54. Consensus-Aware Visual-Semantic Embedding for Image-Text Matching

    Haoran Wang, Ying Zhang, Zhong Ji +2

    cs.CVarXiv:2007.08883v22020
  55. You Only Hypothesize Once: Point Cloud Registration with Rotation-equivariant Descriptors

    Haiping Wang, Yuan Liu, Zhen Dong +1

    cs.CVarXiv:2109.00182v22021
  56. DADA: Driver Attention Prediction in Driving Accident Scenarios

    Jianwu Fang, Dingxin Yan, Jiahuan Qiao +2

    cs.CVarXiv:1912.12148v22019
  57. A Controlled Evaluation of Model Rankings and Input Reliance in Surface Water Segmentation

    Kittipat Phunjanna, Kristóf Karacs, Chayut Ngamkhanong

    cs.CVarXiv:2608.30895v12026
  58. Edge Deep Learning in Computer Vision and Medical Diagnostics: A Comprehensive Survey

    Yiwen Xu, Tariq M. Khan, Yang Song +1

    cs.CVcs.AIarXiv:2605.06714v12026
  59. Stable and robust sampling strategies for compressive imaging

    Felix Krahmer, Rachel Ward

    cs.CVcs.ITmath.NAarXiv:1210.2380v32012
  60. Interpreting Chest X-rays via CNNs that Exploit Hierarchical Disease Dependencies and Uncertainty Labels

    Hieu H. Pham, Tung T. Le, Dat T. Ngo +2

    cs.CVarXiv:2005.12734v12020