Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,341 to 5,400 of 18,955

  1. Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

    Yunxin Li, Zhenyu Liu, Zitao Li +19

    cs.CVcs.CLarXiv:2505.04921v22025
  2. NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors

    Congyue Deng, Chiyu "Max'' Jiang, Charles R. Qi +4

    cs.CVarXiv:2212.03267v12022
  3. Wavelet Diffusion Models are fast and scalable Image Generators

    Hao Phung, Quan Dao, Anh Tran

    cs.CVeess.IVarXiv:2211.16152v22022
  4. VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control

    Yuxuan Bian, Zhaoyang Zhang, Xuan Ju +4

    cs.CVcs.AIcs.MMarXiv:2503.05639v32025
  5. Streaming Long Video Understanding with Large Language Models

    Rui Qian, Xiaoyi Dong, Pan Zhang +4

    cs.CVarXiv:2405.16009v12024
  6. Improving Object Localization with Fitness NMS and Bounded IoU Loss

    Lachlan Tychsen-Smith, Lars Petersson

    cs.CVarXiv:1711.00164v32017
  7. Hessian-based Analysis of Large Batch Training and Robustness to Adversaries

    Zhewei Yao, Amir Gholami, Qi Lei +2

    cs.CVcs.LGstat.MLarXiv:1802.08241v42018
  8. Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key

    Zhihe Yang, Xufang Luo, Dongqi Han +2

    cs.CVarXiv:2501.09695v22025
  9. MASTER: Multi-Aspect Non-local Network for Scene Text Recognition

    Ning Lu, Wenwen Yu, Xianbiao Qi +4

    cs.CVarXiv:1910.02562v32019
  10. 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

    Yipeng Gao, Lei Shu, Genzhi Ye +5

    cs.CVcs.AIcs.GRarXiv:2606.01057v12026
  11. Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

    Nuria Alina Chandra, Hannah Lee, Ryan Murtfeldt +10

    cs.CVcs.AIcs.CYarXiv:2503.02857v52025
  12. LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent

    Jianing Yang, Xuweiyi Chen, Shengyi Qian +4

    cs.CVcs.AIcs.CLarXiv:2309.12311v12023
  13. PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

    Soroush Nasiriany, Fei Xia, Wenhao Yu +20

    cs.ROcs.CLcs.CVarXiv:2402.07872v12024
  14. Pruning and Quantization for Deep Neural Network Acceleration: A Survey

    Tailin Liang, John Glossner, Lei Wang +2

    cs.CVcs.AIarXiv:2101.09671v32021
  15. Transparency of Deep Neural Networks for Medical Image Analysis: A Review of Interpretability Methods

    Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee +1

    eess.IVcs.AIcs.CVarXiv:2111.02398v12021
  16. FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery

    Xian Sun, Peijin Wang, Zhiyuan Yan +11

    cs.CVarXiv:2103.05569v22021
  17. VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction

    Hiroto Nakata, Yawen Zou, Shunsuke Sakai +5

    cs.CVarXiv:2603.13964v12026
  18. Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

    Yumi Lee, Harim Oh, Hyoryung Kim +52

    cs.CVcs.AIarXiv:2609.00866v12026
  19. Video models are zero-shot learners and reasoners

    Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6

    cs.LGcs.AIcs.CVarXiv:2509.20328v22025
  20. World Simulation with Video Foundation Models for Physical AI

    NVIDIA, :, Arslan Ali +87

    cs.CVcs.AIcs.LGarXiv:2511.00062v22025
  21. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

    Junxiang Xu, Ruisi Wang, Fanyi Pu +49

    cs.CVcs.AIcs.LGarXiv:2608.26105v12026
  22. Stitched Value Model for Diffusion Alignment

    Hyojun Go, Hyungjin Chung, Prune Truong +8

    cs.CVcs.AIcs.LGarXiv:2605.19804v12026
  23. Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

    Shuo Liang, Yixing Ma, Pengfei Zhou +32

    cs.CVcs.AIarXiv:2608.14391v12026
  24. ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

    Fan Jiang, Zhaoxu Sun, Mengchao Wang +38

    cs.CVcs.AIcs.LGarXiv:2607.19191v12026
  25. SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

    Yuyang Zhao, Yicheng Pan, Qiyuan He +6

    cs.CVcs.AIarXiv:2605.30409v12026
  26. PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation

    Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2

    cs.CVarXiv:2609.02860v12026
  27. Deep Visual Domain Adaptation: A Survey

    Mei Wang, Weihong Deng

    cs.CVarXiv:1802.03601v42018
  28. Amortized Set Prediction for Inverse IFS Reconstruction from Density Maps

    Yutaka Yamaguti

    cs.CVarXiv:2608.24175v12026
  29. PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

    Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty +1

    cs.IRcs.CLcs.CVarXiv:2608.26091v12026
  30. A Very Big Video Reasoning Suite

    Maijunxian Wang, Ruisi Wang, Juyi Lin +53

    cs.CVcs.AIcs.LGarXiv:2602.20159v22026
  31. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark

    Chang Xu, Jinwang Wang, Wen Yang +3

    cs.CVarXiv:2206.13996v12022
  32. PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

    Selim Kuzucu, Alessio Tonioni, Vasile Lup +3

    cs.CVcs.AIcs.CLarXiv:2605.30126v12026
  33. Cosmos World Foundation Model Platform for Physical AI

    NVIDIA, :, Niket Agarwal +76

    cs.CVcs.AIcs.LGarXiv:2501.03575v32025
  34. Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

    Wenli Xiao, Haotian Lin, Andy Peng +9

    cs.CVcs.ROarXiv:2511.00091v12025
  35. Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction

    Alfonso López-Ruiz, Diego Royo

    cs.DCcs.CVarXiv:2608.28183v12026
  36. Parallel Decoding Distillation for Fast Image and Video Generation

    Neta Shaul, Chao Liu, Arash Vahdat +1

    cs.CVcs.LGarXiv:2607.26004v12026
  37. RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

    Tianxing Chen, Yue Chen, Zixuan Li +41

    cs.ROcs.AIcs.CVarXiv:2607.04434v32026
  38. Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

    Chiara Tappermann, Steffen Renisch, Lars Ole Schwen +3

    cs.CVcs.AIarXiv:2608.16725v12026
  39. Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

    Alex Fang, Gabriel Ilharco, Mitchell Wortsman +4

    cs.CVcs.CLcs.LGarXiv:2205.01397v22022
  40. Deep learning for cardiac image segmentation: A review

    Chen Chen, Chen Qin, Huaqi Qiu +4

    eess.IVcs.CVcs.LGarXiv:1911.03723v12019
  41. EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

    Ryan Punamiya, Simar Kareer, Zeyi Liu +37

    cs.ROcs.CVarXiv:2604.07607v22026
  42. The Evolution of First Person Vision Methods: A Survey

    Alejandro Betancourt, Pietro Morerio, Carlo S. Regazzoni +1

    cs.CVarXiv:1409.1484v32014
  43. Predicting Citywide Crowd Flows in Irregular Regions Using Multi-View Graph Convolutional Networks

    Junkai Sun, Junbo Zhang, Qiaofei Li +3

    cs.CVcs.LGarXiv:1903.07789v22019
  44. Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

    Orcun Cetintas, Guillem Brasó, Tim Meinhardt +1

    cs.CVcs.AIarXiv:2609.00924v12026
  45. Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

    Teresa DiMeola, Charles Walter, Hong Xiao

    cs.CVcs.AIarXiv:2609.00628v12026
  46. VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents

    Ryota Tanaka, Taichi Iki, Taku Hasegawa +3

    cs.CLcs.AIcs.CVarXiv:2504.09795v12025
  47. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

    Chengzhi Liu, Zhongxing Xu, Qingyue Wei +5

    cs.CLcs.AIcs.CVarXiv:2505.21523v32025
  48. Automated and Interpretable Patient ECG Profiles for Disease Detection, Tracking, and Discovery

    Geoffrey H. Tison, Jeffrey Zhang, Francesca N. Delling +1

    cs.CVarXiv:1807.02569v12018
  49. Unpaired Image Captioning via Scene Graph Alignments

    Jiuxiang Gu, Shafiq Joty, Jianfei Cai +3

    cs.CVarXiv:1903.10658v42019
  50. From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network

    Yuxin Wang, Hongtao Xie, Shancheng Fang +3

    cs.CVarXiv:2108.09661v12021
  51. On the Adversarial Robustness of Vision Transformers

    Rulin Shao, Zhouxing Shi, Jinfeng Yi +2

    cs.CVcs.AIcs.LGarXiv:2103.15670v32021
  52. Towards Efficient Model Compression via Learned Global Ranking

    Ting-Wu Chin, Ruizhou Ding, Cha Zhang +1

    cs.CVcs.LGarXiv:1904.12368v22019
  53. Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

    Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot +2

    cs.CVcs.AIcs.LGarXiv:2504.02821v32025
  54. PointVLA: Injecting the 3D World into Vision-Language-Action Models

    Chengmeng Li, Junjie Wen, Yan Peng +3

    cs.ROcs.CVcs.LGarXiv:2503.07511v12025
  55. Stack-Captioning: Coarse-to-Fine Learning for Image Captioning

    Jiuxiang Gu, Jianfei Cai, Gang Wang +1

    cs.CVarXiv:1709.03376v32017
  56. Benchmarking and Error Diagnosis in Multi-Instance Pose Estimation

    Matteo Ruggero Ronchi, Pietro Perona

    cs.CVarXiv:1707.05388v22017
  57. High-Resolution Semantic Labeling with Convolutional Neural Networks

    Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat +1

    cs.CVarXiv:1611.01962v12016
  58. Assessing bikeability with street view imagery and computer vision

    Koichi Ito, Filip Biljecki

    cs.CVarXiv:2105.08499v32021
  59. StreamScout: Learning When to Look Deeper for Streaming Video Understanding

    Ce Zhang, Jing Bi, Jinxi He +9

    cs.CVarXiv:2609.00291v12026
  60. Correlation-Aware Deep Tracking

    Fei Xie, Chunyu Wang, Guangting Wang +3

    cs.CVarXiv:2203.01666v12022