Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,341 to 5,400 of 18,955
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Yunxin Li, Zhenyu Liu, Zitao Li +19
cs.CVcs.CLarXiv:2505.04921v22025NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors
Congyue Deng, Chiyu "Max'' Jiang, Charles R. Qi +4
cs.CVarXiv:2212.03267v12022Wavelet Diffusion Models are fast and scalable Image Generators
Hao Phung, Quan Dao, Anh Tran
cs.CVeess.IVarXiv:2211.16152v22022VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
Yuxuan Bian, Zhaoyang Zhang, Xuan Ju +4
cs.CVcs.AIcs.MMarXiv:2503.05639v32025Streaming Long Video Understanding with Large Language Models
Rui Qian, Xiaoyi Dong, Pan Zhang +4
cs.CVarXiv:2405.16009v12024Improving Object Localization with Fitness NMS and Bounded IoU Loss
Lachlan Tychsen-Smith, Lars Petersson
cs.CVarXiv:1711.00164v32017Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
Zhewei Yao, Amir Gholami, Qi Lei +2
cs.CVcs.LGstat.MLarXiv:1802.08241v42018Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
Zhihe Yang, Xufang Luo, Dongqi Han +2
cs.CVarXiv:2501.09695v22025MASTER: Multi-Aspect Non-local Network for Scene Text Recognition
Ning Lu, Wenwen Yu, Xianbiao Qi +4
cs.CVarXiv:1910.02562v320193DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code
Yipeng Gao, Lei Shu, Genzhi Ye +5
cs.CVcs.AIcs.GRarXiv:2606.01057v12026Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Nuria Alina Chandra, Hannah Lee, Ryan Murtfeldt +10
cs.CVcs.AIcs.CYarXiv:2503.02857v52025LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent
Jianing Yang, Xuweiyi Chen, Shengyi Qian +4
cs.CVcs.AIcs.CLarXiv:2309.12311v12023PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
Soroush Nasiriany, Fei Xia, Wenhao Yu +20
cs.ROcs.CLcs.CVarXiv:2402.07872v12024Pruning and Quantization for Deep Neural Network Acceleration: A Survey
Tailin Liang, John Glossner, Lei Wang +2
cs.CVcs.AIarXiv:2101.09671v32021Transparency of Deep Neural Networks for Medical Image Analysis: A Review of Interpretability Methods
Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee +1
eess.IVcs.AIcs.CVarXiv:2111.02398v12021FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
Xian Sun, Peijin Wang, Zhiyuan Yan +11
cs.CVarXiv:2103.05569v22021VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
Hiroto Nakata, Yawen Zou, Shunsuke Sakai +5
cs.CVarXiv:2603.13964v12026Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
Yumi Lee, Harim Oh, Hyoryung Kim +52
cs.CVcs.AIarXiv:2609.00866v12026Video models are zero-shot learners and reasoners
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol +6
cs.LGcs.AIcs.CVarXiv:2509.20328v22025World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
cs.CVcs.AIcs.LGarXiv:2511.00062v22025VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Junxiang Xu, Ruisi Wang, Fanyi Pu +49
cs.CVcs.AIcs.LGarXiv:2608.26105v12026Stitched Value Model for Diffusion Alignment
Hyojun Go, Hyungjin Chung, Prune Truong +8
cs.CVcs.AIcs.LGarXiv:2605.19804v12026Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Shuo Liang, Yixing Ma, Pengfei Zhou +32
cs.CVcs.AIarXiv:2608.14391v12026ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Fan Jiang, Zhaoxu Sun, Mengchao Wang +38
cs.CVcs.AIcs.LGarXiv:2607.19191v12026SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
Yuyang Zhao, Yicheng Pan, Qiyuan He +6
cs.CVcs.AIarXiv:2605.30409v12026PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation
Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2
cs.CVarXiv:2609.02860v12026Deep Visual Domain Adaptation: A Survey
Mei Wang, Weihong Deng
cs.CVarXiv:1802.03601v42018Amortized Set Prediction for Inverse IFS Reconstruction from Density Maps
Yutaka Yamaguti
cs.CVarXiv:2608.24175v12026PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty +1
cs.IRcs.CLcs.CVarXiv:2608.26091v12026A Very Big Video Reasoning Suite
Maijunxian Wang, Ruisi Wang, Juyi Lin +53
cs.CVcs.AIcs.LGarXiv:2602.20159v22026Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
Chang Xu, Jinwang Wang, Wen Yang +3
cs.CVarXiv:2206.13996v12022PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
Selim Kuzucu, Alessio Tonioni, Vasile Lup +3
cs.CVcs.AIcs.CLarXiv:2605.30126v12026Cosmos World Foundation Model Platform for Physical AI
NVIDIA, :, Niket Agarwal +76
cs.CVcs.AIcs.LGarXiv:2501.03575v32025Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
Wenli Xiao, Haotian Lin, Andy Peng +9
cs.CVcs.ROarXiv:2511.00091v12025Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction
Alfonso López-Ruiz, Diego Royo
cs.DCcs.CVarXiv:2608.28183v12026Parallel Decoding Distillation for Fast Image and Video Generation
Neta Shaul, Chao Liu, Arash Vahdat +1
cs.CVcs.LGarXiv:2607.26004v12026RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Tianxing Chen, Yue Chen, Zixuan Li +41
cs.ROcs.AIcs.CVarXiv:2607.04434v32026Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
Chiara Tappermann, Steffen Renisch, Lars Ole Schwen +3
cs.CVcs.AIarXiv:2608.16725v12026Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)
Alex Fang, Gabriel Ilharco, Mitchell Wortsman +4
cs.CVcs.CLcs.LGarXiv:2205.01397v22022Deep learning for cardiac image segmentation: A review
Chen Chen, Chen Qin, Huaqi Qiu +4
eess.IVcs.CVcs.LGarXiv:1911.03723v12019EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
Ryan Punamiya, Simar Kareer, Zeyi Liu +37
cs.ROcs.CVarXiv:2604.07607v22026The Evolution of First Person Vision Methods: A Survey
Alejandro Betancourt, Pietro Morerio, Carlo S. Regazzoni +1
cs.CVarXiv:1409.1484v32014Predicting Citywide Crowd Flows in Irregular Regions Using Multi-View Graph Convolutional Networks
Junkai Sun, Junbo Zhang, Qiaofei Li +3
cs.CVcs.LGarXiv:1903.07789v22019Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
Orcun Cetintas, Guillem Brasó, Tim Meinhardt +1
cs.CVcs.AIarXiv:2609.00924v12026Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation
Teresa DiMeola, Charles Walter, Hong Xiao
cs.CVcs.AIarXiv:2609.00628v12026VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
Ryota Tanaka, Taichi Iki, Taku Hasegawa +3
cs.CLcs.AIcs.CVarXiv:2504.09795v12025More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
Chengzhi Liu, Zhongxing Xu, Qingyue Wei +5
cs.CLcs.AIcs.CVarXiv:2505.21523v32025Automated and Interpretable Patient ECG Profiles for Disease Detection, Tracking, and Discovery
Geoffrey H. Tison, Jeffrey Zhang, Francesca N. Delling +1
cs.CVarXiv:1807.02569v12018Unpaired Image Captioning via Scene Graph Alignments
Jiuxiang Gu, Shafiq Joty, Jianfei Cai +3
cs.CVarXiv:1903.10658v42019From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network
Yuxin Wang, Hongtao Xie, Shancheng Fang +3
cs.CVarXiv:2108.09661v12021On the Adversarial Robustness of Vision Transformers
Rulin Shao, Zhouxing Shi, Jinfeng Yi +2
cs.CVcs.AIcs.LGarXiv:2103.15670v32021Towards Efficient Model Compression via Learned Global Ranking
Ting-Wu Chin, Ruizhou Ding, Cha Zhang +1
cs.CVcs.LGarXiv:1904.12368v22019Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot +2
cs.CVcs.AIcs.LGarXiv:2504.02821v32025PointVLA: Injecting the 3D World into Vision-Language-Action Models
Chengmeng Li, Junjie Wen, Yan Peng +3
cs.ROcs.CVcs.LGarXiv:2503.07511v12025Stack-Captioning: Coarse-to-Fine Learning for Image Captioning
Jiuxiang Gu, Jianfei Cai, Gang Wang +1
cs.CVarXiv:1709.03376v32017Benchmarking and Error Diagnosis in Multi-Instance Pose Estimation
Matteo Ruggero Ronchi, Pietro Perona
cs.CVarXiv:1707.05388v22017High-Resolution Semantic Labeling with Convolutional Neural Networks
Emmanuel Maggiori, Yuliya Tarabalka, Guillaume Charpiat +1
cs.CVarXiv:1611.01962v12016Assessing bikeability with street view imagery and computer vision
Koichi Ito, Filip Biljecki
cs.CVarXiv:2105.08499v32021StreamScout: Learning When to Look Deeper for Streaming Video Understanding
Ce Zhang, Jing Bi, Jinxi He +9
cs.CVarXiv:2609.00291v12026Correlation-Aware Deep Tracking
Fei Xie, Chunyu Wang, Guangting Wang +3
cs.CVarXiv:2203.01666v12022