Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,241 to 6,300 of 18,904
Deep Learning Object Detection Methods for Ecological Camera Trap Data
Stefan Schneider, Graham W. Taylor, Stefan C. Kremer
cs.CVarXiv:1803.10842v12018VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Zirui Wang, Junyi Zhang, Jiaxin Ge +9
cs.CVarXiv:2601.16973v12026RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
Yuheng Ji, Huajie Tan, Jiayu Shi +14
cs.ROcs.CVarXiv:2502.21257v22025A Generative Appearance Model for End-to-end Video Object Segmentation
Joakim Johnander, Martin Danelljan, Emil Brissman +2
cs.CVarXiv:1811.11611v22018An All-in-One Network for Dehazing and Beyond
Boyi Li, Xiulian Peng, Zhangyang Wang +2
cs.CVcs.AIarXiv:1707.06543v12017Taming Visually Guided Sound Generation
Vladimir Iashin, Esa Rahtu
cs.CVcs.AIcs.LGarXiv:2110.08791v12021A Survey on Diffusion Models for Inverse Problems
Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai +5
cs.LGcs.AIcs.CVarXiv:2410.00083v12024InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
Haofan Wang, Matteo Spinelli, Qixun Wang +3
cs.CVarXiv:2404.02733v22024Rethinking FID: Towards a Better Evaluation Metric for Image Generation
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit +3
cs.CVarXiv:2401.09603v22023Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Shengbang Tong, Zhuang Liu, Yuexiang Zhai +3
cs.CVarXiv:2401.06209v22024Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
cs.CVcs.AIarXiv:2311.18259v42023FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
Shuang Zeng, Xinyuan Chang, Mengwei Xie +6
cs.CVarXiv:2505.17685v32025Point Transformer V3: Simpler, Faster, Stronger
Xiaoyang Wu, Li Jiang, Peng-Shuai Wang +6
cs.CVarXiv:2312.10035v22023Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn +6
cs.CVcs.AIcs.LGarXiv:2312.06662v12023T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Dongzhi Jiang, Ziyu Guo, Renrui Zhang +6
cs.CVcs.AIcs.CLarXiv:2505.00703v22025Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising
Xianlei Han, Shaoyu Wang, Jiancheng Fang +2
cs.CVarXiv:2609.02654v12026Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang +4
cs.CVcs.CLarXiv:2309.10020v12023Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao +6
cs.LGcs.AIcs.CVarXiv:2310.02279v32023Self-Consuming Generative Models Go MAD
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi +5
cs.LGcs.AIcs.CVarXiv:2307.01850v12023Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy
Clare McGenity, Emily L Clarke, Charlotte Jennings +5
physics.med-phcs.AIcs.CVarXiv:2306.07999v32023Generative Diffusion Prior for Unified Image Restoration and Enhancement
Ben Fei, Zhaoyang Lyu, Liang Pan +5
cs.CVarXiv:2304.01247v12023Leapfrog Diffusion Model for Stochastic Trajectory Prediction
Weibo Mao, Chenxin Xu, Qi Zhu +2
cs.CVarXiv:2303.10895v12023Consistency Models
Yang Song, Prafulla Dhariwal, Mark Chen +1
cs.LGcs.CVstat.MLarXiv:2303.01469v22023SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3
cs.CLcs.CVarXiv:2301.04883v12023CREPE: Can Vision-Language Foundation Models Reason Compositionally?
Zixian Ma, Jerry Hong, Mustafa Omer Gul +3
cs.CLcs.CVarXiv:2212.07796v32022Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation
Narek Tumanyan, Michal Geyer, Shai Bagon +1
cs.CVcs.AIarXiv:2211.12572v12022Diffusion Models: A Comprehensive Survey of Methods and Applications
Ling Yang, Zhilong Zhang, Yang Song +6
cs.LGcs.AIcs.CVarXiv:2209.00796v152022MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration
Zhi Jin, Yuwei Qiu, Kaihao Zhang +2
cs.CVarXiv:2501.04486v22025Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution
Alice Lucas, Santiago Lopez Tapia, Rafael Molina +1
cs.CVarXiv:1806.05764v22018State of the Art on Diffusion Models for Visual Computing
Ryan Po, Wang Yifan, Vladislav Golyanik +15
cs.AIcs.CVcs.GRarXiv:2310.07204v12023Streaming 4D Visual Geometry Transformer
Dong Zhuo, Wenzhao Zheng, Jiahe Guo +3
cs.CVcs.AIcs.LGarXiv:2507.11539v22025Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos
Haobo Yuan, Xiangtai Li, Tao Zhang +8
cs.CVarXiv:2501.04001v42025TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
Jakob Wasserthal, Hanns-Christian Breit, Manfred T. Meyer +9
eess.IVcs.CVarXiv:2208.05868v22022Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective
Gabriel Resende Machado, Eugênio Silva, Ronaldo Ribeiro Goldschmidt
cs.CVarXiv:2009.03728v12020Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
Garrick Brazil, Abhinav Kumar, Julian Straub +3
cs.CVarXiv:2207.10660v22022DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
Yingyan Li, Shuyao Shang, Weisong Liu +10
cs.CVcs.AIarXiv:2510.12796v22025Bridging the Domain Gap for Ground-to-Aerial Image Matching
Krishna Regmi, Mubarak Shah
cs.CVarXiv:1904.11045v22019GLIPv2: Unifying Localization and Vision-Language Understanding
Haotian Zhang, Pengchuan Zhang, Xiaowei Hu +7
cs.CVcs.AIcs.CLarXiv:2206.05836v22022Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
Jiwen Yu, Jianhong Bai, Yiran Qin +5
cs.CVarXiv:2506.03141v22025PyDoseRT Proton: A GPU Pencil-Beam Engine with a Convolutional Residual-Correction Network for Fast Proton Dose Calculation
Lukas Zimmermann, Hermann Fuchs, Attila Simkó +1
physics.med-phcs.CVarXiv:2609.01018v12026Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time
Liao Wang, Jiakai Zhang, Xinhang Liu +6
cs.CVcs.GRarXiv:2202.08614v22022Point-to-Voxel Knowledge Distillation for LiDAR Semantic Segmentation
Yuenan Hou, Xinge Zhu, Yuexin Ma +2
cs.CVarXiv:2206.02099v12022MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara +1
cs.CVcs.AIcs.CLarXiv:2502.17422v12025ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen +2
cs.CVcs.AIcs.LGarXiv:2507.16815v22025Deep Multi-modal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review
Can Cui, Haichun Yang, Yaohong Wang +6
cs.LGcs.AIcs.CVarXiv:2203.15588v32022Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade
Daniela Ruiz, Manuel Castellote, Zhongqi Miao +5
cs.SDcs.CVarXiv:2609.01792v12026TWIST: Teleoperated Whole-Body Imitation System
Yanjie Ze, Zixuan Chen, João Pedro Araújo +4
cs.ROcs.CVcs.LGarXiv:2505.02833v12025Using Deep Networks for Drone Detection
Cemal Aker, Sinan Kalkan
cs.CVarXiv:1706.05726v12017Consistency as Regularization for Unsupervised Shadow Removal
Anh-Kiet Duong, Petra Gomez-Krämer, Jean-Michel Carozza
cs.CVarXiv:2609.01806v12026VLP: A Survey on Vision-Language Pre-training
Feilong Chen, Duzhen Zhang, Minglun Han +4
cs.CVcs.CLarXiv:2202.09061v42022MaskGIT: Masked Generative Image Transformer
Huiwen Chang, Han Zhang, Lu Jiang +2
cs.CVarXiv:2202.04200v12022Road Segmentation in SAR Satellite Images with Deep Fully-Convolutional Neural Networks
Corentin Henry, Seyed Majid Azimi, Nina Merkle
cs.CVarXiv:1802.01445v22018QuadTree Attention for Vision Transformers
Shitao Tang, Jiahui Zhang, Siyu Zhu +1
cs.CVarXiv:2201.02767v22022NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
Zihan Zhu, Songyou Peng, Viktor Larsson +5
cs.CVarXiv:2112.12130v22021Dense Depth Priors for Neural Radiance Fields from Sparse Input Views
Barbara Roessle, Jonathan T. Barron, Ben Mildenhall +2
cs.CVarXiv:2112.03288v22021OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
Xingcheng Zhou, Xuyuan Han, Feng Yang +3
cs.CVarXiv:2503.23463v22025IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection
Mingjin Zhang, Yuchun Wang, Jie Guo +3
cs.CVarXiv:2407.07520v12024mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
Jonas Pai, Liam Achenbach, Victoriano Montesinos +3
cs.ROcs.AIcs.CVarXiv:2512.15692v22025Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Saaket Agashe, Kyle Wong, Vincent Tu +3
cs.AIcs.CLcs.CVarXiv:2504.00906v12025Person Re-identification by Saliency Learning
Rui Zhao, Wanli Ouyang, Xiaogang Wang
cs.CVarXiv:1412.1908v12014