Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,241 to 6,300 of 18,904

  1. Deep Learning Object Detection Methods for Ecological Camera Trap Data

    Stefan Schneider, Graham W. Taylor, Stefan C. Kremer

    cs.CVarXiv:1803.10842v12018
  2. VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

    Zirui Wang, Junyi Zhang, Jiaxin Ge +9

    cs.CVarXiv:2601.16973v12026
  3. RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

    Yuheng Ji, Huajie Tan, Jiayu Shi +14

    cs.ROcs.CVarXiv:2502.21257v22025
  4. A Generative Appearance Model for End-to-end Video Object Segmentation

    Joakim Johnander, Martin Danelljan, Emil Brissman +2

    cs.CVarXiv:1811.11611v22018
  5. An All-in-One Network for Dehazing and Beyond

    Boyi Li, Xiulian Peng, Zhangyang Wang +2

    cs.CVcs.AIarXiv:1707.06543v12017
  6. Taming Visually Guided Sound Generation

    Vladimir Iashin, Esa Rahtu

    cs.CVcs.AIcs.LGarXiv:2110.08791v12021
  7. A Survey on Diffusion Models for Inverse Problems

    Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai +5

    cs.LGcs.AIcs.CVarXiv:2410.00083v12024
  8. InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation

    Haofan Wang, Matteo Spinelli, Qixun Wang +3

    cs.CVarXiv:2404.02733v22024
  9. Rethinking FID: Towards a Better Evaluation Metric for Image Generation

    Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit +3

    cs.CVarXiv:2401.09603v22023
  10. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

    Shengbang Tong, Zhuang Liu, Yuexiang Zhai +3

    cs.CVarXiv:2401.06209v22024
  11. Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

    cs.CVcs.AIarXiv:2311.18259v42023
  12. FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

    Shuang Zeng, Xinyuan Chang, Mengwei Xie +6

    cs.CVarXiv:2505.17685v32025
  13. Point Transformer V3: Simpler, Faster, Stronger

    Xiaoyang Wu, Li Jiang, Peng-Shuai Wang +6

    cs.CVarXiv:2312.10035v22023
  14. Photorealistic Video Generation with Diffusion Models

    Agrim Gupta, Lijun Yu, Kihyuk Sohn +6

    cs.CVcs.AIcs.LGarXiv:2312.06662v12023
  15. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

    Dongzhi Jiang, Ziyu Guo, Renrui Zhang +6

    cs.CVcs.AIcs.CLarXiv:2505.00703v22025
  16. Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising

    Xianlei Han, Shaoyu Wang, Jiancheng Fang +2

    cs.CVarXiv:2609.02654v12026
  17. Multimodal Foundation Models: From Specialists to General-Purpose Assistants

    Chunyuan Li, Zhe Gan, Zhengyuan Yang +4

    cs.CVcs.CLarXiv:2309.10020v12023
  18. Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

    Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao +6

    cs.LGcs.AIcs.CVarXiv:2310.02279v32023
  19. Self-Consuming Generative Models Go MAD

    Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi +5

    cs.LGcs.AIcs.CVarXiv:2307.01850v12023
  20. Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy

    Clare McGenity, Emily L Clarke, Charlotte Jennings +5

    physics.med-phcs.AIcs.CVarXiv:2306.07999v32023
  21. Generative Diffusion Prior for Unified Image Restoration and Enhancement

    Ben Fei, Zhaoyang Lyu, Liang Pan +5

    cs.CVarXiv:2304.01247v12023
  22. Leapfrog Diffusion Model for Stochastic Trajectory Prediction

    Weibo Mao, Chenxin Xu, Qi Zhu +2

    cs.CVarXiv:2303.10895v12023
  23. Consistency Models

    Yang Song, Prafulla Dhariwal, Mark Chen +1

    cs.LGcs.CVstat.MLarXiv:2303.01469v22023
  24. SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

    Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3

    cs.CLcs.CVarXiv:2301.04883v12023
  25. CREPE: Can Vision-Language Foundation Models Reason Compositionally?

    Zixian Ma, Jerry Hong, Mustafa Omer Gul +3

    cs.CLcs.CVarXiv:2212.07796v32022
  26. Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

    Narek Tumanyan, Michal Geyer, Shai Bagon +1

    cs.CVcs.AIarXiv:2211.12572v12022
  27. Diffusion Models: A Comprehensive Survey of Methods and Applications

    Ling Yang, Zhilong Zhang, Yang Song +6

    cs.LGcs.AIcs.CVarXiv:2209.00796v152022
  28. MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration

    Zhi Jin, Yuwei Qiu, Kaihao Zhang +2

    cs.CVarXiv:2501.04486v22025
  29. Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution

    Alice Lucas, Santiago Lopez Tapia, Rafael Molina +1

    cs.CVarXiv:1806.05764v22018
  30. State of the Art on Diffusion Models for Visual Computing

    Ryan Po, Wang Yifan, Vladislav Golyanik +15

    cs.AIcs.CVcs.GRarXiv:2310.07204v12023
  31. Streaming 4D Visual Geometry Transformer

    Dong Zhuo, Wenzhao Zheng, Jiahe Guo +3

    cs.CVcs.AIcs.LGarXiv:2507.11539v22025
  32. Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

    Haobo Yuan, Xiangtai Li, Tao Zhang +8

    cs.CVarXiv:2501.04001v42025
  33. TotalSegmentator: robust segmentation of 104 anatomical structures in CT images

    Jakob Wasserthal, Hanns-Christian Breit, Manfred T. Meyer +9

    eess.IVcs.CVarXiv:2208.05868v22022
  34. Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

    Gabriel Resende Machado, Eugênio Silva, Ronaldo Ribeiro Goldschmidt

    cs.CVarXiv:2009.03728v12020
  35. Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

    Garrick Brazil, Abhinav Kumar, Julian Straub +3

    cs.CVarXiv:2207.10660v22022
  36. DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

    Yingyan Li, Shuyao Shang, Weisong Liu +10

    cs.CVcs.AIarXiv:2510.12796v22025
  37. Bridging the Domain Gap for Ground-to-Aerial Image Matching

    Krishna Regmi, Mubarak Shah

    cs.CVarXiv:1904.11045v22019
  38. GLIPv2: Unifying Localization and Vision-Language Understanding

    Haotian Zhang, Pengchuan Zhang, Xiaowei Hu +7

    cs.CVcs.AIcs.CLarXiv:2206.05836v22022
  39. Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

    Jiwen Yu, Jianhong Bai, Yiran Qin +5

    cs.CVarXiv:2506.03141v22025
  40. PyDoseRT Proton: A GPU Pencil-Beam Engine with a Convolutional Residual-Correction Network for Fast Proton Dose Calculation

    Lukas Zimmermann, Hermann Fuchs, Attila Simkó +1

    physics.med-phcs.CVarXiv:2609.01018v12026
  41. Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time

    Liao Wang, Jiakai Zhang, Xinhang Liu +6

    cs.CVcs.GRarXiv:2202.08614v22022
  42. Point-to-Voxel Knowledge Distillation for LiDAR Semantic Segmentation

    Yuenan Hou, Xinge Zhu, Yuexin Ma +2

    cs.CVarXiv:2206.02099v12022
  43. MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

    Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara +1

    cs.CVcs.AIcs.CLarXiv:2502.17422v12025
  44. ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

    Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen +2

    cs.CVcs.AIcs.LGarXiv:2507.16815v22025
  45. Deep Multi-modal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review

    Can Cui, Haichun Yang, Yaohong Wang +6

    cs.LGcs.AIcs.CVarXiv:2203.15588v32022
  46. Efficient Passive Acoustic Monitoring of Killer Whales Using a Two-Stage Detection and Ecotype Classification Cascade

    Daniela Ruiz, Manuel Castellote, Zhongqi Miao +5

    cs.SDcs.CVarXiv:2609.01792v12026
  47. TWIST: Teleoperated Whole-Body Imitation System

    Yanjie Ze, Zixuan Chen, João Pedro Araújo +4

    cs.ROcs.CVcs.LGarXiv:2505.02833v12025
  48. Using Deep Networks for Drone Detection

    Cemal Aker, Sinan Kalkan

    cs.CVarXiv:1706.05726v12017
  49. Consistency as Regularization for Unsupervised Shadow Removal

    Anh-Kiet Duong, Petra Gomez-Krämer, Jean-Michel Carozza

    cs.CVarXiv:2609.01806v12026
  50. VLP: A Survey on Vision-Language Pre-training

    Feilong Chen, Duzhen Zhang, Minglun Han +4

    cs.CVcs.CLarXiv:2202.09061v42022
  51. MaskGIT: Masked Generative Image Transformer

    Huiwen Chang, Han Zhang, Lu Jiang +2

    cs.CVarXiv:2202.04200v12022
  52. Road Segmentation in SAR Satellite Images with Deep Fully-Convolutional Neural Networks

    Corentin Henry, Seyed Majid Azimi, Nina Merkle

    cs.CVarXiv:1802.01445v22018
  53. QuadTree Attention for Vision Transformers

    Shitao Tang, Jiahui Zhang, Siyu Zhu +1

    cs.CVarXiv:2201.02767v22022
  54. NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

    Zihan Zhu, Songyou Peng, Viktor Larsson +5

    cs.CVarXiv:2112.12130v22021
  55. Dense Depth Priors for Neural Radiance Fields from Sparse Input Views

    Barbara Roessle, Jonathan T. Barron, Ben Mildenhall +2

    cs.CVarXiv:2112.03288v22021
  56. OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

    Xingcheng Zhou, Xuyuan Han, Feng Yang +3

    cs.CVarXiv:2503.23463v22025
  57. IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection

    Mingjin Zhang, Yuchun Wang, Jie Guo +3

    cs.CVarXiv:2407.07520v12024
  58. mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

    Jonas Pai, Liam Achenbach, Victoriano Montesinos +3

    cs.ROcs.AIcs.CVarXiv:2512.15692v22025
  59. Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

    Saaket Agashe, Kyle Wong, Vincent Tu +3

    cs.AIcs.CLcs.CVarXiv:2504.00906v12025
  60. Person Re-identification by Saliency Learning

    Rui Zhao, Wanli Ouyang, Xiaogang Wang

    cs.CVarXiv:1412.1908v12014