Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,021 to 4,080 of 18,866

  1. Human Activity Recognition Using Tools of Convolutional Neural Networks: A State of the Art Review, Data Sets, Challenges and Future Prospects

    Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray +1

    eess.SPcs.CVcs.LGarXiv:2202.03274v12022
  2. ReBridge-Flow: Re-Coupling Posterior Bridges in Flow Matching for Image Restoration

    Jiaqi Zhang, Yiqi Wang, Hongjie Wu +6

    cs.CVarXiv:2609.00811v12026
  3. MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network

    Guozhi Qiu, Zhiwei Chen, Zixu Li +4

    cs.CVcs.AIarXiv:2603.29291v12026
  4. Topological Planning with Transformers for Vision-and-Language Navigation

    Kevin Chen, Junshen K. Chen, Jo Chuang +2

    cs.ROcs.AIcs.CLarXiv:2012.05292v12020
  5. ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

    Haian Jin, Rundi Wu, Tianyuan Zhang +4

    cs.CVcs.AIcs.LGarXiv:2603.04385v32026
  6. openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit

    Maximilian Schmitt, Björn W. Schuller

    cs.CVcs.CLcs.IRarXiv:1605.06778v12016
  7. Slice-to-volume medical image registration: a survey

    Enzo Ferrante, Nikos Paragios

    cs.CVarXiv:1702.01636v22017
  8. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval

    Zhiwei Chen, Yupeng Hu, Zixu Li +3

    cs.CVarXiv:2507.05631v22025
  9. Vision Transformers For Weeds and Crops Classification Of High Resolution UAV Images

    Reenul Reedha, Eric Dericquebourg, Raphael Canals +1

    cs.CVarXiv:2109.02716v22021
  10. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

    Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

    cs.CVarXiv:2503.11576v12025
  11. Transitive Invariance for Self-supervised Visual Representation Learning

    Xiaolong Wang, Kaiming He, Abhinav Gupta

    cs.CVarXiv:1708.02901v32017
  12. AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection

    Zongdai Liu, Dingfu Zhou, Feixiang Lu +2

    cs.CVarXiv:2108.11127v12021
  13. A Deep Learning Interpretable Classifier for Diabetic Retinopathy Disease Grading

    Jordi de la Torre, Aida Valls, Domenec Puig

    cs.LGcs.CVstat.MLarXiv:1712.08107v12017
  14. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    Zhexiao Xiong, Xin Ye, Burhan Yaman +5

    cs.CVarXiv:2601.04453v42026
  15. Person Re-identification with Metric Learning using Privileged Information

    Xun Yang, Meng Wang, Dacheng Tao

    cs.CVarXiv:1904.05005v12019
  16. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

    Zhi Wang, Botao He, Kelin Yu +4

    cs.ROcs.AIcs.CVarXiv:2605.24934v22026
  17. The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

    Jiatong Li, Zheng Chen, Kai Liu +91

    cs.CVarXiv:2604.17306v12026
  18. NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

    Xin Li, Yeying Jin, Suhang Yao +95

    cs.CVarXiv:2604.10634v22026
  19. P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement

    Ruoyu Guo, Haonan Zhong, Maurice Pagnucco +1

    cs.CVarXiv:2609.01123v12026
  20. The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

    Xingyu Qiu, Yuqian Fu, Jiawei Geng +71

    cs.CVcs.AIarXiv:2604.11998v12026
  21. Checkerboard artifact free sub-pixel convolution: A note on sub-pixel convolution, resize convolution and convolution resize

    Andrew Aitken, Christian Ledig, Lucas Theis +3

    cs.CVarXiv:1707.02937v12017
  22. NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)

    Guanyi Qin, Jie Liang, Bingbing Zhang +50

    cs.CVcs.AIarXiv:2604.12512v12026
  23. Comparing Different Deep Learning Architectures for Classification of Chest Radiographs

    Keno K. Bressem, Lisa Adams, Christoph Erxleben +3

    cs.LGcs.CVeess.IVarXiv:2002.08991v12020
  24. RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection

    Pei Sun, Weiyue Wang, Yuning Chai +5

    cs.CVarXiv:2106.13365v12021
  25. FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

    Ziyang Fan, Keyu Chen, Ruilong Xing +3

    cs.CVcs.AIcs.CLarXiv:2602.08024v12026
  26. DReSG: Diffusion Residuals for Stylized Gaussian Splatting

    Zhongliang Liu, Wenjie Liu, Yang Li

    cs.CVcs.GRarXiv:2608.29048v22026
  27. HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery

    Jingtao Li, Yingyi Liu, Xinyu Wang +9

    cs.CVarXiv:2503.21841v12025
  28. Hyperspectral Image Classification in the Presence of Noisy Labels

    Junjun Jiang, Jiayi Ma, Zheng Wang +2

    cs.CVarXiv:1809.04212v22018
  29. Localizing Moments in Video with Temporal Language

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman +3

    cs.CVcs.CLarXiv:1809.01337v12018
  30. Diffusion Models for Video Prediction and Infilling

    Tobias Höppe, Arash Mehrjou, Stefan Bauer +2

    cs.CVcs.LGstat.MLarXiv:2206.07696v32022
  31. GUI Agents for Continual Game Generation

    Yixu Huang, Bo Li, Na Li +8

    cs.SEcs.AIcs.CVarXiv:2605.28258v12026
  32. VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

    Sixiao Zheng, Minghao Yin, Wenbo Hu +3

    cs.CVarXiv:2601.05138v22026
  33. Di$^2$CycleSB: Towards High-Quality Unsupervised Nighttime Visibility Enhancement via Schrödinger Bridge Transformer

    Hanting Li, Xin Sun, Wei Ye +2

    cs.CVarXiv:2608.29043v12026
  34. Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking

    You Wu, Xucheng Wang, Xiangyang Yang +4

    cs.CVarXiv:2504.09228v12025
  35. FiLM-GPNet: Geometry-Aware Pseudo-Supervised Phase Restoration with Zero-Shot Generalization for Large Temporal InSAR Stacks

    Getnet Demil, Muhammad Farhan Humayun, Tomi Westerlund +2

    cs.CVcs.LGarXiv:2608.29384v12026
  36. Progressively Normalized Self-Attention Network for Video Polyp Segmentation

    Ge-Peng Ji, Yu-Cheng Chou, Deng-Ping Fan +4

    cs.CVarXiv:2105.08468v22021
  37. Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA Challenge

    Kai Geissler, Raphael Schäfer

    cs.CVcs.AIcs.LGarXiv:2608.29162v12026
  38. ChordEdit: One-Step Low-Energy Transport for Image Editing

    Liangsi Lu, Xuhang Chen, Minzhe Guo +3

    cs.CVarXiv:2602.19083v22026
  39. Towards Universal Modal Tracking with Online Dense Temporal Token Learning

    Yaozong Zheng, Bineng Zhong, Qihua Liang +4

    cs.CVcs.MMarXiv:2507.20177v12025
  40. KeepLoRA: Continual Learning with Residual Gradient Adaptation

    Mao-Lin Luo, Zi-Hao Zhou, Yi-Lin Zhang +3

    cs.CVcs.LGarXiv:2601.19659v12026
  41. Unsupervised Learning for Real-World Super-Resolution

    Andreas Lugmayr, Martin Danelljan, Radu Timofte

    eess.IVcs.CVarXiv:1909.09629v12019
  42. NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results

    Xin Li, Jiachao Gong, Xijun Wang +75

    cs.CVarXiv:2604.10551v12026
  43. Human POSEitioning System (HPS): 3D Human Pose Estimation and Self-localization in Large Scenes from Body-Mounted Sensors

    Vladimir Guzov, Aymen Mir, Torsten Sattler +1

    cs.CVarXiv:2103.17265v12021
  44. FrePGAN: Robust Deepfake Detection Using Frequency-level Perturbations

    Yonghyun Jeong, Doyeon Kim, Youngmin Ro +1

    cs.CVcs.LGeess.IVarXiv:2202.03347v12022
  45. SPP-Net: Deep Absolute Pose Regression with Synthetic Views

    Pulak Purkait, Cheng Zhao, Christopher Zach

    cs.CVarXiv:1712.03452v12017
  46. Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring

    Dongxu Zhang, Yiding Sun, Cheng Tan +4

    cs.MMcs.CLcs.CVarXiv:2601.13879v42026
  47. Towards Image Understanding from Deep Compression without Decoding

    Robert Torfason, Fabian Mentzer, Eirikur Agustsson +3

    cs.CVarXiv:1803.06131v12018
  48. Hierarchy Parsing for Image Captioning

    Ting Yao, Yingwei Pan, Yehao Li +1

    cs.CVcs.CLarXiv:1909.03918v22019
  49. Multitask AET with Orthogonal Tangent Regularity for Dark Object Detection

    Ziteng Cui, Guo-Jun Qi, Lin Gu +3

    cs.CVarXiv:2205.03346v12022
  50. NeRFReN: Neural Radiance Fields with Reflections

    Yuan-Chen Guo, Di Kang, Linchao Bao +2

    cs.CVcs.GRarXiv:2111.15234v22021
  51. Fail2Drive: Benchmarking Closed-Loop Driving Generalization

    Simon Gerstenecker, Andreas Geiger, Katrin Renz

    cs.ROcs.CVarXiv:2604.08535v12026
  52. Visual-Advantage On-Policy Distillation for Vision-Language Models

    Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8

    cs.CVarXiv:2605.21924v12026
  53. What makes instance discrimination good for transfer learning?

    Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau +1

    cs.CVarXiv:2006.06606v22020
  54. Training Vision Transformers for Image Retrieval

    Alaaeldin El-Nouby, Natalia Neverova, Ivan Laptev +1

    cs.CVarXiv:2102.05644v12021
  55. Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

    Mingwen Zhang, Jisheng Dang, Minqiang Yang +3

    cs.CVcs.CLarXiv:2608.28675v12026
  56. DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies

    Tony Tao, Mohan Kumar Srirama, Jason Jingzhou Liu +2

    cs.ROcs.AIcs.CVarXiv:2505.07813v22025
  57. EPNet++: Cascade Bi-directional Fusion for Multi-Modal 3D Object Detection

    Zhe Liu, Tengteng Huang, Bingling Li +3

    cs.CVarXiv:2112.11088v42021
  58. CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

    Chubin Zhang, Jianan Wang, Zifeng Gao +5

    cs.ROcs.CVarXiv:2601.04061v22026
  59. Learning Geometry-Disentangled Representation for Complementary Understanding of 3D Object Point Cloud

    Mutian Xu, Junhao Zhang, Zhipeng Zhou +3

    cs.CVarXiv:2012.10921v32020
  60. Learning Aligned Cross-Modal Representations from Weakly Aligned Data

    Lluis Castrejon, Yusuf Aytar, Carl Vondrick +2

    cs.CVarXiv:1607.07295v12016