Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

4,081 to 4,140 of 18,955

  1. Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding

    Xun Long Ng, Kian Eng Ong, Qichen Zheng +3

    cs.CVarXiv:2204.08129v22022
  2. Long-term Tracking in the Wild: A Benchmark

    Jack Valmadre, Luca Bertinetto, João F. Henriques +5

    cs.CVarXiv:1803.09502v32018
  3. Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

    Simian Luo, Chuanhao Yan, Chenxu Hu +1

    cs.SDcs.CVcs.LGarXiv:2306.17203v12023
  4. LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood

    Abhinav Kumar, Tim K. Marks, Wenxuan Mou +6

    cs.CVcs.LGeess.IVarXiv:2004.02980v12020
  5. SPair-71k: A Large-scale Benchmark for Semantic Correspondence

    Juhong Min, Jongmin Lee, Jean Ponce +1

    cs.CVarXiv:1908.10543v12019
    Summaries:한국어
  6. OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

    Lingyi Hong, Shilin Yan, Renrui Zhang +8

    cs.CVarXiv:2403.09634v12024
  7. Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual Features

    Simone Palazzo, Concetto Spampinato, Isaak Kavasidis +3

    cs.CVcs.LGq-bio.NCarXiv:1810.10974v22018
  8. Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

    Jason Y. Zhang, Sam Pepose, Hanbyul Joo +3

    cs.CVarXiv:2007.15649v22020
  9. Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

    Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu +12

    cs.CVarXiv:2601.23224v22026
  10. Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks

    Weihua Chen, Xianzhe Xu, Jian Jia +5

    cs.CVarXiv:2303.17602v12023
  11. NTIRE 2026 Challenge on Single Image Reflection Removal in the Wild: Datasets, Results, and Methods

    Jie Cai, Kangning Yang, Zhiyuan Li +50

    cs.CVarXiv:2604.10321v32026
  12. LEDITS++: Limitless Image Editing using Text-to-Image Models

    Manuel Brack, Felix Friedrich, Katharina Kornmeier +4

    cs.CVcs.AIcs.HCarXiv:2311.16711v22023
  13. Person Image Synthesis via Denoising Diffusion Model

    Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal +4

    cs.CVarXiv:2211.12500v22022
  14. CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation

    Qinghe Wang, Yawen Luo, Xiaoyu Shi +7

    cs.CVarXiv:2502.08639v12025
  15. Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization

    Junting Pan, Siyu Chen, Mike Zheng Shou +3

    cs.CVcs.LGeess.IVarXiv:2006.07976v32020
  16. QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

    Shuxiang Cao, Zijian Zhang, Abhishek Agarwal +29

    quant-phcs.CVarXiv:2604.25884v12026
  17. COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

    Zixu Li, Yupeng Hu, Zhiwei Chen +3

    cs.CVarXiv:2606.04604v12026
  18. Manifold Preserving Guided Diffusion

    Yutong He, Naoki Murata, Chieh-Hsin Lai +8

    cs.LGcs.AIcs.CVarXiv:2311.16424v12023
  19. Investigating Tradeoffs in Real-World Video Super-Resolution

    Kelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu +1

    cs.CVarXiv:2111.12704v12021
  20. The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results

    Jingkai Wang, Jue Gong, Zheng Chen +50

    cs.CVarXiv:2604.10532v22026
  21. SD-CNN: a Shallow-Deep CNN for Improved Breast Cancer Diagnosis

    Fei Gao, Teresa Wu, Jing Li +4

    cs.CVarXiv:1803.00663v22018
  22. AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

    Ruoxuan Feng, Yuxuan Zhou, Siyu Mei +6

    cs.ROcs.AIcs.CVarXiv:2602.09617v12026
  23. A survey on deep learning in medical image registration: new technologies, uncertainty, evaluation metrics, and beyond

    Junyu Chen, Yihao Liu, Shuwen Wei +5

    eess.IVcs.CVarXiv:2307.15615v42023
  24. Ten Years of Generative Adversarial Nets (GANs): A survey of the state-of-the-art

    Tanujit Chakraborty, Ujjwal Reddy K S, Shraddha M. Naik +2

    cs.LGcs.CVarXiv:2308.16316v12023
  25. VGR: Visual Grounded Reasoning

    Jiacong Wang, Zijian Kang, Haochen Wang +8

    cs.CVcs.AIcs.CLarXiv:2506.11991v32025
  26. AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control

    Peng Xu, Zhengnan Deng, Jiayan Deng +2

    cs.CVcs.AIcs.ROarXiv:2603.14363v12026
  27. Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models

    Yang Liu, Dingkang Yang, Yan Wang +5

    cs.CVcs.MMarXiv:2302.05087v32023
  28. OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

    Xueyu Hu, Tao Xiong, Biao Yi +26

    cs.AIcs.CLcs.CVarXiv:2508.04482v12025
  29. IT-TextFusion: Iterative Text-Image Interaction with Text-Guided Residual Refinement for Degradation-Aware Image Fusion

    Siyang Liu, Peiyi Zhou, Tianle Jin +3

    cs.CVarXiv:2609.01092v12026
  30. Human Activity Recognition Using Tools of Convolutional Neural Networks: A State of the Art Review, Data Sets, Challenges and Future Prospects

    Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray +1

    eess.SPcs.CVcs.LGarXiv:2202.03274v12022
  31. ReBridge-Flow: Re-Coupling Posterior Bridges in Flow Matching for Image Restoration

    Jiaqi Zhang, Yiqi Wang, Hongjie Wu +6

    cs.CVarXiv:2609.00811v12026
  32. MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network

    Guozhi Qiu, Zhiwei Chen, Zixu Li +4

    cs.CVcs.AIarXiv:2603.29291v12026
  33. Topological Planning with Transformers for Vision-and-Language Navigation

    Kevin Chen, Junshen K. Chen, Jo Chuang +2

    cs.ROcs.AIcs.CLarXiv:2012.05292v12020
  34. ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

    Haian Jin, Rundi Wu, Tianyuan Zhang +4

    cs.CVcs.AIcs.LGarXiv:2603.04385v32026
  35. openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit

    Maximilian Schmitt, Björn W. Schuller

    cs.CVcs.CLcs.IRarXiv:1605.06778v12016
  36. Slice-to-volume medical image registration: a survey

    Enzo Ferrante, Nikos Paragios

    cs.CVarXiv:1702.01636v22017
  37. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval

    Zhiwei Chen, Yupeng Hu, Zixu Li +3

    cs.CVarXiv:2507.05631v22025
  38. Vision Transformers For Weeds and Crops Classification Of High Resolution UAV Images

    Reenul Reedha, Eric Dericquebourg, Raphael Canals +1

    cs.CVarXiv:2109.02716v22021
  39. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

    Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

    cs.CVarXiv:2503.11576v12025
  40. Transitive Invariance for Self-supervised Visual Representation Learning

    Xiaolong Wang, Kaiming He, Abhinav Gupta

    cs.CVarXiv:1708.02901v32017
  41. AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection

    Zongdai Liu, Dingfu Zhou, Feixiang Lu +2

    cs.CVarXiv:2108.11127v12021
  42. A Deep Learning Interpretable Classifier for Diabetic Retinopathy Disease Grading

    Jordi de la Torre, Aida Valls, Domenec Puig

    cs.LGcs.CVstat.MLarXiv:1712.08107v12017
  43. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    Zhexiao Xiong, Xin Ye, Burhan Yaman +5

    cs.CVarXiv:2601.04453v42026
  44. Person Re-identification with Metric Learning using Privileged Information

    Xun Yang, Meng Wang, Dacheng Tao

    cs.CVarXiv:1904.05005v12019
  45. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

    Zhi Wang, Botao He, Kelin Yu +4

    cs.ROcs.AIcs.CVarXiv:2605.24934v22026
  46. The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

    Jiatong Li, Zheng Chen, Kai Liu +91

    cs.CVarXiv:2604.17306v12026
  47. NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

    Xin Li, Yeying Jin, Suhang Yao +95

    cs.CVarXiv:2604.10634v22026
  48. P-PatchDiff: Progressive Patch Diffusion Models for Low-light Image Enhancement

    Ruoyu Guo, Haonan Zhong, Maurice Pagnucco +1

    cs.CVarXiv:2609.01123v12026
  49. The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

    Xingyu Qiu, Yuqian Fu, Jiawei Geng +71

    cs.CVcs.AIarXiv:2604.11998v12026
  50. Checkerboard artifact free sub-pixel convolution: A note on sub-pixel convolution, resize convolution and convolution resize

    Andrew Aitken, Christian Ledig, Lucas Theis +3

    cs.CVarXiv:1707.02937v12017
  51. NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)

    Guanyi Qin, Jie Liang, Bingbing Zhang +50

    cs.CVcs.AIarXiv:2604.12512v12026
  52. Comparing Different Deep Learning Architectures for Classification of Chest Radiographs

    Keno K. Bressem, Lisa Adams, Christoph Erxleben +3

    cs.LGcs.CVeess.IVarXiv:2002.08991v12020
  53. RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection

    Pei Sun, Weiyue Wang, Yuning Chai +5

    cs.CVarXiv:2106.13365v12021
  54. FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

    Ziyang Fan, Keyu Chen, Ruilong Xing +3

    cs.CVcs.AIcs.CLarXiv:2602.08024v12026
  55. DReSG: Diffusion Residuals for Stylized Gaussian Splatting

    Zhongliang Liu, Wenjie Liu, Yang Li

    cs.CVcs.GRarXiv:2608.29048v22026
  56. HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery

    Jingtao Li, Yingyi Liu, Xinyu Wang +9

    cs.CVarXiv:2503.21841v12025
  57. Hyperspectral Image Classification in the Presence of Noisy Labels

    Junjun Jiang, Jiayi Ma, Zheng Wang +2

    cs.CVarXiv:1809.04212v22018
  58. Localizing Moments in Video with Temporal Language

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman +3

    cs.CVcs.CLarXiv:1809.01337v12018
  59. Diffusion Models for Video Prediction and Infilling

    Tobias Höppe, Arash Mehrjou, Stefan Bauer +2

    cs.CVcs.LGstat.MLarXiv:2206.07696v32022
  60. GUI Agents for Continual Game Generation

    Yixu Huang, Bo Li, Na Li +8

    cs.SEcs.AIcs.CVarXiv:2605.28258v12026