Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,961 to 4,020 of 18,849

  1. Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction

    Jiaxin Huang, Yuanbo Yang, Bangbang Yang +3

    cs.CVarXiv:2601.04090v22026
  2. MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models

    Hanyang Yu, Haitao Lin, Jingbo Zhang +4

    cs.CVcs.LGcs.ROarXiv:2606.13515v12026
  3. Predicting Visual Exemplars of Unseen Classes for Zero-Shot Learning

    Soravit Changpinyo, Wei-Lun Chao, Fei Sha

    cs.CVarXiv:1605.08151v22016
  4. Connecting Touch and Vision via Cross-Modal Prediction

    Yunzhu Li, Jun-Yan Zhu, Russ Tedrake +1

    cs.CVcs.LGcs.ROarXiv:1906.06322v12019
  5. SCULPT: Training Edge Vision Models for Post-Training Quantization Readiness

    Bharadwaj Kavuri, Sourav Babu-PK, Varadhraj Ellapan +2

    cs.CVarXiv:2609.01743v12026
  6. Flow caching for autoregressive video generation

    Yuexiao Ma, Xuzhe Zheng, Jing Xu +9

    cs.CVcs.AIarXiv:2602.10825v12026
  7. NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Multi-Exposure Image Fusion in Dynamic Scenes (Track 2)

    Lishen Qu, Yao Liu, Jie Liang +32

    cs.CVarXiv:2604.09030v12026
  8. The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation

    Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi +3

    cs.CVarXiv:1905.00737v12019
  9. Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

    Yiding Sun, Jihua Zhu, Haozhe Cheng +4

    cs.CVarXiv:2602.23069v22026
  10. Deep Fashion3D: A Dataset and Benchmark for 3D Garment Reconstruction from Single Images

    Heming Zhu, Yu Cao, Hang Jin +5

    cs.CVarXiv:2003.12753v22020
  11. NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge Report

    Andrei Dumitriu, Aakash Ralhan, Florin Miron +32

    cs.CVarXiv:2604.17070v22026
  12. Self-Critical Reasoning for Robust Visual Question Answering

    Jialin Wu, Raymond J. Mooney

    cs.CVcs.CLarXiv:1905.09998v32019
  13. A Comprehensive Review of U-Net and Its Variants: Advances and Applications in Medical Image Segmentation

    Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan

    eess.IVcs.CVarXiv:2502.06895v12025
  14. PuLID: Pure and Lightning ID Customization via Contrastive Alignment

    Zinan Guo, Yanze Wu, Zhuowei Chen +3

    cs.CVarXiv:2404.16022v22024
  15. Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding

    Xun Long Ng, Kian Eng Ong, Qichen Zheng +3

    cs.CVarXiv:2204.08129v22022
  16. Long-term Tracking in the Wild: A Benchmark

    Jack Valmadre, Luca Bertinetto, João F. Henriques +5

    cs.CVarXiv:1803.09502v32018
  17. Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

    Simian Luo, Chuanhao Yan, Chenxu Hu +1

    cs.SDcs.CVcs.LGarXiv:2306.17203v12023
  18. LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood

    Abhinav Kumar, Tim K. Marks, Wenxuan Mou +6

    cs.CVcs.LGeess.IVarXiv:2004.02980v12020
  19. SPair-71k: A Large-scale Benchmark for Semantic Correspondence

    Juhong Min, Jongmin Lee, Jean Ponce +1

    cs.CVarXiv:1908.10543v12019
    Summaries:한국어
  20. OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

    Lingyi Hong, Shilin Yan, Renrui Zhang +8

    cs.CVarXiv:2403.09634v12024
  21. Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual Features

    Simone Palazzo, Concetto Spampinato, Isaak Kavasidis +3

    cs.CVcs.LGq-bio.NCarXiv:1810.10974v22018
  22. Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

    Jason Y. Zhang, Sam Pepose, Hanbyul Joo +3

    cs.CVarXiv:2007.15649v22020
  23. Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

    Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu +12

    cs.CVarXiv:2601.23224v22026
  24. Beyond Appearance: a Semantic Controllable Self-Supervised Learning Framework for Human-Centric Visual Tasks

    Weihua Chen, Xianzhe Xu, Jian Jia +5

    cs.CVarXiv:2303.17602v12023
  25. NTIRE 2026 Challenge on Single Image Reflection Removal in the Wild: Datasets, Results, and Methods

    Jie Cai, Kangning Yang, Zhiyuan Li +50

    cs.CVarXiv:2604.10321v32026
  26. LEDITS++: Limitless Image Editing using Text-to-Image Models

    Manuel Brack, Felix Friedrich, Katharina Kornmeier +4

    cs.CVcs.AIcs.HCarXiv:2311.16711v22023
  27. Person Image Synthesis via Denoising Diffusion Model

    Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal +4

    cs.CVarXiv:2211.12500v22022
  28. CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation

    Qinghe Wang, Yawen Luo, Xiaoyu Shi +7

    cs.CVarXiv:2502.08639v12025
  29. Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization

    Junting Pan, Siyu Chen, Mike Zheng Shou +3

    cs.CVcs.LGeess.IVarXiv:2006.07976v32020
  30. QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

    Shuxiang Cao, Zijian Zhang, Abhishek Agarwal +29

    quant-phcs.CVarXiv:2604.25884v12026
  31. COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

    Zixu Li, Yupeng Hu, Zhiwei Chen +3

    cs.CVarXiv:2606.04604v12026
  32. Manifold Preserving Guided Diffusion

    Yutong He, Naoki Murata, Chieh-Hsin Lai +8

    cs.LGcs.AIcs.CVarXiv:2311.16424v12023
  33. Investigating Tradeoffs in Real-World Video Super-Resolution

    Kelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu +1

    cs.CVarXiv:2111.12704v12021
  34. The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results

    Jingkai Wang, Jue Gong, Zheng Chen +50

    cs.CVarXiv:2604.10532v22026
  35. SD-CNN: a Shallow-Deep CNN for Improved Breast Cancer Diagnosis

    Fei Gao, Teresa Wu, Jing Li +4

    cs.CVarXiv:1803.00663v22018
  36. AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception

    Ruoxuan Feng, Yuxuan Zhou, Siyu Mei +6

    cs.ROcs.AIcs.CVarXiv:2602.09617v12026
  37. A survey on deep learning in medical image registration: new technologies, uncertainty, evaluation metrics, and beyond

    Junyu Chen, Yihao Liu, Shuwen Wei +5

    eess.IVcs.CVarXiv:2307.15615v42023
  38. Ten Years of Generative Adversarial Nets (GANs): A survey of the state-of-the-art

    Tanujit Chakraborty, Ujjwal Reddy K S, Shraddha M. Naik +2

    cs.LGcs.CVarXiv:2308.16316v12023
  39. VGR: Visual Grounded Reasoning

    Jiacong Wang, Zijian Kang, Haochen Wang +8

    cs.CVcs.AIcs.CLarXiv:2506.11991v32025
  40. AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control

    Peng Xu, Zhengnan Deng, Jiayan Deng +2

    cs.CVcs.AIcs.ROarXiv:2603.14363v12026
  41. Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models

    Yang Liu, Dingkang Yang, Yan Wang +5

    cs.CVcs.MMarXiv:2302.05087v32023
  42. OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

    Xueyu Hu, Tao Xiong, Biao Yi +26

    cs.AIcs.CLcs.CVarXiv:2508.04482v12025
  43. IT-TextFusion: Iterative Text-Image Interaction with Text-Guided Residual Refinement for Degradation-Aware Image Fusion

    Siyang Liu, Peiyi Zhou, Tianle Jin +3

    cs.CVarXiv:2609.01092v12026
  44. Human Activity Recognition Using Tools of Convolutional Neural Networks: A State of the Art Review, Data Sets, Challenges and Future Prospects

    Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray +1

    eess.SPcs.CVcs.LGarXiv:2202.03274v12022
  45. ReBridge-Flow: Re-Coupling Posterior Bridges in Flow Matching for Image Restoration

    Jiaqi Zhang, Yiqi Wang, Hongjie Wu +6

    cs.CVarXiv:2609.00811v12026
  46. MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network

    Guozhi Qiu, Zhiwei Chen, Zixu Li +4

    cs.CVcs.AIarXiv:2603.29291v12026
  47. Topological Planning with Transformers for Vision-and-Language Navigation

    Kevin Chen, Junshen K. Chen, Jo Chuang +2

    cs.ROcs.AIcs.CLarXiv:2012.05292v12020
  48. ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

    Haian Jin, Rundi Wu, Tianyuan Zhang +4

    cs.CVcs.AIcs.LGarXiv:2603.04385v32026
  49. openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit

    Maximilian Schmitt, Björn W. Schuller

    cs.CVcs.CLcs.IRarXiv:1605.06778v12016
  50. Slice-to-volume medical image registration: a survey

    Enzo Ferrante, Nikos Paragios

    cs.CVarXiv:1702.01636v22017
  51. OFFSET: Segmentation-based Focus Shift Revision for Composed Image Retrieval

    Zhiwei Chen, Yupeng Hu, Zixu Li +3

    cs.CVarXiv:2507.05631v22025
  52. Vision Transformers For Weeds and Crops Classification Of High Resolution UAV Images

    Reenul Reedha, Eric Dericquebourg, Raphael Canals +1

    cs.CVarXiv:2109.02716v22021
  53. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

    Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

    cs.CVarXiv:2503.11576v12025
  54. Transitive Invariance for Self-supervised Visual Representation Learning

    Xiaolong Wang, Kaiming He, Abhinav Gupta

    cs.CVarXiv:1708.02901v32017
  55. AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection

    Zongdai Liu, Dingfu Zhou, Feixiang Lu +2

    cs.CVarXiv:2108.11127v12021
  56. A Deep Learning Interpretable Classifier for Diabetic Retinopathy Disease Grading

    Jordi de la Torre, Aida Valls, Domenec Puig

    cs.LGcs.CVstat.MLarXiv:1712.08107v12017
  57. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    Zhexiao Xiong, Xin Ye, Burhan Yaman +5

    cs.CVarXiv:2601.04453v42026
  58. Person Re-identification with Metric Learning using Privileged Information

    Xun Yang, Meng Wang, Dacheng Tao

    cs.CVarXiv:1904.05005v12019
  59. HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

    Zhi Wang, Botao He, Kelin Yu +4

    cs.ROcs.AIcs.CVarXiv:2605.24934v22026
  60. The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

    Jiatong Li, Zheng Chen, Kai Liu +91

    cs.CVarXiv:2604.17306v12026