Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

5,761 to 5,820 of 18,867

  1. Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization

    Ruyi Ji, Longyin Wen, Libo Zhang +5

    cs.CVarXiv:1909.11378v22019
  2. X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

    Zigang Geng, Yibing Wang, Yeyao Ma +10

    cs.CVarXiv:2507.22058v12025
  3. Uncertainty-aware Score Distribution Learning for Action Quality Assessment

    Yansong Tang, Zanlin Ni, Jiahuan Zhou +4

    cs.CVarXiv:2006.07665v12020
  4. Single Image Reflection Removal Exploiting Misaligned Training Data and Network Enhancements

    Kaixuan Wei, Jiaolong Yang, Ying Fu +2

    cs.CVarXiv:1904.00637v12019
  5. Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

    Weiqi Li, Xuanyu Zhang, Shijie Zhao +4

    cs.CVarXiv:2503.22679v22025
  6. Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

    Ranjan Sapkota, Manoj Karkee

    cs.CVcs.AIarXiv:2510.09653v32025
  7. ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation

    Yongzhi Su, Mahdi Saleh, Torben Fetzer +5

    cs.CVarXiv:2203.09418v22022
  8. Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

    Guoqing Ma, Haoyang Huang, Kun Yan +112

    cs.CVcs.CLarXiv:2502.10248v32025
  9. Controllable Image Captioning with Prompt-Conditioned Scene Rewards

    Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim

    cs.CVcs.CLcs.LGarXiv:2609.00709v12026
  10. XAttention: Block Sparse Attention with Antidiagonal Scoring

    Ruyi Xu, Guangxuan Xiao, Haofeng Huang +2

    cs.CLcs.CVarXiv:2503.16428v12025
  11. Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation

    Qianhao Yuan, Jie Lou, Xing Yu +4

    cs.CVcs.AIcs.CLarXiv:2605.18740v42026
  12. A Unified RGB-T Saliency Detection Benchmark: Dataset, Baselines, Analysis and A Novel Approach

    Chenglong Li, Guizhao Wang, Yunpeng Ma +3

    cs.CVarXiv:1701.02829v12017
  13. PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures

    Dan Hendrycks, Andy Zou, Mantas Mazeika +4

    cs.LGcs.CVarXiv:2112.05135v32021
  14. Deep Exemplar-based Video Colorization

    Bo Zhang, Mingming He, Jing Liao +4

    cs.CVcs.AIcs.LGarXiv:1906.09909v12019
  15. Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout

    Hidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan +2

    cs.CVarXiv:2511.20649v32025
  16. EV-SegNet: Semantic Segmentation for Event-based Cameras

    Iñigo Alonso, Ana C. Murillo

    cs.CVarXiv:1811.12039v12018
  17. BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation

    Yu Yao, Ella Atkins, Matthew Johnson-Roberson +2

    cs.CVcs.ROarXiv:2007.14558v22020
  18. ArtEmis: Affective Language for Visual Art

    Panos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov +2

    cs.CVcs.CLarXiv:2101.07396v12021
  19. Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

    Yunzhuo Hao, Jiawei Gu, Huichen Will Wang +4

    cs.CVarXiv:2501.05444v12025
  20. Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

    Ryan Burgert, Yuancheng Xu, Wenqi Xian +10

    cs.CVarXiv:2501.08331v52025
  21. SPIn-NeRF: Multiview Segmentation and Perceptual Inpainting with Neural Radiance Fields

    Ashkan Mirzaei, Tristan Aumentado-Armstrong, Konstantinos G. Derpanis +4

    cs.CVarXiv:2211.12254v22022
  22. Self-Supervised Learning for Videos: A Survey

    Madeline C. Schiappa, Yogesh S. Rawat, Mubarak Shah

    cs.CVcs.MMarXiv:2207.00419v32022
  23. Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

    Zeyuan Yang, Xueyang Yu, Delin Chen +2

    cs.CVcs.AIarXiv:2506.17218v12025
  24. MmWave Radar and Vision Fusion for Object Detection in Autonomous Driving: A Review

    Zhiqing Wei, Fengkai Zhang, Shuo Chang +3

    cs.CVarXiv:2108.03004v32021
  25. Towards Understanding Regularization in Batch Normalization

    Ping Luo, Xinjiang Wang, Wenqi Shao +1

    cs.LGcs.CVeess.SYarXiv:1809.00846v42018
  26. Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

    Shangzhe Di, Zhelun Yu, Guanghao Zhang +7

    cs.CVarXiv:2503.00540v12025
  27. FG-CLIP: Fine-Grained Visual and Textual Alignment

    Chunyu Xie, Bin Wang, Fanjing Kong +5

    cs.CVcs.AIarXiv:2505.05071v32025
  28. Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning

    Haokui Zhang, Ying Li, Yenan Jiang +3

    cs.CVarXiv:2012.03439v12020
  29. Striking the Right Balance with Uncertainty

    Salman Khan, Munawar Hayat, Waqas Zamir +2

    cs.CVarXiv:1901.07590v32019
  30. Deep learning analysis of the myocardium in coronary CT angiography for identification of patients with functionally significant coronary artery stenosis

    Majd Zreik, Nikolas Lessmann, Robbert W. van Hamersvelt +5

    cs.CVarXiv:1711.08917v22017
  31. Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

    Chetwin Low, Weimin Wang, Calder Katyal

    cs.MMcs.CVcs.SDarXiv:2510.01284v12025
  32. NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

    Chia-Yu Hung, Qi Sun, Pengfei Hong +5

    cs.ROcs.AIcs.CVarXiv:2504.19854v12025
  33. Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

    Shuang Wu, Youtian Lin, Feihu Zhang +8

    cs.CVarXiv:2505.17412v22025
  34. Fully-Convolutional Point Networks for Large-Scale Point Clouds

    Dario Rethage, Johanna Wald, Jürgen Sturm +2

    cs.CVarXiv:1808.06840v12018
  35. SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting

    Haozheng Yu, Xinyu Yang, Rundong Luo +2

    cs.CVarXiv:2608.31023v12026
  36. Visual prompt engineering for video models

    Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7

    cs.CVcs.AIarXiv:2607.25537v12026
  37. WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

    Runsheng Xu, Hubert Lin, Wonseok Jeon +11

    cs.CVcs.AIarXiv:2510.26125v32025
  38. Bilateral Attention Network for RGB-D Salient Object Detection

    Zhao Zhang, Zheng Lin, Jun Xu +3

    cs.CVarXiv:2004.14582v12020
  39. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

    Bingnan Li, Haozhe Wang, Haozhong Xiong +5

    cs.CVcs.AIcs.LGarXiv:2607.24731v22026
  40. CVM-Cervix: A Hybrid Cervical Pap-Smear Image Classification Framework Using CNN, Visual Transformer and Multilayer Perceptron

    Wanli Liu, Chen Li, Ning Xu +9

    cs.CVarXiv:2206.00971v12022
  41. Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

    Jiho Choi, Seonho Lee, Seojeong Park +1

    cs.CVarXiv:2606.23557v12026
  42. Detection of Christmas tree plantations from high-resolution aerial imagery. A case study in the French Morvan

    Francesca Razzano, Emanuele Dalsasso, Adrien Baysse-Lainé +3

    cs.CVarXiv:2608.27290v12026
  43. A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

    Ranjan Sapkota, William Bu, Chen Chen +2

    cs.CVcs.AIarXiv:2608.24935v12026
  44. GeoWAM: Visual Geometry World Action Models for Autonomous Driving

    Yiren Lu, Xin Ye, Jiaming Liu +9

    cs.CVcs.ROarXiv:2608.23486v12026
  45. FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

    Dingyun Zhang, Lixue Gong, Wei Liu

    cs.CVarXiv:2607.18227v12026
  46. Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing

    Keyan Hu, Mingtao Wang, Ziyu Zhou +4

    cs.CVarXiv:2608.28517v12026
  47. GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

    Arka Mukherjee, Soham Roy, Kartikeya Trivedi +1

    cs.CVcs.CLarXiv:2608.29483v12026
  48. Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

    Songtao Jiang, Yuan Wang, Sibo Song +22

    cs.CVarXiv:2510.08668v22025
  49. Beyond Accuracy: Quantifying Pulmonary Attribution in Anatomy-Guided Chest X-Ray Classification Under Domain Shift

    Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash +2

    cs.CVarXiv:2608.30467v12026
  50. Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

    Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi +5

    cs.CVcs.LGarXiv:2608.26500v12026
  51. Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

    Cong Cao, Huanjing Yue, Xin Liu +1

    cs.CVarXiv:2608.26476v12026
  52. SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images

    Zhishe Wang, Yanlin Chen, Wenyu Shao +2

    cs.CVarXiv:2204.11436v12022
  53. When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

    Ruoqi Hu, Chulin Zhao, Jiashuo Chang +2

    cs.CVcs.AIarXiv:2608.25933v12026
  54. CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

    Rana Muhammad Ahmed, Sabahat Abbas

    cs.CVcs.LGarXiv:2608.25539v12026
  55. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

    Zhaochen Su, Linjie Li, Mingyang Song +8

    cs.CVarXiv:2505.08617v22025
  56. TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

    Mark YU, Wenbo Hu, Jinbo Xing +1

    cs.CVcs.AIcs.GRarXiv:2503.05638v12025
  57. Bridging Adversarial and Collaborative Learning for AI-Generated Image Quality Assessment

    Baoliang Chen, Qing Lin, Sijie Mai

    cs.CVarXiv:2608.24372v12026
  58. Primate vision reveals a missing principle for robust dynamic AI

    Matteo Dunnhofer, Christian Micheloni, Kohitij Kar

    cs.CVq-bio.NCarXiv:2608.23790v12026
  59. Mover360: Controllable Object Manipulation in 360° Panoramic Images

    Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers +1

    cs.CVarXiv:2608.23238v12026
  60. Hybrid Generative-Discriminative Object Placement

    Siyuan Zhou, Li Niu

    cs.CVarXiv:2608.22692v12026