Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,761 to 5,820 of 18,867
Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
Ruyi Ji, Longyin Wen, Libo Zhang +5
cs.CVarXiv:1909.11378v22019X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Zigang Geng, Yibing Wang, Yeyao Ma +10
cs.CVarXiv:2507.22058v12025Uncertainty-aware Score Distribution Learning for Action Quality Assessment
Yansong Tang, Zanlin Ni, Jiahuan Zhou +4
cs.CVarXiv:2006.07665v12020Single Image Reflection Removal Exploiting Misaligned Training Data and Network Enhancements
Kaixuan Wei, Jiaolong Yang, Ying Fu +2
cs.CVarXiv:1904.00637v12019Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
Weiqi Li, Xuanyu Zhang, Shijie Zhao +4
cs.CVarXiv:2503.22679v22025Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
Ranjan Sapkota, Manoj Karkee
cs.CVcs.AIarXiv:2510.09653v32025ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose Estimation
Yongzhi Su, Mahdi Saleh, Torben Fetzer +5
cs.CVarXiv:2203.09418v22022Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
cs.CVcs.CLarXiv:2502.10248v32025Controllable Image Captioning with Prompt-Conditioned Scene Rewards
Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim
cs.CVcs.CLcs.LGarXiv:2609.00709v12026XAttention: Block Sparse Attention with Antidiagonal Scoring
Ruyi Xu, Guangxuan Xiao, Haofeng Huang +2
cs.CLcs.CVarXiv:2503.16428v12025Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
Qianhao Yuan, Jie Lou, Xing Yu +4
cs.CVcs.AIcs.CLarXiv:2605.18740v42026A Unified RGB-T Saliency Detection Benchmark: Dataset, Baselines, Analysis and A Novel Approach
Chenglong Li, Guizhao Wang, Yunpeng Ma +3
cs.CVarXiv:1701.02829v12017PixMix: Dreamlike Pictures Comprehensively Improve Safety Measures
Dan Hendrycks, Andy Zou, Mantas Mazeika +4
cs.LGcs.CVarXiv:2112.05135v32021Deep Exemplar-based Video Colorization
Bo Zhang, Mingming He, Jing Liao +4
cs.CVcs.AIcs.LGarXiv:1906.09909v12019Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
Hidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan +2
cs.CVarXiv:2511.20649v32025EV-SegNet: Semantic Segmentation for Event-based Cameras
Iñigo Alonso, Ana C. Murillo
cs.CVarXiv:1811.12039v12018BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation
Yu Yao, Ella Atkins, Matthew Johnson-Roberson +2
cs.CVcs.ROarXiv:2007.14558v22020ArtEmis: Affective Language for Visual Art
Panos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov +2
cs.CVcs.CLarXiv:2101.07396v12021Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang +4
cs.CVarXiv:2501.05444v12025Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise
Ryan Burgert, Yuancheng Xu, Wenqi Xian +10
cs.CVarXiv:2501.08331v52025SPIn-NeRF: Multiview Segmentation and Perceptual Inpainting with Neural Radiance Fields
Ashkan Mirzaei, Tristan Aumentado-Armstrong, Konstantinos G. Derpanis +4
cs.CVarXiv:2211.12254v22022Self-Supervised Learning for Videos: A Survey
Madeline C. Schiappa, Yogesh S. Rawat, Mubarak Shah
cs.CVcs.MMarXiv:2207.00419v32022Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Zeyuan Yang, Xueyang Yu, Delin Chen +2
cs.CVcs.AIarXiv:2506.17218v12025MmWave Radar and Vision Fusion for Object Detection in Autonomous Driving: A Review
Zhiqing Wei, Fengkai Zhang, Shuo Chang +3
cs.CVarXiv:2108.03004v32021Towards Understanding Regularization in Batch Normalization
Ping Luo, Xinjiang Wang, Wenqi Shao +1
cs.LGcs.CVeess.SYarXiv:1809.00846v42018Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
Shangzhe Di, Zhelun Yu, Guanghao Zhang +7
cs.CVarXiv:2503.00540v12025FG-CLIP: Fine-Grained Visual and Textual Alignment
Chunyu Xie, Bin Wang, Fanjing Kong +5
cs.CVcs.AIarXiv:2505.05071v32025Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning
Haokui Zhang, Ying Li, Yenan Jiang +3
cs.CVarXiv:2012.03439v12020Striking the Right Balance with Uncertainty
Salman Khan, Munawar Hayat, Waqas Zamir +2
cs.CVarXiv:1901.07590v32019Deep learning analysis of the myocardium in coronary CT angiography for identification of patients with functionally significant coronary artery stenosis
Majd Zreik, Nikolas Lessmann, Robbert W. van Hamersvelt +5
cs.CVarXiv:1711.08917v22017Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
Chetwin Low, Weimin Wang, Calder Katyal
cs.MMcs.CVcs.SDarXiv:2510.01284v12025NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
Chia-Yu Hung, Qi Sun, Pengfei Hong +5
cs.ROcs.AIcs.CVarXiv:2504.19854v12025Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Shuang Wu, Youtian Lin, Feihu Zhang +8
cs.CVarXiv:2505.17412v22025Fully-Convolutional Point Networks for Large-Scale Point Clouds
Dario Rethage, Johanna Wald, Jürgen Sturm +2
cs.CVarXiv:1808.06840v12018SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
Haozheng Yu, Xinyu Yang, Rundong Luo +2
cs.CVarXiv:2608.31023v12026Visual prompt engineering for video models
Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7
cs.CVcs.AIarXiv:2607.25537v12026WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
Runsheng Xu, Hubert Lin, Wonseok Jeon +11
cs.CVcs.AIarXiv:2510.26125v32025Bilateral Attention Network for RGB-D Salient Object Detection
Zhao Zhang, Zheng Lin, Jun Xu +3
cs.CVarXiv:2004.14582v12020Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Bingnan Li, Haozhe Wang, Haozhong Xiong +5
cs.CVcs.AIcs.LGarXiv:2607.24731v22026CVM-Cervix: A Hybrid Cervical Pap-Smear Image Classification Framework Using CNN, Visual Transformer and Multilayer Perceptron
Wanli Liu, Chen Li, Ning Xu +9
cs.CVarXiv:2206.00971v12022Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
Jiho Choi, Seonho Lee, Seojeong Park +1
cs.CVarXiv:2606.23557v12026Detection of Christmas tree plantations from high-resolution aerial imagery. A case study in the French Morvan
Francesca Razzano, Emanuele Dalsasso, Adrien Baysse-Lainé +3
cs.CVarXiv:2608.27290v12026A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
Ranjan Sapkota, William Bu, Chen Chen +2
cs.CVcs.AIarXiv:2608.24935v12026GeoWAM: Visual Geometry World Action Models for Autonomous Driving
Yiren Lu, Xin Ye, Jiaming Liu +9
cs.CVcs.ROarXiv:2608.23486v12026FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
Dingyun Zhang, Lixue Gong, Wei Liu
cs.CVarXiv:2607.18227v12026Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing
Keyan Hu, Mingtao Wang, Ziyu Zhou +4
cs.CVarXiv:2608.28517v12026GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
Arka Mukherjee, Soham Roy, Kartikeya Trivedi +1
cs.CVcs.CLarXiv:2608.29483v12026Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song +22
cs.CVarXiv:2510.08668v22025Beyond Accuracy: Quantifying Pulmonary Attribution in Anatomy-Guided Chest X-Ray Classification Under Domain Shift
Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash +2
cs.CVarXiv:2608.30467v12026Systematic Literature Review of Machine Learning Models and Applications for Text Recognition
Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi +5
cs.CVcs.LGarXiv:2608.26500v12026Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References
Cong Cao, Huanjing Yue, Xin Liu +1
cs.CVarXiv:2608.26476v12026SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images
Zhishe Wang, Yanlin Chen, Wenyu Shao +2
cs.CVarXiv:2204.11436v12022When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
Ruoqi Hu, Chulin Zhao, Jiashuo Chang +2
cs.CVcs.AIarXiv:2608.25933v12026CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact
Rana Muhammad Ahmed, Sabahat Abbas
cs.CVcs.LGarXiv:2608.25539v12026OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Zhaochen Su, Linjie Li, Mingyang Song +8
cs.CVarXiv:2505.08617v22025TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
Mark YU, Wenbo Hu, Jinbo Xing +1
cs.CVcs.AIcs.GRarXiv:2503.05638v12025Bridging Adversarial and Collaborative Learning for AI-Generated Image Quality Assessment
Baoliang Chen, Qing Lin, Sijie Mai
cs.CVarXiv:2608.24372v12026Primate vision reveals a missing principle for robust dynamic AI
Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
cs.CVq-bio.NCarXiv:2608.23790v12026Mover360: Controllable Object Manipulation in 360° Panoramic Images
Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers +1
cs.CVarXiv:2608.23238v12026Hybrid Generative-Discriminative Object Placement
Siyuan Zhou, Li Niu
cs.CVarXiv:2608.22692v12026