Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
5,701 to 5,760 of 18,794
Deep Exemplar-based Video Colorization
Bo Zhang, Mingming He, Jing Liao +4
cs.CVcs.AIcs.LGarXiv:1906.09909v12019Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
Hidir Yesiltepe, Tuna Han Salih Meral, Adil Kaan Akan +2
cs.CVarXiv:2511.20649v32025EV-SegNet: Semantic Segmentation for Event-based Cameras
Iñigo Alonso, Ana C. Murillo
cs.CVarXiv:1811.12039v12018BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation
Yu Yao, Ella Atkins, Matthew Johnson-Roberson +2
cs.CVcs.ROarXiv:2007.14558v22020ArtEmis: Affective Language for Visual Art
Panos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov +2
cs.CVcs.CLarXiv:2101.07396v12021Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang +4
cs.CVarXiv:2501.05444v12025Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise
Ryan Burgert, Yuancheng Xu, Wenqi Xian +10
cs.CVarXiv:2501.08331v52025SPIn-NeRF: Multiview Segmentation and Perceptual Inpainting with Neural Radiance Fields
Ashkan Mirzaei, Tristan Aumentado-Armstrong, Konstantinos G. Derpanis +4
cs.CVarXiv:2211.12254v22022Self-Supervised Learning for Videos: A Survey
Madeline C. Schiappa, Yogesh S. Rawat, Mubarak Shah
cs.CVcs.MMarXiv:2207.00419v32022Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Zeyuan Yang, Xueyang Yu, Delin Chen +2
cs.CVcs.AIarXiv:2506.17218v12025MmWave Radar and Vision Fusion for Object Detection in Autonomous Driving: A Review
Zhiqing Wei, Fengkai Zhang, Shuo Chang +3
cs.CVarXiv:2108.03004v32021Towards Understanding Regularization in Batch Normalization
Ping Luo, Xinjiang Wang, Wenqi Shao +1
cs.LGcs.CVeess.SYarXiv:1809.00846v42018Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
Shangzhe Di, Zhelun Yu, Guanghao Zhang +7
cs.CVarXiv:2503.00540v12025FG-CLIP: Fine-Grained Visual and Textual Alignment
Chunyu Xie, Bin Wang, Fanjing Kong +5
cs.CVcs.AIarXiv:2505.05071v32025Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning
Haokui Zhang, Ying Li, Yenan Jiang +3
cs.CVarXiv:2012.03439v12020Striking the Right Balance with Uncertainty
Salman Khan, Munawar Hayat, Waqas Zamir +2
cs.CVarXiv:1901.07590v32019Deep learning analysis of the myocardium in coronary CT angiography for identification of patients with functionally significant coronary artery stenosis
Majd Zreik, Nikolas Lessmann, Robbert W. van Hamersvelt +5
cs.CVarXiv:1711.08917v22017Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
Chetwin Low, Weimin Wang, Calder Katyal
cs.MMcs.CVcs.SDarXiv:2510.01284v12025NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
Chia-Yu Hung, Qi Sun, Pengfei Hong +5
cs.ROcs.AIcs.CVarXiv:2504.19854v12025Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Shuang Wu, Youtian Lin, Feihu Zhang +8
cs.CVarXiv:2505.17412v22025Fully-Convolutional Point Networks for Large-Scale Point Clouds
Dario Rethage, Johanna Wald, Jürgen Sturm +2
cs.CVarXiv:1808.06840v12018SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
Haozheng Yu, Xinyu Yang, Rundong Luo +2
cs.CVarXiv:2608.31023v12026Visual prompt engineering for video models
Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7
cs.CVcs.AIarXiv:2607.25537v12026WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
Runsheng Xu, Hubert Lin, Wonseok Jeon +11
cs.CVcs.AIarXiv:2510.26125v32025Bilateral Attention Network for RGB-D Salient Object Detection
Zhao Zhang, Zheng Lin, Jun Xu +3
cs.CVarXiv:2004.14582v12020Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Bingnan Li, Haozhe Wang, Haozhong Xiong +5
cs.CVcs.AIcs.LGarXiv:2607.24731v22026CVM-Cervix: A Hybrid Cervical Pap-Smear Image Classification Framework Using CNN, Visual Transformer and Multilayer Perceptron
Wanli Liu, Chen Li, Ning Xu +9
cs.CVarXiv:2206.00971v12022Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
Jiho Choi, Seonho Lee, Seojeong Park +1
cs.CVarXiv:2606.23557v12026Detection of Christmas tree plantations from high-resolution aerial imagery. A case study in the French Morvan
Francesca Razzano, Emanuele Dalsasso, Adrien Baysse-Lainé +3
cs.CVarXiv:2608.27290v12026A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
Ranjan Sapkota, William Bu, Chen Chen +2
cs.CVcs.AIarXiv:2608.24935v12026GeoWAM: Visual Geometry World Action Models for Autonomous Driving
Yiren Lu, Xin Ye, Jiaming Liu +9
cs.CVcs.ROarXiv:2608.23486v12026FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
Dingyun Zhang, Lixue Gong, Wei Liu
cs.CVarXiv:2607.18227v12026Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing
Keyan Hu, Mingtao Wang, Ziyu Zhou +4
cs.CVarXiv:2608.28517v12026GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation
Arka Mukherjee, Soham Roy, Kartikeya Trivedi +1
cs.CVcs.CLarXiv:2608.29483v12026Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song +22
cs.CVarXiv:2510.08668v22025Beyond Accuracy: Quantifying Pulmonary Attribution in Anatomy-Guided Chest X-Ray Classification Under Domain Shift
Abdullah Al Mamun, Md. Nasif Osman Khansur, Md Ashraful Hossen Akash +2
cs.CVarXiv:2608.30467v12026Systematic Literature Review of Machine Learning Models and Applications for Text Recognition
Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi +5
cs.CVcs.LGarXiv:2608.26500v12026Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References
Cong Cao, Huanjing Yue, Xin Liu +1
cs.CVarXiv:2608.26476v12026SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images
Zhishe Wang, Yanlin Chen, Wenyu Shao +2
cs.CVarXiv:2204.11436v12022When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
Ruoqi Hu, Chulin Zhao, Jiashuo Chang +2
cs.CVcs.AIarXiv:2608.25933v12026CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact
Rana Muhammad Ahmed, Sabahat Abbas
cs.CVcs.LGarXiv:2608.25539v12026OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Zhaochen Su, Linjie Li, Mingyang Song +8
cs.CVarXiv:2505.08617v22025TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
Mark YU, Wenbo Hu, Jinbo Xing +1
cs.CVcs.AIcs.GRarXiv:2503.05638v12025Bridging Adversarial and Collaborative Learning for AI-Generated Image Quality Assessment
Baoliang Chen, Qing Lin, Sijie Mai
cs.CVarXiv:2608.24372v12026Primate vision reveals a missing principle for robust dynamic AI
Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
cs.CVq-bio.NCarXiv:2608.23790v12026Mover360: Controllable Object Manipulation in 360° Panoramic Images
Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers +1
cs.CVarXiv:2608.23238v12026Hybrid Generative-Discriminative Object Placement
Siyuan Zhou, Li Niu
cs.CVarXiv:2608.22692v12026ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation
Qianqian Tang, Jiayi Gao, Ting Lei +1
cs.MMcs.CVarXiv:2608.22329v12026Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
Rujin Liang, Zhongpu Chen, Yuhao Lei +1
cs.CVcs.AIarXiv:2608.20756v12026Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores
Xiang Fu, Jixiang Ma, Xinpeng Zhang +3
cs.DCcs.CVarXiv:2608.20725v12026Pre-training Visual Dexterity in Simulation
Sarthak Kamat, Adam Rashid, Satvik Sharma +4
cs.ROcs.AIcs.CVarXiv:2608.15917v12026The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang
cs.AIcs.CVarXiv:2608.14558v12026Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
Haibo Wang, Zihao Lin, Zhiyang Xu +1
cs.CVcs.AIarXiv:2604.00528v22026HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
cs.AIcs.CLcs.CVarXiv:2606.31179v12026Summaries:한국어INSID3: Training-Free In-Context Segmentation with DINOv3
Claudia Cuttano, Gabriele Trivigno, Christoph Reich +3
cs.CVarXiv:2603.28480v12026CLIP4Caption: CLIP for Video Caption
Mingkang Tang, Zhanyu Wang, Zhenhua Liu +3
cs.CVarXiv:2110.06615v12021DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects
Tianshan Zhang, Yijia Duan, Yanjun Li +2
cs.ROcs.CVarXiv:2606.15133v12026SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Seokju Cho, Ryo Hachiuma, Abhishek Badki +8
cs.CVcs.AIarXiv:2606.13673v12026LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
cs.CVcs.CLarXiv:2603.27538v12026OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
cs.CVarXiv:2209.07526v22022