Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,961 to 7,020 of 18,817
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Wenxuan Huang, Bohan Jia, Zijie Zhai +7
cs.CVcs.AIcs.CLarXiv:2503.06749v42025Grounded Human-Object Interaction Hotspots from Video
Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman
cs.CVarXiv:1812.04558v22018Comprehensive Graph-conditional Similarity Preserving Network for Unsupervised Cross-modal Hashing
Jun Yu, Hao Zhou, Yibing Zhan +1
cs.IRcs.CVarXiv:2012.13538v12020Understanding metric-related pitfalls in image analysis validation
Annika Reinke, Minu D. Tizabi, Michael Baumgartner +75
cs.CVarXiv:2302.01790v42023MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Yubo Ma, Yuhang Zang, Liangyu Chen +13
cs.CVcs.CLarXiv:2407.01523v32024Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
cs.CLcs.AIcs.CVarXiv:2509.17765v12025A Benchmark for Lidar Sensors in Fog: Is Detection Breaking Down?
Mario Bijelic, Tobias Gruber, Werner Ritter
cs.CVarXiv:1912.03251v12019Spot the conversation: speaker diarisation in the wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani +2
cs.SDcs.CVeess.ASarXiv:2007.01216v32020SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection
Rui Fan, Hengli Wang, Peide Cai +1
cs.CVcs.ROeess.IVarXiv:2008.11351v12020InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation
Vanshika Vats, Ashwani Rathee, James Davis
cs.CVcs.AIarXiv:2609.02002v12026Continuous 3D Perception Model with Persistent State
Qianqian Wang, Yifei Zhang, Aleksander Holynski +2
cs.CVarXiv:2501.12387v12025Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment
Li Ding, Chenliang Xu
cs.CVarXiv:1803.10699v12018Video-R1: Reinforcing Video Reasoning in MLLMs
Kaituo Feng, Kaixiong Gong, Bohao Li +7
cs.CVarXiv:2503.21776v42025Sparse Instance Activation for Real-Time Instance Segmentation
Tianheng Cheng, Xinggang Wang, Shaoyu Chen +5
cs.CVarXiv:2203.12827v12022Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
Royi Rassin, Eran Hirsch, Daniel Glickman +3
cs.CLcs.CVarXiv:2306.08877v32023VACE: All-in-One Video Creation and Editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao +3
cs.CVarXiv:2503.07598v22025Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks
Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz
cs.LGcs.CRcs.CVarXiv:1906.10908v22019VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
Xuan He, Dongfu Jiang, Ge Zhang +16
cs.CVcs.AIarXiv:2406.15252v32024SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Tianzhe Chu, Yuexiang Zhai, Jihan Yang +6
cs.AIcs.CVcs.LGarXiv:2501.17161v22025UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Yujia Qin, Yining Ye, Junjie Fang +32
cs.AIcs.CLcs.CVarXiv:2501.12326v12025RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
Tianxing Chen, Zanxin Chen, Baijun Chen +23
cs.ROcs.AIcs.CLarXiv:2506.18088v22025Motion-Aware Feature for Improved Video Anomaly Detection
Yi Zhu, Shawn Newsam
cs.CVcs.LGeess.IVarXiv:1907.10211v12019Deep Image Spatial Transformation for Person Image Generation
Yurui Ren, Xiaoming Yu, Junming Chen +2
cs.CVcs.AIarXiv:2003.00696v22020Mean Flows for One-step Generative Modeling
Zhengyang Geng, Mingyang Deng, Xingjian Bai +2
cs.LGcs.CVarXiv:2505.13447v12025Flow-GRPO: Training Flow Matching Models via Online RL
Jie Liu, Gongye Liu, Jiajun Liang +6
cs.CVcs.AIarXiv:2505.05470v52025V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan +27
cs.AIcs.CVcs.LGarXiv:2506.09985v12025MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5
cs.CLcs.CVarXiv:2609.01772v12026Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG
Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel +5
cs.CVcs.HCarXiv:2609.03569v12026Abnormal respiratory patterns classifier may contribute to large-scale screening of people infected with COVID-19 in an accurate and unobtrusive manner
Yunlu Wang, Menghan Hu, Qingli Li +3
cs.LGcs.CVeess.SParXiv:2002.05534v22020Conditional Generative Neural System for Probabilistic Trajectory Prediction
Jiachen Li, Hengbo Ma, Masayoshi Tomizuka
cs.CVcs.AIcs.LGarXiv:1905.01631v22019Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief
Robert Engel
eess.IVcs.CVcs.HCarXiv:2609.03095v12026Towards Interpretable Semantic Segmentation via Gradient-weighted Class Activation Mapping
Kira Vinogradova, Alexandr Dibrov, Gene Myers
cs.CVcs.LGeess.IVarXiv:2002.11434v12020Supervised Classification Performance of Multispectral Images
K. Perumal, R. Bhaskaran
cs.LGcs.CVarXiv:1002.4046v12010Scene Parsing with Multiscale Feature Learning, Purity Trees, and Optimal Covers
Clément Farabet, Camille Couprie, Laurent Najman +1
cs.CVcs.LGarXiv:1202.2160v22012YOLOv5, YOLOv8 and YOLOv10: The Go-To Detectors for Real-time Vision
Muhammad Hussain
cs.CVarXiv:2407.02988v12024Subjective and Objective Quality Assessment of Image: A Survey
Pedram Mohammadi, Abbas Ebrahimi-Moghadam, Shahram Shirani
cs.MMcs.CVarXiv:1406.7799v12014Tensor-based Brain Surface Modeling and Analysis
Moo K. Chung, Keith J. Worsley, Steve Robbins +1
cs.CVq-bio.NCarXiv:2609.03302v12026HELIOS: From midnight to noon, continuous outdoor urban scene relighting
Hala Djeghim, Nathan Piasco, Luis Roldão +4
cs.CVarXiv:2609.00901v22026Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments
Kai Niu, Yan Huang, Wanli Ouyang +1
cs.CVarXiv:1906.09610v12019Image Fusion Transformer
Vibashan VS, Jeya Maria Jose Valanarasu, Poojan Oza +1
cs.CVarXiv:2107.09011v42021GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation
Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10
eess.IVcs.AIcs.CVarXiv:2609.01310v12026Benchmarking Spatial, Spectral, and Self-Supervised Cues for Face Forgery Detection under Realistic Degradation
Lucas Cunha, Lucas Sotomaior, Lucas Gasperin +3
cs.CVarXiv:2609.01511v12026Prostate Cancer Detection using Deep Convolutional Neural Networks
Sunghwan Yoo, Isha Gujrathi, Masoom A. Haider +1
cs.CVeess.IVq-bio.QMarXiv:1905.13145v12019Do Satellites See Commuters? A Critical Benchmark of Vision Foundation Models
Ashiq Shukoor Iqbal, Wilson Wongso, Flora D. Salim
cs.CVarXiv:2609.00661v12026Robust Unsupervised Video Anomaly Detection by Multi-Path Frame Prediction
Xuanzhao Wang, Zhengping Che, Bo Jiang +6
cs.CVcs.LGarXiv:2011.02763v22020Analyzing Classifiers: Fisher Vectors and Deep Neural Networks
Sebastian Bach, Alexander Binder, Grégoire Montavon +2
cs.CVarXiv:1512.00172v12015Harvesting Multiple Views for Marker-less 3D Human Pose Annotations
Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis +1
cs.CVarXiv:1704.04793v12017Prompting for Multi-Modal Tracking
Jinyu Yang, Zhe Li, Feng Zheng +2
cs.CVarXiv:2207.14571v22022Hierarchical Recurrent Neural Network for Video Summarization
Bin Zhao, Xuelong Li, Xiaoqiang Lu
cs.CVarXiv:1904.12251v12019I-ViT: Integer-only Quantization for Efficient Vision Transformer Inference
Zhikai Li, Qingyi Gu
cs.CVarXiv:2207.01405v42022BoWFire: Detection of Fire in Still Images by Integrating Pixel Color and Texture Analysis
Daniel Y. T. Chino, Letricia P. S. Avalhais, Jose F. Rodrigues +1
cs.CVarXiv:1506.03495v12015DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation
Atif Belal, Lilian Hollard, Marco Pedersoli +1
cs.CVarXiv:2609.01795v12026Hetero-Center Loss for Cross-Modality Person Re-Identification
Yuanxin Zhu, Zhao Yang, Li Wang +3
cs.CVeess.IVarXiv:1910.09830v12019SAM 3: Segment Anything with Concepts
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35
cs.CVcs.AIarXiv:2511.16719v22025Noise Flow: Noise Modeling with Conditional Normalizing Flows
Abdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. Brown
cs.CVcs.LGeess.IVarXiv:1908.08453v12019Emerging Properties in Unified Multimodal Pretraining
Chaorui Deng, Deyao Zhu, Kunchang Li +9
cs.CVarXiv:2505.14683v32025DADA: Depth-aware Domain Adaptation in Semantic Segmentation
Tuan-Hung Vu, Himalaya Jain, Maxime Bucher +2
cs.CVarXiv:1904.01886v32019Convolutional Neural Network Pruning with Structural Redundancy Reduction
Zi Wang, Chengcheng Li, Xiangyang Wang
cs.CVcs.LGarXiv:2104.03438v12021Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Chin-Yang Lin, Yang-Che Sun, Cheng Sun +5
cs.CVarXiv:2609.04201v12026Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
Xun Huang, Zhengqi Li, Guande He +2
cs.CVcs.AIcs.LGarXiv:2506.08009v22025