Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,961 to 7,020 of 18,817

  1. Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

    Wenxuan Huang, Bohan Jia, Zijie Zhai +7

    cs.CVcs.AIcs.CLarXiv:2503.06749v42025
  2. Grounded Human-Object Interaction Hotspots from Video

    Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

    cs.CVarXiv:1812.04558v22018
  3. Comprehensive Graph-conditional Similarity Preserving Network for Unsupervised Cross-modal Hashing

    Jun Yu, Hao Zhou, Yibing Zhan +1

    cs.IRcs.CVarXiv:2012.13538v12020
  4. Understanding metric-related pitfalls in image analysis validation

    Annika Reinke, Minu D. Tizabi, Michael Baumgartner +75

    cs.CVarXiv:2302.01790v42023
  5. MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

    Yubo Ma, Yuhang Zang, Liangyu Chen +13

    cs.CVcs.CLarXiv:2407.01523v32024
  6. Qwen3-Omni Technical Report

    Jin Xu, Zhifang Guo, Hangrui Hu +35

    cs.CLcs.AIcs.CVarXiv:2509.17765v12025
  7. A Benchmark for Lidar Sensors in Fog: Is Detection Breaking Down?

    Mario Bijelic, Tobias Gruber, Werner Ritter

    cs.CVarXiv:1912.03251v12019
  8. Spot the conversation: speaker diarisation in the wild

    Joon Son Chung, Jaesung Huh, Arsha Nagrani +2

    cs.SDcs.CVeess.ASarXiv:2007.01216v32020
  9. SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection

    Rui Fan, Hengli Wang, Peide Cai +1

    cs.CVcs.ROeess.IVarXiv:2008.11351v12020
  10. InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation

    Vanshika Vats, Ashwani Rathee, James Davis

    cs.CVcs.AIarXiv:2609.02002v12026
  11. Continuous 3D Perception Model with Persistent State

    Qianqian Wang, Yifei Zhang, Aleksander Holynski +2

    cs.CVarXiv:2501.12387v12025
  12. Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment

    Li Ding, Chenliang Xu

    cs.CVarXiv:1803.10699v12018
  13. Video-R1: Reinforcing Video Reasoning in MLLMs

    Kaituo Feng, Kaixiong Gong, Bohao Li +7

    cs.CVarXiv:2503.21776v42025
  14. Sparse Instance Activation for Real-Time Instance Segmentation

    Tianheng Cheng, Xinggang Wang, Shaoyu Chen +5

    cs.CVarXiv:2203.12827v12022
  15. Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment

    Royi Rassin, Eran Hirsch, Daniel Glickman +3

    cs.CLcs.CVarXiv:2306.08877v32023
  16. VACE: All-in-One Video Creation and Editing

    Zeyinzi Jiang, Zhen Han, Chaojie Mao +3

    cs.CVarXiv:2503.07598v22025
  17. Prediction Poisoning: Towards Defenses Against DNN Model Stealing Attacks

    Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

    cs.LGcs.CRcs.CVarXiv:1906.10908v22019
  18. VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

    Xuan He, Dongfu Jiang, Ge Zhang +16

    cs.CVcs.AIarXiv:2406.15252v32024
  19. SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

    Tianzhe Chu, Yuexiang Zhai, Jihan Yang +6

    cs.AIcs.CVcs.LGarXiv:2501.17161v22025
  20. UI-TARS: Pioneering Automated GUI Interaction with Native Agents

    Yujia Qin, Yining Ye, Junjie Fang +32

    cs.AIcs.CLcs.CVarXiv:2501.12326v12025
  21. RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

    Tianxing Chen, Zanxin Chen, Baijun Chen +23

    cs.ROcs.AIcs.CLarXiv:2506.18088v22025
  22. Motion-Aware Feature for Improved Video Anomaly Detection

    Yi Zhu, Shawn Newsam

    cs.CVcs.LGeess.IVarXiv:1907.10211v12019
  23. Deep Image Spatial Transformation for Person Image Generation

    Yurui Ren, Xiaoming Yu, Junming Chen +2

    cs.CVcs.AIarXiv:2003.00696v22020
  24. Mean Flows for One-step Generative Modeling

    Zhengyang Geng, Mingyang Deng, Xingjian Bai +2

    cs.LGcs.CVarXiv:2505.13447v12025
  25. Flow-GRPO: Training Flow Matching Models via Online RL

    Jie Liu, Gongye Liu, Jiajun Liang +6

    cs.CVcs.AIarXiv:2505.05470v52025
  26. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

    Mido Assran, Adrien Bardes, David Fan +27

    cs.AIcs.CVcs.LGarXiv:2506.09985v12025
  27. MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

    Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5

    cs.CLcs.CVarXiv:2609.01772v12026
  28. Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

    Birgit Nierula, Karam Tomotaki-Dawoud, Mert Akguel +5

    cs.CVcs.HCarXiv:2609.03569v12026
  29. Abnormal respiratory patterns classifier may contribute to large-scale screening of people infected with COVID-19 in an accurate and unobtrusive manner

    Yunlu Wang, Menghan Hu, Qingli Li +3

    cs.LGcs.CVeess.SParXiv:2002.05534v22020
  30. Conditional Generative Neural System for Probabilistic Trajectory Prediction

    Jiachen Li, Hengbo Ma, Masayoshi Tomizuka

    cs.CVcs.AIcs.LGarXiv:1905.01631v22019
  31. Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief

    Robert Engel

    eess.IVcs.CVcs.HCarXiv:2609.03095v12026
  32. Towards Interpretable Semantic Segmentation via Gradient-weighted Class Activation Mapping

    Kira Vinogradova, Alexandr Dibrov, Gene Myers

    cs.CVcs.LGeess.IVarXiv:2002.11434v12020
  33. Supervised Classification Performance of Multispectral Images

    K. Perumal, R. Bhaskaran

    cs.LGcs.CVarXiv:1002.4046v12010
  34. Scene Parsing with Multiscale Feature Learning, Purity Trees, and Optimal Covers

    Clément Farabet, Camille Couprie, Laurent Najman +1

    cs.CVcs.LGarXiv:1202.2160v22012
  35. YOLOv5, YOLOv8 and YOLOv10: The Go-To Detectors for Real-time Vision

    Muhammad Hussain

    cs.CVarXiv:2407.02988v12024
  36. Subjective and Objective Quality Assessment of Image: A Survey

    Pedram Mohammadi, Abbas Ebrahimi-Moghadam, Shahram Shirani

    cs.MMcs.CVarXiv:1406.7799v12014
  37. Tensor-based Brain Surface Modeling and Analysis

    Moo K. Chung, Keith J. Worsley, Steve Robbins +1

    cs.CVq-bio.NCarXiv:2609.03302v12026
  38. HELIOS: From midnight to noon, continuous outdoor urban scene relighting

    Hala Djeghim, Nathan Piasco, Luis Roldão +4

    cs.CVarXiv:2609.00901v22026
  39. Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments

    Kai Niu, Yan Huang, Wanli Ouyang +1

    cs.CVarXiv:1906.09610v12019
  40. Image Fusion Transformer

    Vibashan VS, Jeya Maria Jose Valanarasu, Poojan Oza +1

    cs.CVarXiv:2107.09011v42021
  41. GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

    Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri +10

    eess.IVcs.AIcs.CVarXiv:2609.01310v12026
  42. Benchmarking Spatial, Spectral, and Self-Supervised Cues for Face Forgery Detection under Realistic Degradation

    Lucas Cunha, Lucas Sotomaior, Lucas Gasperin +3

    cs.CVarXiv:2609.01511v12026
  43. Prostate Cancer Detection using Deep Convolutional Neural Networks

    Sunghwan Yoo, Isha Gujrathi, Masoom A. Haider +1

    cs.CVeess.IVq-bio.QMarXiv:1905.13145v12019
  44. Do Satellites See Commuters? A Critical Benchmark of Vision Foundation Models

    Ashiq Shukoor Iqbal, Wilson Wongso, Flora D. Salim

    cs.CVarXiv:2609.00661v12026
  45. Robust Unsupervised Video Anomaly Detection by Multi-Path Frame Prediction

    Xuanzhao Wang, Zhengping Che, Bo Jiang +6

    cs.CVcs.LGarXiv:2011.02763v22020
  46. Analyzing Classifiers: Fisher Vectors and Deep Neural Networks

    Sebastian Bach, Alexander Binder, Grégoire Montavon +2

    cs.CVarXiv:1512.00172v12015
  47. Harvesting Multiple Views for Marker-less 3D Human Pose Annotations

    Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis +1

    cs.CVarXiv:1704.04793v12017
  48. Prompting for Multi-Modal Tracking

    Jinyu Yang, Zhe Li, Feng Zheng +2

    cs.CVarXiv:2207.14571v22022
  49. Hierarchical Recurrent Neural Network for Video Summarization

    Bin Zhao, Xuelong Li, Xiaoqiang Lu

    cs.CVarXiv:1904.12251v12019
  50. I-ViT: Integer-only Quantization for Efficient Vision Transformer Inference

    Zhikai Li, Qingyi Gu

    cs.CVarXiv:2207.01405v42022
  51. BoWFire: Detection of Fire in Still Images by Integrating Pixel Color and Texture Analysis

    Daniel Y. T. Chino, Letricia P. S. Avalhais, Jose F. Rodrigues +1

    cs.CVarXiv:1506.03495v12015
  52. DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation

    Atif Belal, Lilian Hollard, Marco Pedersoli +1

    cs.CVarXiv:2609.01795v12026
  53. Hetero-Center Loss for Cross-Modality Person Re-Identification

    Yuanxin Zhu, Zhao Yang, Li Wang +3

    cs.CVeess.IVarXiv:1910.09830v12019
  54. SAM 3: Segment Anything with Concepts

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35

    cs.CVcs.AIarXiv:2511.16719v22025
  55. Noise Flow: Noise Modeling with Conditional Normalizing Flows

    Abdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. Brown

    cs.CVcs.LGeess.IVarXiv:1908.08453v12019
  56. Emerging Properties in Unified Multimodal Pretraining

    Chaorui Deng, Deyao Zhu, Kunchang Li +9

    cs.CVarXiv:2505.14683v32025
  57. DADA: Depth-aware Domain Adaptation in Semantic Segmentation

    Tuan-Hung Vu, Himalaya Jain, Maxime Bucher +2

    cs.CVarXiv:1904.01886v32019
  58. Convolutional Neural Network Pruning with Structural Redundancy Reduction

    Zi Wang, Chengcheng Li, Xiangyang Wang

    cs.CVcs.LGarXiv:2104.03438v12021
  59. Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

    Chin-Yang Lin, Yang-Che Sun, Cheng Sun +5

    cs.CVarXiv:2609.04201v12026
  60. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

    Xun Huang, Zhengqi Li, Guande He +2

    cs.CVcs.AIcs.LGarXiv:2506.08009v22025