Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

781 to 840 of 18,866

  1. Video-Based Reward Modeling for Computer-Use Agents

    Linxin Song, Jieyu Zhang, Huanxin Sheng +6

    cs.CVcs.CLarXiv:2603.10178v12026
  2. Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch

    Hossein Souri, Liam Fowl, Rama Chellappa +2

    cs.LGcs.CRcs.CVarXiv:2106.08970v32021
  3. A path following algorithm for the graph matching problem

    Mikhail Zaslavskiy, Francis Bach, Jean-Philippe Vert

    cs.CVcs.DMarXiv:0801.3654v22008
  4. Anticipative Video Transformer

    Rohit Girdhar, Kristen Grauman

    cs.CVcs.AIcs.LGarXiv:2106.02036v22021
  5. Intriguing Properties of Vision Transformers

    Muzammal Naseer, Kanchana Ranasinghe, Salman Khan +3

    cs.CVcs.AIcs.LGarXiv:2105.10497v32021
  6. MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

    Baorong Shi, Bo Cui, Boyuan Jiang +17

    cs.CLcs.AIcs.CVarXiv:2602.12705v42026
  7. Applications of Deep Learning Techniques for Automated Multiple Sclerosis Detection Using Magnetic Resonance Imaging: A Review

    Afshin Shoeibi, Marjane Khodatars, Mahboobeh Jafari +9

    eess.IVcs.CVarXiv:2105.04881v22021
  8. Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search

    Tianming Liang, Qirui Du, Jian-Fang Hu +3

    cs.CVarXiv:2602.04454v12026
  9. Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion

    Hanmo Chen, Chenghao Xu, Xu Yang +2

    cs.CVarXiv:2601.21896v32026
  10. Stable and Scalable Bundle Adjustment of Holistic 3D Structures

    Shaohui Liu, Rémi Pautrat, Daniel Barath +3

    cs.CVarXiv:2609.04026v12026
  11. Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

    Denis M. Akola, David F. Fouhey

    cs.CVarXiv:2609.04174v12026
  12. Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets

    Jiashi Feng, Xiu Li, Jing Lin +25

    eess.IVcs.CVarXiv:2510.19944v12025
  13. Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

    Yikang Ding, Jiwen Liu, Wenyuan Zhang +11

    cs.CVarXiv:2509.09595v22025
  14. CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images

    Chengqi Duan, Kaiyue Sun, Rongyao Fang +10

    cs.CVcs.AIarXiv:2510.11718v12025
  15. CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

    Long Xing, Xiaoyi Dong, Yuhang Zang +6

    cs.CVcs.AIcs.CLarXiv:2509.22647v12025
  16. Boosting of Image Denoising Algorithms

    Yaniv Romano, Michael Elad

    cs.CVmath.NAarXiv:1502.06220v22015
  17. Ovis2.5 Technical Report

    Shiyin Lu, Yang Li, Yu Xia +39

    cs.CVcs.AIcs.CLarXiv:2508.11737v12025
  18. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    Haohan Chi, Huan-ang Gao, Ziming Liu +12

    cs.CVarXiv:2505.23757v12025
  19. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

    Yi-Fan Zhang, Tao Yu, Haochen Tian +17

    cs.CLcs.CVarXiv:2502.10391v12025
  20. HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images

    Chengxi Han, Chen Wu, Haonan Guo +2

    cs.CVarXiv:2404.09178v12024
  21. Explainable artificial intelligence (XAI) in deep learning-based medical image analysis

    Bas H. M. van der Velden, Hugo J. Kuijf, Kenneth G. A. Gilhuijs +1

    eess.IVcs.CVarXiv:2107.10912v12021
  22. Common Limitations of Image Processing Metrics: A Picture Story

    Annika Reinke, Minu D. Tizabi, Carole H. Sudre +90

    eess.IVcs.CVarXiv:2104.05642v82021
  23. Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image

    Zhengqin Li, Kalyan Sunkavalli, Manmohan Chandraker

    cs.CVarXiv:1804.05790v12018
  24. Image biomarker standardisation initiative

    Alex Zwanenburg, Stefan Leger, Martin Vallières +1

    cs.CVeess.IVarXiv:1612.07003v112016
  25. RGB-D-based Action Recognition Datasets: A Survey

    Jing Zhang, Wanqing Li, Philip O. Ogunbona +2

    cs.CVarXiv:1601.05511v12016
  26. Moving Object Detection by Detecting Contiguous Outliers in the Low-Rank Representation

    Xiaowei Zhou, Can Yang, Weichuan Yu

    cs.CVarXiv:1109.0882v22011
  27. Faster and better: a machine learning approach to corner detection

    Edward Rosten, Reid Porter, Tom Drummond

    cs.CVcs.LGarXiv:0810.2434v12008
  28. EditThinker: Unlocking Iterative Reasoning for Any Image Editor

    Hongyu Li, Manyuan Zhang, Dian Zheng +11

    cs.CVarXiv:2512.05965v12025
  29. SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

    Federico Putamorsi, Leonardo Zini, Marcella Cornia +1

    cs.CVarXiv:2609.03813v12026
  30. Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery

    Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai +6

    cs.CVarXiv:2510.15869v42025
  31. TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation

    Adeela Islam, Zorah Lähner, Vittorio Murino +1

    cs.CVarXiv:2609.04202v12026
  32. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

    Zhihong Chen, Xuehai Bai, Yang Shi +9

    cs.CVcs.AIarXiv:2509.24900v22025
  33. DA$^{2}$: Depth Anything in Any Direction

    Haodong Li, Wangguangdong Zheng, Jing He +5

    cs.CVarXiv:2509.26618v52025
  34. More Grounded Image Captioning by Distilling Image-Text Matching Model

    Yuanen Zhou, Meng Wang, Daqing Liu +2

    cs.CVcs.CLarXiv:2004.00390v12020
  35. Sparse auto-regressive modeling for scene generation from multi-view images

    Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4

    cs.CVcs.LGarXiv:2609.03931v12026
  36. OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot

    Junhan Zhu, Hesong Wang, Mingluo Su +2

    cs.CVarXiv:2510.06751v32025
  37. Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling

    Spyros Gidaris, Nikos Komodakis

    cs.CVcs.LGarXiv:1612.04770v12016
  38. When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

    Xuehao Wang, Jiaxin Hua, Runmei Li +4

    cs.CVarXiv:2609.03378v12026
  39. PointGT: Simultaneous Geometry and Texture Editing for Point-Based Representations

    Yanshu Zhang, George Shramko, Pratul P. Srinivasan +1

    cs.CVcs.GRarXiv:2609.03341v12026
  40. Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++

    Antonio Scardace, Francesco Guarnera, Sebastiano Battiato +1

    cs.CVarXiv:2609.03615v12026
  41. LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning

    Shenghao Fu, Qize Yang, Yuan-Ming Li +3

    cs.CVarXiv:2509.24786v12025
  42. Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

    Huanyu Zhang, Wenshan Wu, Chengzu Li +9

    cs.CVcs.CLarXiv:2510.24514v12025
  43. Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

    Shilong Zhang, He Zhang, Zhifei Zhang +11

    cs.CVarXiv:2512.17909v12025
  44. DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions

    Zhenjia Xu, Jiajun Wu, Andy Zeng +2

    cs.ROcs.AIcs.CVarXiv:1906.03853v22019
  45. A Survey on Efficient Vision-Language-Action Models

    Zhaoshu Yu, Bo Wang, Pengpeng Zeng +7

    cs.CVcs.AIcs.LGarXiv:2510.24795v22025
  46. RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning

    Sicheng Feng, Kaiwen Tuo, Song Wang +3

    cs.CVcs.AIarXiv:2510.02240v22025
  47. VABench: A Comprehensive Benchmark for Audio-Video Generation

    Daili Hua, Xizhi Wang, Bohan Zeng +6

    cs.CVcs.SDarXiv:2512.09299v22025
  48. An End-to-End Network for Panoptic Segmentation

    Huanyu Liu, Chao Peng, Changqian Yu +4

    cs.CVarXiv:1903.05027v22019
  49. Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models

    Runqian Wang, Yilun Du

    cs.LGcs.AIcs.CVarXiv:2510.02300v32025
  50. Salient ImageNet: How to discover spurious features in Deep Learning?

    Sahil Singla, Soheil Feizi

    cs.LGcs.CVarXiv:2110.04301v42021
  51. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead

    Chaojun Ni, Cheng Chen, Xiaofeng Wang +12

    cs.CVcs.ROarXiv:2512.00903v12025
  52. Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models

    Pu Jian, Junhong Wu, Wei Sun +3

    cs.CVcs.CLarXiv:2509.12132v12025
  53. UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation

    Yibin Wang, Zhimin Li, Yuhang Zang +8

    cs.CVarXiv:2510.18701v22025
  54. MedQA-MM: Shortcuts Behind Medical Visual Reasoning

    Benlu Wang, Yifan Zhang, Jiaqing Yu +7

    cs.CVcs.CLarXiv:2609.03261v22026
  55. WorldGen: From Text to Traversable and Interactive 3D Worlds

    Dilin Wang, Hyunyoung Jung, Tom Monnier +22

    cs.CVcs.AIarXiv:2511.16825v12025
  56. Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

    Onat Şahin, Mohammad Altillawi, George Eskandar +2

    cs.CVarXiv:2609.03657v12026
  57. ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

    Jiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen +5

    cs.CVarXiv:2510.06208v12025
  58. Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

    Cheng Yang, Haiyuan Wan, Yiran Peng +8

    cs.CVcs.AIarXiv:2511.15065v22025
  59. Sharp Monocular View Synthesis in Less Than a Second

    Lars Mescheder, Wei Dong, Shiwei Li +10

    cs.CVcs.LGarXiv:2512.10685v22025
  60. Observation-Conditioned Latent Energy Priors for Sparse Implicit Neural Shape Completion

    Paul Büschl, Ezequiel de la Rosa, Julia Wolleb +3

    cs.CVarXiv:2609.03694v12026