Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

721 to 780 of 18,810

  1. Towards Generalizable Robotic Manipulation in Dynamic Environments

    Heng Fang, Shangru Li, Shuhan Wang +3

    cs.CVcs.ROarXiv:2603.15620v32026
  2. Efficient Interactive Annotation of Segmentation Datasets with Polygon-RNN++

    David Acuna, Huan Ling, Amlan Kar +1

    cs.CVarXiv:1803.09693v12018
  3. MixStyle Neural Networks for Domain Generalization and Adaptation

    Kaiyang Zhou, Yongxin Yang, Yu Qiao +1

    cs.CVcs.AIcs.LGarXiv:2107.02053v22021
  4. Visual-ERM: Reward Modeling for Visual Equivalence

    Ziyu Liu, Shengyuan Ding, Xinyu Fang +7

    cs.CVcs.AIarXiv:2603.13224v22026
  5. Video-Based Reward Modeling for Computer-Use Agents

    Linxin Song, Jieyu Zhang, Huanxin Sheng +6

    cs.CVcs.CLarXiv:2603.10178v12026
  6. Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch

    Hossein Souri, Liam Fowl, Rama Chellappa +2

    cs.LGcs.CRcs.CVarXiv:2106.08970v32021
  7. A path following algorithm for the graph matching problem

    Mikhail Zaslavskiy, Francis Bach, Jean-Philippe Vert

    cs.CVcs.DMarXiv:0801.3654v22008
  8. Anticipative Video Transformer

    Rohit Girdhar, Kristen Grauman

    cs.CVcs.AIcs.LGarXiv:2106.02036v22021
  9. Intriguing Properties of Vision Transformers

    Muzammal Naseer, Kanchana Ranasinghe, Salman Khan +3

    cs.CVcs.AIcs.LGarXiv:2105.10497v32021
  10. MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

    Baorong Shi, Bo Cui, Boyuan Jiang +17

    cs.CLcs.AIcs.CVarXiv:2602.12705v42026
  11. Applications of Deep Learning Techniques for Automated Multiple Sclerosis Detection Using Magnetic Resonance Imaging: A Review

    Afshin Shoeibi, Marjane Khodatars, Mahboobeh Jafari +9

    eess.IVcs.CVarXiv:2105.04881v22021
  12. Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search

    Tianming Liang, Qirui Du, Jian-Fang Hu +3

    cs.CVarXiv:2602.04454v12026
  13. Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion

    Hanmo Chen, Chenghao Xu, Xu Yang +2

    cs.CVarXiv:2601.21896v32026
  14. Stable and Scalable Bundle Adjustment of Holistic 3D Structures

    Shaohui Liu, Rémi Pautrat, Daniel Barath +3

    cs.CVarXiv:2609.04026v12026
  15. Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

    Denis M. Akola, David F. Fouhey

    cs.CVarXiv:2609.04174v12026
  16. Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets

    Jiashi Feng, Xiu Li, Jing Lin +25

    eess.IVcs.CVarXiv:2510.19944v12025
  17. Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

    Yikang Ding, Jiwen Liu, Wenyuan Zhang +11

    cs.CVarXiv:2509.09595v22025
  18. CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images

    Chengqi Duan, Kaiyue Sun, Rongyao Fang +10

    cs.CVcs.AIarXiv:2510.11718v12025
  19. CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

    Long Xing, Xiaoyi Dong, Yuhang Zang +6

    cs.CVcs.AIcs.CLarXiv:2509.22647v12025
  20. Boosting of Image Denoising Algorithms

    Yaniv Romano, Michael Elad

    cs.CVmath.NAarXiv:1502.06220v22015
  21. Ovis2.5 Technical Report

    Shiyin Lu, Yang Li, Yu Xia +39

    cs.CVcs.AIcs.CLarXiv:2508.11737v12025
  22. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    Haohan Chi, Huan-ang Gao, Ziming Liu +12

    cs.CVarXiv:2505.23757v12025
  23. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

    Yi-Fan Zhang, Tao Yu, Haochen Tian +17

    cs.CLcs.CVarXiv:2502.10391v12025
  24. HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images

    Chengxi Han, Chen Wu, Haonan Guo +2

    cs.CVarXiv:2404.09178v12024
  25. Explainable artificial intelligence (XAI) in deep learning-based medical image analysis

    Bas H. M. van der Velden, Hugo J. Kuijf, Kenneth G. A. Gilhuijs +1

    eess.IVcs.CVarXiv:2107.10912v12021
  26. Common Limitations of Image Processing Metrics: A Picture Story

    Annika Reinke, Minu D. Tizabi, Carole H. Sudre +90

    eess.IVcs.CVarXiv:2104.05642v82021
  27. Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image

    Zhengqin Li, Kalyan Sunkavalli, Manmohan Chandraker

    cs.CVarXiv:1804.05790v12018
  28. Image biomarker standardisation initiative

    Alex Zwanenburg, Stefan Leger, Martin Vallières +1

    cs.CVeess.IVarXiv:1612.07003v112016
  29. RGB-D-based Action Recognition Datasets: A Survey

    Jing Zhang, Wanqing Li, Philip O. Ogunbona +2

    cs.CVarXiv:1601.05511v12016
  30. Moving Object Detection by Detecting Contiguous Outliers in the Low-Rank Representation

    Xiaowei Zhou, Can Yang, Weichuan Yu

    cs.CVarXiv:1109.0882v22011
  31. Faster and better: a machine learning approach to corner detection

    Edward Rosten, Reid Porter, Tom Drummond

    cs.CVcs.LGarXiv:0810.2434v12008
  32. EditThinker: Unlocking Iterative Reasoning for Any Image Editor

    Hongyu Li, Manyuan Zhang, Dian Zheng +11

    cs.CVarXiv:2512.05965v12025
  33. SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

    Federico Putamorsi, Leonardo Zini, Marcella Cornia +1

    cs.CVarXiv:2609.03813v12026
  34. Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery

    Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai +6

    cs.CVarXiv:2510.15869v42025
  35. TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation

    Adeela Islam, Zorah Lähner, Vittorio Murino +1

    cs.CVarXiv:2609.04202v12026
  36. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing

    Zhihong Chen, Xuehai Bai, Yang Shi +9

    cs.CVcs.AIarXiv:2509.24900v22025
  37. DA$^{2}$: Depth Anything in Any Direction

    Haodong Li, Wangguangdong Zheng, Jing He +5

    cs.CVarXiv:2509.26618v52025
  38. More Grounded Image Captioning by Distilling Image-Text Matching Model

    Yuanen Zhou, Meng Wang, Daqing Liu +2

    cs.CVcs.CLarXiv:2004.00390v12020
  39. Sparse auto-regressive modeling for scene generation from multi-view images

    Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4

    cs.CVcs.LGarXiv:2609.03931v12026
  40. OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot

    Junhan Zhu, Hesong Wang, Mingluo Su +2

    cs.CVarXiv:2510.06751v32025
  41. Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling

    Spyros Gidaris, Nikos Komodakis

    cs.CVcs.LGarXiv:1612.04770v12016
  42. When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

    Xuehao Wang, Jiaxin Hua, Runmei Li +4

    cs.CVarXiv:2609.03378v12026
  43. PointGT: Simultaneous Geometry and Texture Editing for Point-Based Representations

    Yanshu Zhang, George Shramko, Pratul P. Srinivasan +1

    cs.CVcs.GRarXiv:2609.03341v12026
  44. Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++

    Antonio Scardace, Francesco Guarnera, Sebastiano Battiato +1

    cs.CVarXiv:2609.03615v12026
  45. LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning

    Shenghao Fu, Qize Yang, Yuan-Ming Li +3

    cs.CVarXiv:2509.24786v12025
  46. Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs

    Huanyu Zhang, Wenshan Wu, Chengzu Li +9

    cs.CVcs.CLarXiv:2510.24514v12025
  47. Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

    Shilong Zhang, He Zhang, Zhifei Zhang +11

    cs.CVarXiv:2512.17909v12025
  48. DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions

    Zhenjia Xu, Jiajun Wu, Andy Zeng +2

    cs.ROcs.AIcs.CVarXiv:1906.03853v22019
  49. A Survey on Efficient Vision-Language-Action Models

    Zhaoshu Yu, Bo Wang, Pengpeng Zeng +7

    cs.CVcs.AIcs.LGarXiv:2510.24795v22025
  50. RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning

    Sicheng Feng, Kaiwen Tuo, Song Wang +3

    cs.CVcs.AIarXiv:2510.02240v22025
  51. VABench: A Comprehensive Benchmark for Audio-Video Generation

    Daili Hua, Xizhi Wang, Bohan Zeng +6

    cs.CVcs.SDarXiv:2512.09299v22025
  52. An End-to-End Network for Panoptic Segmentation

    Huanyu Liu, Chao Peng, Changqian Yu +4

    cs.CVarXiv:1903.05027v22019
  53. Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models

    Runqian Wang, Yilun Du

    cs.LGcs.AIcs.CVarXiv:2510.02300v32025
  54. Salient ImageNet: How to discover spurious features in Deep Learning?

    Sahil Singla, Soheil Feizi

    cs.LGcs.CVarXiv:2110.04301v42021
  55. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead

    Chaojun Ni, Cheng Chen, Xiaofeng Wang +12

    cs.CVcs.ROarXiv:2512.00903v12025
  56. Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models

    Pu Jian, Junhong Wu, Wei Sun +3

    cs.CVcs.CLarXiv:2509.12132v12025
  57. UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation

    Yibin Wang, Zhimin Li, Yuhang Zang +8

    cs.CVarXiv:2510.18701v22025
  58. MedQA-MM: Shortcuts Behind Medical Visual Reasoning

    Benlu Wang, Yifan Zhang, Jiaqing Yu +7

    cs.CVcs.CLarXiv:2609.03261v22026
  59. WorldGen: From Text to Traversable and Interactive 3D Worlds

    Dilin Wang, Hyunyoung Jung, Tom Monnier +22

    cs.CVcs.AIarXiv:2511.16825v12025
  60. Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

    Onat Şahin, Mohammad Altillawi, George Eskandar +2

    cs.CVarXiv:2609.03657v12026