Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
721 to 780 of 18,810
Towards Generalizable Robotic Manipulation in Dynamic Environments
Heng Fang, Shangru Li, Shuhan Wang +3
cs.CVcs.ROarXiv:2603.15620v32026Efficient Interactive Annotation of Segmentation Datasets with Polygon-RNN++
David Acuna, Huan Ling, Amlan Kar +1
cs.CVarXiv:1803.09693v12018MixStyle Neural Networks for Domain Generalization and Adaptation
Kaiyang Zhou, Yongxin Yang, Yu Qiao +1
cs.CVcs.AIcs.LGarXiv:2107.02053v22021Visual-ERM: Reward Modeling for Visual Equivalence
Ziyu Liu, Shengyuan Ding, Xinyu Fang +7
cs.CVcs.AIarXiv:2603.13224v22026Video-Based Reward Modeling for Computer-Use Agents
Linxin Song, Jieyu Zhang, Huanxin Sheng +6
cs.CVcs.CLarXiv:2603.10178v12026Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch
Hossein Souri, Liam Fowl, Rama Chellappa +2
cs.LGcs.CRcs.CVarXiv:2106.08970v32021A path following algorithm for the graph matching problem
Mikhail Zaslavskiy, Francis Bach, Jean-Philippe Vert
cs.CVcs.DMarXiv:0801.3654v22008Anticipative Video Transformer
Rohit Girdhar, Kristen Grauman
cs.CVcs.AIcs.LGarXiv:2106.02036v22021Intriguing Properties of Vision Transformers
Muzammal Naseer, Kanchana Ranasinghe, Salman Khan +3
cs.CVcs.AIcs.LGarXiv:2105.10497v32021MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
Baorong Shi, Bo Cui, Boyuan Jiang +17
cs.CLcs.AIcs.CVarXiv:2602.12705v42026Applications of Deep Learning Techniques for Automated Multiple Sclerosis Detection Using Magnetic Resonance Imaging: A Review
Afshin Shoeibi, Marjane Khodatars, Mahboobeh Jafari +9
eess.IVcs.CVarXiv:2105.04881v22021Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search
Tianming Liang, Qirui Du, Jian-Fang Hu +3
cs.CVarXiv:2602.04454v12026Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
Hanmo Chen, Chenghao Xu, Xu Yang +2
cs.CVarXiv:2601.21896v32026Stable and Scalable Bundle Adjustment of Holistic 3D Structures
Shaohui Liu, Rémi Pautrat, Daniel Barath +3
cs.CVarXiv:2609.04026v12026Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations
Denis M. Akola, David F. Fouhey
cs.CVarXiv:2609.04174v12026Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
Jiashi Feng, Xiu Li, Jing Lin +25
eess.IVcs.CVarXiv:2510.19944v12025Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
Yikang Ding, Jiwen Liu, Wenyuan Zhang +11
cs.CVarXiv:2509.09595v22025CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
Chengqi Duan, Kaiyue Sun, Rongyao Fang +10
cs.CVcs.AIarXiv:2510.11718v12025CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
Long Xing, Xiaoyi Dong, Yuhang Zang +6
cs.CVcs.AIcs.CLarXiv:2509.22647v12025Boosting of Image Denoising Algorithms
Yaniv Romano, Michael Elad
cs.CVmath.NAarXiv:1502.06220v22015Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
cs.CVcs.AIcs.CLarXiv:2508.11737v12025Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
Haohan Chi, Huan-ang Gao, Ziming Liu +12
cs.CVarXiv:2505.23757v12025MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Yi-Fan Zhang, Tao Yu, Haochen Tian +17
cs.CLcs.CVarXiv:2502.10391v12025HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
Chengxi Han, Chen Wu, Haonan Guo +2
cs.CVarXiv:2404.09178v12024Explainable artificial intelligence (XAI) in deep learning-based medical image analysis
Bas H. M. van der Velden, Hugo J. Kuijf, Kenneth G. A. Gilhuijs +1
eess.IVcs.CVarXiv:2107.10912v12021Common Limitations of Image Processing Metrics: A Picture Story
Annika Reinke, Minu D. Tizabi, Carole H. Sudre +90
eess.IVcs.CVarXiv:2104.05642v82021Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image
Zhengqin Li, Kalyan Sunkavalli, Manmohan Chandraker
cs.CVarXiv:1804.05790v12018Image biomarker standardisation initiative
Alex Zwanenburg, Stefan Leger, Martin Vallières +1
cs.CVeess.IVarXiv:1612.07003v112016RGB-D-based Action Recognition Datasets: A Survey
Jing Zhang, Wanqing Li, Philip O. Ogunbona +2
cs.CVarXiv:1601.05511v12016Moving Object Detection by Detecting Contiguous Outliers in the Low-Rank Representation
Xiaowei Zhou, Can Yang, Weichuan Yu
cs.CVarXiv:1109.0882v22011Faster and better: a machine learning approach to corner detection
Edward Rosten, Reid Porter, Tom Drummond
cs.CVcs.LGarXiv:0810.2434v12008EditThinker: Unlocking Iterative Reasoning for Any Image Editor
Hongyu Li, Manyuan Zhang, Dian Zheng +11
cs.CVarXiv:2512.05965v12025SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution
Federico Putamorsi, Leonardo Zini, Marcella Cornia +1
cs.CVarXiv:2609.03813v12026Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai +6
cs.CVarXiv:2510.15869v42025TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation
Adeela Islam, Zorah Lähner, Vittorio Murino +1
cs.CVarXiv:2609.04202v12026OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
Zhihong Chen, Xuehai Bai, Yang Shi +9
cs.CVcs.AIarXiv:2509.24900v22025DA$^{2}$: Depth Anything in Any Direction
Haodong Li, Wangguangdong Zheng, Jing He +5
cs.CVarXiv:2509.26618v52025More Grounded Image Captioning by Distilling Image-Text Matching Model
Yuanen Zhou, Meng Wang, Daqing Liu +2
cs.CVcs.CLarXiv:2004.00390v12020Sparse auto-regressive modeling for scene generation from multi-view images
Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4
cs.CVcs.LGarXiv:2609.03931v12026OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
Junhan Zhu, Hesong Wang, Mingluo Su +2
cs.CVarXiv:2510.06751v32025Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling
Spyros Gidaris, Nikos Komodakis
cs.CVcs.LGarXiv:1612.04770v12016When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection
Xuehao Wang, Jiaxin Hua, Runmei Li +4
cs.CVarXiv:2609.03378v12026PointGT: Simultaneous Geometry and Texture Editing for Point-Based Representations
Yanshu Zhang, George Shramko, Pratul P. Srinivasan +1
cs.CVcs.GRarXiv:2609.03341v12026Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++
Antonio Scardace, Francesco Guarnera, Sebastiano Battiato +1
cs.CVarXiv:2609.03615v12026LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
Shenghao Fu, Qize Yang, Yuan-Ming Li +3
cs.CVarXiv:2509.24786v12025Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
Huanyu Zhang, Wenshan Wu, Chengzu Li +9
cs.CVcs.CLarXiv:2510.24514v12025Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Shilong Zhang, He Zhang, Zhifei Zhang +11
cs.CVarXiv:2512.17909v12025DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions
Zhenjia Xu, Jiajun Wu, Andy Zeng +2
cs.ROcs.AIcs.CVarXiv:1906.03853v22019A Survey on Efficient Vision-Language-Action Models
Zhaoshu Yu, Bo Wang, Pengpeng Zeng +7
cs.CVcs.AIcs.LGarXiv:2510.24795v22025RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
Sicheng Feng, Kaiwen Tuo, Song Wang +3
cs.CVcs.AIarXiv:2510.02240v22025VABench: A Comprehensive Benchmark for Audio-Video Generation
Daili Hua, Xizhi Wang, Bohan Zeng +6
cs.CVcs.SDarXiv:2512.09299v22025An End-to-End Network for Panoptic Segmentation
Huanyu Liu, Chao Peng, Changqian Yu +4
cs.CVarXiv:1903.05027v22019Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
Runqian Wang, Yilun Du
cs.LGcs.AIcs.CVarXiv:2510.02300v32025Salient ImageNet: How to discover spurious features in Deep Learning?
Sahil Singla, Soheil Feizi
cs.LGcs.CVarXiv:2110.04301v42021SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
Chaojun Ni, Cheng Chen, Xiaofeng Wang +12
cs.CVcs.ROarXiv:2512.00903v12025Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
Pu Jian, Junhong Wu, Wei Sun +3
cs.CVcs.CLarXiv:2509.12132v12025UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
Yibin Wang, Zhimin Li, Yuhang Zang +8
cs.CVarXiv:2510.18701v22025MedQA-MM: Shortcuts Behind Medical Visual Reasoning
Benlu Wang, Yifan Zhang, Jiaqing Yu +7
cs.CVcs.CLarXiv:2609.03261v22026WorldGen: From Text to Traversable and Interactive 3D Worlds
Dilin Wang, Hyunyoung Jung, Tom Monnier +22
cs.CVcs.AIarXiv:2511.16825v12025Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations
Onat Şahin, Mohammad Altillawi, George Eskandar +2
cs.CVarXiv:2609.03657v12026