Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
18,721 to 18,780 of 18,811
A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
Zixuan Fu, Chong Wang, Lanqing Guo +3
cs.CVarXiv:2607.29122v12026Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle
Jiaming Zhang, Boyang Chen, Zherui Li +14
cs.CRcs.CVarXiv:2608.04314v12026Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis
Lingwei Dang, Ziyan Qiu, Jiajia Cheng +9
cs.ROcs.CVarXiv:2608.01973v12026AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
Guiyu Zhao, Longteng Guo, Yanghong Mei +7
cs.ROcs.CVarXiv:2608.06729v12026Decoding Children's Gait Behavior
Yifan Shen, Boyi Li, Meihuan Huang +12
cs.CVarXiv:2608.00371v12026SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers
Kiran Nair, Rodrigue Rizk, KC Santosh
cs.LGcs.AIcs.CVarXiv:2608.13702v12026MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Wenjie Zhu, Yabin Zhang, Wenjun Zeng +1
cs.CVarXiv:2607.27637v22026What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Zhijing Zhang, Jinpeng Yu, Xin Song +6
cs.CVcs.AIarXiv:2608.07565v12026The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents
Weiwei Li, Junzhuo Liu, Tong Chu +2
cs.CVarXiv:2608.06065v12026TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection
Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo +3
eess.IVcs.CVcs.LGarXiv:2608.13711v120263DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
Changwoo Baek, Kyeongbo Kong
cs.CVcs.LGarXiv:2608.01185v12026DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12
cs.CVcs.AIarXiv:2608.01827v12026MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
Rafi Ibn Sultan, Hui Zhu, Chengyin Li +1
cs.CVcs.AIarXiv:2608.13690v12026Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
Siming Fu, Zheming Fu, Ruizhe He +7
cs.LGcs.CVarXiv:2608.03316v12026Uncertainty-Aware World Model for Aerial Image-Goal Navigation
Deyi Zhu, Haoyu Fan, Yinan Zhu +4
cs.CVarXiv:2608.05597v12026Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
Jiazhen Liu, Mingkuan Feng, Long Chen
cs.CVarXiv:2608.02791v12026Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Zhen Fang, Yu Zeng, Wenxuan Huang +17
cs.CVcs.AIarXiv:2608.03979v12026WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Senyu Fei, Xiaopeng Yu, Siyin Wang +3
cs.ROcs.CLcs.CVarXiv:2607.29613v12026Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure
Gauthier Avité, Maxime Sanchez-Renauld, Nicolas Bourriez +1
cs.CVcs.LGarXiv:2608.14293v12026Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara +2
cs.CVcs.AIcs.CLarXiv:2608.14286v12026ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
cs.CVarXiv:2608.04956v12026What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation
Amal Saqib, Tausifa Jan Saleem, Numan Saeed +1
cs.CVcs.LGarXiv:2608.13660v12026UniWorld-Design: From Pixel Generation to Layer-Native Design
Zongjian Li, Zhiyuan Yan, Chenxu Bai +9
cs.CVarXiv:2608.03971v12026Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
Siming Fu, Haojun Xu, Ruizhe He +9
cs.CVarXiv:2608.04349v12026Invisible Shortcuts: Why Vision Encoders Know Your Camera
Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos +2
cs.CVcs.LGarXiv:2608.05424v12026VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao +9
cs.CVcs.CLarXiv:2607.28590v12026Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
Zhixue Fang, Zhimin Zhang, Bi'an Du +6
cs.CVarXiv:2608.01628v22026Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans
Tarandeep Singh Mandhiratta, ANK Zaman, Abdul-Rahman Mawlood-Yunis
cs.CVcs.AIcs.HCarXiv:2608.14317v12026CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang
cs.CVarXiv:2608.02589v12026iFAN: Inference-Aware Learning for Plain Mask Transformers
Fang Li, Yu He, Haoyang Tong +7
cs.CVarXiv:2608.03216v22026EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
Feier Wu, Wanke Xia, Xu He +8
cs.CVarXiv:2608.05565v12026CADENA: Stepwise CAD Reverse Engineering
Soslan Kabisov, Gennadiy Savrasov, Maksim Elistratov +9
cs.CVarXiv:2608.00799v12026Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use
Yi Ding, Yanzhao Yu, Xili Dai +5
cs.ROcs.AIcs.CVarXiv:2608.14047v12026GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection
Yingjie Ma, Zitong Yu, Wei Jia +2
cs.CVcs.AIarXiv:2608.14389v12026OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao +5
cs.CVcs.AIarXiv:2608.05131v22026MiniWorld: Democratizing the Training of Video World Models from Scratch
Yian Zhao, Ruochong Zheng, Hongcan Guo +3
cs.CVarXiv:2608.01127v22026Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
Gunjan Paul, Senthil Palanisamy, Satpal Singh Rathore +3
cs.CVcs.ARcs.ROarXiv:2608.08285v22026An AI4AI Framework for Visual Token Pruning
Zhen Liu, Wenli Huang, Wei Song +3
cs.LGcs.CVarXiv:2608.07193v12026Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Weili Zeng, Yitong Xing, Fulong Liu +10
cs.ROcs.CVarXiv:2607.26657v32026WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Yuxue Yang, Shuyao Shang, Jiahe Wang +13
cs.CVarXiv:2608.02603v12026ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5
cs.CVarXiv:2608.04436v12026ChronoVision: Temporal Reasoning via Latent State Reconstruction
Yifan Shen, Jian Xu, Boyi Li +6
cs.CVarXiv:2608.05631v12026CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification
Bingxin Yu, Xueli Wang, Jerry Zhou +6
cs.CVcs.AIarXiv:2608.13939v12026WorldClaw: Agentic 3D Open-World Generation at Scale
Chunchao Guo, Jinpeng Li, Yang Li +1
cs.AIcs.CVarXiv:2608.05248v12026GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Qifeng Zhang, Kaixiang Huang, Heng Dong +6
cs.CVarXiv:2608.05747v12026ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Xinye Li, Lingshuai Lin, Lei Wang +8
cs.CVcs.AIarXiv:2608.14022v12026Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation
Michael R. Martin, Joseph Insley, Victor A. Mateevitsi +2
cs.CVcs.AIcs.CEarXiv:2608.14112v12026OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
Wanshun Su, Yang Shi, Feihu Liu +10
cs.CVarXiv:2608.03812v12026DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
DreamX Team, Rui Chen, Xiangxiang Chu +7
cs.CVcs.ROarXiv:2608.13489v12026Summaries:한국어Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Junliang Ye, Kenkun Liu, Guocun Wang +13
cs.CVarXiv:2608.02711v32026Summaries:한국어Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Haoqi Yuan, Zhixuan Liang, Anzhe Chen +20
cs.ROcs.CVcs.LGarXiv:2606.17846v22026Summaries:한국어Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten +1
cs.CVcs.LGarXiv:1608.06993v52016Summaries:한국어U-Net: Convolutional Networks for Biomedical Image Segmentation
Olaf Ronneberger, Philipp Fischer, Thomas Brox
cs.CVarXiv:1505.04597v12015Summaries:한국어A Survey of Large Models in Sports
Yichen Xu, Jianzhe Ma, Chuhan Wang +4
cs.CLcs.CVarXiv:2608.14377v12026Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation
Nikolai Röhrich, Isabell Hans, Felix Krause +1
cs.CVcs.AIcs.LGarXiv:2608.14172v12026Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk +2
cs.CLcs.AIcs.CVarXiv:2608.13760v12026AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Yuqing Wen, Yukai Huang, Qianqian Xie +6
cs.MMcs.CVcs.SDarXiv:2607.24821v12026SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Jinsheng Quan, Jianhua Li, Siyi Xie +7
cs.CVcs.AIarXiv:2608.14138v12026A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade +9
cs.AIcs.CVcs.DLarXiv:2608.14075v12026RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
Kabila Haile Soboka
cs.CVarXiv:2608.06914v22026