Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,721 to 18,780 of 18,811

  1. A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

    Zixuan Fu, Chong Wang, Lanqing Guo +3

    cs.CVarXiv:2607.29122v12026
  2. Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

    Jiaming Zhang, Boyang Chen, Zherui Li +14

    cs.CRcs.CVarXiv:2608.04314v12026
  3. Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

    Lingwei Dang, Ziyan Qiu, Jiajia Cheng +9

    cs.ROcs.CVarXiv:2608.01973v12026
  4. AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

    Guiyu Zhao, Longteng Guo, Yanghong Mei +7

    cs.ROcs.CVarXiv:2608.06729v12026
  5. Decoding Children's Gait Behavior

    Yifan Shen, Boyi Li, Meihuan Huang +12

    cs.CVarXiv:2608.00371v12026
  6. SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

    Kiran Nair, Rodrigue Rizk, KC Santosh

    cs.LGcs.AIcs.CVarXiv:2608.13702v12026
  7. MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

    Wenjie Zhu, Yabin Zhang, Wenjun Zeng +1

    cs.CVarXiv:2607.27637v22026
  8. What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

    Zhijing Zhang, Jinpeng Yu, Xin Song +6

    cs.CVcs.AIarXiv:2608.07565v12026
  9. The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

    Weiwei Li, Junzhuo Liu, Tong Chu +2

    cs.CVarXiv:2608.06065v12026
  10. TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

    Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo +3

    eess.IVcs.CVcs.LGarXiv:2608.13711v12026
  11. 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

    Changwoo Baek, Kyeongbo Kong

    cs.CVcs.LGarXiv:2608.01185v12026
  12. DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

    Huanyao Zhang, Jiepeng Zhou, Runhao Zhao +12

    cs.CVcs.AIarXiv:2608.01827v12026
  13. MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

    Rafi Ibn Sultan, Hui Zhu, Chengyin Li +1

    cs.CVcs.AIarXiv:2608.13690v12026
  14. Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

    Siming Fu, Zheming Fu, Ruizhe He +7

    cs.LGcs.CVarXiv:2608.03316v12026
  15. Uncertainty-Aware World Model for Aerial Image-Goal Navigation

    Deyi Zhu, Haoyu Fan, Yinan Zhu +4

    cs.CVarXiv:2608.05597v12026
  16. Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

    Jiazhen Liu, Mingkuan Feng, Long Chen

    cs.CVarXiv:2608.02791v12026
  17. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

    Zhen Fang, Yu Zeng, Wenxuan Huang +17

    cs.CVcs.AIarXiv:2608.03979v12026
  18. WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

    Senyu Fei, Xiaopeng Yu, Siyin Wang +3

    cs.ROcs.CLcs.CVarXiv:2607.29613v12026
  19. Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure

    Gauthier Avité, Maxime Sanchez-Renauld, Nicolas Bourriez +1

    cs.CVcs.LGarXiv:2608.14293v12026
  20. Seeing Red, Thinking Bad: Color Bias in Vision Language Models

    Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara +2

    cs.CVcs.AIcs.CLarXiv:2608.14286v12026
  21. ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

    Xu Guo, Zhengxuan Wei, Xinghui Li +11

    cs.CVarXiv:2608.04956v12026
  22. What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation

    Amal Saqib, Tausifa Jan Saleem, Numan Saeed +1

    cs.CVcs.LGarXiv:2608.13660v12026
  23. UniWorld-Design: From Pixel Generation to Layer-Native Design

    Zongjian Li, Zhiyuan Yan, Chenxu Bai +9

    cs.CVarXiv:2608.03971v12026
  24. Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models

    Siming Fu, Haojun Xu, Ruizhe He +9

    cs.CVarXiv:2608.04349v12026
  25. Invisible Shortcuts: Why Vision Encoders Know Your Camera

    Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos +2

    cs.CVcs.LGarXiv:2608.05424v12026
  26. VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

    Kangning Zhang, Yixing Li, Shuai Shao +9

    cs.CVcs.CLarXiv:2607.28590v12026
  27. Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

    Zhixue Fang, Zhimin Zhang, Bi'an Du +6

    cs.CVarXiv:2608.01628v22026
  28. Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans

    Tarandeep Singh Mandhiratta, ANK Zaman, Abdul-Rahman Mawlood-Yunis

    cs.CVcs.AIcs.HCarXiv:2608.14317v12026
  29. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang

    cs.CVarXiv:2608.02589v12026
  30. iFAN: Inference-Aware Learning for Plain Mask Transformers

    Fang Li, Yu He, Haoyang Tong +7

    cs.CVarXiv:2608.03216v22026
  31. EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

    Feier Wu, Wanke Xia, Xu He +8

    cs.CVarXiv:2608.05565v12026
  32. CADENA: Stepwise CAD Reverse Engineering

    Soslan Kabisov, Gennadiy Savrasov, Maksim Elistratov +9

    cs.CVarXiv:2608.00799v12026
  33. Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

    Yi Ding, Yanzhao Yu, Xili Dai +5

    cs.ROcs.AIcs.CVarXiv:2608.14047v12026
  34. GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection

    Yingjie Ma, Zitong Yu, Wei Jia +2

    cs.CVcs.AIarXiv:2608.14389v12026
  35. OPD-V: Visual On-Policy Self-Distillation with Modality Balance

    Aniri, Jinhe Bi, Peng Liao +5

    cs.CVcs.AIarXiv:2608.05131v22026
  36. MiniWorld: Democratizing the Training of Video World Models from Scratch

    Yian Zhao, Ruochong Zheng, Hongcan Guo +3

    cs.CVarXiv:2608.01127v22026
  37. Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

    Gunjan Paul, Senthil Palanisamy, Satpal Singh Rathore +3

    cs.CVcs.ARcs.ROarXiv:2608.08285v22026
  38. An AI4AI Framework for Visual Token Pruning

    Zhen Liu, Wenli Huang, Wei Song +3

    cs.LGcs.CVarXiv:2608.07193v12026
  39. Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

    Weili Zeng, Yitong Xing, Fulong Liu +10

    cs.ROcs.CVarXiv:2607.26657v32026
  40. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    Yuxue Yang, Shuyao Shang, Jiahe Wang +13

    cs.CVarXiv:2608.02603v12026
  41. ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

    Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5

    cs.CVarXiv:2608.04436v12026
  42. ChronoVision: Temporal Reasoning via Latent State Reconstruction

    Yifan Shen, Jian Xu, Boyi Li +6

    cs.CVarXiv:2608.05631v12026
  43. CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

    Bingxin Yu, Xueli Wang, Jerry Zhou +6

    cs.CVcs.AIarXiv:2608.13939v12026
  44. WorldClaw: Agentic 3D Open-World Generation at Scale

    Chunchao Guo, Jinpeng Li, Yang Li +1

    cs.AIcs.CVarXiv:2608.05248v12026
  45. GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    Qifeng Zhang, Kaixiang Huang, Heng Dong +6

    cs.CVarXiv:2608.05747v12026
  46. ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

    Xinye Li, Lingshuai Lin, Lei Wang +8

    cs.CVcs.AIarXiv:2608.14022v12026
  47. Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation

    Michael R. Martin, Joseph Insley, Victor A. Mateevitsi +2

    cs.CVcs.AIcs.CEarXiv:2608.14112v12026
  48. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    Wanshun Su, Yang Shi, Feihu Liu +10

    cs.CVarXiv:2608.03812v12026
  49. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

    DreamX Team, Rui Chen, Xiangxiang Chu +7

    cs.CVcs.ROarXiv:2608.13489v12026
    Summaries:한국어
  50. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

    Junliang Ye, Kenkun Liu, Guocun Wang +13

    cs.CVarXiv:2608.02711v32026
    Summaries:한국어
  51. Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

    Haoqi Yuan, Zhixuan Liang, Anzhe Chen +20

    cs.ROcs.CVcs.LGarXiv:2606.17846v22026
    Summaries:한국어
  52. Densely Connected Convolutional Networks

    Gao Huang, Zhuang Liu, Laurens van der Maaten +1

    cs.CVcs.LGarXiv:1608.06993v52016
    Summaries:한국어
  53. U-Net: Convolutional Networks for Biomedical Image Segmentation

    Olaf Ronneberger, Philipp Fischer, Thomas Brox

    cs.CVarXiv:1505.04597v12015
    Summaries:한국어
  54. A Survey of Large Models in Sports

    Yichen Xu, Jianzhe Ma, Chuhan Wang +4

    cs.CLcs.CVarXiv:2608.14377v12026
  55. Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

    Nikolai Röhrich, Isabell Hans, Felix Krause +1

    cs.CVcs.AIcs.LGarXiv:2608.14172v12026
  56. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

    Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk +2

    cs.CLcs.AIcs.CVarXiv:2608.13760v12026
  57. AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

    Yuqing Wen, Yukai Huang, Qianqian Xie +6

    cs.MMcs.CVcs.SDarXiv:2607.24821v12026
  58. SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

    Jinsheng Quan, Jianhua Li, Siyi Xie +7

    cs.CVcs.AIarXiv:2608.14138v12026
  59. A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

    Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade +9

    cs.AIcs.CVcs.DLarXiv:2608.14075v12026
  60. RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

    Kabila Haile Soboka

    cs.CVarXiv:2608.06914v22026