Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,901 to 18,957 of 18,957

  1. OPD-V: Visual On-Policy Self-Distillation with Modality Balance

    Aniri, Jinhe Bi, Peng Liao +5

    cs.CVcs.AIarXiv:2608.05131v22026
  2. MiniWorld: Democratizing the Training of Video World Models from Scratch

    Yian Zhao, Ruochong Zheng, Hongcan Guo +3

    cs.CVarXiv:2608.01127v22026
  3. Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

    Gunjan Paul, Senthil Palanisamy, Satpal Singh Rathore +3

    cs.CVcs.ARcs.ROarXiv:2608.08285v22026
  4. An AI4AI Framework for Visual Token Pruning

    Zhen Liu, Wenli Huang, Wei Song +3

    cs.LGcs.CVarXiv:2608.07193v12026
  5. Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

    Weili Zeng, Yitong Xing, Fulong Liu +10

    cs.ROcs.CVarXiv:2607.26657v32026
  6. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    Yuxue Yang, Shuyao Shang, Jiahe Wang +13

    cs.CVarXiv:2608.02603v12026
  7. ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

    Jiahao Zhao, Xiaomin Yu, Zhongxiang Sun +5

    cs.CVarXiv:2608.04436v12026
  8. ChronoVision: Temporal Reasoning via Latent State Reconstruction

    Yifan Shen, Jian Xu, Boyi Li +6

    cs.CVarXiv:2608.05631v12026
  9. CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

    Bingxin Yu, Xueli Wang, Jerry Zhou +6

    cs.CVcs.AIarXiv:2608.13939v12026
  10. WorldClaw: Agentic 3D Open-World Generation at Scale

    Chunchao Guo, Jinpeng Li, Yang Li +1

    cs.AIcs.CVarXiv:2608.05248v12026
  11. GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    Qifeng Zhang, Kaixiang Huang, Heng Dong +6

    cs.CVarXiv:2608.05747v12026
  12. ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

    Xinye Li, Lingshuai Lin, Lei Wang +8

    cs.CVcs.AIarXiv:2608.14022v12026
  13. Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation

    Michael R. Martin, Joseph Insley, Victor A. Mateevitsi +2

    cs.CVcs.AIcs.CEarXiv:2608.14112v12026
  14. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

    Wanshun Su, Yang Shi, Feihu Liu +10

    cs.CVarXiv:2608.03812v12026
  15. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

    DreamX Team, Rui Chen, Xiangxiang Chu +7

    cs.CVcs.ROarXiv:2608.13489v12026
    Summaries:한국어
  16. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

    Junliang Ye, Kenkun Liu, Guocun Wang +13

    cs.CVarXiv:2608.02711v32026
    Summaries:한국어
  17. Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

    Haoqi Yuan, Zhixuan Liang, Anzhe Chen +20

    cs.ROcs.CVcs.LGarXiv:2606.17846v22026
    Summaries:한국어
  18. Densely Connected Convolutional Networks

    Gao Huang, Zhuang Liu, Laurens van der Maaten +1

    cs.CVcs.LGarXiv:1608.06993v52016
    Summaries:한국어
  19. U-Net: Convolutional Networks for Biomedical Image Segmentation

    Olaf Ronneberger, Philipp Fischer, Thomas Brox

    cs.CVarXiv:1505.04597v12015
    Summaries:한국어
  20. A Survey of Large Models in Sports

    Yichen Xu, Jianzhe Ma, Chuhan Wang +4

    cs.CLcs.CVarXiv:2608.14377v12026
  21. Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

    Nikolai Röhrich, Isabell Hans, Felix Krause +1

    cs.CVcs.AIcs.LGarXiv:2608.14172v12026
  22. Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

    Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk +2

    cs.CLcs.AIcs.CVarXiv:2608.13760v12026
  23. AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

    Yuqing Wen, Yukai Huang, Qianqian Xie +6

    cs.MMcs.CVcs.SDarXiv:2607.24821v12026
  24. SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

    Jinsheng Quan, Jianhua Li, Siyi Xie +7

    cs.CVcs.AIarXiv:2608.14138v12026
  25. A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

    Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade +9

    cs.AIcs.CVcs.DLarXiv:2608.14075v12026
  26. RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

    Kabila Haile Soboka

    cs.CVarXiv:2608.06914v22026
  27. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    Haiyang Zhou, Wangbo Yu, Chaoran Feng +3

    cs.CVarXiv:2608.04701v12026
  28. OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    Hongyu Liu, Chun Wang, Feng Gao +6

    cs.CVarXiv:2607.08766v12026
  29. HelloWorld: Enabling Socially Interactive Characters in Video World Models

    Liangyang Ouyang, Ruicong Liu, Xuangeng Chu +2

    cs.CVarXiv:2608.05070v12026
  30. Scalable Visual Pretraining for Language Intelligence

    Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14

    cs.CVcs.AIcs.MMarXiv:2607.09657v22026
  31. LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

    Hang Long, Tianhao Zhao, Junkai Lin +8

    cs.GRcs.CVarXiv:2607.10623v22026
  32. Video Generation Models are General-Purpose Vision Learners

    Letian Wang, Chuhan Zhang, Rishabh Kabra +9

    cs.CVcs.AIarXiv:2607.09024v12026
  33. InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

    Jiawei Wang, Hao Yu, Yongzhen Hu +7

    cs.CVarXiv:2608.02437v22026
  34. Self-Supervised Visual On-Policy Distillation

    Yijiang Li, Yijun Liang, Yunjie Tian +6

    cs.CVcs.AIarXiv:2608.14144v12026
    Summaries:한국어
  35. Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

    Kaiwen Zheng, Guande He, Min Zhao +7

    cs.CVcs.LGarXiv:2606.25473v12026
  36. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik +3

    cs.CVcs.GRarXiv:2003.08934v22020
  37. Unlimited OCR Works

    Youyang Yin, Huanhuan Liu, YY +14

    cs.CVcs.CLarXiv:2606.23050v12026
  38. Vidu S1: A Real-Time Interactive Video Generation Model

    Jintao Zhang, Kai Jiang, Jintao Chen +24

    cs.CVcs.LGarXiv:2607.03118v22026
  39. KVAE: Family of Tokenizers for Multimodal Generative Models

    Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11

    cs.CVcs.LGcs.SDarXiv:2608.05798v12026
  40. DanceOPD: On-Policy Generative Field Distillation

    Wei Zhou, Xiongwei Zhu, Zelin Xu +8

    cs.CVcs.CLcs.LGarXiv:2606.27377v22026
  41. JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

    Yicheng Xiao, Wenxun Dai, Xinran Qin +22

    cs.CVarXiv:2608.03974v12026
  42. SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

    Zongchuang Zhao, Xin Zhou, Tianyang Xu +5

    cs.CVarXiv:2608.07468v22026
  43. CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

    Qinye Zhou, Jun Zheng, Yongchao Du +17

    cs.CVarXiv:2608.14546v12026
  44. Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

    Hmrishav Bandyopadhyay, Xuanchi Ren, Zijian Huang +7

    cs.CVarXiv:2608.13391v12026
  45. Infinite Worlds with Versatile Interactions

    Zelin Gao, Qiuyu Wang, Jiapeng Zhu +17

    cs.CVarXiv:2607.07534v12026
  46. PixSDS: Why Latent SDS Makes Noisy Pixels

    Vsevolod Skorokhodov

    cs.CVarXiv:2608.12997v12026
  47. Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

    Ryosei Hara, Masashi Hatano, Rintaro Yanagi +3

    cs.CVarXiv:2608.11574v12026
  48. StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

    Yuyang Yin, Zixiang Li, Longxuan Deng +11

    cs.CVarXiv:2608.12314v12026
  49. Intern-S2-Preview: Scientific Agentic Foundation Model

    Lei Bai, Jiaqi Cao, Chiyu Chen +122

    cs.LGcs.CLcs.CVarXiv:2608.13505v12026
  50. PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

    Kaixin Ding, Xi Chen, Minghong Cai +9

    cs.CVarXiv:2608.13552v12026
  51. From SRA to Self-Flow: Data Augmentation or Self-Supervision?

    Dengyang Jiang, Mengmeng Wang, Harry Yang +1

    cs.CVarXiv:2607.02508v12026
  52. Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

    Ling Xu, Chuyu Han, Borui Li +6

    cs.ROcs.CVcs.OSarXiv:2607.02501v22026
  53. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

    Ronghan Chen, Yandan Yang, Zuojin Tang +18

    cs.CVcs.ROarXiv:2607.00678v22026
  54. GEAR: Guided End-to-End AutoRegression for Image Synthesis

    Bin Lin, Zheyuan Liu, Chenguo Lin +8

    cs.CVarXiv:2606.32039v12026
  55. AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer

    Xinyu Hou, Xiaoming Li, Zongsheng Yue +1

    cs.CVarXiv:2606.31959v12026
  56. InstanceControl: Controllable Complex Image Generation without Instance Labeling

    Xiaoyu Liu, Huan Wang, Fan Li +4

    cs.CVarXiv:2606.31924v12026
  57. You Only Look Once: Unified, Real-Time Object Detection

    Joseph Redmon, Santosh Divvala, Ross Girshick +1

    cs.CVarXiv:1506.02640v52015