Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

18,361 to 18,420 of 18,808

  1. OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer

    Boyang Wang, Guangyi Xu, Jiahui Zhang +2

    cs.CVarXiv:2604.24762v22026
  2. S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

    Yalun Dai, Hao Li, Shulin Tian +8

    cs.CVarXiv:2606.20515v32026
  3. Flash-WAM: Modality-Aware Distillation for World Action Models

    Arman Akbari, Ci Zhang, Arash Akbari +6

    cs.LGcs.CVcs.ROarXiv:2606.05254v12026
  4. Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

    Chenming Zhu, Jingli Lin, Yilin Long +4

    cs.CVarXiv:2606.06476v12026
  5. UniSHARP: Universal Sharp Monocular View Synthesis

    Meixi Song, Dizhe Zhang, Hao Ren +4

    cs.CVarXiv:2606.07514v12026
  6. Watch, Remember, Reason: Human-View Video Understanding with MLLMs

    Jiahao Meng, Yue Tan, Qi Xu +12

    cs.CVcs.AIcs.MMarXiv:2606.07433v12026
  7. Latent Spatial Memory for Video World Models

    Weijie Wang, Haoyu Zhao, Yifan Yang +7

    cs.CVarXiv:2606.09828v12026
  8. WorldOlympiad: Can Your World Model Survive a Triathlon?

    Yuke Zhao, Wangbo Zhao, Weijie Wang +8

    cs.CVarXiv:2606.11129v22026
  9. UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer

    Shuai Wang, Liang Li, Yang Chen +3

    cs.CVarXiv:2606.16255v12026
  10. VisualClaw: A Real-Time, Personalized Agent for the Physical World

    Haoqin Tu, Jianwen Chen, Zijun Wang +14

    cs.CVcs.CLarXiv:2606.16295v12026
  11. Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

    Wujian Peng, Lingchen Meng, Yuxuan Cai +7

    cs.CVarXiv:2606.18249v22026
  12. EgoCS-400K: An Egocentric Gameplay Dataset for World Models

    Rongjin Guo, Dong Liang, Yuhao Liu +4

    cs.CVarXiv:2606.18180v12026
  13. MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

    Lichen Bai, Tianhao Zhang, Shitong Shao +14

    cs.CVarXiv:2606.17800v12026
  14. Meta-CoT: Enhancing Granularity and Generalization in Image Editing

    Shiyi Zhang, Yiji Cheng, Tiankai Hang +8

    cs.CVcs.AIcs.LGarXiv:2604.24625v12026
  15. Diffusion Model as a Generalist Segmentation Learner

    Haoxiao Wang, Antao Xiang, Haiyang Sun +8

    cs.CVarXiv:2604.24575v12026
  16. A Systematic Post-Train Framework for Video Generation

    Zeyue Xue, Siming Fu, Jie Huang +9

    cs.CVarXiv:2604.25427v12026
  17. LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention

    Shitong Shao, Zikai Zhou, Haopeng Li +4

    cs.CVarXiv:2605.04569v22026
  18. Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine

    Marta Sumyk, Oleksandr Kosovan, Iryna Voitsitska

    cs.CVcs.AIarXiv:2608.16380v12026
  19. FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

    Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim +4

    cs.DCcs.AIcs.CVarXiv:2608.15410v12026
  20. Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning

    Yuxing Long, Lei Kang, Ziyan Yu +8

    cs.ROcs.AIcs.CLarXiv:2608.15863v12026
  21. LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting

    Louen Pottier

    cs.CVcs.LGarXiv:2608.16324v12026
  22. GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

    Ming Qian, Zijian Wang, Minchao Sun +5

    cs.CVarXiv:2608.17988v12026
  23. Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology

    Hadi Hasan, Safaa Salman, Lama Sleem +2

    cs.CVcs.AIarXiv:2608.15915v12026
  24. Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

    Bingxin Xu, Yuzhang Shang, Emilio Ferrara

    cs.ROcs.AIcs.CVarXiv:2608.16889v12026
  25. Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

    Siyuan Huang, Xiaoye Qu, Yafu Li +6

    cs.CVcs.AIarXiv:2605.00814v22026
  26. WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

    Junzhe Huang, Xiaoxiao Sun, Yan Yang +6

    cs.CVarXiv:2605.01018v22026
  27. Asymmetric Flow Models

    Hansheng Chen, Jan Ackermann, Minseo Kim +2

    cs.CVarXiv:2605.12964v22026
  28. OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

    Yiren Song, Xiyao Deng, Pei Yang +2

    cs.CVarXiv:2605.12038v12026
  29. Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

    Yuanzhi Xu, Qian Gao, Jun Fan +4

    cs.CVcs.AIarXiv:2608.16805v12026
  30. JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

    Lin Song, Wenbo Li, Guoqing Ma +16

    cs.GRcs.AIcs.CLarXiv:2605.04128v22026
  31. Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

    Ye Lu, Shen Wang, Zhaoyang Zhang +4

    cs.CVcs.AIcs.CRarXiv:2608.16791v12026
  32. HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes

    Yujia Li, Yiqun Zhang, Zihan Cheng +7

    cs.CVcs.AIarXiv:2608.16622v12026
  33. AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

    Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi +9

    cs.CVcs.LGarXiv:2608.14778v12026
  34. Native Audio-Visual Alignment for Generation

    Longbin Ji, Guan Wang, Xuan Wei +6

    cs.CVarXiv:2605.30073v12026
  35. Delta Attention Residuals

    Cheng Luo, Zefan Cai, Junjie Hu

    cs.LGcs.CVarXiv:2605.18855v12026
    Summaries:한국어
  36. Geometry-Aware Image Flow Matching

    Junho Lee, Kwanseok Kim, Joonseok Lee

    cs.CVarXiv:2605.25294v12026
  37. Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos

    Mohamed Afham, Christoph Reich, Oliver Hahn +2

    cs.CVcs.AIarXiv:2608.16457v12026
  38. Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System

    Alam Noor, Luis Almeida, Kai Li +3

    cs.CVcs.AIarXiv:2608.16142v12026
  39. P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

    Yikang Yang, Zhanpeng Hu, Youtian Lin +5

    cs.CVarXiv:2606.11152v22026
  40. Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

    Che Liu, Lichao Ma, Xiangyu Tony Zhang +4

    cs.MMcs.AIcs.CVarXiv:2605.12034v22026
  41. CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

    Dongsheng Ma, Jiayu Li, Zhengren Wang +8

    cs.CLcs.CVarXiv:2605.12882v12026
  42. FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

    Quanjian Song, Yefeng Shen, Mengting Chen +5

    cs.CVarXiv:2605.15824v22026
  43. WavFlow: Audio Generation in Waveform Space

    Feiyan Zhou, Luyuan Wang, Shoufa Chen +6

    cs.SDcs.CVarXiv:2605.18749v12026
  44. From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

    Xingjian Wang, Zhao Wang, Taihang Hu +14

    cs.CVcs.AIarXiv:2608.18076v12026
  45. Convolution-Free Holistic Multivariance Decomposition Layer for Efficient Hyperspectral Image Classification Tensor Networks

    Süha Tuna, Ülker Başar

    cs.CVcs.LGmath.OCarXiv:2608.16241v12026
  46. ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

    Shangpin Peng, Gengluo Li, Xingyu Wan +10

    cs.CVarXiv:2606.01348v32026
  47. ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

    Mary Lynn Martin, Yifei Zhang, Martha Palmer +1

    cs.CLcs.CVarXiv:2608.16011v12026
  48. Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth

    Zhaocen Liu, Satvik Praveen, Yi Sheng

    cs.LGcs.CVarXiv:2608.16010v12026
  49. Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

    Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani +2

    cs.CVcs.AIcs.CLarXiv:2608.16514v12026
  50. Unsupervised Learning of Cell Instances with Generative Routing Pyramids

    Ziwen Liu, Martin Weigert

    cs.CVcs.LGq-bio.QMarXiv:2608.16810v12026
  51. Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation

    Qinghui Liu, Jon André Ottesen, Atle Bjørnerud +1

    cs.CVcs.AIarXiv:2608.16377v12026
  52. MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter

    Vignesh Nagarajan, Sriram Venkatapathy

    cs.CVcs.AIcs.LGarXiv:2608.16709v12026
  53. A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

    Siyuan Ma, Liang He, Mengying Zhu +14

    eess.IVcs.AIcs.CVarXiv:2608.16233v12026
  54. Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

    Fabian Gröger, Marco Weishaupt, Philippe Gottfrois +6

    cs.CVcs.AIarXiv:2608.16198v12026
  55. DriveCache: Action-Aware Caching for Driving World Model Inference

    Jianchun Yang, Jian Liang, Xianda Guo +5

    cs.AIcs.CVarXiv:2608.16354v12026
  56. X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

    Zichao Zeng, Weijia Fan, Yufan Chen +7

    cs.CVcs.AIcs.ROarXiv:2608.16658v12026
  57. Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

    Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar +4

    cs.CVcs.AIcs.CLarXiv:2608.15869v12026
  58. Turning spectra into images improves plant trait retrieval with 2D-CNNs

    Javier Lopatin, Teja Kattenborn, Eya Cherif +1

    cs.CVcs.LGarXiv:2608.16661v12026
  59. OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

    Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola +3

    cs.LGcs.AIcs.CVarXiv:2608.16373v22026
  60. FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion

    Yu Lu, Junjie Yang, Piotr Koniusz +2

    cs.CVarXiv:2606.10671v22026