Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

6,181 to 6,240 of 18,867

  1. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

    Peiwen Zhang, Yufan Deng, Shangkun Sun +11

    cs.CVcs.AIcs.ROarXiv:2606.28128v12026
    Summaries:한국어
  2. FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

    Haorui Ji, Weizhe Liu, Hongdong Li +1

    cs.CVcs.AIarXiv:2606.24874v12026
  3. $μ_0$: A Scalable 3D Interaction-Trace World Model

    Seungjae Lee, Yoonkyo Jung, Jusuk Lee +6

    cs.ROcs.CVcs.LGarXiv:2606.13769v22026
  4. Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

    Shangwen Zhu, Qianyu Peng, Zhao Pu +12

    cs.CVarXiv:2605.18601v22026
  5. Map2World: Segment Map Conditioned Text to 3D World Generation

    Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang +2

    cs.CVarXiv:2605.00781v12026
  6. World2Minecraft: Occupancy-Driven Simulated Scenes Construction

    Lechao Zhang, Haoran Xu, Jingyu Gong +3

    cs.CVarXiv:2604.27578v12026
  7. 3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis

    Stefan Schulz, Fernando Edelstein, Hannah Dröge +2

    cs.CVcs.LGcs.MMarXiv:2604.11211v12026
  8. POS-ISP: Pipeline Optimization at the Sequence Level for Task-aware ISP

    Jiyun Won, Heemin Yang, Woohyeok Kim +2

    cs.CVarXiv:2604.06938v12026
  9. Video-Oasis: Rethinking Evaluation of Video Understanding

    Geuntaek Lim, Sungjune Park, Jaeyun Lee +5

    cs.CVarXiv:2603.29616v22026
  10. Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation

    Donald Shenaj, Federico Errica, Antonio Carta

    cs.CVcs.AIcs.LGarXiv:2603.21884v12026
  11. Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

    Yuanhao Cai, Jing Lin, Xiaowan Hu +5

    cs.CVarXiv:2203.04845v32022
  12. Beyond Single Tokens: Distilling Discrete Diffusion Models via Discrete MMD

    Emiel Hoogeboom, David Ruhe, Jonathan Heek +2

    cs.LGcs.CVstat.MLarXiv:2603.20155v12026
  13. Fully Convolutional Networks for Panoptic Segmentation

    Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +4

    cs.CVarXiv:2012.00720v22020
  14. Scale Space Diffusion

    Soumik Mukhopadhyay, Prateksha Udhayanan, Abhinav Shrivastava

    cs.CVcs.AIarXiv:2603.08709v12026
  15. DDFlow: Learning Optical Flow with Unlabeled Data Distillation

    Pengpeng Liu, Irwin King, Michael R. Lyu +1

    cs.CVarXiv:1902.09145v12019
  16. Denoising Diffusion Generative Models Secretly Calculate Attentions

    Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3

    cs.AIcs.CVcs.LGarXiv:2609.00885v12026
  17. MIBURI: Towards Expressive Interactive Gesture Synthesis

    M. Hamza Mughal, Rishabh Dabral, Vera Demberg +1

    cs.CVcs.GRcs.HCarXiv:2603.03282v22026
  18. Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models

    Arnas Uselis, Andrea Dittadi, Seong Joon Oh

    cs.CVcs.LGarXiv:2602.24264v22026
  19. Learning Graph Embeddings for Compositional Zero-shot Learning

    Muhammad Ferjad Naeem, Yongqin Xian, Federico Tombari +1

    cs.CVarXiv:2102.01987v32021
  20. See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis

    Jaehyun Park, Minyoung Ahn, Minkyu Kim +3

    cs.CVcs.AIarXiv:2602.20951v22026
  21. Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

    Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

    cs.CVcs.CLarXiv:1811.10652v32018
  22. LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

    Xueyang Zhou, Yangming Xu, Guiyao Tie +5

    cs.CVcs.ROarXiv:2510.03827v22025
  23. Spatial-Angular Interaction for Light Field Image Super-Resolution

    Yingqian Wang, Longguang Wang, Jungang Yang +3

    eess.IVcs.CVarXiv:1912.07849v32019
  24. Deep Learning Object Detection Methods for Ecological Camera Trap Data

    Stefan Schneider, Graham W. Taylor, Stefan C. Kremer

    cs.CVarXiv:1803.10842v12018
  25. VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

    Zirui Wang, Junyi Zhang, Jiaxin Ge +9

    cs.CVarXiv:2601.16973v12026
  26. RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

    Yuheng Ji, Huajie Tan, Jiayu Shi +14

    cs.ROcs.CVarXiv:2502.21257v22025
  27. A Generative Appearance Model for End-to-end Video Object Segmentation

    Joakim Johnander, Martin Danelljan, Emil Brissman +2

    cs.CVarXiv:1811.11611v22018
  28. An All-in-One Network for Dehazing and Beyond

    Boyi Li, Xiulian Peng, Zhangyang Wang +2

    cs.CVcs.AIarXiv:1707.06543v12017
  29. Taming Visually Guided Sound Generation

    Vladimir Iashin, Esa Rahtu

    cs.CVcs.AIcs.LGarXiv:2110.08791v12021
  30. A Survey on Diffusion Models for Inverse Problems

    Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai +5

    cs.LGcs.AIcs.CVarXiv:2410.00083v12024
  31. InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation

    Haofan Wang, Matteo Spinelli, Qixun Wang +3

    cs.CVarXiv:2404.02733v22024
  32. Rethinking FID: Towards a Better Evaluation Metric for Image Generation

    Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit +3

    cs.CVarXiv:2401.09603v22023
  33. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

    Shengbang Tong, Zhuang Liu, Yuexiang Zhai +3

    cs.CVarXiv:2401.06209v22024
  34. Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

    cs.CVcs.AIarXiv:2311.18259v42023
  35. FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

    Shuang Zeng, Xinyuan Chang, Mengwei Xie +6

    cs.CVarXiv:2505.17685v32025
  36. Point Transformer V3: Simpler, Faster, Stronger

    Xiaoyang Wu, Li Jiang, Peng-Shuai Wang +6

    cs.CVarXiv:2312.10035v22023
  37. Photorealistic Video Generation with Diffusion Models

    Agrim Gupta, Lijun Yu, Kihyuk Sohn +6

    cs.CVcs.AIcs.LGarXiv:2312.06662v12023
  38. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

    Dongzhi Jiang, Ziyu Guo, Renrui Zhang +6

    cs.CVcs.AIcs.CLarXiv:2505.00703v22025
  39. Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising

    Xianlei Han, Shaoyu Wang, Jiancheng Fang +2

    cs.CVarXiv:2609.02654v12026
  40. Multimodal Foundation Models: From Specialists to General-Purpose Assistants

    Chunyuan Li, Zhe Gan, Zhengyuan Yang +4

    cs.CVcs.CLarXiv:2309.10020v12023
  41. Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

    Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao +6

    cs.LGcs.AIcs.CVarXiv:2310.02279v32023
  42. Self-Consuming Generative Models Go MAD

    Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi +5

    cs.LGcs.AIcs.CVarXiv:2307.01850v12023
  43. Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy

    Clare McGenity, Emily L Clarke, Charlotte Jennings +5

    physics.med-phcs.AIcs.CVarXiv:2306.07999v32023
  44. Generative Diffusion Prior for Unified Image Restoration and Enhancement

    Ben Fei, Zhaoyang Lyu, Liang Pan +5

    cs.CVarXiv:2304.01247v12023
  45. Leapfrog Diffusion Model for Stochastic Trajectory Prediction

    Weibo Mao, Chenxin Xu, Qi Zhu +2

    cs.CVarXiv:2303.10895v12023
  46. Consistency Models

    Yang Song, Prafulla Dhariwal, Mark Chen +1

    cs.LGcs.CVstat.MLarXiv:2303.01469v22023
  47. SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

    Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3

    cs.CLcs.CVarXiv:2301.04883v12023
  48. CREPE: Can Vision-Language Foundation Models Reason Compositionally?

    Zixian Ma, Jerry Hong, Mustafa Omer Gul +3

    cs.CLcs.CVarXiv:2212.07796v32022
  49. Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

    Narek Tumanyan, Michal Geyer, Shai Bagon +1

    cs.CVcs.AIarXiv:2211.12572v12022
  50. Diffusion Models: A Comprehensive Survey of Methods and Applications

    Ling Yang, Zhilong Zhang, Yang Song +6

    cs.LGcs.AIcs.CVarXiv:2209.00796v152022
  51. MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration

    Zhi Jin, Yuwei Qiu, Kaihao Zhang +2

    cs.CVarXiv:2501.04486v22025
  52. Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution

    Alice Lucas, Santiago Lopez Tapia, Rafael Molina +1

    cs.CVarXiv:1806.05764v22018
  53. State of the Art on Diffusion Models for Visual Computing

    Ryan Po, Wang Yifan, Vladislav Golyanik +15

    cs.AIcs.CVcs.GRarXiv:2310.07204v12023
  54. Streaming 4D Visual Geometry Transformer

    Dong Zhuo, Wenzhao Zheng, Jiahe Guo +3

    cs.CVcs.AIcs.LGarXiv:2507.11539v22025
  55. Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

    Haobo Yuan, Xiangtai Li, Tao Zhang +8

    cs.CVarXiv:2501.04001v42025
  56. TotalSegmentator: robust segmentation of 104 anatomical structures in CT images

    Jakob Wasserthal, Hanns-Christian Breit, Manfred T. Meyer +9

    eess.IVcs.CVarXiv:2208.05868v22022
  57. Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective

    Gabriel Resende Machado, Eugênio Silva, Ronaldo Ribeiro Goldschmidt

    cs.CVarXiv:2009.03728v12020
  58. Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

    Garrick Brazil, Abhinav Kumar, Julian Straub +3

    cs.CVarXiv:2207.10660v22022
  59. DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

    Yingyan Li, Shuyao Shang, Weisong Liu +10

    cs.CVcs.AIarXiv:2510.12796v22025
  60. Bridging the Domain Gap for Ground-to-Aerial Image Matching

    Krishna Regmi, Mubarak Shah

    cs.CVarXiv:1904.11045v22019