Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
6,181 to 6,240 of 18,867
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Peiwen Zhang, Yufan Deng, Shangkun Sun +11
cs.CVcs.AIcs.ROarXiv:2606.28128v12026Summaries:한국어FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
Haorui Ji, Weizhe Liu, Hongdong Li +1
cs.CVcs.AIarXiv:2606.24874v12026$μ_0$: A Scalable 3D Interaction-Trace World Model
Seungjae Lee, Yoonkyo Jung, Jusuk Lee +6
cs.ROcs.CVcs.LGarXiv:2606.13769v22026Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Shangwen Zhu, Qianyu Peng, Zhao Pu +12
cs.CVarXiv:2605.18601v22026Map2World: Segment Map Conditioned Text to 3D World Generation
Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang +2
cs.CVarXiv:2605.00781v12026World2Minecraft: Occupancy-Driven Simulated Scenes Construction
Lechao Zhang, Haoran Xu, Jingyu Gong +3
cs.CVarXiv:2604.27578v120263DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
Stefan Schulz, Fernando Edelstein, Hannah Dröge +2
cs.CVcs.LGcs.MMarXiv:2604.11211v12026POS-ISP: Pipeline Optimization at the Sequence Level for Task-aware ISP
Jiyun Won, Heemin Yang, Woohyeok Kim +2
cs.CVarXiv:2604.06938v12026Video-Oasis: Rethinking Evaluation of Video Understanding
Geuntaek Lim, Sungjune Park, Jaeyun Lee +5
cs.CVarXiv:2603.29616v22026Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation
Donald Shenaj, Federico Errica, Antonio Carta
cs.CVcs.AIcs.LGarXiv:2603.21884v12026Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction
Yuanhao Cai, Jing Lin, Xiaowan Hu +5
cs.CVarXiv:2203.04845v32022Beyond Single Tokens: Distilling Discrete Diffusion Models via Discrete MMD
Emiel Hoogeboom, David Ruhe, Jonathan Heek +2
cs.LGcs.CVstat.MLarXiv:2603.20155v12026Fully Convolutional Networks for Panoptic Segmentation
Yanwei Li, Hengshuang Zhao, Xiaojuan Qi +4
cs.CVarXiv:2012.00720v22020Scale Space Diffusion
Soumik Mukhopadhyay, Prateksha Udhayanan, Abhinav Shrivastava
cs.CVcs.AIarXiv:2603.08709v12026DDFlow: Learning Optical Flow with Unlabeled Data Distillation
Pengpeng Liu, Irwin King, Michael R. Lyu +1
cs.CVarXiv:1902.09145v12019Denoising Diffusion Generative Models Secretly Calculate Attentions
Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3
cs.AIcs.CVcs.LGarXiv:2609.00885v12026MIBURI: Towards Expressive Interactive Gesture Synthesis
M. Hamza Mughal, Rishabh Dabral, Vera Demberg +1
cs.CVcs.GRcs.HCarXiv:2603.03282v22026Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
Arnas Uselis, Andrea Dittadi, Seong Joon Oh
cs.CVcs.LGarXiv:2602.24264v22026Learning Graph Embeddings for Compositional Zero-shot Learning
Muhammad Ferjad Naeem, Yongqin Xian, Federico Tombari +1
cs.CVarXiv:2102.01987v32021See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
Jaehyun Park, Minyoung Ahn, Minkyu Kim +3
cs.CVcs.AIarXiv:2602.20951v22026Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
cs.CVcs.CLarXiv:1811.10652v32018LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Xueyang Zhou, Yangming Xu, Guiyao Tie +5
cs.CVcs.ROarXiv:2510.03827v22025Spatial-Angular Interaction for Light Field Image Super-Resolution
Yingqian Wang, Longguang Wang, Jungang Yang +3
eess.IVcs.CVarXiv:1912.07849v32019Deep Learning Object Detection Methods for Ecological Camera Trap Data
Stefan Schneider, Graham W. Taylor, Stefan C. Kremer
cs.CVarXiv:1803.10842v12018VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Zirui Wang, Junyi Zhang, Jiaxin Ge +9
cs.CVarXiv:2601.16973v12026RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
Yuheng Ji, Huajie Tan, Jiayu Shi +14
cs.ROcs.CVarXiv:2502.21257v22025A Generative Appearance Model for End-to-end Video Object Segmentation
Joakim Johnander, Martin Danelljan, Emil Brissman +2
cs.CVarXiv:1811.11611v22018An All-in-One Network for Dehazing and Beyond
Boyi Li, Xiulian Peng, Zhangyang Wang +2
cs.CVcs.AIarXiv:1707.06543v12017Taming Visually Guided Sound Generation
Vladimir Iashin, Esa Rahtu
cs.CVcs.AIcs.LGarXiv:2110.08791v12021A Survey on Diffusion Models for Inverse Problems
Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai +5
cs.LGcs.AIcs.CVarXiv:2410.00083v12024InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
Haofan Wang, Matteo Spinelli, Qixun Wang +3
cs.CVarXiv:2404.02733v22024Rethinking FID: Towards a Better Evaluation Metric for Image Generation
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit +3
cs.CVarXiv:2401.09603v22023Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Shengbang Tong, Zhuang Liu, Yuexiang Zhai +3
cs.CVarXiv:2401.06209v22024Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
cs.CVcs.AIarXiv:2311.18259v42023FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
Shuang Zeng, Xinyuan Chang, Mengwei Xie +6
cs.CVarXiv:2505.17685v32025Point Transformer V3: Simpler, Faster, Stronger
Xiaoyang Wu, Li Jiang, Peng-Shuai Wang +6
cs.CVarXiv:2312.10035v22023Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn +6
cs.CVcs.AIcs.LGarXiv:2312.06662v12023T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Dongzhi Jiang, Ziyu Guo, Renrui Zhang +6
cs.CVcs.AIcs.CLarXiv:2505.00703v22025Physics-Driven Independent Pair Generation for Iterative Self-Supervised Low-Dose CT Denoising
Xianlei Han, Shaoyu Wang, Jiancheng Fang +2
cs.CVarXiv:2609.02654v12026Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang +4
cs.CVcs.CLarXiv:2309.10020v12023Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao +6
cs.LGcs.AIcs.CVarXiv:2310.02279v32023Self-Consuming Generative Models Go MAD
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi +5
cs.LGcs.AIcs.CVarXiv:2307.01850v12023Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy
Clare McGenity, Emily L Clarke, Charlotte Jennings +5
physics.med-phcs.AIcs.CVarXiv:2306.07999v32023Generative Diffusion Prior for Unified Image Restoration and Enhancement
Ben Fei, Zhaoyang Lyu, Liang Pan +5
cs.CVarXiv:2304.01247v12023Leapfrog Diffusion Model for Stochastic Trajectory Prediction
Weibo Mao, Chenxin Xu, Qi Zhu +2
cs.CVarXiv:2303.10895v12023Consistency Models
Yang Song, Prafulla Dhariwal, Mark Chen +1
cs.LGcs.CVstat.MLarXiv:2303.01469v22023SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida +3
cs.CLcs.CVarXiv:2301.04883v12023CREPE: Can Vision-Language Foundation Models Reason Compositionally?
Zixian Ma, Jerry Hong, Mustafa Omer Gul +3
cs.CLcs.CVarXiv:2212.07796v32022Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation
Narek Tumanyan, Michal Geyer, Shai Bagon +1
cs.CVcs.AIarXiv:2211.12572v12022Diffusion Models: A Comprehensive Survey of Methods and Applications
Ling Yang, Zhilong Zhang, Yang Song +6
cs.LGcs.AIcs.CVarXiv:2209.00796v152022MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration
Zhi Jin, Yuwei Qiu, Kaihao Zhang +2
cs.CVarXiv:2501.04486v22025Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution
Alice Lucas, Santiago Lopez Tapia, Rafael Molina +1
cs.CVarXiv:1806.05764v22018State of the Art on Diffusion Models for Visual Computing
Ryan Po, Wang Yifan, Vladislav Golyanik +15
cs.AIcs.CVcs.GRarXiv:2310.07204v12023Streaming 4D Visual Geometry Transformer
Dong Zhuo, Wenzhao Zheng, Jiahe Guo +3
cs.CVcs.AIcs.LGarXiv:2507.11539v22025Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos
Haobo Yuan, Xiangtai Li, Tao Zhang +8
cs.CVarXiv:2501.04001v42025TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
Jakob Wasserthal, Hanns-Christian Breit, Manfred T. Meyer +9
eess.IVcs.CVarXiv:2208.05868v22022Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective
Gabriel Resende Machado, Eugênio Silva, Ronaldo Ribeiro Goldschmidt
cs.CVarXiv:2009.03728v12020Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
Garrick Brazil, Abhinav Kumar, Julian Straub +3
cs.CVarXiv:2207.10660v22022DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
Yingyan Li, Shuyao Shang, Weisong Liu +10
cs.CVcs.AIarXiv:2510.12796v22025Bridging the Domain Gap for Ground-to-Aerial Image Matching
Krishna Regmi, Mubarak Shah
cs.CVarXiv:1904.11045v22019