Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,061 to 15,120 of 18,839
Automated brain extraction of multi-sequence MRI using artificial neural networks
Fabian Isensee, Marianne Schell, Irada Tursunova +10
cs.CVarXiv:1901.11341v22019Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh +1
cs.CVcs.AIcs.CLarXiv:1712.00377v22017A Corpus for Reasoning About Natural Language Grounded in Photographs
Alane Suhr, Stephanie Zhou, Ally Zhang +3
cs.CLcs.CVarXiv:1811.00491v32018Summaries:한국어HDINO: A Concise and Efficient Open-Vocabulary Detector
Hao Zhang, Yiqun Wang, Qinran Lin +2
cs.CVarXiv:2603.02924v12026OmniCAD: A Large-Scale Benchmark for 3D Spatial Reasoning in Robotics Assemblies
Mingjia Wang, Taiting Lu, Ziwei Dong +25
cs.CVarXiv:2608.22637v12026A Probabilistic U-Net for Segmentation of Ambiguous Images
Simon A. A. Kohl, Bernardino Romera-Paredes, Clemens Meyer +6
cs.CVcs.LGcs.NEarXiv:1806.05034v42018The Event-Camera Dataset and Simulator: Event-based Data for Pose Estimation, Visual Odometry, and SLAM
Elias Mueggler, Henri Rebecq, Guillermo Gallego +2
cs.ROcs.CVarXiv:1610.08336v42016Invertible Residual Networks
Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen +2
cs.LGcs.AIcs.CVarXiv:1811.00995v32018Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning
Richard J. Chen, Chengkuan Chen, Yicong Li +4
cs.CVarXiv:2206.02647v12022SLER-IR: Spherical Layer-wise Expert Routing for All-in-One Image Restoration
Peng Shurui, Xin Lin, Shi Luo +5
cs.CVarXiv:2603.05940v12026BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
Zihao Zhu, Ruotong Wang, Siwei Lyu +2
cs.CVcs.AIarXiv:2603.02816v22026T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun +5
cs.CVarXiv:2301.06052v42023EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
Seungjun Lee, Zihan Wang, Yunsong Wang +1
cs.CVarXiv:2603.04254v12026TNT: Target-driveN Trajectory Prediction
Hang Zhao, Jiyang Gao, Tian Lan +9
cs.CVcs.ROarXiv:2008.08294v22020Locality-Attending Vision Transformer
Sina Hajimiri, Farzad Beizaee, Fereshteh Shakeri +3
cs.CVarXiv:2603.04892v12026TAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events
Jiaxiong Liu, Zhen Tan, Jinpu Zhang +4
cs.CVarXiv:2603.04989v22026Layer by layer, module by module: Choose both for optimal OOD probing of ViT
Ambroise Odonnat, Vasilii Feofanov, Laetitia Chapel +2
cs.CVcs.LGstat.MLarXiv:2603.05280v12026DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias +4
cs.CVcs.AIarXiv:2603.05607v22026PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
Xiang Zhang, Sohyun Yoo, Hongrui Wu +3
cs.CVcs.GRcs.LGarXiv:2603.05888v12026DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
Akash Haridas, Utkarsh Saxena, Parsa Ashrafi Fashi +3
cs.CVcs.AIcs.LGarXiv:2603.06351v22026FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models
Haoyang Li, Liang Wang, Siyu Zhou +5
cs.CVarXiv:2603.08708v12026Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
Yehonatan Elisha, Oren Barkan, Noam Koenigstein
cs.CVcs.AIcs.LGarXiv:2603.08309v22026Diffusion Models for Adversarial Purification
Weili Nie, Brandon Guo, Yujia Huang +3
cs.LGcs.CRcs.CVarXiv:2205.07460v12022Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation
Jose Caballero, Christian Ledig, Andrew Aitken +4
cs.CVarXiv:1611.05250v22016Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction
Yao Zhang, Zhuchenyang Liu, Yanlan He +2
cs.CVcs.IRarXiv:2603.09930v22026Visualizing Deep Neural Network Decisions: Prediction Difference Analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel +1
cs.CVcs.AIarXiv:1702.04595v12017OpenScene: 3D Scene Understanding with Open Vocabularies
Songyou Peng, Kyle Genova, Chiyu "Max" Jiang +3
cs.CVarXiv:2211.15654v22022CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing
Yucheng Wang, Zedong Wang, Yuetong Wu +2
cs.CVarXiv:2603.08589v12026Semi-supervised Domain Adaptation via Minimax Entropy
Kuniaki Saito, Donghyun Kim, Stan Sclaroff +2
cs.CVarXiv:1904.06487v52019BiCLIP: Domain Canonicalization via Structured Geometric Transformation
Pranav Mantini, Shishir K. Shah
cs.CVcs.AIcs.CLarXiv:2603.08942v22026PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
Danfei Xu, Dragomir Anguelov, Ashesh Jain
cs.CVarXiv:1711.10871v22017Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation
Xin Chen, Lingxi Xie, Jun Wu +1
cs.CVcs.LGarXiv:1904.12760v120194DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
Jin Lyu, Liang An, Pujin Cheng +2
cs.CVarXiv:2603.10125v12026Spatiotemporal Residual Networks for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
cs.CVarXiv:1611.02155v12016Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks
Vineet Kosaraju, Amir Sadeghian, Roberto Martín-Martín +3
cs.CVcs.LGarXiv:1907.03395v22019UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
Yaqi Zhao, Wang Lin, Zijian Zhang +5
cs.CVarXiv:2603.10702v12026Inverting Visual Representations with Convolutional Networks
Alexey Dosovitskiy, Thomas Brox
cs.NEcs.CVcs.LGarXiv:1506.02753v42015MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
Chenyang Zhu, Hongxiang Li, Xiu Li +1
cs.CVcs.AIcs.CLarXiv:2603.12743v12026SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
Anbang Wang, Yuzhuo Ao, Shangzhe Wu +1
cs.CVarXiv:2603.14152v22026VersaDB: A High-Performance AI Storage Database for Unifying Mutimodal Datasets
Cong Wang, Zelin Liu, Yang Luo Ran Zhang +6
cs.CVarXiv:2608.22795v12026SCoCCA: Multi-modal Sparse Concept Decomposition via Canonical Correlation Analysis
Ehud Gordon, Meir Yossef Levi, Guy Gilboa
cs.CVarXiv:2603.13884v12026Image-Conditioned Diffusion Models for Quality Assurance of Organ-at-Risk Segmentations in Radiotherapy
Clea Dronne, Catharine H Clark, Xavier Loizeau +3
cs.CVarXiv:2608.23432v12026V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Yan-Bo Lin, Jonah Casebeer, Long Mai +3
cs.CVcs.AIcs.LGarXiv:2603.11042v22026Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Jinheng Xie, Weijia Mao, Zechen Bai +7
cs.CVarXiv:2408.12528v72024Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline
Eugenia Moris, José Ignacio Orlando
cs.CVarXiv:2608.23336v12026Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
Umberto Cappellazzo, Stavros Petridis, Maja Pantic
eess.AScs.CVcs.SDarXiv:2603.12046v22026Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks
Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou +5
cs.ROcs.CVcs.LGarXiv:1809.10732v22018Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
Chengqun Yang, Liang Xu, Yanping Li +4
cs.CVarXiv:2608.23279v12026Anabranch Network for Camouflaged Object Segmentation
Trung-Nghia Le, Tam V. Nguyen, Zhongliang Nie +2
cs.CVarXiv:2105.09451v12021Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging
Amir Rezaei, Wen-Xin Pan, Giuseppe Caire
cs.CVeess.SParXiv:2608.23249v12026Revisiting Skeleton-based Action Recognition
Haodong Duan, Yue Zhao, Kai Chen +2
cs.CVarXiv:2104.13586v22021Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Large Language Models
Fan Xu, Luis A. Leiva
cs.CVcs.IRarXiv:2608.23102v12026Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski +2
cs.NEcs.AIcs.CVarXiv:1605.09304v52016Learning a Predictable and Generative Vector Representation for Objects
Rohit Girdhar, David F. Fouhey, Mikel Rodriguez +1
cs.CVarXiv:1603.08637v22016WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models
Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz +4
cs.CVarXiv:2608.22950v12026Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan +1
cs.CVarXiv:1708.01471v12017Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation
Yanliang Qi, Kexi Chen, Muchao Ye +1
cs.CVarXiv:2608.22819v12026Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan +5
cs.CVcs.GRarXiv:2011.02523v52020Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Junyao Hu, Zhongwei Cheng, Waikeung Wong +1
cs.CVarXiv:2603.14153v12026Rethinking Pre-training and Self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin +4
cs.CVcs.LGstat.MLarXiv:2006.06882v22020