Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
15,121 to 15,180 of 18,866
Spatiotemporal Residual Networks for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
cs.CVarXiv:1611.02155v12016Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks
Vineet Kosaraju, Amir Sadeghian, Roberto Martín-Martín +3
cs.CVcs.LGarXiv:1907.03395v22019UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
Yaqi Zhao, Wang Lin, Zijian Zhang +5
cs.CVarXiv:2603.10702v12026Inverting Visual Representations with Convolutional Networks
Alexey Dosovitskiy, Thomas Brox
cs.NEcs.CVcs.LGarXiv:1506.02753v42015MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
Chenyang Zhu, Hongxiang Li, Xiu Li +1
cs.CVcs.AIcs.CLarXiv:2603.12743v12026SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
Anbang Wang, Yuzhuo Ao, Shangzhe Wu +1
cs.CVarXiv:2603.14152v22026VersaDB: A High-Performance AI Storage Database for Unifying Mutimodal Datasets
Cong Wang, Zelin Liu, Yang Luo Ran Zhang +6
cs.CVarXiv:2608.22795v12026SCoCCA: Multi-modal Sparse Concept Decomposition via Canonical Correlation Analysis
Ehud Gordon, Meir Yossef Levi, Guy Gilboa
cs.CVarXiv:2603.13884v12026Image-Conditioned Diffusion Models for Quality Assurance of Organ-at-Risk Segmentations in Radiotherapy
Clea Dronne, Catharine H Clark, Xavier Loizeau +3
cs.CVarXiv:2608.23432v12026V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Yan-Bo Lin, Jonah Casebeer, Long Mai +3
cs.CVcs.AIcs.LGarXiv:2603.11042v22026Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Jinheng Xie, Weijia Mao, Zechen Bai +7
cs.CVarXiv:2408.12528v72024Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline
Eugenia Moris, José Ignacio Orlando
cs.CVarXiv:2608.23336v12026Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
Umberto Cappellazzo, Stavros Petridis, Maja Pantic
eess.AScs.CVcs.SDarXiv:2603.12046v22026Multimodal Trajectory Predictions for Autonomous Driving using Deep Convolutional Networks
Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou +5
cs.ROcs.CVcs.LGarXiv:1809.10732v22018Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
Chengqun Yang, Liang Xu, Yanping Li +4
cs.CVarXiv:2608.23279v12026Anabranch Network for Camouflaged Object Segmentation
Trung-Nghia Le, Tam V. Nguyen, Zhongliang Nie +2
cs.CVarXiv:2105.09451v12021Semantic Reconstruction and 3-D Detection via Learned Multi-Pair Fusion in RF Imaging
Amir Rezaei, Wen-Xin Pan, Giuseppe Caire
cs.CVeess.SParXiv:2608.23249v12026Revisiting Skeleton-based Action Recognition
Haodong Duan, Yue Zhao, Kai Chen +2
cs.CVarXiv:2104.13586v22021Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Large Language Models
Fan Xu, Luis A. Leiva
cs.CVcs.IRarXiv:2608.23102v12026Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski +2
cs.NEcs.AIcs.CVarXiv:1605.09304v52016Learning a Predictable and Generative Vector Representation for Objects
Rohit Girdhar, David F. Fouhey, Mikel Rodriguez +1
cs.CVarXiv:1603.08637v22016WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models
Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz +4
cs.CVarXiv:2608.22950v12026Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan +1
cs.CVarXiv:1708.01471v12017Direct, Parallel, or Sequential? A Comparative Study of Training-Free Multi-Subject Image-to-Video Generation
Yanliang Qi, Kexi Chen, Muchao Ye +1
cs.CVarXiv:2608.22819v12026Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan +5
cs.CVcs.GRarXiv:2011.02523v52020Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Junyao Hu, Zhongwei Cheng, Waikeung Wong +1
cs.CVarXiv:2603.14153v12026Rethinking Pre-training and Self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin +4
cs.CVcs.LGstat.MLarXiv:2006.06882v22020PREDATOR: Registration of 3D Point Clouds with Low Overlap
Shengyu Huang, Zan Gojcic, Mikhail Usvyatsov +2
cs.CVeess.IVarXiv:2011.13005v32020SEEDS: Superpixels Extracted via Energy-Driven Sampling
Michael Van den Bergh, Xavier Boix, Gemma Roig +1
cs.CVarXiv:1309.3848v12013WiT: Waypoint Diffusion Transformers via Trajectory Conflict Navigation
Hainuo Wang, Mingjia Li, Xiaojie Guo
cs.CVarXiv:2603.15132v22026DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang +31
cs.CVcs.CLcs.LGarXiv:2304.14108v52023A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
Mateusz Malinowski, Mario Fritz
cs.AIcs.CLcs.CVarXiv:1410.0210v42014ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
Ruonan Yu, Zhenxiong Tan, Zigeng Chen +2
cs.CVarXiv:2603.15478v12026Fooling automated surveillance cameras: adversarial patches to attack person detection
Simen Thys, Wiebe Van Ranst, Toon Goedemé
cs.CVarXiv:1904.08653v12019Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen +8
cs.CVarXiv:2603.15614v12026DeMoN: Depth and Motion Network for Learning Monocular Stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig +4
cs.CVarXiv:1612.02401v22016Rotation Equivariant CNNs for Digital Pathology
Bastiaan S. Veeling, Jasper Linmans, Jim Winkens +2
cs.CVcs.LGstat.MLarXiv:1806.03962v12018Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin Qu, Haozhe Qi, Mihai Dusmanu +3
cs.CVcs.AIcs.CLarXiv:2603.18002v12026Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks
Guotai Wang, Wenqi Li, Michael Aertsen +3
cs.CVarXiv:1807.07356v32018Prompt-Free Universal Region Proposal Network
Qihong Tang, Changhan Liu, Shaofeng Zhang +3
cs.CVarXiv:2603.17554v12026Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu +3
cs.CVarXiv:1406.2031v12014Chain-of-Trajectories: Unlocking the Intrinsic Generative Optimality of Diffusion Models via Graph-Theoretic Planning
Ping Chen, Xiang Liu, Xingpeng Zhang +7
cs.LGcs.CVstat.MLarXiv:2603.14704v12026Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Andy Zeng, Maria Attarian, Brian Ichter +10
cs.CVcs.AIcs.CLarXiv:2204.00598v22022AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers
Salim Khazem
cs.CVcs.AIcs.LGarXiv:2603.14706v12026ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
Cheng Luo, Bizhu Wu, Bing Li +5
cs.CVcs.AIcs.HCarXiv:2603.15083v12026Latent Embeddings for Zero-shot Classification
Yongqin Xian, Zeynep Akata, Gaurav Sharma +3
cs.CVarXiv:1603.08895v22016Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching
Fangran Miao, Jian Huang, Ting Li
cs.CVstat.MLarXiv:2603.15016v12026SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation
Shufan Li, Jiuxiang Gu, Kangning Liu +3
cs.CVarXiv:2603.15150v12026Efficient Document Parsing via Parallel Token Prediction
Lei Li, Ze Zhao, Meng Li +6
cs.CLcs.CVarXiv:2603.15206v12026Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
Lexiang Xiong, Qi Li, Jingwen Ye +1
cs.CVarXiv:2603.15557v12026ViT-AdaLA: Adapting Vision Transformers with Linear Attention
Yifan Li, Seunghyun Yoon, Viet Dac Lai +4
cs.CVarXiv:2603.16063v12026Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
Peng Sun, Jun Xie, Tao Lin
cs.CVarXiv:2603.16139v12026Mixture of Style Experts for Diverse Image Stylization
Shihao Zhu, Ziheng Ouyang, Yijia Kang +5
cs.CVarXiv:2603.16649v32026ACE-LoRA: Graph-Attentive Context Enhancement for Parameter-Efficient Adaptation of Medical Vision-Language Models
M. Arda Aydın, Melih B. Yilmaz, Aykut Koç +1
cs.CVarXiv:2603.17079v12026AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time
Hao-Shu Fang, Jiefeng Li, Hongyang Tang +5
cs.CVarXiv:2211.03375v12022Multi-Task Feature Learning Via Efficient l2,1-Norm Minimization
Jun Liu, Shuiwang Ji, Jieping Ye
cs.LGcs.CVstat.MLarXiv:1205.2631v12012Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views
Hao Su, Charles R. Qi, Yangyan Li +1
cs.CVarXiv:1505.05641v12015ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
Thomas De Min, Subhankar Roy, Stéphane Lathuilière +2
cs.CVarXiv:2603.19466v22026EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings
Md Thamed Bin Zaman Chowdhury, Moazzem Hossain
cs.CVcs.AIarXiv:2608.23563v12026AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting
Ye Yuan, Xinshuo Weng, Yanglan Ou +1
cs.AIcs.CVcs.LGarXiv:2103.14023v32021