Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,881 to 14,940 of 18,868
EfficientFormer: Vision Transformers at MobileNet Speed
Yanyu Li, Geng Yuan, Yang Wen +5
cs.CVarXiv:2206.01191v52022Deep learning with noisy labels: exploring techniques and remedies in medical image analysis
Davood Karimi, Haoran Dou, Simon K. Warfield +1
cs.CVcs.LGeess.IVarXiv:1912.02911v42019MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Bocheng Zou, Mu Cai, Mark Stanley +2
cs.CVarXiv:2603.25744v22026DeepStereo: Learning to Predict New Views from the World's Imagery
John Flynn, Ivan Neulander, James Philbin +1
cs.CVarXiv:1506.06825v12015Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
Dangwei Li, Xiaotang Chen, Zhang Zhang +1
cs.CVarXiv:1710.06555v12017Unified Number-Free Text-to-Motion Generation Via Flow Matching
Guanhe Huang, Oya Celiktutan
cs.CVarXiv:2603.27040v12026CHIMERA Challenge: Biochemical Recurrence Prediction in Prostate Cancer Patients using multimodal datasets
Robert N. Spaans, Catherine Chia, Tongjie Wang +7
eess.IVcs.CVarXiv:2608.21497v12026Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
cs.CVarXiv:2604.22280v32026Summaries:한국어Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
Jie Hu, Li Shen, Samuel Albanie +2
cs.CVarXiv:1810.12348v32018VirtualHome: Simulating Household Activities via Programs
Xavier Puig, Kevin Ra, Marko Boben +4
cs.CVcs.AIcs.LGarXiv:1806.07011v12018MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes
Mingzhe Du, Thong Thanh Nguyen, Nguyen Tran Cong Duy +2
cs.CVcs.AIarXiv:2608.24107v12026Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis
Ming Cheng, Hongyu Sun, Zhaolin Chen +3
cs.CVcs.MMarXiv:2608.23974v120266-DOF GraspNet: Variational Grasp Generation for Object Manipulation
Arsalan Mousavian, Clemens Eppner, Dieter Fox
cs.CVcs.ROarXiv:1905.10520v22019InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis
Prateek Mittal, Ayush Srivastava, Joohi Chauhan
q-bio.QMcs.CVcs.ITarXiv:2608.23574v12026Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking
Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le +10
cs.CVcs.AIarXiv:2608.24130v12026Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection
Long Hoang Pham, Quoc Pham-Nam Ho, Huy-Hung Nguyen +10
cs.CVcs.AIarXiv:2608.24154v12026Infant Care Video Dataset for Classification of Interventions Using Transformers
Igor Bogdanov, James Green
cs.CVcs.AIcs.LGarXiv:2608.23838v12026Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil +2
cs.CVarXiv:2603.27183v22026A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
Kirill Skobelev, Eric Fithian, Yegor Baranovski +9
cs.AIcs.CVcs.LGarXiv:2603.27341v42026Towards Universal Fake Image Detectors that Generalize Across Generative Models
Utkarsh Ojha, Yuheng Li, Yong Jae Lee
cs.CVcs.LGarXiv:2302.10174v22023When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri +2
cs.CVcs.AIcs.CLarXiv:2210.01936v32022Falcon Perception
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou +6
cs.CVarXiv:2603.27365v12026Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
Yuhe Liu, Zhenxiong Tan, Yujia Hu +2
cs.CVarXiv:2603.27666v12026RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching
Lahav Lipson, Zachary Teed, Jia Deng
cs.CVarXiv:2109.07547v12021Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang +10
cs.CVarXiv:2608.24138v12026CNN: Single-label to Multi-label
Yunchao Wei, Wei Xia, Junshi Huang +4
cs.CVarXiv:1406.5726v32014An Unsupervised Learning Model for Deformable Medical Image Registration
Guha Balakrishnan, Amy Zhao, Mert R. Sabuncu +2
cs.CVarXiv:1802.02604v32018A Survey of Deep Learning Applications to Autonomous Vehicle Control
Sampo Kuutti, Richard Bowden, Yaochu Jin +2
cs.LGcs.CVeess.SYarXiv:1912.10773v12019Deep Sliding Shapes for Amodal 3D Object Detection in RGB-D Images
Shuran Song, Jianxiong Xiao
cs.CVarXiv:1511.02300v22015A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
Vasiliki Kondyli, Jakob Suchan, Mehul Bhatt
q-bio.NCcs.AIcs.CVarXiv:2608.23572v12026Fast-SCNN: Fast Semantic Segmentation Network
Rudra P K Poudel, Stephan Liwicki, Roberto Cipolla
cs.CVarXiv:1902.04502v12019LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Haoning Wu, Dongxu Li, Bei Chen +1
cs.CVcs.CLcs.LGarXiv:2407.15754v12024PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
Lorenza Prospero, Orest Kupyn, Ostap Viniavskyi +2
cs.CVarXiv:2603.28763v12026Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
Tommaso Galliena, Stefano Rosa, Tommaso Apicella +3
cs.CVarXiv:2603.24257v22026SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
Wenli Li, Kai Zhao, Haoran Jiang +3
cs.CVarXiv:2603.29437v12026Asymmetric Non-local Neural Networks for Semantic Segmentation
Zhen Zhu, Mengde Xu, Song Bai +2
cs.CVcs.LGarXiv:1908.07678v52019ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding
Yuanhao Sun, Huawei Ji, Jiaxin Ding +2
cs.CVarXiv:2608.22996v12026Robust and Efficient Subspace Segmentation via Least Squares Regression
Can-Yi Lu, Hai Min, Zhong-Qiu Zhao +3
cs.CVarXiv:1404.6736v12014CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
Oier Mees, Lukas Hermann, Erick Rosete-Beas +1
cs.ROcs.AIcs.CLarXiv:2112.03227v42021LERF: Language Embedded Radiance Fields
Justin Kerr, Chung Min Kim, Ken Goldberg +2
cs.CVcs.GRarXiv:2303.09553v12023TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
Yunseok Jang, Yale Song, Youngjae Yu +2
cs.CVarXiv:1704.04497v32017MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
Bharath Krishnamurthy, Ajita Rattani
cs.CVcs.AIarXiv:2603.29029v12026NeRV: Neural Reflectance and Visibility Fields for Relighting and View Synthesis
Pratul P. Srinivasan, Boyang Deng, Xiuming Zhang +3
cs.CVcs.GRarXiv:2012.03927v12020StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets
Axel Sauer, Katja Schwarz, Andreas Geiger
cs.LGcs.CVarXiv:2202.00273v22022Contrastive learning of global and local features for medical image segmentation with limited annotations
Krishna Chaitanya, Ertunc Erdil, Neerav Karani +1
cs.CVcs.LGeess.IVarXiv:2006.10511v22020RawGen: Learning Camera Raw Image Generation
Dongyoung Kim, Junyong Lee, Abhijith Punnappurath +4
cs.CVarXiv:2604.00093v12026Improving Diffusion Models for Inverse Problems using Manifold Constraints
Hyungjin Chung, Byeongsu Sim, Dohoon Ryu +1
cs.LGcs.AIcs.CVarXiv:2206.00941v32022HP-UniIF: Hierarchical Prompt Learning for Unified Image Fusion
Xingxin Xu, Siqi Zhao, Xin Li +3
cs.CVarXiv:2608.21786v12026Which Tasks Should Be Learned Together in Multi-task Learning?
Trevor Standley, Amir R. Zamir, Dawn Chen +3
cs.CVarXiv:1905.07553v42019PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
Yuhui Xu, Lingxi Xie, Xiaopeng Zhang +4
cs.CVcs.LGarXiv:1907.05737v42019Learning by Cheating
Dian Chen, Brady Zhou, Vladlen Koltun +1
cs.ROcs.AIcs.CVarXiv:1912.12294v12019Measuring Robustness to Natural Distribution Shifts in Image Classification
Rohan Taori, Achal Dave, Vaishaal Shankar +3
cs.LGcs.CVstat.MLarXiv:2007.00644v22020SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
Nanye Ma, Mark Goldstein, Michael S. Albergo +3
cs.CVcs.LGarXiv:2401.08740v22024Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
Hongxu Yin, Pavlo Molchanov, Zhizhong Li +5
cs.LGcs.CVstat.MLarXiv:1912.08795v22019InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs
Yujun Shen, Ceyuan Yang, Xiaoou Tang +1
cs.CVcs.LGeess.IVarXiv:2005.09635v22020ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology
Duncan Stothers, Ren-Chin Wu, William Lotter
cs.CVcs.AIcs.LGarXiv:2608.22066v12026Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection
Debidatta Dwibedi, Ishan Misra, Martial Hebert
cs.CVarXiv:1708.01642v12017Gate Voltage Effect on Pulse Detection Efficiency of Perimeter-Gated SPADs
Hunter Guthrie, Md Sakibur Sajal, Zexi Liu +1
physics.ins-detcs.CVarXiv:2608.21371v12026VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen +8
cs.CVarXiv:2305.11175v22023Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
Cheng Chen, Yifan Zhao, Jia Li
cs.CVarXiv:2608.22313v12026