Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,281 to 17,340 of 18,785
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
Xudong Lu, Xueying Li, Annan Wang +8
cs.CVcs.CLarXiv:2605.26485v12026MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu, Peixian Chen, Yunhang Shen +11
cs.CVarXiv:2306.13394v52023Inverting Gradients -- How easy is it to break privacy in federated learning?
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge +1
cs.CVcs.CRcs.LGarXiv:2003.14053v22020QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor +8
cs.LGcs.AIcs.CVarXiv:1806.10293v32018ResNeSt: Split-Attention Networks
Hang Zhang, Chongruo Wu, Zhongyue Zhang +9
cs.CVarXiv:2004.08955v22020Panoptic Segmentation
Alexander Kirillov, Kaiming He, Ross Girshick +2
cs.CVarXiv:1801.00868v32018Grounded Language-Image Pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang +9
cs.CVcs.AIcs.CLarXiv:2112.03857v22021Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
Yuke Zhu, Roozbeh Mottaghi, Eric Kolve +4
cs.CVarXiv:1609.05143v12016MemNet: A Persistent Memory Network for Image Restoration
Ying Tai, Jian Yang, Xiaoming Liu +1
cs.CVarXiv:1708.02209v12017Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
Zifeng Wu, Chunhua Shen, Anton van den Hengel
cs.CVarXiv:1611.10080v12016Interpretable Explanations of Black Boxes by Meaningful Perturbation
Ruth Fong, Andrea Vedaldi
cs.CVcs.AIcs.LGarXiv:1704.03296v42017GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training
Samet Akcay, Amir Atapour-Abarghouei, Toby P. Breckon
cs.CVarXiv:1805.06725v32018Hierarchical Question-Image Co-Attention for Visual Question Answering
Jiasen Lu, Jianwei Yang, Dhruv Batra +1
cs.CVcs.CLarXiv:1606.00061v52016SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control
Geonung Kim, Jeongeun Park, Nuri Ryu +2
cs.GRcs.CVarXiv:2607.04540v12026Summaries:한국어ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
Anindya Mondal, Sauradip Nag, Anjan Dutta
cs.CVeess.IVarXiv:2606.23835v12026Summaries:한국어PhyCo: Learning Controllable Physical Priors for Generative Motion
Sriram Narayanan, Ziyu Jiang, Srinivasa Narasimhan +1
cs.CVcs.AIcs.LGarXiv:2604.28169v12026Summaries:한국어LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Chunyuan Li, Cliff Wong, Sheng Zhang +6
cs.CVcs.CLarXiv:2306.00890v12023The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos +6
cs.CVarXiv:1604.01685v22016Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes +3
cs.CVcs.CLarXiv:1505.04870v42015Video Enhancement with Task-Oriented Flow
Tianfan Xue, Baian Chen, Jiajun Wu +2
cs.CVarXiv:1711.09078v32017Large-Scale Evolution of Image Classifiers
Esteban Real, Sherry Moore, Andrew Selle +5
cs.NEcs.AIcs.CVarXiv:1703.01041v22017Generation and Comprehension of Unambiguous Object Descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev +3
cs.CVcs.CLcs.LGarXiv:1511.02283v32015MesoNet: a Compact Facial Video Forgery Detection Network
Darius Afchar, Vincent Nozick, Junichi Yamagishi +1
cs.CVeess.IVarXiv:1809.00888v12018Deep Metric Learning via Lifted Structured Feature Embedding
Hyun Oh Song, Yu Xiang, Stefanie Jegelka +1
cs.CVcs.LGarXiv:1511.06452v12015Objaverse: A Universe of Annotated 3D Objects
Matt Deitke, Dustin Schwenk, Jordi Salvador +7
cs.CVcs.AIcs.GRarXiv:2212.08051v12022Image Captioning with Semantic Attention
Quanzeng You, Hailin Jin, Zhaowen Wang +2
cs.CVarXiv:1603.03925v12016Fully Convolutional Siamese Networks for Change Detection
Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch
cs.CVcs.LGarXiv:1810.08462v12018NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
Jun Liu, Amir Shahroudy, Mauricio Perez +3
cs.CVarXiv:1905.04757v22019RMPE: Regional Multi-person Pose Estimation
Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai +1
cs.CVarXiv:1612.00137v52016Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection
Dong Gong, Lingqiao Liu, Vuong Le +4
cs.CVarXiv:1904.02639v22019TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
Xingkui Zhu, Shuchang Lyu, Xu Wang +1
cs.CVcs.AIarXiv:2108.11539v12021Kindling the Darkness: A Practical Low-light Image Enhancer
Yonghua Zhang, Jiawan Zhang, Xiaojie Guo
cs.CVarXiv:1905.04161v12019Towards Total Recall in Industrial Anomaly Detection
Karsten Roth, Latha Pemula, Joaquin Zepeda +3
cs.CVarXiv:2106.08265v22021A Survey on Contrastive Self-supervised Learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh +2
cs.CVarXiv:2011.00362v32020Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1
cs.CVarXiv:2605.28132v12026Deep Anomaly Detection with Outlier Exposure
Dan Hendrycks, Mantas Mazeika, Thomas Dietterich
cs.LGcs.CLcs.CVarXiv:1812.04606v32018Deeper, Broader and Artier Domain Generalization
Da Li, Yongxin Yang, Yi-Zhe Song +1
cs.CVarXiv:1710.03077v12017LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
Feng Han, Zhixiong Zhang, Zheming Liang +2
cs.CVcs.CLarXiv:2605.30265v12026Learning a Variational Network for Reconstruction of Accelerated MRI Data
Kerstin Hammernik, Teresa Klatzer, Erich Kobler +4
cs.CVarXiv:1704.00447v12017Network Dissection: Quantifying Interpretability of Deep Visual Representations
David Bau, Bolei Zhou, Aditya Khosla +2
cs.CVcs.AIarXiv:1704.05796v12017Model-Contrastive Federated Learning
Qinbin Li, Bingsheng He, Dawn Song
cs.LGcs.AIcs.CVarXiv:2103.16257v12021NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection
Golnaz Ghiasi, Tsung-Yi Lin, Ruoming Pang +1
cs.CVcs.LGarXiv:1904.07392v12019Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang +2
cs.CVcs.CLarXiv:1608.00272v32016CoCa: Contrastive Captioners are Image-Text Foundation Models
Jiahui Yu, Zirui Wang, Vijay Vasudevan +3
cs.CVcs.LGcs.MMarXiv:2205.01917v22022Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
Artur Jesslen, Olaf Dünkel, Adam Kortylewski
cs.CVarXiv:2605.30093v12026WIDER FACE: A Face Detection Benchmark
Shuo Yang, Ping Luo, Chen Change Loy +1
cs.CVarXiv:1511.06523v12015WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14
cs.CVcs.CLarXiv:2605.29341v22026TensoRF: Tensorial Radiance Fields
Anpei Chen, Zexiang Xu, Andreas Geiger +2
cs.CVarXiv:2203.09517v22022Zero-1-to-3: Zero-shot One Image to 3D Object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick +3
cs.CVcs.GRcs.ROarXiv:2303.11328v12023T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
Chong Mou, Xintao Wang, Liangbin Xie +5
cs.CVcs.AIcs.LGarXiv:2302.08453v22023One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation
Sanghyun Jo, Seo Jin Lee, Seohyung Hong +4
cs.CVarXiv:2605.29429v22026Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
Federica Bogo, Angjoo Kanazawa, Christoph Lassner +3
cs.CVarXiv:1607.08128v12016Generalizing to Unseen Domains: A Survey on Domain Generalization
Jindong Wang, Cuiling Lan, Chang Liu +6
cs.LGcs.AIcs.CVarXiv:2103.03097v72021RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
Ulrich Prestel, Stefan Andreas Baumann, Nick Stracke +1
cs.CVcs.AIcs.LGarXiv:2605.31535v12026Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
Zhaofan Qiu, Ting Yao, Tao Mei
cs.CVarXiv:1711.10305v12017Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2
cs.CVarXiv:2605.29198v22026LVIS: A Dataset for Large Vocabulary Instance Segmentation
Agrim Gupta, Piotr Dollár, Ross Girshick
cs.CVarXiv:1908.03195v22019Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation
Remi Denton, Wojciech Zaremba, Joan Bruna +2
cs.CVcs.LGarXiv:1404.0736v22014MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Pan Lu, Hritik Bansal, Tony Xia +7
cs.CVcs.AIcs.CLarXiv:2310.02255v32023LLNet: A Deep Autoencoder Approach to Natural Low-light Image Enhancement
Kin Gwn Lore, Adedotun Akintayo, Soumik Sarkar
cs.CVarXiv:1511.03995v32015