Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,281 to 17,340 of 18,785

  1. OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

    Xudong Lu, Xueying Li, Annan Wang +8

    cs.CVcs.CLarXiv:2605.26485v12026
  2. MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

    Chaoyou Fu, Peixian Chen, Yunhang Shen +11

    cs.CVarXiv:2306.13394v52023
  3. Inverting Gradients -- How easy is it to break privacy in federated learning?

    Jonas Geiping, Hartmut Bauermeister, Hannah Dröge +1

    cs.CVcs.CRcs.LGarXiv:2003.14053v22020
  4. QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

    Dmitry Kalashnikov, Alex Irpan, Peter Pastor +8

    cs.LGcs.AIcs.CVarXiv:1806.10293v32018
  5. ResNeSt: Split-Attention Networks

    Hang Zhang, Chongruo Wu, Zhongyue Zhang +9

    cs.CVarXiv:2004.08955v22020
  6. Panoptic Segmentation

    Alexander Kirillov, Kaiming He, Ross Girshick +2

    cs.CVarXiv:1801.00868v32018
  7. Grounded Language-Image Pre-training

    Liunian Harold Li, Pengchuan Zhang, Haotian Zhang +9

    cs.CVcs.AIcs.CLarXiv:2112.03857v22021
  8. Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning

    Yuke Zhu, Roozbeh Mottaghi, Eric Kolve +4

    cs.CVarXiv:1609.05143v12016
  9. MemNet: A Persistent Memory Network for Image Restoration

    Ying Tai, Jian Yang, Xiaoming Liu +1

    cs.CVarXiv:1708.02209v12017
  10. Wider or Deeper: Revisiting the ResNet Model for Visual Recognition

    Zifeng Wu, Chunhua Shen, Anton van den Hengel

    cs.CVarXiv:1611.10080v12016
  11. Interpretable Explanations of Black Boxes by Meaningful Perturbation

    Ruth Fong, Andrea Vedaldi

    cs.CVcs.AIcs.LGarXiv:1704.03296v42017
  12. GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training

    Samet Akcay, Amir Atapour-Abarghouei, Toby P. Breckon

    cs.CVarXiv:1805.06725v32018
  13. Hierarchical Question-Image Co-Attention for Visual Question Answering

    Jiasen Lu, Jianwei Yang, Dhruv Batra +1

    cs.CVcs.CLarXiv:1606.00061v52016
  14. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

    Geonung Kim, Jeongeun Park, Nuri Ryu +2

    cs.GRcs.CVarXiv:2607.04540v12026
    Summaries:한국어
  15. ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

    Anindya Mondal, Sauradip Nag, Anjan Dutta

    cs.CVeess.IVarXiv:2606.23835v12026
    Summaries:한국어
  16. PhyCo: Learning Controllable Physical Priors for Generative Motion

    Sriram Narayanan, Ziyu Jiang, Srinivasa Narasimhan +1

    cs.CVcs.AIcs.LGarXiv:2604.28169v12026
    Summaries:한국어
  17. LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

    Chunyuan Li, Cliff Wong, Sheng Zhang +6

    cs.CVcs.CLarXiv:2306.00890v12023
  18. The Cityscapes Dataset for Semantic Urban Scene Understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos +6

    cs.CVarXiv:1604.01685v22016
  19. Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

    Bryan A. Plummer, Liwei Wang, Chris M. Cervantes +3

    cs.CVcs.CLarXiv:1505.04870v42015
  20. Video Enhancement with Task-Oriented Flow

    Tianfan Xue, Baian Chen, Jiajun Wu +2

    cs.CVarXiv:1711.09078v32017
  21. Large-Scale Evolution of Image Classifiers

    Esteban Real, Sherry Moore, Andrew Selle +5

    cs.NEcs.AIcs.CVarXiv:1703.01041v22017
  22. Generation and Comprehension of Unambiguous Object Descriptions

    Junhua Mao, Jonathan Huang, Alexander Toshev +3

    cs.CVcs.CLcs.LGarXiv:1511.02283v32015
  23. MesoNet: a Compact Facial Video Forgery Detection Network

    Darius Afchar, Vincent Nozick, Junichi Yamagishi +1

    cs.CVeess.IVarXiv:1809.00888v12018
  24. Deep Metric Learning via Lifted Structured Feature Embedding

    Hyun Oh Song, Yu Xiang, Stefanie Jegelka +1

    cs.CVcs.LGarXiv:1511.06452v12015
  25. Objaverse: A Universe of Annotated 3D Objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador +7

    cs.CVcs.AIcs.GRarXiv:2212.08051v12022
  26. Image Captioning with Semantic Attention

    Quanzeng You, Hailin Jin, Zhaowen Wang +2

    cs.CVarXiv:1603.03925v12016
  27. Fully Convolutional Siamese Networks for Change Detection

    Rodrigo Caye Daudt, Bertrand Le Saux, Alexandre Boulch

    cs.CVcs.LGarXiv:1810.08462v12018
  28. NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding

    Jun Liu, Amir Shahroudy, Mauricio Perez +3

    cs.CVarXiv:1905.04757v22019
  29. RMPE: Regional Multi-person Pose Estimation

    Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai +1

    cs.CVarXiv:1612.00137v52016
  30. Memorizing Normality to Detect Anomaly: Memory-augmented Deep Autoencoder for Unsupervised Anomaly Detection

    Dong Gong, Lingqiao Liu, Vuong Le +4

    cs.CVarXiv:1904.02639v22019
  31. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios

    Xingkui Zhu, Shuchang Lyu, Xu Wang +1

    cs.CVcs.AIarXiv:2108.11539v12021
  32. Kindling the Darkness: A Practical Low-light Image Enhancer

    Yonghua Zhang, Jiawan Zhang, Xiaojie Guo

    cs.CVarXiv:1905.04161v12019
  33. Towards Total Recall in Industrial Anomaly Detection

    Karsten Roth, Latha Pemula, Joaquin Zepeda +3

    cs.CVarXiv:2106.08265v22021
  34. A Survey on Contrastive Self-supervised Learning

    Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh +2

    cs.CVarXiv:2011.00362v32020
  35. Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

    Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1

    cs.CVarXiv:2605.28132v12026
  36. Deep Anomaly Detection with Outlier Exposure

    Dan Hendrycks, Mantas Mazeika, Thomas Dietterich

    cs.LGcs.CLcs.CVarXiv:1812.04606v32018
  37. Deeper, Broader and Artier Domain Generalization

    Da Li, Yongxin Yang, Yi-Zhe Song +1

    cs.CVarXiv:1710.03077v12017
  38. LoMo: Local Modality Substitution for Deeper Vision-Language Fusion

    Feng Han, Zhixiong Zhang, Zheming Liang +2

    cs.CVcs.CLarXiv:2605.30265v12026
  39. Learning a Variational Network for Reconstruction of Accelerated MRI Data

    Kerstin Hammernik, Teresa Klatzer, Erich Kobler +4

    cs.CVarXiv:1704.00447v12017
  40. Network Dissection: Quantifying Interpretability of Deep Visual Representations

    David Bau, Bolei Zhou, Aditya Khosla +2

    cs.CVcs.AIarXiv:1704.05796v12017
  41. Model-Contrastive Federated Learning

    Qinbin Li, Bingsheng He, Dawn Song

    cs.LGcs.AIcs.CVarXiv:2103.16257v12021
  42. NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection

    Golnaz Ghiasi, Tsung-Yi Lin, Ruoming Pang +1

    cs.CVcs.LGarXiv:1904.07392v12019
  43. Modeling Context in Referring Expressions

    Licheng Yu, Patrick Poirson, Shan Yang +2

    cs.CVcs.CLarXiv:1608.00272v32016
  44. CoCa: Contrastive Captioners are Image-Text Foundation Models

    Jiahui Yu, Zirui Wang, Vijay Vasudevan +3

    cs.CVcs.LGcs.MMarXiv:2205.01917v22022
  45. Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

    Artur Jesslen, Olaf Dünkel, Adam Kortylewski

    cs.CVarXiv:2605.30093v12026
  46. WIDER FACE: A Face Detection Benchmark

    Shuo Yang, Ping Luo, Chen Change Loy +1

    cs.CVarXiv:1511.06523v12015
  47. WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

    Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14

    cs.CVcs.CLarXiv:2605.29341v22026
  48. TensoRF: Tensorial Radiance Fields

    Anpei Chen, Zexiang Xu, Andreas Geiger +2

    cs.CVarXiv:2203.09517v22022
  49. Zero-1-to-3: Zero-shot One Image to 3D Object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick +3

    cs.CVcs.GRcs.ROarXiv:2303.11328v12023
  50. T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

    Chong Mou, Xintao Wang, Liangbin Xie +5

    cs.CVcs.AIcs.LGarXiv:2302.08453v22023
  51. One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation

    Sanghyun Jo, Seo Jin Lee, Seohyung Hong +4

    cs.CVarXiv:2605.29429v22026
  52. Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image

    Federica Bogo, Angjoo Kanazawa, Christoph Lassner +3

    cs.CVarXiv:1607.08128v12016
  53. Generalizing to Unseen Domains: A Survey on Domain Generalization

    Jindong Wang, Cuiling Lan, Chang Liu +6

    cs.LGcs.AIcs.CVarXiv:2103.03097v72021
  54. RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

    Ulrich Prestel, Stefan Andreas Baumann, Nick Stracke +1

    cs.CVcs.AIcs.LGarXiv:2605.31535v12026
  55. Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks

    Zhaofan Qiu, Ting Yao, Tao Mei

    cs.CVarXiv:1711.10305v12017
  56. Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

    Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2

    cs.CVarXiv:2605.29198v22026
  57. LVIS: A Dataset for Large Vocabulary Instance Segmentation

    Agrim Gupta, Piotr Dollár, Ross Girshick

    cs.CVarXiv:1908.03195v22019
  58. Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation

    Remi Denton, Wojciech Zaremba, Joan Bruna +2

    cs.CVcs.LGarXiv:1404.0736v22014
  59. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

    Pan Lu, Hritik Bansal, Tony Xia +7

    cs.CVcs.AIcs.CLarXiv:2310.02255v32023
  60. LLNet: A Deep Autoencoder Approach to Natural Low-light Image Enhancement

    Kin Gwn Lore, Adedotun Akintayo, Soumik Sarkar

    cs.CVarXiv:1511.03995v32015