Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

17,881 to 17,940 of 18,866

  1. Supervised Contrastive Learning

    Prannay Khosla, Piotr Teterwak, Chen Wang +6

    cs.LGcs.CVstat.MLarXiv:2004.11362v52020
  2. A Simple Framework for Contrastive Learning of Visual Representations

    Ting Chen, Simon Kornblith, Mohammad Norouzi +1

    cs.LGcs.CVstat.MLarXiv:2002.05709v32020
  3. Momentum Contrast for Unsupervised Visual Representation Learning

    Kaiming He, Haoqi Fan, Yuxin Wu +2

    cs.CVarXiv:1911.05722v32019
  4. Deep learning in agriculture: A survey

    Andreas Kamilaris, Francesc X. Prenafeta-Boldu

    cs.LGcs.CVstat.MLarXiv:1807.11809v12018
  5. CBAM: Convolutional Block Attention Module

    Sanghyun Woo, Jongchan Park, Joon-Young Lee +1

    cs.CVarXiv:1807.06521v22018
  6. BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

    Fisher Yu, Haofeng Chen, Xin Wang +5

    cs.CVarXiv:1805.04687v22018
  7. SphereFace: Deep Hypersphere Embedding for Face Recognition

    Weiyang Liu, Yandong Wen, Zhiding Yu +3

    cs.CVarXiv:1704.08063v42017
  8. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks

    Chelsea Finn, Pieter Abbeel, Sergey Levine

    cs.LGcs.AIcs.CVarXiv:1703.03400v32017
  9. Fully-Convolutional Siamese Networks for Object Tracking

    Luca Bertinetto, Jack Valmadre, João F. Henriques +2

    cs.CVarXiv:1606.09549v32016
  10. Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks

    Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li +1

    cs.CVarXiv:1604.02878v12016
  11. Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images

    Anh Nguyen, Jason Yosinski, Jeff Clune

    cs.CVcs.AIcs.NEarXiv:1412.1897v42014
  12. Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

    Jiawei Wang, Ke Rui, Yushen Zuo +2

    cs.LGcs.CVarXiv:2608.18746v12026
  13. H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

    Dingyi Rong, Yue Shi, Chaofan Ma +6

    cs.ROcs.CVarXiv:2608.13049v12026
  14. Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

    Woojung Han, Seil Kang, Youngjun Jun +3

    cs.CVarXiv:2606.06361v22026
  15. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

    Hojun Choi, Jaeyo Shin, Suin Lee +1

    cs.AIcs.CVcs.LGarXiv:2608.11616v22026
  16. Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder

    Sachin Deb, Harshit Sharma, Asif Salekin

    cs.CVcs.LGarXiv:2608.16087v12026
  17. The Limits of Binding in Dual Encoders

    Kin Ian Lo

    cs.LGcs.CLcs.CVarXiv:2608.15971v12026
  18. Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

    Taewook Kang, Taeheon Kim, Donghyun Shin +1

    cs.ROcs.CVcs.LGarXiv:2607.00666v12026
  19. Continuous Latent Diffusion Language Model

    Hongcan Guo, Qinyu Zhao, Yian Zhao +8

    cs.CLcs.AIcs.CVarXiv:2605.06548v12026
    Summaries:한국어
  20. LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

    Jianzong Wu, Hao Lian, Jiongfan Yang +12

    cs.CVarXiv:2606.06042v22026
  21. What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

    Zhengrong Yue, Taihang Hu, Mengting Chen +8

    cs.CVarXiv:2605.07915v12026
  22. Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

    Mohammad Javad Ahmadi, Hamid D. Taghirad

    cs.CVcs.AIarXiv:2608.17522v12026
  23. Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer

    Sergey Zagoruyko, Nikos Komodakis

    cs.CVarXiv:1612.03928v32016
  24. Is Space-Time Attention All You Need for Video Understanding?

    Gedas Bertasius, Heng Wang, Lorenzo Torresani

    cs.CVarXiv:2102.05095v42021
  25. VGGFace2: A dataset for recognising faces across pose and age

    Qiong Cao, Li Shen, Weidi Xie +2

    cs.CVarXiv:1710.08092v22017
  26. NetVLAD: CNN architecture for weakly supervised place recognition

    Relja Arandjelović, Petr Gronat, Akihiko Torii +2

    cs.CVcs.LGarXiv:1511.07247v32015
  27. Multi-View 3D Object Detection Network for Autonomous Driving

    Xiaozhi Chen, Huimin Ma, Ji Wan +2

    cs.CVarXiv:1611.07759v32016
  28. The Effectiveness of Data Augmentation in Image Classification using Deep Learning

    Luis Perez, Jason Wang

    cs.CVarXiv:1712.04621v12017
  29. DeepPose: Human Pose Estimation via Deep Neural Networks

    Alexander Toshev, Christian Szegedy

    cs.CVarXiv:1312.4659v32013
  30. Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking

    Ergys Ristani, Francesco Solera, Roger S. Zou +2

    cs.CVarXiv:1609.01775v22016
  31. CLIPScore: A Reference-free Evaluation Metric for Image Captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes +2

    cs.CVcs.CLarXiv:2104.08718v32021
  32. Semantic Image Synthesis with Spatially-Adaptive Normalization

    Taesung Park, Ming-Yu Liu, Ting-Chun Wang +1

    cs.CVcs.AIcs.GRarXiv:1903.07291v22019
  33. ViViT: A Video Vision Transformer

    Anurag Arnab, Mostafa Dehghani, Georg Heigold +3

    cs.CVarXiv:2103.15691v22021
  34. Empirical Evaluation of Rectified Activations in Convolutional Network

    Bing Xu, Naiyan Wang, Tianqi Chen +1

    cs.LGcs.CVstat.MLarXiv:1505.00853v22015
  35. Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles

    Mehdi Noroozi, Paolo Favaro

    cs.CVarXiv:1603.09246v32016
  36. KPConv: Flexible and Deformable Convolution for Point Clouds

    Hugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud +3

    cs.CVarXiv:1904.08889v22019
  37. BinaryConnect: Training Deep Neural Networks with binary weights during propagations

    Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David

    cs.LGcs.CVcs.NEarXiv:1511.00363v32015
  38. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang +12

    cs.CVarXiv:2312.14238v32023
  39. InstructPix2Pix: Learning to Follow Image Editing Instructions

    Tim Brooks, Aleksander Holynski, Alexei A. Efros

    cs.CVcs.AIcs.CLarXiv:2211.09800v22022
  40. FaceForensics++: Learning to Detect Manipulated Facial Images

    Andreas Rössler, Davide Cozzolino, Luisa Verdoliva +3

    cs.CVarXiv:1901.08971v32019
  41. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis +3

    cs.CVcs.AIcs.LGarXiv:1811.12231v32018
  42. Class-Balanced Loss Based on Effective Number of Samples

    Yin Cui, Menglin Jia, Tsung-Yi Lin +2

    cs.CVarXiv:1901.05555v12019
  43. CyCADA: Cycle-Consistent Adversarial Domain Adaptation

    Judy Hoffman, Eric Tzeng, Taesung Park +5

    cs.CVarXiv:1711.03213v32017
  44. Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels

    Zhilu Zhang, Mert R. Sabuncu

    cs.LGcs.CVstat.MLarXiv:1805.07836v42018
  45. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

    Deyao Zhu, Jun Chen, Xiaoqian Shen +2

    cs.CVarXiv:2304.10592v22023
  46. Shortcut Learning in Deep Neural Networks

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis +4

    cs.CVcs.AIcs.LGarXiv:2004.07780v52020
  47. YOLOv11: An Overview of the Key Architectural Enhancements

    Rahima Khanam, Muhammad Hussain

    cs.CVarXiv:2410.17725v12024
  48. Regularized Evolution for Image Classifier Architecture Search

    Esteban Real, Alok Aggarwal, Yanping Huang +1

    cs.NEcs.AIcs.CVarXiv:1802.01548v72018
  49. Generative Adversarial Text to Image Synthesis

    Scott Reed, Zeynep Akata, Xinchen Yan +3

    cs.NEcs.CVarXiv:1605.05396v22016
  50. CenterNet: Keypoint Triplets for Object Detection

    Kaiwen Duan, Song Bai, Lingxi Xie +3

    cs.CVarXiv:1904.08189v32019
  51. CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning

    Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu +9

    cs.CVcs.LGstat.MLarXiv:1711.05225v32017
  52. Occupancy Networks: Learning 3D Reconstruction in Function Space

    Lars Mescheder, Michael Oechsle, Michael Niemeyer +2

    cs.CVarXiv:1812.03828v22018
  53. MMDetection: Open MMLab Detection Toolbox and Benchmark

    Kai Chen, Jiaqi Wang, Jiangmiao Pang +22

    cs.CVcs.LGeess.IVarXiv:1906.07155v12019
  54. Unsupervised Deep Embedding for Clustering Analysis

    Junyuan Xie, Ross Girshick, Ali Farhadi

    cs.LGcs.CVarXiv:1511.06335v22015
  55. Object Detection in 20 Years: A Survey

    Zhengxia Zou, Keyan Chen, Zhenwei Shi +2

    cs.CVarXiv:1905.05055v32019
  56. Adversarial Machine Learning at Scale

    Alexey Kurakin, Ian Goodfellow, Samy Bengio

    cs.CVcs.CRcs.LGarXiv:1611.01236v22016
  57. Accelerating the Super-Resolution Convolutional Neural Network

    Chao Dong, Chen Change Loy, Xiaoou Tang

    cs.CVarXiv:1608.00367v12016
  58. Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

    Ze Zhang, Yang Zhang

    cs.CVcs.LGarXiv:2608.14717v12026
  59. MnasNet: Platform-Aware Neural Architecture Search for Mobile

    Mingxing Tan, Bo Chen, Ruoming Pang +4

    cs.CVcs.LGarXiv:1807.11626v32018
  60. PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

    Siddharth Patel

    cs.CVcs.AIarXiv:2608.14741v12026