Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
17,881 to 17,940 of 18,866
Supervised Contrastive Learning
Prannay Khosla, Piotr Teterwak, Chen Wang +6
cs.LGcs.CVstat.MLarXiv:2004.11362v52020A Simple Framework for Contrastive Learning of Visual Representations
Ting Chen, Simon Kornblith, Mohammad Norouzi +1
cs.LGcs.CVstat.MLarXiv:2002.05709v32020Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu +2
cs.CVarXiv:1911.05722v32019Deep learning in agriculture: A survey
Andreas Kamilaris, Francesc X. Prenafeta-Boldu
cs.LGcs.CVstat.MLarXiv:1807.11809v12018CBAM: Convolutional Block Attention Module
Sanghyun Woo, Jongchan Park, Joon-Young Lee +1
cs.CVarXiv:1807.06521v22018BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
Fisher Yu, Haofeng Chen, Xin Wang +5
cs.CVarXiv:1805.04687v22018SphereFace: Deep Hypersphere Embedding for Face Recognition
Weiyang Liu, Yandong Wen, Zhiding Yu +3
cs.CVarXiv:1704.08063v42017Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Chelsea Finn, Pieter Abbeel, Sergey Levine
cs.LGcs.AIcs.CVarXiv:1703.03400v32017Fully-Convolutional Siamese Networks for Object Tracking
Luca Bertinetto, Jack Valmadre, João F. Henriques +2
cs.CVarXiv:1606.09549v32016Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks
Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li +1
cs.CVarXiv:1604.02878v12016Deep Neural Networks are Easily Fooled: High Confidence Predictions for Unrecognizable Images
Anh Nguyen, Jason Yosinski, Jeff Clune
cs.CVcs.AIcs.NEarXiv:1412.1897v42014Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
Jiawei Wang, Ke Rui, Yushen Zuo +2
cs.LGcs.CVarXiv:2608.18746v12026H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Dingyi Rong, Yue Shi, Chaofan Ma +6
cs.ROcs.CVarXiv:2608.13049v12026Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
Woojung Han, Seil Kang, Youngjun Jun +3
cs.CVarXiv:2606.06361v22026MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Hojun Choi, Jaeyo Shin, Suin Lee +1
cs.AIcs.CVcs.LGarXiv:2608.11616v22026Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder
Sachin Deb, Harshit Sharma, Asif Salekin
cs.CVcs.LGarXiv:2608.16087v12026The Limits of Binding in Dual Encoders
Kin Ian Lo
cs.LGcs.CLcs.CVarXiv:2608.15971v12026Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
Taewook Kang, Taeheon Kim, Donghyun Shin +1
cs.ROcs.CVcs.LGarXiv:2607.00666v12026Continuous Latent Diffusion Language Model
Hongcan Guo, Qinyu Zhao, Yian Zhao +8
cs.CLcs.AIcs.CVarXiv:2605.06548v12026Summaries:한국어LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing
Jianzong Wu, Hao Lian, Jiongfan Yang +12
cs.CVarXiv:2606.06042v22026What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
Zhengrong Yue, Taihang Hu, Mengting Chen +8
cs.CVarXiv:2605.07915v12026Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
Mohammad Javad Ahmadi, Hamid D. Taghirad
cs.CVcs.AIarXiv:2608.17522v12026Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
Sergey Zagoruyko, Nikos Komodakis
cs.CVarXiv:1612.03928v32016Is Space-Time Attention All You Need for Video Understanding?
Gedas Bertasius, Heng Wang, Lorenzo Torresani
cs.CVarXiv:2102.05095v42021VGGFace2: A dataset for recognising faces across pose and age
Qiong Cao, Li Shen, Weidi Xie +2
cs.CVarXiv:1710.08092v22017NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelović, Petr Gronat, Akihiko Torii +2
cs.CVcs.LGarXiv:1511.07247v32015Multi-View 3D Object Detection Network for Autonomous Driving
Xiaozhi Chen, Huimin Ma, Ji Wan +2
cs.CVarXiv:1611.07759v32016The Effectiveness of Data Augmentation in Image Classification using Deep Learning
Luis Perez, Jason Wang
cs.CVarXiv:1712.04621v12017DeepPose: Human Pose Estimation via Deep Neural Networks
Alexander Toshev, Christian Szegedy
cs.CVarXiv:1312.4659v32013Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking
Ergys Ristani, Francesco Solera, Roger S. Zou +2
cs.CVarXiv:1609.01775v22016CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes +2
cs.CVcs.CLarXiv:2104.08718v32021Semantic Image Synthesis with Spatially-Adaptive Normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang +1
cs.CVcs.AIcs.GRarXiv:1903.07291v22019ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold +3
cs.CVarXiv:2103.15691v22021Empirical Evaluation of Rectified Activations in Convolutional Network
Bing Xu, Naiyan Wang, Tianqi Chen +1
cs.LGcs.CVstat.MLarXiv:1505.00853v22015Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
Mehdi Noroozi, Paolo Favaro
cs.CVarXiv:1603.09246v32016KPConv: Flexible and Deformable Convolution for Point Clouds
Hugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud +3
cs.CVarXiv:1904.08889v22019BinaryConnect: Training Deep Neural Networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David
cs.LGcs.CVcs.NEarXiv:1511.00363v32015InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Zhe Chen, Jiannan Wu, Wenhai Wang +12
cs.CVarXiv:2312.14238v32023InstructPix2Pix: Learning to Follow Image Editing Instructions
Tim Brooks, Aleksander Holynski, Alexei A. Efros
cs.CVcs.AIcs.CLarXiv:2211.09800v22022FaceForensics++: Learning to Detect Manipulated Facial Images
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva +3
cs.CVarXiv:1901.08971v32019ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis +3
cs.CVcs.AIcs.LGarXiv:1811.12231v32018Class-Balanced Loss Based on Effective Number of Samples
Yin Cui, Menglin Jia, Tsung-Yi Lin +2
cs.CVarXiv:1901.05555v12019CyCADA: Cycle-Consistent Adversarial Domain Adaptation
Judy Hoffman, Eric Tzeng, Taesung Park +5
cs.CVarXiv:1711.03213v32017Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
Zhilu Zhang, Mert R. Sabuncu
cs.LGcs.CVstat.MLarXiv:1805.07836v42018MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Deyao Zhu, Jun Chen, Xiaoqian Shen +2
cs.CVarXiv:2304.10592v22023Shortcut Learning in Deep Neural Networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis +4
cs.CVcs.AIcs.LGarXiv:2004.07780v52020YOLOv11: An Overview of the Key Architectural Enhancements
Rahima Khanam, Muhammad Hussain
cs.CVarXiv:2410.17725v12024Regularized Evolution for Image Classifier Architecture Search
Esteban Real, Alok Aggarwal, Yanping Huang +1
cs.NEcs.AIcs.CVarXiv:1802.01548v72018Generative Adversarial Text to Image Synthesis
Scott Reed, Zeynep Akata, Xinchen Yan +3
cs.NEcs.CVarXiv:1605.05396v22016CenterNet: Keypoint Triplets for Object Detection
Kaiwen Duan, Song Bai, Lingxi Xie +3
cs.CVarXiv:1904.08189v32019CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning
Pranav Rajpurkar, Jeremy Irvin, Kaylie Zhu +9
cs.CVcs.LGstat.MLarXiv:1711.05225v32017Occupancy Networks: Learning 3D Reconstruction in Function Space
Lars Mescheder, Michael Oechsle, Michael Niemeyer +2
cs.CVarXiv:1812.03828v22018MMDetection: Open MMLab Detection Toolbox and Benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang +22
cs.CVcs.LGeess.IVarXiv:1906.07155v12019Unsupervised Deep Embedding for Clustering Analysis
Junyuan Xie, Ross Girshick, Ali Farhadi
cs.LGcs.CVarXiv:1511.06335v22015Object Detection in 20 Years: A Survey
Zhengxia Zou, Keyan Chen, Zhenwei Shi +2
cs.CVarXiv:1905.05055v32019Adversarial Machine Learning at Scale
Alexey Kurakin, Ian Goodfellow, Samy Bengio
cs.CVcs.CRcs.LGarXiv:1611.01236v22016Accelerating the Super-Resolution Convolutional Neural Network
Chao Dong, Chen Change Loy, Xiaoou Tang
cs.CVarXiv:1608.00367v12016Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders
Ze Zhang, Yang Zhang
cs.CVcs.LGarXiv:2608.14717v12026MnasNet: Platform-Aware Neural Architecture Search for Mobile
Mingxing Tan, Bo Chen, Ruoming Pang +4
cs.CVcs.LGarXiv:1807.11626v32018PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models
Siddharth Patel
cs.CVcs.AIarXiv:2608.14741v12026