Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,901 to 9,960 of 18,819
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Wentao Yuan, Jiafei Duan, Valts Blukis +5
cs.ROcs.AIcs.CVarXiv:2406.10721v12024ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT
Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol +1
cs.CVcs.AIarXiv:2608.28455v12026Constrained-CNN losses for weakly supervised segmentation
Hoel Kervadec, Jose Dolz, Meng Tang +3
cs.CVarXiv:1805.04628v22018DeepISP: Towards Learning an End-to-End Image Processing Pipeline
Eli Schwartz, Raja Giryes, Alex M. Bronstein
eess.IVcs.CVarXiv:1801.06724v22018ArcFace: Additive Angular Margin Loss for Deep Face Recognition
Jiankang Deng, Jia Guo, Jing Yang +3
cs.CVarXiv:1801.07698v42018NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks
Yandong Li, Lijun Li, Liqiang Wang +2
cs.LGcs.CRcs.CVarXiv:1905.00441v32019Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging
Eric L. Wisotzky, Jost Triller, Simon W. Härtl +3
cs.CVcs.AIarXiv:2608.28341v12026DKM: Dense Kernelized Feature Matching for Geometry Estimation
Johan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck +1
cs.CVcs.LGarXiv:2202.00667v32022Query2Label: A Simple Transformer Way to Multi-Label Classification
Shilong Liu, Lei Zhang, Xiao Yang +2
cs.CVarXiv:2107.10834v12021Unrestricted Facial Geometry Reconstruction Using Image-to-Image Translation
Matan Sela, Elad Richardson, Ron Kimmel
cs.CVarXiv:1703.10131v22017Token Merging for Fast Stable Diffusion
Daniel Bolya, Judy Hoffman
cs.CVarXiv:2303.17604v12023Machine Learning Techniques for Biomedical Image Segmentation: An Overview of Technical Aspects and Introduction to State-of-Art Applications
Hyunseok Seo, Masoud Badiei Khuzani, Varun Vasudevan +5
eess.IVcs.CVcs.LGarXiv:1911.02521v12019Physically Grounded Vision-Language Models for Robotic Manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia +5
cs.ROcs.AIcs.CVarXiv:2309.02561v42023Learning Semantic Segmentation from Synthetic Data: A Geometrically Guided Input-Output Adaptation Approach
Yuhua Chen, Wen Li, Xiaoran Chen +1
cs.CVarXiv:1812.05040v22018Multi-level Semantic Feature Augmentation for One-shot Learning
Zitian Chen, Yanwei Fu, Yinda Zhang +3
cs.CVarXiv:1804.05298v42018Short-term traffic flow forecasting with spatial-temporal correlation in a hybrid deep learning framework
Yuankai Wu, Huachun Tan
cs.CVarXiv:1612.01022v12016Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction
Jason Ku, Alex D. Pon, Steven L. Waslander
cs.CVarXiv:1904.01690v12019Learning from All Vehicles
Dian Chen, Philipp Krähenbühl
cs.ROcs.CVcs.LGarXiv:2203.11934v32022TimeLens: Event-based Video Frame Interpolation
Stepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis +4
cs.CVarXiv:2106.07286v12021General Facial Representation Learning in a Visual-Linguistic Manner
Yinglin Zheng, Hao Yang, Ting Zhang +7
cs.CVcs.CLarXiv:2112.03109v32021Multimodal Industrial Anomaly Detection via Hybrid Fusion
Yue Wang, Jinlong Peng, Jiangning Zhang +3
cs.CVarXiv:2303.00601v22023A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
Tadej Tomanič, Alice Baudhuin, Jan Sotošek +4
cs.CVcs.AIarXiv:2608.28247v12026Intriguing Findings of Frequency Selection for Image Deblurring
Xintian Mao, Yiming Liu, Fengze Liu +3
cs.CVarXiv:2111.11745v22021Automatic segmentation of the spinal cord and intramedullary multiple sclerosis lesions with convolutional neural networks
Charley Gros, Benjamin De Leener, Atef Badji +49
cs.CVarXiv:1805.06349v22018Real-time Joint Tracking of a Hand Manipulating an Object from RGB-D Input
Srinath Sridhar, Franziska Mueller, Michael Zollhöfer +3
cs.CVarXiv:1610.04889v12016ATISS: Autoregressive Transformers for Indoor Scene Synthesis
Despoina Paschalidou, Amlan Kar, Maria Shugrina +3
cs.CVarXiv:2110.03675v12021$ N^4 $-Fields: Neural Network Nearest Neighbor Fields for Image Transforms
Yaroslav Ganin, Victor Lempitsky
cs.CVarXiv:1406.6558v22014OneLLM: One Framework to Align All Modalities with Language
Jiaming Han, Kaixiong Gong, Yiyuan Zhang +6
cs.CVcs.AIcs.CLarXiv:2312.03700v22023Dual Attention Suppression Attack: Generate Adversarial Camouflage in Physical World
Jiakai Wang, Aishan Liu, Zixin Yin +3
cs.CVarXiv:2103.01050v12021End-to-end optimization of nonlinear transform codes for perceptual quality
Johannes Ballé, Valero Laparra, Eero P. Simoncelli
cs.ITcs.CVarXiv:1607.05006v22016Ranked List Loss for Deep Metric Learning
Xinshao Wang, Yang Hua, Elyor Kodirov +1
cs.CVarXiv:1903.03238v82019MVImgNet: A Large-scale Dataset of Multi-view Images
Xianggang Yu, Mutian Xu, Yidan Zhang +10
cs.CVarXiv:2303.06042v12023Few-Shot Segmentation via Cycle-Consistent Transformer
Gengwei Zhang, Guoliang Kang, Yi Yang +1
cs.CVarXiv:2106.02320v42021Adversarially Robust Distillation
Micah Goldblum, Liam Fowl, Soheil Feizi +1
cs.LGcs.CVstat.MLarXiv:1905.09747v22019DuDoNet: Dual Domain Network for CT Metal Artifact Reduction
Wei-An Lin, Haofu Liao, Cheng Peng +5
eess.IVcs.CVarXiv:1907.00273v12019Effectively Unbiased FID and Inception Score and where to find them
Min Jin Chong, David Forsyth
cs.CVcs.LGarXiv:1911.07023v32019Improve Unsupervised Domain Adaptation with Mixup Training
Shen Yan, Huan Song, Nanxiang Li +2
stat.MLcs.CVcs.LGarXiv:2001.00677v12020Adaptive Token Sampling For Efficient Vision Transformers
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari +5
cs.CVarXiv:2111.15667v32021Continental-Scale Building Detection from High Resolution Satellite Imagery
Wojciech Sirko, Sergii Kashubin, Marvin Ritter +7
cs.CVarXiv:2107.12283v22021Contrastive Embedding for Generalized Zero-Shot Learning
Zongyan Han, Zhenyong Fu, Shuo Chen +1
cs.CVcs.AIarXiv:2103.16173v12021Hierarchical Discrete Distribution Decomposition for Match Density Estimation
Zhichao Yin, Trevor Darrell, Fisher Yu
cs.CVarXiv:1812.06264v32018OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
Wenzhao Zheng, Weiliang Chen, Yuanhui Huang +3
cs.CVcs.AIcs.LGarXiv:2311.16038v12023Recurrent Neural Networks for Driver Activity Anticipation via Sensory-Fusion Architecture
Ashesh Jain, Avi Singh, Hema S Koppula +2
cs.CVcs.AIcs.ROarXiv:1509.05016v12015Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Kunchang Li, Yali Wang, Yizhuo Li +4
cs.CVarXiv:2303.16058v22023SoftMatch: Addressing the Quantity-Quality Trade-off in Semi-supervised Learning
Hao Chen, Ran Tao, Yue Fan +6
cs.LGcs.AIcs.CVarXiv:2301.10921v22023Cross-Batch Memory for Embedding Learning
Xun Wang, Haozhi Zhang, Weilin Huang +1
cs.LGcs.CVarXiv:1912.06798v32019CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs
Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy
eess.IVcs.AIcs.CVarXiv:2608.28137v12026QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Xiuying Wei, Ruihao Gong, Yuhang Li +2
cs.CVcs.AIarXiv:2203.05740v22022Skeleton Aware Multi-modal Sign Language Recognition
Songyao Jiang, Bin Sun, Lichen Wang +3
cs.CVarXiv:2103.08833v52021Style Aligned Image Generation via Shared Attention
Amir Hertz, Andrey Voynov, Shlomi Fruchter +1
cs.CVcs.GRcs.LGarXiv:2312.02133v22023BIRNet: Brain Image Registration Using Dual-Supervised Fully Convolutional Networks
Jingfan Fan, Xiaohuan Cao, Pew-Thian Yap +1
cs.CVarXiv:1802.04692v12018Car Detection using Unmanned Aerial Vehicles: Comparison between Faster R-CNN and YOLOv3
Bilel Benjdira, Taha Khursheed, Anis Koubaa +2
cs.ROcs.CVcs.LGarXiv:1812.10968v12018Transferable Adversarial Attacks for Image and Video Object Detection
Xingxing Wei, Siyuan Liang, Ning Chen +1
cs.CVarXiv:1811.12641v52018Self-supervised learning methods and applications in medical imaging analysis: A survey
Saeed Shurrab, Rehab Duwairi
eess.IVcs.CVcs.LGarXiv:2109.08685v32021Extremely Simple Activation Shaping for Out-of-Distribution Detection
Andrija Djurisic, Nebojsa Bozanic, Arjun Ashok +1
cs.LGcs.CVarXiv:2209.09858v22022DABNet: Depth-wise Asymmetric Bottleneck for Real-time Semantic Segmentation
Gen Li, Inyoung Yun, Jonghyun Kim +1
cs.CVarXiv:1907.11357v22019Dive into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition
Jiahui She, Yibo Hu, Hailin Shi +3
cs.CVarXiv:2104.00232v12021Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations
Naren Akash, Neeraja Ramanan
eess.IVcs.AIcs.CVarXiv:2608.28092v12026VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
Ruijie Su, Lingxiao Yang, Xiaohua Xie +1
cs.CVcs.AIarXiv:2608.28069v12026Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
Kairong Yu, Zixin Zhu, Le Yu +1
cs.CVcs.AIarXiv:2608.28058v12026