Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,901 to 9,960 of 18,819

  1. RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

    Wentao Yuan, Jiafei Duan, Valts Blukis +5

    cs.ROcs.AIcs.CVarXiv:2406.10721v12024
  2. ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

    Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol +1

    cs.CVcs.AIarXiv:2608.28455v12026
  3. Constrained-CNN losses for weakly supervised segmentation

    Hoel Kervadec, Jose Dolz, Meng Tang +3

    cs.CVarXiv:1805.04628v22018
  4. DeepISP: Towards Learning an End-to-End Image Processing Pipeline

    Eli Schwartz, Raja Giryes, Alex M. Bronstein

    eess.IVcs.CVarXiv:1801.06724v22018
  5. ArcFace: Additive Angular Margin Loss for Deep Face Recognition

    Jiankang Deng, Jia Guo, Jing Yang +3

    cs.CVarXiv:1801.07698v42018
  6. NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks

    Yandong Li, Lijun Li, Liqiang Wang +2

    cs.LGcs.CRcs.CVarXiv:1905.00441v32019
  7. Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging

    Eric L. Wisotzky, Jost Triller, Simon W. Härtl +3

    cs.CVcs.AIarXiv:2608.28341v12026
  8. DKM: Dense Kernelized Feature Matching for Geometry Estimation

    Johan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck +1

    cs.CVcs.LGarXiv:2202.00667v32022
  9. Query2Label: A Simple Transformer Way to Multi-Label Classification

    Shilong Liu, Lei Zhang, Xiao Yang +2

    cs.CVarXiv:2107.10834v12021
  10. Unrestricted Facial Geometry Reconstruction Using Image-to-Image Translation

    Matan Sela, Elad Richardson, Ron Kimmel

    cs.CVarXiv:1703.10131v22017
  11. Token Merging for Fast Stable Diffusion

    Daniel Bolya, Judy Hoffman

    cs.CVarXiv:2303.17604v12023
  12. Machine Learning Techniques for Biomedical Image Segmentation: An Overview of Technical Aspects and Introduction to State-of-Art Applications

    Hyunseok Seo, Masoud Badiei Khuzani, Varun Vasudevan +5

    eess.IVcs.CVcs.LGarXiv:1911.02521v12019
  13. Physically Grounded Vision-Language Models for Robotic Manipulation

    Jensen Gao, Bidipta Sarkar, Fei Xia +5

    cs.ROcs.AIcs.CVarXiv:2309.02561v42023
  14. Learning Semantic Segmentation from Synthetic Data: A Geometrically Guided Input-Output Adaptation Approach

    Yuhua Chen, Wen Li, Xiaoran Chen +1

    cs.CVarXiv:1812.05040v22018
  15. Multi-level Semantic Feature Augmentation for One-shot Learning

    Zitian Chen, Yanwei Fu, Yinda Zhang +3

    cs.CVarXiv:1804.05298v42018
  16. Short-term traffic flow forecasting with spatial-temporal correlation in a hybrid deep learning framework

    Yuankai Wu, Huachun Tan

    cs.CVarXiv:1612.01022v12016
  17. Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction

    Jason Ku, Alex D. Pon, Steven L. Waslander

    cs.CVarXiv:1904.01690v12019
  18. Learning from All Vehicles

    Dian Chen, Philipp Krähenbühl

    cs.ROcs.CVcs.LGarXiv:2203.11934v32022
  19. TimeLens: Event-based Video Frame Interpolation

    Stepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis +4

    cs.CVarXiv:2106.07286v12021
  20. General Facial Representation Learning in a Visual-Linguistic Manner

    Yinglin Zheng, Hao Yang, Ting Zhang +7

    cs.CVcs.CLarXiv:2112.03109v32021
  21. Multimodal Industrial Anomaly Detection via Hybrid Fusion

    Yue Wang, Jinlong Peng, Jiangning Zhang +3

    cs.CVarXiv:2303.00601v22023
  22. A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation

    Tadej Tomanič, Alice Baudhuin, Jan Sotošek +4

    cs.CVcs.AIarXiv:2608.28247v12026
  23. Intriguing Findings of Frequency Selection for Image Deblurring

    Xintian Mao, Yiming Liu, Fengze Liu +3

    cs.CVarXiv:2111.11745v22021
  24. Automatic segmentation of the spinal cord and intramedullary multiple sclerosis lesions with convolutional neural networks

    Charley Gros, Benjamin De Leener, Atef Badji +49

    cs.CVarXiv:1805.06349v22018
  25. Real-time Joint Tracking of a Hand Manipulating an Object from RGB-D Input

    Srinath Sridhar, Franziska Mueller, Michael Zollhöfer +3

    cs.CVarXiv:1610.04889v12016
  26. ATISS: Autoregressive Transformers for Indoor Scene Synthesis

    Despoina Paschalidou, Amlan Kar, Maria Shugrina +3

    cs.CVarXiv:2110.03675v12021
  27. $ N^4 $-Fields: Neural Network Nearest Neighbor Fields for Image Transforms

    Yaroslav Ganin, Victor Lempitsky

    cs.CVarXiv:1406.6558v22014
  28. OneLLM: One Framework to Align All Modalities with Language

    Jiaming Han, Kaixiong Gong, Yiyuan Zhang +6

    cs.CVcs.AIcs.CLarXiv:2312.03700v22023
  29. Dual Attention Suppression Attack: Generate Adversarial Camouflage in Physical World

    Jiakai Wang, Aishan Liu, Zixin Yin +3

    cs.CVarXiv:2103.01050v12021
  30. End-to-end optimization of nonlinear transform codes for perceptual quality

    Johannes Ballé, Valero Laparra, Eero P. Simoncelli

    cs.ITcs.CVarXiv:1607.05006v22016
  31. Ranked List Loss for Deep Metric Learning

    Xinshao Wang, Yang Hua, Elyor Kodirov +1

    cs.CVarXiv:1903.03238v82019
  32. MVImgNet: A Large-scale Dataset of Multi-view Images

    Xianggang Yu, Mutian Xu, Yidan Zhang +10

    cs.CVarXiv:2303.06042v12023
  33. Few-Shot Segmentation via Cycle-Consistent Transformer

    Gengwei Zhang, Guoliang Kang, Yi Yang +1

    cs.CVarXiv:2106.02320v42021
  34. Adversarially Robust Distillation

    Micah Goldblum, Liam Fowl, Soheil Feizi +1

    cs.LGcs.CVstat.MLarXiv:1905.09747v22019
  35. DuDoNet: Dual Domain Network for CT Metal Artifact Reduction

    Wei-An Lin, Haofu Liao, Cheng Peng +5

    eess.IVcs.CVarXiv:1907.00273v12019
  36. Effectively Unbiased FID and Inception Score and where to find them

    Min Jin Chong, David Forsyth

    cs.CVcs.LGarXiv:1911.07023v32019
  37. Improve Unsupervised Domain Adaptation with Mixup Training

    Shen Yan, Huan Song, Nanxiang Li +2

    stat.MLcs.CVcs.LGarXiv:2001.00677v12020
  38. Adaptive Token Sampling For Efficient Vision Transformers

    Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari +5

    cs.CVarXiv:2111.15667v32021
  39. Continental-Scale Building Detection from High Resolution Satellite Imagery

    Wojciech Sirko, Sergii Kashubin, Marvin Ritter +7

    cs.CVarXiv:2107.12283v22021
  40. Contrastive Embedding for Generalized Zero-Shot Learning

    Zongyan Han, Zhenyong Fu, Shuo Chen +1

    cs.CVcs.AIarXiv:2103.16173v12021
  41. Hierarchical Discrete Distribution Decomposition for Match Density Estimation

    Zhichao Yin, Trevor Darrell, Fisher Yu

    cs.CVarXiv:1812.06264v32018
  42. OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

    Wenzhao Zheng, Weiliang Chen, Yuanhui Huang +3

    cs.CVcs.AIcs.LGarXiv:2311.16038v12023
  43. Recurrent Neural Networks for Driver Activity Anticipation via Sensory-Fusion Architecture

    Ashesh Jain, Avi Singh, Hema S Koppula +2

    cs.CVcs.AIcs.ROarXiv:1509.05016v12015
  44. Unmasked Teacher: Towards Training-Efficient Video Foundation Models

    Kunchang Li, Yali Wang, Yizhuo Li +4

    cs.CVarXiv:2303.16058v22023
  45. SoftMatch: Addressing the Quantity-Quality Trade-off in Semi-supervised Learning

    Hao Chen, Ran Tao, Yue Fan +6

    cs.LGcs.AIcs.CVarXiv:2301.10921v22023
  46. Cross-Batch Memory for Embedding Learning

    Xun Wang, Haozhi Zhang, Weilin Huang +1

    cs.LGcs.CVarXiv:1912.06798v32019
  47. CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs

    Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy

    eess.IVcs.AIcs.CVarXiv:2608.28137v12026
  48. QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization

    Xiuying Wei, Ruihao Gong, Yuhang Li +2

    cs.CVcs.AIarXiv:2203.05740v22022
  49. Skeleton Aware Multi-modal Sign Language Recognition

    Songyao Jiang, Bin Sun, Lichen Wang +3

    cs.CVarXiv:2103.08833v52021
  50. Style Aligned Image Generation via Shared Attention

    Amir Hertz, Andrey Voynov, Shlomi Fruchter +1

    cs.CVcs.GRcs.LGarXiv:2312.02133v22023
  51. BIRNet: Brain Image Registration Using Dual-Supervised Fully Convolutional Networks

    Jingfan Fan, Xiaohuan Cao, Pew-Thian Yap +1

    cs.CVarXiv:1802.04692v12018
  52. Car Detection using Unmanned Aerial Vehicles: Comparison between Faster R-CNN and YOLOv3

    Bilel Benjdira, Taha Khursheed, Anis Koubaa +2

    cs.ROcs.CVcs.LGarXiv:1812.10968v12018
  53. Transferable Adversarial Attacks for Image and Video Object Detection

    Xingxing Wei, Siyuan Liang, Ning Chen +1

    cs.CVarXiv:1811.12641v52018
  54. Self-supervised learning methods and applications in medical imaging analysis: A survey

    Saeed Shurrab, Rehab Duwairi

    eess.IVcs.CVcs.LGarXiv:2109.08685v32021
  55. Extremely Simple Activation Shaping for Out-of-Distribution Detection

    Andrija Djurisic, Nebojsa Bozanic, Arjun Ashok +1

    cs.LGcs.CVarXiv:2209.09858v22022
  56. DABNet: Depth-wise Asymmetric Bottleneck for Real-time Semantic Segmentation

    Gen Li, Inyoung Yun, Jonghyun Kim +1

    cs.CVarXiv:1907.11357v22019
  57. Dive into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition

    Jiahui She, Yibo Hu, Hailin Shi +3

    cs.CVarXiv:2104.00232v12021
  58. Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

    Naren Akash, Neeraja Ramanan

    eess.IVcs.AIcs.CVarXiv:2608.28092v12026
  59. VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians

    Ruijie Su, Lingxiao Yang, Xiaohua Xie +1

    cs.CVcs.AIarXiv:2608.28069v12026
  60. Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

    Kairong Yu, Zixin Zhu, Le Yu +1

    cs.CVcs.AIarXiv:2608.28058v12026