Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,381 to 10,440 of 18,867

  1. Single-Path NAS: Designing Hardware-Efficient ConvNets in less than 4 Hours

    Dimitrios Stamoulis, Ruizhou Ding, Di Wang +4

    cs.LGcs.CVstat.MLarXiv:1904.02877v12019
  2. Deep Learning Predicts Hip Fracture using Confounding Patient and Healthcare Variables

    Marcus A. Badgeley, John R. Zech, Luke Oakden-Rayner +7

    cs.CVarXiv:1811.03695v12018
  3. Multi-graph Fusion for Multi-view Spectral Clustering

    Zhao Kang, Guoxin Shi, Shudong Huang +4

    cs.LGcs.CVstat.MLarXiv:1909.06940v12019
  4. Centralized Feature Pyramid for Object Detection

    Yu Quan, Dong Zhang, Liyan Zhang +1

    cs.CVarXiv:2210.02093v12022
  5. Label Efficient Learning of Transferable Representations across Domains and Tasks

    Zelun Luo, Yuliang Zou, Judy Hoffman +1

    stat.MLcs.CVarXiv:1712.00123v12017
  6. Hyperbolic Deep Neural Networks: A Survey

    Wei Peng, Tuomas Varanka, Abdelrahman Mostafa +2

    cs.LGcs.CVarXiv:2101.04562v32021
  7. You Only Look Yourself: Unsupervised and Untrained Single Image Dehazing Neural Network

    Boyun Li, Yuanbiao Gou, Shuhang Gu +3

    cs.CVarXiv:2006.16829v12020
  8. Unsupervised Detection of Lesions in Brain MRI using constrained adversarial auto-encoders

    Xiaoran Chen, Ender Konukoglu

    cs.CVarXiv:1806.04972v12018
  9. Fast Interactive Object Annotation with Curve-GCN

    Huan Ling, Jun Gao, Amlan Kar +2

    cs.CVcs.LGarXiv:1903.06874v12019
  10. What does CLIP know about a red circle? Visual prompt engineering for VLMs

    Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi

    cs.CVarXiv:2304.06712v22023
  11. M2DGR: A Multi-sensor and Multi-scenario SLAM Dataset for Ground Robots

    Jie Yin, Ang Li, Tao Li +2

    cs.ROcs.CVarXiv:2112.13659v12021
  12. ROSE: A Retinal OCT-Angiography Vessel Segmentation Dataset and New Model

    Yuhui Ma, Huaying Hao, Huazhu Fu +5

    eess.IVcs.CVarXiv:2007.05201v22020
  13. More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity

    Shiwei Liu, Tianlong Chen, Xiaohan Chen +7

    cs.CVarXiv:2207.03620v32022
  14. 3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment

    Ziyu Zhu, Xiaojian Ma, Yixin Chen +3

    cs.CVarXiv:2308.04352v12023
  15. MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

    Reiner Birkl, Diana Wofk, Matthias Müller

    cs.CVarXiv:2307.14460v12023
  16. Monocular Expressive Body Regression through Body-Driven Attention

    Vasileios Choutas, Georgios Pavlakos, Timo Bolkart +2

    cs.CVcs.GRarXiv:2008.09062v12020
  17. Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos

    Yitian Yuan, Lin Ma, Jingwen Wang +2

    cs.CVarXiv:1910.14303v12019
  18. MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

    Yann Labbé, Lucas Manuelli, Arsalan Mousavian +7

    cs.CVcs.ROarXiv:2212.06870v12022
  19. Adaptive Diffusion Priors for Accelerated MRI Reconstruction

    Alper Güngör, Salman UH Dar, Şaban Öztürk +4

    eess.IVcs.CVarXiv:2207.05876v32022
  20. Towards Ghost-free Shadow Removal via Dual Hierarchical Aggregation Network and Shadow Matting GAN

    Xiaodong Cun, Chi-Man Pun, Cheng Shi

    cs.CVarXiv:1911.08718v22019
  21. Robust Graph Learning from Noisy Data

    Zhao Kang, Haiqi Pan, Steven C. H. Hoi +1

    cs.CVcs.AIcs.LGarXiv:1812.06673v12018
  22. DreamBooth3D: Subject-Driven Text-to-3D Generation

    Amit Raj, Srinivas Kaza, Ben Poole +9

    cs.CVcs.AIcs.GRarXiv:2303.13508v22023
  23. Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

    Haoning Wu, Zicheng Zhang, Erli Zhang +8

    cs.CVcs.AIcs.MMarXiv:2309.14181v32023
  24. Prototype Rectification for Few-Shot Learning

    Jinlu Liu, Liang Song, Yongqiang Qin

    cs.CVarXiv:1911.10713v42019
  25. Long-term Human Motion Prediction with Scene Context

    Zhe Cao, Hang Gao, Karttikeya Mangalam +3

    cs.CVarXiv:2007.03672v32020
  26. Emotion Recognition in Speech using Cross-Modal Transfer in the Wild

    Samuel Albanie, Arsha Nagrani, Andrea Vedaldi +1

    cs.CVarXiv:1808.05561v12018
  27. Single Image 3D Interpreter Network

    Jiajun Wu, Tianfan Xue, Joseph J. Lim +4

    cs.CVcs.LGarXiv:1604.08685v22016
  28. DepthSplat: Connecting Gaussian Splatting and Depth

    Haofei Xu, Songyou Peng, Fangjinhua Wang +4

    cs.CVarXiv:2410.13862v32024
  29. Incremental Learning Techniques for Semantic Segmentation

    Umberto Michieli, Pietro Zanuttigh

    cs.CVcs.LGeess.IVarXiv:1907.13372v42019
  30. Towards Vision-Based Deep Reinforcement Learning for Robotic Motion Control

    Fangyi Zhang, Jürgen Leitner, Michael Milford +2

    cs.LGcs.CVcs.ROarXiv:1511.03791v22015
  31. Squeeze-and-Attention Networks for Semantic Segmentation

    Zilong Zhong, Zhong Qiu Lin, Rene Bidart +6

    cs.CVarXiv:1909.03402v42019
  32. SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds

    Yanyu Li, Huan Wang, Qing Jin +6

    cs.CVcs.AIcs.LGarXiv:2306.00980v32023
  33. Polarized Self-Attention: Towards High-quality Pixel-wise Regression

    Huajun Liu, Fuqiang Liu, Xinyi Fan +1

    cs.CVarXiv:2107.00782v22021
  34. A Fast and Accurate Unconstrained Face Detector

    Shengcai Liao, Anil K. Jain, Stan Z. Li

    cs.CVarXiv:1408.1656v32014
  35. Unsupervised Person Re-identification by Soft Multilabel Learning

    Hong-Xing Yu, Wei-Shi Zheng, Ancong Wu +3

    cs.CVarXiv:1903.06325v22019
  36. Stepwise Feature Fusion: Local Guides Global

    Jinfeng Wang, Qiming Huang, Feilong Tang +3

    eess.IVcs.CVarXiv:2203.03635v32022
  37. TIDE: A General Toolbox for Identifying Object Detection Errors

    Daniel Bolya, Sean Foley, James Hays +1

    cs.CVarXiv:2008.08115v22020
  38. Generalized Nonconvex Nonsmooth Low-Rank Minimization

    Canyi Lu, Jinhui Tang, Shuicheng Yan +1

    cs.CVcs.LGstat.MLarXiv:1404.7306v12014
  39. PCRNet: Point Cloud Registration Network using PointNet Encoding

    Vinit Sarode, Xueqian Li, Hunter Goforth +4

    cs.CVarXiv:1908.07906v22019
  40. Peak-Piloted Deep Network for Facial Expression Recognition

    Xiangyun Zhao, Xiaodan Liang, Luoqi Liu +4

    cs.CVarXiv:1607.06997v22016
  41. DA-TransUNet: Integrating Spatial and Channel Dual Attention with Transformer U-Net for Medical Image Segmentation

    Guanqun Sun, Yizhi Pan, Weikun Kong +5

    eess.IVcs.CVcs.GRarXiv:2310.12570v22023
  42. MTU-Net: Multi-level TransUNet for Space-based Infrared Tiny Ship Detection

    Tianhao Wu, Boyang Li, Yihang Luo +6

    cs.CVarXiv:2209.13756v12022
  43. Learn2Reg: comprehensive multi-task medical image registration challenge, dataset and evaluation in the era of deep learning

    Alessa Hering, Lasse Hansen, Tony C. W. Mok +50

    eess.IVcs.CVarXiv:2112.04489v32021
  44. Adversarial Self-Supervised Contrastive Learning

    Minseon Kim, Jihoon Tack, Sung Ju Hwang

    cs.LGcs.CVstat.MLarXiv:2006.07589v22020
  45. Adversarial Latent Autoencoders

    Stanislav Pidhorskyi, Donald Adjeroh, Gianfranco Doretto

    cs.LGcs.CVarXiv:2004.04467v12020
  46. Chasing Sparsity in Vision Transformers: An End-to-End Exploration

    Tianlong Chen, Yu Cheng, Zhe Gan +3

    cs.CVcs.AIarXiv:2106.04533v32021
  47. TripoSR: Fast 3D Object Reconstruction from a Single Image

    Dmitry Tochilkin, David Pankratz, Zexiang Liu +7

    cs.CVarXiv:2403.02151v12024
  48. Deep Compositional Captioning: Describing Novel Object Categories without Paired Training Data

    Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach +3

    cs.CVcs.CLarXiv:1511.05284v22015
  49. Local Learning with Deep and Handcrafted Features for Facial Expression Recognition

    Mariana-Iuliana Georgescu, Radu Tudor Ionescu, Marius Popescu

    cs.CVarXiv:1804.10892v72018
  50. A Survey on Neural Architecture Search

    Martin Wistuba, Ambrish Rawat, Tejaswini Pedapati

    cs.LGcs.CVcs.NEarXiv:1905.01392v22019
  51. Sketch-Guided Text-to-Image Diffusion Models

    Andrey Voynov, Kfir Aberman, Daniel Cohen-Or

    cs.CVcs.GRcs.LGarXiv:2211.13752v12022
  52. Deep Learning for Photoacoustic Tomography from Sparse Data

    Stephan Antholzer, Markus Haltmeier, Johannes Schwab

    cs.CVcs.LGarXiv:1704.04587v32017
  53. Pixel-Perfect Structure-from-Motion with Featuremetric Refinement

    Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson +1

    cs.CVarXiv:2108.08291v12021
  54. Learning Generalisable Omni-Scale Representations for Person Re-Identification

    Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro +1

    cs.CVarXiv:1910.06827v52019
  55. Cross-Image Relational Knowledge Distillation for Semantic Segmentation

    Chuanguang Yang, Helong Zhou, Zhulin An +3

    cs.CVarXiv:2204.06986v22022
  56. Invertible Image Rescaling

    Mingqing Xiao, Shuxin Zheng, Chang Liu +6

    eess.IVcs.CVcs.LGarXiv:2005.05650v12020
  57. Partial FC: Training 10 Million Identities on a Single Machine

    Xiang An, Xuhan Zhu, Yang Xiao +7

    cs.CVcs.DCarXiv:2010.05222v42020
  58. Online and Offline Handwritten Chinese Character Recognition: A Comprehensive Study and New Benchmark

    Xu-Yao Zhang, Yoshua Bengio, Cheng-Lin Liu

    cs.CVarXiv:1606.05763v12016
  59. Island Loss for Learning Discriminative Features in Facial Expression Recognition

    Jie Cai, Zibo Meng, Ahmed Shehab Khan +3

    cs.CVarXiv:1710.03144v32017
  60. Language Models for Image Captioning: The Quirks and What Works

    Jacob Devlin, Hao Cheng, Hao Fang +5

    cs.CLcs.AIcs.CVarXiv:1505.01809v32015