Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,181 to 12,240 of 18,841

  1. Hands Deep in Deep Learning for Hand Pose Estimation

    Markus Oberweger, Paul Wohlhart, Vincent Lepetit

    cs.CVarXiv:1502.06807v22015
  2. Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models

    Jiuxiang Gu, Jianfei Cai, Shafiq Joty +2

    cs.CVarXiv:1711.06420v22017
  3. Deep Direct Regression for Multi-Oriented Scene Text Detection

    Wenhao He, Xu-Yao Zhang, Fei Yin +1

    cs.CVarXiv:1703.08289v12017
  4. Implicit Diffusion Models for Continuous Super-Resolution

    Sicheng Gao, Xuhui Liu, Bohan Zeng +6

    cs.CVarXiv:2303.16491v22023
  5. Exploring the Landscape of Spatial Robustness

    Logan Engstrom, Brandon Tran, Dimitris Tsipras +2

    cs.LGcs.CVcs.NEarXiv:1712.02779v42017
  6. LEDNet: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation

    Yu Wang, Quan Zhou, Jia Liu +4

    cs.CVarXiv:1905.02423v32019
  7. GridMask Data Augmentation

    Pengguang Chen, Shu Liu, Hengshuang Zhao +2

    cs.CVarXiv:2001.04086v32020
  8. Delta-encoder: an effective sample synthesis method for few-shot object recognition

    Eli Schwartz, Leonid Karlinsky, Joseph Shtok +6

    cs.CVarXiv:1806.04734v32018
  9. Learning the Model Update for Siamese Trackers

    Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer +2

    cs.CVarXiv:1908.00855v22019
  10. 3DFeat-Net: Weakly Supervised Local 3D Features for Point Cloud Registration

    Zi Jian Yew, Gim Hee Lee

    cs.CVarXiv:1807.09413v12018
  11. Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs

    Ruigang Fu, Qingyong Hu, Xiaohu Dong +3

    cs.CVcs.AIcs.LGarXiv:2008.02312v42020
  12. One Million Scenes for Autonomous Driving: ONCE Dataset

    Jiageng Mao, Minzhe Niu, Chenhan Jiang +10

    cs.CVarXiv:2106.11037v32021
  13. Identifying Mislabeled Data using the Area Under the Margin Ranking

    Geoff Pleiss, Tianyi Zhang, Ethan R. Elenberg +1

    cs.LGcs.CVstat.MLarXiv:2001.10528v42020
  14. FaceNet2ExpNet: Regularizing a Deep Face Recognition Net for Expression Recognition

    Hui Ding, Shaohua Kevin Zhou, Rama Chellappa

    cs.CVarXiv:1609.06591v22016
  15. PSCC-Net: Progressive Spatio-Channel Correlation Network for Image Manipulation Detection and Localization

    Xiaohong Liu, Yaojie Liu, Jun Chen +1

    cs.CVarXiv:2103.10596v22021
  16. How much data is needed to train a medical image deep learning system to achieve necessary high accuracy?

    Junghwan Cho, Kyewook Lee, Ellie Shin +2

    cs.LGcs.CVcs.NEarXiv:1511.06348v22015
  17. ReconFusion: 3D Reconstruction with Diffusion Priors

    Rundi Wu, Ben Mildenhall, Philipp Henzler +8

    cs.CVarXiv:2312.02981v12023
  18. Rethinking Transformer-based Set Prediction for Object Detection

    Zhiqing Sun, Shengcao Cao, Yiming Yang +1

    cs.CVcs.LGarXiv:2011.10881v22020
  19. RS-Mamba for Large Remote Sensing Image Dense Prediction

    Sijie Zhao, Hao Chen, Xueliang Zhang +3

    cs.CVarXiv:2404.02668v22024
  20. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

    Frederik Berenz

    cs.CVcs.AIarXiv:2608.27367v12026
  21. MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking

    Patrick Dendorfer, Aljoša Ošep, Anton Milan +5

    cs.CVarXiv:2010.07548v22020
  22. BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs

    Valentin Bazarevsky, Yury Kartynnik, Andrey Vakunov +2

    cs.CVarXiv:1907.05047v22019
  23. 3D Shape Segmentation with Projective Convolutional Networks

    Evangelos Kalogerakis, Melinos Averkiou, Subhransu Maji +1

    cs.CVcs.GRarXiv:1612.02808v32016
  24. Generalized Decoding for Pixel, Image, and Language

    Xueyan Zou, Zi-Yi Dou, Jianwei Yang +11

    cs.CVcs.CLarXiv:2212.11270v12022
  25. Fast Optical Flow using Dense Inverse Search

    Till Kroeger, Radu Timofte, Dengxin Dai +1

    cs.CVcs.ROarXiv:1603.03590v12016
  26. Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks

    Pichao Wang, Zhaoyang Li, Yonghong Hou +1

    cs.CVarXiv:1611.02447v22016
  27. Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

    Javier Marin, Aritro Biswas, Ferda Ofli +5

    cs.CVarXiv:1810.06553v22018
  28. Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

    Yan Zeng, Xinsong Zhang, Hang Li

    cs.CLcs.CVarXiv:2111.08276v32021
  29. Reducing Network Agnostophobia

    Akshay Raj Dhamija, Manuel Günther, Terrance E. Boult

    cs.CVarXiv:1811.04110v22018
  30. Nonlinear Transform Source-Channel Coding for Semantic Communications

    Jincheng Dai, Sixian Wang, Kailin Tan +4

    cs.ITcs.CVcs.LGarXiv:2112.10961v32021
  31. Neural Topological SLAM for Visual Navigation

    Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta +1

    cs.CVcs.AIcs.LGarXiv:2005.12256v22020
  32. Free View Synthesis

    Gernot Riegler, Vladlen Koltun

    cs.CVarXiv:2008.05511v12020
  33. Multi-view 3D Models from Single Images with a Convolutional Network

    Maxim Tatarchenko, Alexey Dosovitskiy, Thomas Brox

    cs.CVarXiv:1511.06702v22015
  34. Autoencoders for Unsupervised Anomaly Segmentation in Brain MR Images: A Comparative Study

    Christoph Baur, Stefan Denner, Benedikt Wiestler +2

    eess.IVcs.CVcs.LGarXiv:2004.03271v22020
  35. IONet: Learning to Cure the Curse of Drift in Inertial Odometry

    Changhao Chen, Xiaoxuan Lu, Andrew Markham +1

    cs.ROcs.AIcs.CVarXiv:1802.02209v12018
  36. Efficient parametrization of multi-domain deep neural networks

    Sylvestre-Alvise Rebuffi, Hakan Bilen, Andrea Vedaldi

    cs.CVstat.MLarXiv:1803.10082v12018
  37. MIC: Masked Image Consistency for Context-Enhanced Domain Adaptation

    Lukas Hoyer, Dengxin Dai, Haoran Wang +1

    cs.CVarXiv:2212.01322v22022
  38. Anatomy-Guided Foundation Model Adaptation with Within-Case Prototype Supervision for Standard Plane Detection in Fetal Ultrasound Blind Sweeps

    Yuzhe Zhao

    cs.CVeess.IVarXiv:2608.27051v12026
  39. CheXclusion: Fairness gaps in deep chest X-ray classifiers

    Laleh Seyyed-Kalantari, Guanxiong Liu, Matthew McDermott +2

    cs.CVcs.AIcs.LGarXiv:2003.00827v22020
  40. How good is my GAN?

    Konstantin Shmelkov, Cordelia Schmid, Karteek Alahari

    cs.CVcs.LGarXiv:1807.09499v12018
  41. Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion

    Xu Yan, Jiantao Gao, Jie Li +4

    cs.CVarXiv:2012.03762v12020
  42. Cross-Layer Distillation with Semantic Calibration

    Defang Chen, Jian-Ping Mei, Yuan Zhang +3

    cs.CVcs.AIcs.LGarXiv:2012.03236v22020
  43. LaserNet: An Efficient Probabilistic 3D Object Detector for Autonomous Driving

    Gregory P. Meyer, Ankit Laddha, Eric Kee +2

    cs.CVcs.LGcs.ROarXiv:1903.08701v12019
  44. Multi-modal Cycle-consistent Generalized Zero-Shot Learning

    Rafael Felix, B. G. Vijay Kumar, Ian Reid +1

    cs.CVarXiv:1808.00136v22018
  45. TempJail: Temporal Jailbreak Attacks against Image-to-Video Generation Models

    Qi Lu, Zehui Guo, David Yuanda Gan +4

    cs.CVcs.MMarXiv:2608.26971v12026
  46. Filtered Channel Features for Pedestrian Detection

    Shanshan Zhang, Rodrigo Benenson, Bernt Schiele

    cs.CVarXiv:1501.05759v12015
  47. Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models

    Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney +1

    cs.CVcs.CLcs.IRarXiv:2108.04024v12021
  48. Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification

    Feng Zhu, Hongsheng Li, Wanli Ouyang +2

    cs.CVarXiv:1702.05891v22017
  49. Object Detection from Video Tubelets with Convolutional Neural Networks

    Kai Kang, Wanli Ouyang, Hongsheng Li +1

    cs.CVarXiv:1604.04053v12016
  50. On Evaluating Adversarial Robustness of Large Vision-Language Models

    Yunqing Zhao, Tianyu Pang, Chao Du +4

    cs.CVcs.CLcs.CRarXiv:2305.16934v22023
  51. Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature

    Babak Saleh, Ahmed Elgammal

    cs.CVcs.IRcs.LGarXiv:1505.00855v12015
  52. TEXTure: Text-Guided Texturing of 3D Shapes

    Elad Richardson, Gal Metzer, Yuval Alaluf +2

    cs.CVcs.GRarXiv:2302.01721v12023
  53. Free Lunch for Few-shot Learning: Distribution Calibration

    Shuo Yang, Lu Liu, Min Xu

    cs.LGcs.CVarXiv:2101.06395v32021
  54. Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

    Shizhe Chen, Yida Zhao, Qin Jin +1

    cs.CVcs.AIarXiv:2003.00392v12020
  55. PENet: Towards Precise and Efficient Image Guided Depth Completion

    Mu Hu, Shuling Wang, Bin Li +3

    cs.CVarXiv:2103.00783v32021
  56. Review of Visual Saliency Detection with Comprehensive Information

    Runmin Cong, Jianjun Lei, Huazhu Fu +3

    cs.CVarXiv:1803.03391v22018
  57. Charting the Right Manifold: Manifold Mixup for Few-shot Learning

    Puneet Mangla, Mayank Singh, Abhishek Sinha +3

    cs.LGcs.CVstat.MLarXiv:1907.12087v42019
  58. Generating 3D Adversarial Point Clouds

    Chong Xiang, Charles R. Qi, Bo Li

    cs.CRcs.CVcs.LGarXiv:1809.07016v42018
  59. One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning

    Tianhe Yu, Chelsea Finn, Annie Xie +4

    cs.LGcs.AIcs.CVarXiv:1802.01557v12018
  60. TransPose: Keypoint Localization via Transformer

    Sen Yang, Zhibin Quan, Mu Nie +1

    cs.CVarXiv:2012.14214v52020