Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

13,321 to 13,380 of 18,970

  1. BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

    Chenyu Yang, Yuntao Chen, Hao Tian +9

    cs.CVarXiv:2211.10439v12022
  2. Semi-Supervised Semantic Segmentation with High- and Low-level Consistency

    Sudhanshu Mittal, Maxim Tatarchenko, Thomas Brox

    cs.CVarXiv:1908.05724v12019
  3. QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

    Chenhongyi Yang, Zehao Huang, Naiyan Wang

    cs.CVarXiv:2103.09136v22021
  4. Denoising Diffusion Models for Plug-and-Play Image Restoration

    Yuanzhi Zhu, Kai Zhang, Jingyun Liang +4

    cs.CVeess.IVarXiv:2305.08995v12023
  5. Robotic Grasp Detection using Deep Convolutional Neural Networks

    Sulabh Kumra, Christopher Kanan

    cs.ROcs.CVarXiv:1611.08036v42016
  6. A Neural Approach to Blind Motion Deblurring

    Ayan Chakrabarti

    cs.CVarXiv:1603.04771v22016
  7. Text Line Segmentation of Historical Documents: a Survey

    Laurence Likforman-Sulem, Abderrazak Zahour, Bruno Taconet

    cs.CVarXiv:0704.1267v12007
  8. Low-bit Quantization of Neural Networks for Efficient Inference

    Yoni Choukroun, Eli Kravchik, Fan Yang +1

    cs.LGcs.CVstat.MLarXiv:1902.06822v22019
  9. Multispectral Deep Neural Networks for Pedestrian Detection

    Jingjing Liu, Shaoting Zhang, Shu Wang +1

    cs.CVarXiv:1611.02644v12016
  10. MoMask: Generative Masked Modeling of 3D Human Motions

    Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed +2

    cs.CVarXiv:2312.00063v12023
  11. Salient Object Detection via Integrity Learning

    Mingchen Zhuge, Deng-Ping Fan, Nian Liu +3

    cs.CVarXiv:2101.07663v72021
  12. A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers

    Tianyun Zhang, Shaokai Ye, Kaiqi Zhang +4

    cs.NEcs.CVcs.LGarXiv:1804.03294v32018
  13. Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning

    Yinbo Chen, Zhuang Liu, Huijuan Xu +2

    cs.CVcs.LGarXiv:2003.04390v42020
  14. Fast Dynamic Radiance Fields with Time-Aware Neural Voxels

    Jiemin Fang, Taoran Yi, Xinggang Wang +5

    cs.CVcs.GRarXiv:2205.15285v22022
  15. ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

    Lin Chen, Xilin Wei, Jinsong Li +12

    cs.CVarXiv:2406.04325v12024
  16. Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos

    Serena Yeung, Olga Russakovsky, Ning Jin +3

    cs.CVarXiv:1507.05738v32015
  17. HDMapNet: An Online HD Map Construction and Evaluation Framework

    Qi Li, Yue Wang, Yilun Wang +1

    cs.CVcs.AIarXiv:2107.06307v42021
  18. Learning Not to Learn: Training Deep Neural Networks with Biased Data

    Byungju Kim, Hyunwoo Kim, Kyungsu Kim +2

    cs.CVarXiv:1812.10352v22018
  19. A scoping review of transfer learning research on medical image analysis using ImageNet

    Mohammad Amin Morid, Alireza Borjali, Guilherme Del Fiol

    eess.IVcs.CVcs.LGarXiv:2004.13175v52020
  20. Range Loss for Deep Face Recognition with Long-tail

    Xiao Zhang, Zhiyuan Fang, Yandong Wen +2

    cs.CVarXiv:1611.08976v12016
  21. Side Adapter Network for Open-Vocabulary Semantic Segmentation

    Mengde Xu, Zheng Zhang, Fangyun Wei +2

    cs.CVcs.AIarXiv:2302.12242v22023
  22. Efficient Decision-based Black-box Adversarial Attacks on Face Recognition

    Yinpeng Dong, Hang Su, Baoyuan Wu +4

    cs.CVcs.CRcs.LGarXiv:1904.04433v12019
  23. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics

    Tianyi Xie, Zeshun Zong, Yuxing Qiu +4

    cs.GRcs.AIcs.CVarXiv:2311.12198v32023
  24. Causality matters in medical imaging

    Daniel C. Castro, Ian Walker, Ben Glocker

    eess.IVcs.AIcs.CVarXiv:1912.08142v12019
  25. Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds

    Qi Chen, Sihai Tang, Qing Yang +1

    cs.CVarXiv:1905.05265v12019
  26. Learn To Pay Attention

    Saumya Jetley, Nicholas A. Lord, Namhoon Lee +1

    cs.CVcs.AIarXiv:1804.02391v22018
  27. Focal Modulation Networks

    Jianwei Yang, Chunyuan Li, Xiyang Dai +2

    cs.CVcs.AIcs.LGarXiv:2203.11926v32022
  28. A Recurrent Vision-and-Language BERT for Navigation

    Yicong Hong, Qi Wu, Yuankai Qi +2

    cs.CVarXiv:2011.13922v22020
  29. Wide-Area Image Geolocalization with Aerial Reference Imagery

    Scott Workman, Richard Souvenir, Nathan Jacobs

    cs.CVarXiv:1510.03743v12015
  30. Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

    Zhicheng Huang, Zhaoyang Zeng, Bei Liu +2

    cs.CVcs.CLcs.LGarXiv:2004.00849v22020
  31. PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

    Xiaomeng Fan, Yueran Liu, Shengyu Zhou +6

    cs.CVcs.AIcs.IRarXiv:2608.26658v12026
  32. ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space Model

    Hongruixuan Chen, Jian Song, Chengxi Han +2

    eess.IVcs.AIcs.CVarXiv:2404.03425v72024
  33. Co-Scale Conv-Attentional Image Transformers

    Weijian Xu, Yifan Xu, Tyler Chang +1

    cs.CVcs.LGcs.NEarXiv:2104.06399v22021
  34. Generative OpenMax for Multi-Class Open Set Classification

    ZongYuan Ge, Sergey Demyanov, Zetao Chen +1

    cs.CVarXiv:1707.07418v12017
  35. DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

    Ying Fan, Olivia Watkins, Yuqing Du +7

    cs.LGcs.CVarXiv:2305.16381v32023
  36. Hierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss

    Lele Chen, Ross K. Maddox, Zhiyao Duan +1

    cs.CVarXiv:1905.03820v12019
  37. Beyond Sharing Weights for Deep Domain Adaptation

    Artem Rozantsev, Mathieu Salzmann, Pascal Fua

    cs.CVarXiv:1603.06432v22016
  38. Recalibrating Fully Convolutional Networks with Spatial and Channel 'Squeeze & Excitation' Blocks

    Abhijit Guha Roy, Nassir Navab, Christian Wachinger

    cs.CVarXiv:1808.08127v12018
  39. nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation

    Fabian Isensee, Tassilo Wald, Constantin Ulrich +4

    cs.CVarXiv:2404.09556v22024
  40. Pose-Normalized Image Generation for Person Re-identification

    Xuelin Qian, Yanwei Fu, Tao Xiang +5

    cs.CVcs.AIcs.MMarXiv:1712.02225v62017
  41. Masked Face Recognition Dataset and Application

    Zhongyuan Wang, Guangcheng Wang, Baojin Huang +11

    cs.CVarXiv:2003.09093v22020
  42. Triple Generative Adversarial Nets

    Chongxuan Li, Kun Xu, Jun Zhu +1

    cs.LGcs.CVarXiv:1703.02291v42017
  43. MixConv: Mixed Depthwise Convolutional Kernels

    Mingxing Tan, Quoc V. Le

    cs.CVcs.LGarXiv:1907.09595v32019
  44. Neural Photo Editing with Introspective Adversarial Networks

    Andrew Brock, Theodore Lim, J. M. Ritchie +1

    cs.LGcs.CVcs.NEarXiv:1609.07093v32016
  45. Image Matching across Wide Baselines: From Paper to Practice

    Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk +4

    cs.CVarXiv:2003.01587v52020
  46. YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss

    Debapriya Maji, Soyeb Nagori, Manu Mathew +1

    cs.CVcs.AIcs.LGarXiv:2204.06806v12022
  47. Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal Mixup

    Jang-Hyun Kim, Wonho Choo, Hyun Oh Song

    cs.LGcs.AIcs.CVarXiv:2009.06962v22020
  48. UNITER: UNiversal Image-TExt Representation Learning

    Yen-Chun Chen, Linjie Li, Licheng Yu +5

    cs.CVcs.CLcs.LGarXiv:1909.11740v32019
  49. Dimensionality-Driven Learning with Noisy Labels

    Xingjun Ma, Yisen Wang, Michael E. Houle +5

    cs.CVcs.LGstat.MLarXiv:1806.02612v22018
  50. POI: Multiple Object Tracking with High Performance Detection and Appearance Feature

    Fengwei Yu, Wenbo Li, Quanquan Li +3

    cs.CVarXiv:1610.06136v12016
  51. SpinQuant: LLM quantization with learned rotations

    Zechun Liu, Changsheng Zhao, Igor Fedorov +6

    cs.LGcs.AIcs.CLarXiv:2405.16406v42024
  52. Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data

    Xiangyu Yue, Yang Zhang, Sicheng Zhao +3

    cs.CVarXiv:1909.00889v22019
  53. ESPNetv2: A Light-weight, Power Efficient, and General Purpose Convolutional Neural Network

    Sachin Mehta, Mohammad Rastegari, Linda Shapiro +1

    cs.CVarXiv:1811.11431v32018
  54. Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval

    Adam W. Harley, Alex Ufkes, Konstantinos G. Derpanis

    cs.CVcs.IRcs.LGarXiv:1502.07058v12015
  55. Message Passing Neural PDE Solvers

    Johannes Brandstetter, Daniel Worrall, Max Welling

    cs.LGcs.CVmath.NAarXiv:2202.03376v32022
  56. Procedura: Agentic 3D Modeling with Procedural Control

    Youtian Lin, Yikang Yang, Zhanpeng Hu +5

    cs.CVcs.GRarXiv:2608.26238v12026
  57. Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

    Chenyang Wu, Fuchen Long, Binyuan Huang +4

    cs.CVcs.MMarXiv:2608.26809v12026
  58. VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance

    Katherine Crowson, Stella Biderman, Daniel Kornis +4

    cs.CVarXiv:2204.08583v22022
  59. AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

    Qihang Zhou, Guansong Pang, Yu Tian +2

    cs.CVarXiv:2310.18961v122023
  60. GameWAM: A World Action Model for Video Games

    Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1

    cs.AIcs.CVcs.LGarXiv:2608.26200v12026