Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
13,321 to 13,380 of 18,970
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision
Chenyu Yang, Yuntao Chen, Hao Tian +9
cs.CVarXiv:2211.10439v12022Semi-Supervised Semantic Segmentation with High- and Low-level Consistency
Sudhanshu Mittal, Maxim Tatarchenko, Thomas Brox
cs.CVarXiv:1908.05724v12019QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
Chenhongyi Yang, Zehao Huang, Naiyan Wang
cs.CVarXiv:2103.09136v22021Denoising Diffusion Models for Plug-and-Play Image Restoration
Yuanzhi Zhu, Kai Zhang, Jingyun Liang +4
cs.CVeess.IVarXiv:2305.08995v12023Robotic Grasp Detection using Deep Convolutional Neural Networks
Sulabh Kumra, Christopher Kanan
cs.ROcs.CVarXiv:1611.08036v42016A Neural Approach to Blind Motion Deblurring
Ayan Chakrabarti
cs.CVarXiv:1603.04771v22016Text Line Segmentation of Historical Documents: a Survey
Laurence Likforman-Sulem, Abderrazak Zahour, Bruno Taconet
cs.CVarXiv:0704.1267v12007Low-bit Quantization of Neural Networks for Efficient Inference
Yoni Choukroun, Eli Kravchik, Fan Yang +1
cs.LGcs.CVstat.MLarXiv:1902.06822v22019Multispectral Deep Neural Networks for Pedestrian Detection
Jingjing Liu, Shaoting Zhang, Shu Wang +1
cs.CVarXiv:1611.02644v12016MoMask: Generative Masked Modeling of 3D Human Motions
Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed +2
cs.CVarXiv:2312.00063v12023Salient Object Detection via Integrity Learning
Mingchen Zhuge, Deng-Ping Fan, Nian Liu +3
cs.CVarXiv:2101.07663v72021A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers
Tianyun Zhang, Shaokai Ye, Kaiqi Zhang +4
cs.NEcs.CVcs.LGarXiv:1804.03294v32018Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning
Yinbo Chen, Zhuang Liu, Huijuan Xu +2
cs.CVcs.LGarXiv:2003.04390v42020Fast Dynamic Radiance Fields with Time-Aware Neural Voxels
Jiemin Fang, Taoran Yi, Xinggang Wang +5
cs.CVcs.GRarXiv:2205.15285v22022ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Lin Chen, Xilin Wei, Jinsong Li +12
cs.CVarXiv:2406.04325v12024Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
Serena Yeung, Olga Russakovsky, Ning Jin +3
cs.CVarXiv:1507.05738v32015HDMapNet: An Online HD Map Construction and Evaluation Framework
Qi Li, Yue Wang, Yilun Wang +1
cs.CVcs.AIarXiv:2107.06307v42021Learning Not to Learn: Training Deep Neural Networks with Biased Data
Byungju Kim, Hyunwoo Kim, Kyungsu Kim +2
cs.CVarXiv:1812.10352v22018A scoping review of transfer learning research on medical image analysis using ImageNet
Mohammad Amin Morid, Alireza Borjali, Guilherme Del Fiol
eess.IVcs.CVcs.LGarXiv:2004.13175v52020Range Loss for Deep Face Recognition with Long-tail
Xiao Zhang, Zhiyuan Fang, Yandong Wen +2
cs.CVarXiv:1611.08976v12016Side Adapter Network for Open-Vocabulary Semantic Segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei +2
cs.CVcs.AIarXiv:2302.12242v22023Efficient Decision-based Black-box Adversarial Attacks on Face Recognition
Yinpeng Dong, Hang Su, Baoyuan Wu +4
cs.CVcs.CRcs.LGarXiv:1904.04433v12019PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
Tianyi Xie, Zeshun Zong, Yuxing Qiu +4
cs.GRcs.AIcs.CVarXiv:2311.12198v32023Causality matters in medical imaging
Daniel C. Castro, Ian Walker, Ben Glocker
eess.IVcs.AIcs.CVarXiv:1912.08142v12019Cooper: Cooperative Perception for Connected Autonomous Vehicles based on 3D Point Clouds
Qi Chen, Sihai Tang, Qing Yang +1
cs.CVarXiv:1905.05265v12019Learn To Pay Attention
Saumya Jetley, Nicholas A. Lord, Namhoon Lee +1
cs.CVcs.AIarXiv:1804.02391v22018Focal Modulation Networks
Jianwei Yang, Chunyuan Li, Xiyang Dai +2
cs.CVcs.AIcs.LGarXiv:2203.11926v32022A Recurrent Vision-and-Language BERT for Navigation
Yicong Hong, Qi Wu, Yuankai Qi +2
cs.CVarXiv:2011.13922v22020Wide-Area Image Geolocalization with Aerial Reference Imagery
Scott Workman, Richard Souvenir, Nathan Jacobs
cs.CVarXiv:1510.03743v12015Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu +2
cs.CVcs.CLcs.LGarXiv:2004.00849v22020PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
Xiaomeng Fan, Yueran Liu, Shengyu Zhou +6
cs.CVcs.AIcs.IRarXiv:2608.26658v12026ChangeMamba: Remote Sensing Change Detection With Spatiotemporal State Space Model
Hongruixuan Chen, Jian Song, Chengxi Han +2
eess.IVcs.AIcs.CVarXiv:2404.03425v72024Co-Scale Conv-Attentional Image Transformers
Weijian Xu, Yifan Xu, Tyler Chang +1
cs.CVcs.LGcs.NEarXiv:2104.06399v22021Generative OpenMax for Multi-Class Open Set Classification
ZongYuan Ge, Sergey Demyanov, Zetao Chen +1
cs.CVarXiv:1707.07418v12017DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Ying Fan, Olivia Watkins, Yuqing Du +7
cs.LGcs.CVarXiv:2305.16381v32023Hierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss
Lele Chen, Ross K. Maddox, Zhiyao Duan +1
cs.CVarXiv:1905.03820v12019Beyond Sharing Weights for Deep Domain Adaptation
Artem Rozantsev, Mathieu Salzmann, Pascal Fua
cs.CVarXiv:1603.06432v22016Recalibrating Fully Convolutional Networks with Spatial and Channel 'Squeeze & Excitation' Blocks
Abhijit Guha Roy, Nassir Navab, Christian Wachinger
cs.CVarXiv:1808.08127v12018nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation
Fabian Isensee, Tassilo Wald, Constantin Ulrich +4
cs.CVarXiv:2404.09556v22024Pose-Normalized Image Generation for Person Re-identification
Xuelin Qian, Yanwei Fu, Tao Xiang +5
cs.CVcs.AIcs.MMarXiv:1712.02225v62017Masked Face Recognition Dataset and Application
Zhongyuan Wang, Guangcheng Wang, Baojin Huang +11
cs.CVarXiv:2003.09093v22020Triple Generative Adversarial Nets
Chongxuan Li, Kun Xu, Jun Zhu +1
cs.LGcs.CVarXiv:1703.02291v42017MixConv: Mixed Depthwise Convolutional Kernels
Mingxing Tan, Quoc V. Le
cs.CVcs.LGarXiv:1907.09595v32019Neural Photo Editing with Introspective Adversarial Networks
Andrew Brock, Theodore Lim, J. M. Ritchie +1
cs.LGcs.CVcs.NEarXiv:1609.07093v32016Image Matching across Wide Baselines: From Paper to Practice
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk +4
cs.CVarXiv:2003.01587v52020YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss
Debapriya Maji, Soyeb Nagori, Manu Mathew +1
cs.CVcs.AIcs.LGarXiv:2204.06806v12022Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal Mixup
Jang-Hyun Kim, Wonho Choo, Hyun Oh Song
cs.LGcs.AIcs.CVarXiv:2009.06962v22020UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen, Linjie Li, Licheng Yu +5
cs.CVcs.CLcs.LGarXiv:1909.11740v32019Dimensionality-Driven Learning with Noisy Labels
Xingjun Ma, Yisen Wang, Michael E. Houle +5
cs.CVcs.LGstat.MLarXiv:1806.02612v22018POI: Multiple Object Tracking with High Performance Detection and Appearance Feature
Fengwei Yu, Wenbo Li, Quanquan Li +3
cs.CVarXiv:1610.06136v12016SpinQuant: LLM quantization with learned rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov +6
cs.LGcs.AIcs.CLarXiv:2405.16406v42024Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data
Xiangyu Yue, Yang Zhang, Sicheng Zhao +3
cs.CVarXiv:1909.00889v22019ESPNetv2: A Light-weight, Power Efficient, and General Purpose Convolutional Neural Network
Sachin Mehta, Mohammad Rastegari, Linda Shapiro +1
cs.CVarXiv:1811.11431v32018Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Adam W. Harley, Alex Ufkes, Konstantinos G. Derpanis
cs.CVcs.IRcs.LGarXiv:1502.07058v12015Message Passing Neural PDE Solvers
Johannes Brandstetter, Daniel Worrall, Max Welling
cs.LGcs.CVmath.NAarXiv:2202.03376v32022Procedura: Agentic 3D Modeling with Procedural Control
Youtian Lin, Yikang Yang, Zhanpeng Hu +5
cs.CVcs.GRarXiv:2608.26238v12026Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
Chenyang Wu, Fuchen Long, Binyuan Huang +4
cs.CVcs.MMarXiv:2608.26809v12026VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
Katherine Crowson, Stella Biderman, Daniel Kornis +4
cs.CVarXiv:2204.08583v22022AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Qihang Zhou, Guansong Pang, Yu Tian +2
cs.CVarXiv:2310.18961v122023GameWAM: A World Action Model for Video Games
Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo +1
cs.AIcs.CVcs.LGarXiv:2608.26200v12026