Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
16,741 to 16,800 of 18,830
Describing Videos by Exploiting Temporal Structure
Li Yao, Atousa Torabi, Kyunghyun Cho +4
stat.MLcs.AIcs.CLarXiv:1502.08029v52015Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity
Junlong Shen
cs.CVcs.AIarXiv:2608.20587v12026Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising
Merve Gülle, Junno Yun, Yaşar Utku Alçalar +1
eess.IVcs.AIcs.CVarXiv:2608.20561v12026BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
Jifeng Dai, Kaiming He, Jian Sun
cs.CVarXiv:1503.01640v22015Training Deep Neural Networks on Noisy Labels with Bootstrapping
Scott Reed, Honglak Lee, Dragomir Anguelov +3
cs.CVcs.LGcs.NEarXiv:1412.6596v32014What's the Point: Semantic Segmentation with Point Supervision
Amy Bearman, Olga Russakovsky, Vittorio Ferrari +1
cs.CVarXiv:1506.02106v520153D Bounding Box Estimation Using Deep Learning and Geometry
Arsalan Mousavian, Dragomir Anguelov, John Flynn +1
cs.CVarXiv:1612.00496v22016EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Enjun Du, Siyi Liu, Zirong Chen +8
cs.CVcs.LGarXiv:2608.20886v12026Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle +1
cs.CVcs.LGeess.IVarXiv:1912.07372v22019Summaries:한국어Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation
David P. Stonko
cs.AIcs.CVcs.ROarXiv:2608.21332v12026Resnet in Resnet: Generalizing Residual Architectures
Sasha Targ, Diogo Almeida, Kevin Lyman
cs.LGcs.CVcs.NEarXiv:1603.08029v12016Unsupervised Learning for Physical Interaction through Video Prediction
Chelsea Finn, Ian Goodfellow, Sergey Levine
cs.LGcs.AIcs.CVarXiv:1605.07157v42016TrackFormer: Multi-Object Tracking with Transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe +1
cs.CVarXiv:2101.02702v32021Multi-scale Orderless Pooling of Deep Convolutional Activation Features
Yunchao Gong, Liwei Wang, Ruiqi Guo +1
cs.CVarXiv:1403.1840v32014Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning
Eric Arazo, Diego Ortego, Paul Albert +2
cs.CVarXiv:1908.02983v52019Meshed-Memory Transformer for Image Captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi +1
cs.CVcs.CLarXiv:1912.08226v22019Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry
Madina Kojanazarova, Sidaty El Hadramy, Philippe C. Cattin
cs.AIcs.CGcs.CVarXiv:2608.20967v12026TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Yibo Hu, Yu Qian, Mao Gu +6
cs.AIcs.CVarXiv:2608.20958v12026Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Keyu Tian, Yi Jiang, Zehuan Yuan +2
cs.CVcs.AIarXiv:2404.02905v22024Towards Real-Time Multi-Object Tracking
Zhongdao Wang, Liang Zheng, Yixuan Liu +2
cs.CVarXiv:1909.12605v22019Toward Convolutional Blind Denoising of Real Photographs
Shi Guo, Zifei Yan, Kai Zhang +2
cs.CVarXiv:1807.04686v22018Activating More Pixels in Image Super-Resolution Transformer
Xiangyu Chen, Xintao Wang, Jiantao Zhou +2
eess.IVcs.CVarXiv:2205.04437v32022NVAE: A Deep Hierarchical Variational Autoencoder
Arash Vahdat, Jan Kautz
stat.MLcs.CVcs.LGarXiv:2007.03898v32020PointPainting: Sequential Fusion for 3D Object Detection
Sourabh Vora, Alex H. Lang, Bassam Helou +1
cs.CVcs.LGeess.IVarXiv:1911.10150v22019OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
Xianyun Sun, Chaoyou Fu, Zhengye Zhang +6
cs.CVarXiv:2608.21360v12026Multimodal Learning with Transformers: A Survey
Peng Xu, Xiatian Zhu, David A. Clifton
cs.CVcs.LGarXiv:2206.06488v22022Block-NeRF: Scalable Large Scene Neural View Synthesis
Matthew Tancik, Vincent Casser, Xinchen Yan +5
cs.CVcs.GRarXiv:2202.05263v12022Token Merging: Your ViT But Faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +3
cs.CVarXiv:2210.09461v32022Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal +10
cs.CVcs.AIcs.LGarXiv:2301.00493v12023Contrastive Learning of Medical Visual Representations from Paired Images and Text
Yuhao Zhang, Hang Jiang, Yasuhide Miura +2
cs.CVcs.CLcs.LGarXiv:2010.00747v22020HAQ: Hardware-Aware Automated Quantization with Mixed Precision
Kuan Wang, Zhijian Liu, Yujun Lin +2
cs.CVarXiv:1811.08886v32018Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
Aojun Zhou, Anbang Yao, Yiwen Guo +2
cs.CVcs.AIcs.NEarXiv:1702.03044v22017DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images
Ilke Demir, Krzysztof Koperski, David Lindenbaum +6
cs.CVarXiv:1805.06561v12018PointRend: Image Segmentation as Rendering
Alexander Kirillov, Yuxin Wu, Kaiming He +1
cs.CVarXiv:1912.08193v22019The RSNA-ASNR-MICCAI BraTS 2021 Benchmark on Brain Tumor Segmentation and Radiogenomic Classification
Ujjwal Baid, Satyam Ghodasara, Suyash Mohan +100
cs.CVarXiv:2107.02314v22021LightGlue: Local Feature Matching at Light Speed
Philipp Lindenberger, Paul-Edouard Sarlin, Marc Pollefeys
cs.CVarXiv:2306.13643v12023StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
Michelle Lin
cs.AIcs.CVarXiv:2608.20414v12026TrackingNet: A Large-Scale Dataset and Benchmark for Object Tracking in the Wild
Matthias Müller, Adel Bibi, Silvio Giancola +2
cs.CVcs.ROarXiv:1803.10794v12018Go-ICP: A Globally Optimal Solution to 3D ICP Point-Set Registration
Jiaolong Yang, Hongdong Li, Dylan Campbell +1
cs.CVarXiv:1605.03344v12016PACT: Parameterized Clipping Activation for Quantized Neural Networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani +3
cs.CVcs.AIarXiv:1805.06085v22018Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding
Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla
cs.CVcs.NEarXiv:1511.02680v22015Inf-Net: Automatic COVID-19 Lung Infection Segmentation from CT Images
Deng-Ping Fan, Tao Zhou, Ge-Peng Ji +5
eess.IVcs.CVcs.LGarXiv:2004.14133v42020ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers
Guray Ozgur, Tahar Chettaoui, Eduarda Caldeira +5
cs.CVeess.IVarXiv:2604.22841v12026DeepFakes and Beyond: A Survey of Face Manipulation and Fake Detection
Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez +2
cs.CVcs.MMarXiv:2001.00179v32020Convolutional Occupancy Networks
Songyou Peng, Michael Niemeyer, Lars Mescheder +2
cs.CVarXiv:2003.04618v22020Exposing Deep Fakes Using Inconsistent Head Poses
Xin Yang, Yuezun Li, Siwei Lyu
cs.CVarXiv:1811.00661v22018MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Greenbaum +7
cs.CVcs.LGeess.IVarXiv:1901.07042v52019Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen +20
cs.CVcs.AIcs.LGarXiv:2111.11432v12021Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks
Yinpeng Dong, Tianyu Pang, Hang Su +1
cs.CVcs.CRcs.LGarXiv:1904.02884v12019The Importance of Skip Connections in Biomedical Image Segmentation
Michal Drozdzal, Eugene Vorontsov, Gabriel Chartrand +2
cs.CVarXiv:1608.04117v22016ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic Segmentation
Di Lin, Jifeng Dai, Jiaya Jia +2
cs.CVarXiv:1604.05144v12016Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing
Yixuan Ding, Wei Huang, Ruijie Quan +2
cs.HCcs.CVarXiv:2606.05172v12026Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Mahmoud Assran, Quentin Duval, Ishan Misra +5
cs.CVcs.AIcs.LGarXiv:2301.08243v32023Temporal Relational Reasoning in Videos
Bolei Zhou, Alex Andonian, Aude Oliva +1
cs.CVarXiv:1711.08496v22017UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
Hong Jiang, Wensong Song, Zongxin Yang +2
cs.CVarXiv:2604.17565v42026Learning Background-Aware Correlation Filters for Visual Tracking
Hamed Kiani Galoogahi, Ashton Fagg, Simon Lucey
cs.CVarXiv:1703.04590v22017VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation
Jiyang Gao, Chen Sun, Hang Zhao +4
cs.CVcs.LGstat.MLarXiv:2005.04259v12020DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better
Orest Kupyn, Tetiana Martyniuk, Junru Wu +1
cs.CVcs.LGarXiv:1908.03826v12019Universal Style Transfer via Feature Transforms
Yijun Li, Chen Fang, Jimei Yang +3
cs.CVarXiv:1705.08086v22017Transformer Interpretability Beyond Attention Visualization
Hila Chefer, Shir Gur, Lior Wolf
cs.CVarXiv:2012.09838v22020