Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

16,741 to 16,800 of 18,830

  1. Describing Videos by Exploiting Temporal Structure

    Li Yao, Atousa Torabi, Kyunghyun Cho +4

    stat.MLcs.AIcs.CLarXiv:1502.08029v52015
  2. Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity

    Junlong Shen

    cs.CVcs.AIarXiv:2608.20587v12026
  3. Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising

    Merve Gülle, Junno Yun, Yaşar Utku Alçalar +1

    eess.IVcs.AIcs.CVarXiv:2608.20561v12026
  4. BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation

    Jifeng Dai, Kaiming He, Jian Sun

    cs.CVarXiv:1503.01640v22015
  5. Training Deep Neural Networks on Noisy Labels with Bootstrapping

    Scott Reed, Honglak Lee, Dragomir Anguelov +3

    cs.CVcs.LGcs.NEarXiv:1412.6596v32014
  6. What's the Point: Semantic Segmentation with Point Supervision

    Amy Bearman, Olga Russakovsky, Vittorio Ferrari +1

    cs.CVarXiv:1506.02106v52015
  7. 3D Bounding Box Estimation Using Deep Learning and Geometry

    Arsalan Mousavian, Dragomir Anguelov, John Flynn +1

    cs.CVarXiv:1612.00496v22016
  8. EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

    Enjun Du, Siyi Liu, Zirong Chen +8

    cs.CVcs.LGarXiv:2608.20886v12026
  9. Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision

    Michael Niemeyer, Lars Mescheder, Michael Oechsle +1

    cs.CVcs.LGeess.IVarXiv:1912.07372v22019
    Summaries:한국어
  10. Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation

    David P. Stonko

    cs.AIcs.CVcs.ROarXiv:2608.21332v12026
  11. Resnet in Resnet: Generalizing Residual Architectures

    Sasha Targ, Diogo Almeida, Kevin Lyman

    cs.LGcs.CVcs.NEarXiv:1603.08029v12016
  12. Unsupervised Learning for Physical Interaction through Video Prediction

    Chelsea Finn, Ian Goodfellow, Sergey Levine

    cs.LGcs.AIcs.CVarXiv:1605.07157v42016
  13. TrackFormer: Multi-Object Tracking with Transformers

    Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe +1

    cs.CVarXiv:2101.02702v32021
  14. Multi-scale Orderless Pooling of Deep Convolutional Activation Features

    Yunchao Gong, Liwei Wang, Ruiqi Guo +1

    cs.CVarXiv:1403.1840v32014
  15. Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning

    Eric Arazo, Diego Ortego, Paul Albert +2

    cs.CVarXiv:1908.02983v52019
  16. Meshed-Memory Transformer for Image Captioning

    Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi +1

    cs.CVcs.CLarXiv:1912.08226v22019
  17. Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry

    Madina Kojanazarova, Sidaty El Hadramy, Philippe C. Cattin

    cs.AIcs.CGcs.CVarXiv:2608.20967v12026
  18. TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

    Yibo Hu, Yu Qian, Mao Gu +6

    cs.AIcs.CVarXiv:2608.20958v12026
  19. Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

    Keyu Tian, Yi Jiang, Zehuan Yuan +2

    cs.CVcs.AIarXiv:2404.02905v22024
  20. Towards Real-Time Multi-Object Tracking

    Zhongdao Wang, Liang Zheng, Yixuan Liu +2

    cs.CVarXiv:1909.12605v22019
  21. Toward Convolutional Blind Denoising of Real Photographs

    Shi Guo, Zifei Yan, Kai Zhang +2

    cs.CVarXiv:1807.04686v22018
  22. Activating More Pixels in Image Super-Resolution Transformer

    Xiangyu Chen, Xintao Wang, Jiantao Zhou +2

    eess.IVcs.CVarXiv:2205.04437v32022
  23. NVAE: A Deep Hierarchical Variational Autoencoder

    Arash Vahdat, Jan Kautz

    stat.MLcs.CVcs.LGarXiv:2007.03898v32020
  24. PointPainting: Sequential Fusion for 3D Object Detection

    Sourabh Vora, Alex H. Lang, Bassam Helou +1

    cs.CVcs.LGeess.IVarXiv:1911.10150v22019
  25. OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

    Xianyun Sun, Chaoyou Fu, Zhengye Zhang +6

    cs.CVarXiv:2608.21360v12026
  26. Multimodal Learning with Transformers: A Survey

    Peng Xu, Xiatian Zhu, David A. Clifton

    cs.CVcs.LGarXiv:2206.06488v22022
  27. Block-NeRF: Scalable Large Scene Neural View Synthesis

    Matthew Tancik, Vincent Casser, Xinchen Yan +5

    cs.CVcs.GRarXiv:2202.05263v12022
  28. Token Merging: Your ViT But Faster

    Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai +3

    cs.CVarXiv:2210.09461v32022
  29. Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

    Benjamin Wilson, William Qi, Tanmay Agarwal +10

    cs.CVcs.AIcs.LGarXiv:2301.00493v12023
  30. Contrastive Learning of Medical Visual Representations from Paired Images and Text

    Yuhao Zhang, Hang Jiang, Yasuhide Miura +2

    cs.CVcs.CLcs.LGarXiv:2010.00747v22020
  31. HAQ: Hardware-Aware Automated Quantization with Mixed Precision

    Kuan Wang, Zhijian Liu, Yujun Lin +2

    cs.CVarXiv:1811.08886v32018
  32. Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights

    Aojun Zhou, Anbang Yao, Yiwen Guo +2

    cs.CVcs.AIcs.NEarXiv:1702.03044v22017
  33. DeepGlobe 2018: A Challenge to Parse the Earth through Satellite Images

    Ilke Demir, Krzysztof Koperski, David Lindenbaum +6

    cs.CVarXiv:1805.06561v12018
  34. PointRend: Image Segmentation as Rendering

    Alexander Kirillov, Yuxin Wu, Kaiming He +1

    cs.CVarXiv:1912.08193v22019
  35. The RSNA-ASNR-MICCAI BraTS 2021 Benchmark on Brain Tumor Segmentation and Radiogenomic Classification

    Ujjwal Baid, Satyam Ghodasara, Suyash Mohan +100

    cs.CVarXiv:2107.02314v22021
  36. LightGlue: Local Feature Matching at Light Speed

    Philipp Lindenberger, Paul-Edouard Sarlin, Marc Pollefeys

    cs.CVarXiv:2306.13643v12023
  37. StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

    Michelle Lin

    cs.AIcs.CVarXiv:2608.20414v12026
  38. TrackingNet: A Large-Scale Dataset and Benchmark for Object Tracking in the Wild

    Matthias Müller, Adel Bibi, Silvio Giancola +2

    cs.CVcs.ROarXiv:1803.10794v12018
  39. Go-ICP: A Globally Optimal Solution to 3D ICP Point-Set Registration

    Jiaolong Yang, Hongdong Li, Dylan Campbell +1

    cs.CVarXiv:1605.03344v12016
  40. PACT: Parameterized Clipping Activation for Quantized Neural Networks

    Jungwook Choi, Zhuo Wang, Swagath Venkataramani +3

    cs.CVcs.AIarXiv:1805.06085v22018
  41. Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding

    Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla

    cs.CVcs.NEarXiv:1511.02680v22015
  42. Inf-Net: Automatic COVID-19 Lung Infection Segmentation from CT Images

    Deng-Ping Fan, Tao Zhou, Ge-Peng Ji +5

    eess.IVcs.CVcs.LGarXiv:2004.14133v42020
  43. ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers

    Guray Ozgur, Tahar Chettaoui, Eduarda Caldeira +5

    cs.CVeess.IVarXiv:2604.22841v12026
  44. DeepFakes and Beyond: A Survey of Face Manipulation and Fake Detection

    Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez +2

    cs.CVcs.MMarXiv:2001.00179v32020
  45. Convolutional Occupancy Networks

    Songyou Peng, Michael Niemeyer, Lars Mescheder +2

    cs.CVarXiv:2003.04618v22020
  46. Exposing Deep Fakes Using Inconsistent Head Poses

    Xin Yang, Yuezun Li, Siwei Lyu

    cs.CVarXiv:1811.00661v22018
  47. MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs

    Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Greenbaum +7

    cs.CVcs.LGeess.IVarXiv:1901.07042v52019
  48. Florence: A New Foundation Model for Computer Vision

    Lu Yuan, Dongdong Chen, Yi-Ling Chen +20

    cs.CVcs.AIcs.LGarXiv:2111.11432v12021
  49. Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks

    Yinpeng Dong, Tianyu Pang, Hang Su +1

    cs.CVcs.CRcs.LGarXiv:1904.02884v12019
  50. The Importance of Skip Connections in Biomedical Image Segmentation

    Michal Drozdzal, Eugene Vorontsov, Gabriel Chartrand +2

    cs.CVarXiv:1608.04117v22016
  51. ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic Segmentation

    Di Lin, Jifeng Dai, Jiaya Jia +2

    cs.CVarXiv:1604.05144v12016
  52. Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

    Yixuan Ding, Wei Huang, Ruijie Quan +2

    cs.HCcs.CVarXiv:2606.05172v12026
  53. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra +5

    cs.CVcs.AIcs.LGarXiv:2301.08243v32023
  54. Temporal Relational Reasoning in Videos

    Bolei Zhou, Alex Andonian, Aude Oliva +1

    cs.CVarXiv:1711.08496v22017
  55. UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

    Hong Jiang, Wensong Song, Zongxin Yang +2

    cs.CVarXiv:2604.17565v42026
  56. Learning Background-Aware Correlation Filters for Visual Tracking

    Hamed Kiani Galoogahi, Ashton Fagg, Simon Lucey

    cs.CVarXiv:1703.04590v22017
  57. VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation

    Jiyang Gao, Chen Sun, Hang Zhao +4

    cs.CVcs.LGstat.MLarXiv:2005.04259v12020
  58. DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better

    Orest Kupyn, Tetiana Martyniuk, Junru Wu +1

    cs.CVcs.LGarXiv:1908.03826v12019
  59. Universal Style Transfer via Feature Transforms

    Yijun Li, Chen Fang, Jimei Yang +3

    cs.CVarXiv:1705.08086v22017
  60. Transformer Interpretability Beyond Attention Visualization

    Hila Chefer, Shir Gur, Lior Wolf

    cs.CVarXiv:2012.09838v22020