Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

14,461 to 14,520 of 18,817

  1. Decomposing Motion and Content for Natural Video Sequence Prediction

    Ruben Villegas, Jimei Yang, Seunghoon Hong +2

    cs.CVarXiv:1706.08033v22017
  2. Learning Depth from Monocular Videos using Direct Methods

    Chaoyang Wang, Jose Miguel Buenaposada, Rui Zhu +1

    cs.CVarXiv:1712.00175v12017
  3. Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector

    Qi Fan, Wei Zhuo, Chi-Keung Tang +1

    cs.CVarXiv:1908.01998v42019
  4. GAIA-1: A Generative World Model for Autonomous Driving

    Anthony Hu, Lloyd Russell, Hudson Yeo +5

    cs.CVcs.AIcs.ROarXiv:2309.17080v12023
  5. Med-Flamingo: a Multimodal Medical Few-shot Learner

    Michael Moor, Qian Huang, Shirley Wu +6

    cs.CVcs.AIarXiv:2307.15189v12023
  6. GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images

    Jun Gao, Tianchang Shen, Zian Wang +6

    cs.CVarXiv:2209.11163v12022
  7. U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation

    Junho Kim, Minjae Kim, Hyeonwoo Kang +1

    cs.CVeess.IVarXiv:1907.10830v42019
  8. Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation

    Anurag Ranjan, Varun Jampani, Lukas Balles +4

    cs.CVarXiv:1805.09806v32018
  9. Video Instance Segmentation

    Linjie Yang, Yuchen Fan, Ning Xu

    cs.CVarXiv:1905.04804v42019
  10. Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

    Sixiang Chen, Jiaming Liu, Jixian Wu +7

    cs.ROcs.CVarXiv:2608.24885v12026
  11. Scaling Laws for Autoregressive Generative Modeling

    Tom Henighan, Jared Kaplan, Mor Katz +16

    cs.LGcs.CLcs.CVarXiv:2010.14701v22020
  12. KLTNet: Learning Sparse Feature Tracking for Robust and Accurate Monocular Visual-Inertial Odometry

    Renbiao Jin, Danping Zou, Wenxian Yu

    cs.CVarXiv:2608.24544v12026
  13. Going Deeper into Action Recognition: A Survey

    Samitha Herath, Mehrtash Harandi, Fatih Porikli

    cs.CVarXiv:1605.04988v22016
  14. Speaker-Follower Models for Vision-and-Language Navigation

    Daniel Fried, Ronghang Hu, Volkan Cirik +7

    cs.CVcs.CLarXiv:1806.02724v22018
  15. Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes

    Minghui Liao, Pengyuan Lyu, Minghang He +3

    cs.CVarXiv:1908.08207v12019
  16. OCNet: Object Context Network for Scene Parsing

    Yuhui Yuan, Lang Huang, Jianyuan Guo +3

    cs.CVarXiv:1809.00916v42018
  17. Vision GNN: An Image is Worth Graph of Nodes

    Kai Han, Yunhe Wang, Jianyuan Guo +2

    cs.CVarXiv:2206.00272v32022
  18. DeepViT: Towards Deeper Vision Transformer

    Daquan Zhou, Bingyi Kang, Xiaojie Jin +5

    cs.CVarXiv:2103.11886v42021
  19. DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object Detection

    Haibao Yu, Yizhen Luo, Mao Shu +8

    cs.CVcs.AIarXiv:2204.05575v12022
  20. SLIP: Self-supervision meets Language-Image Pre-training

    Norman Mu, Alexander Kirillov, David Wagner +1

    cs.CVarXiv:2112.12750v12021
  21. Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation

    Xinning Yao, Jingjing Wang, Jinghua Yue +3

    cs.CVarXiv:2608.24541v12026
  22. Vision Language Model Fusion for Explainable Face Recognition

    Ana Estrada-Real, Lydia Alapatt, Christoph Busch +1

    cs.CVarXiv:2608.24430v12026
  23. All are Worth Words: A ViT Backbone for Diffusion Models

    Fan Bao, Shen Nie, Kaiwen Xue +4

    cs.CVcs.AIcs.LGarXiv:2209.12152v42022
  24. GANimation: Anatomically-aware Facial Animation from a Single Image

    Albert Pumarola, Antonio Agudo, Aleix M. Martinez +2

    cs.CVarXiv:1807.09251v22018
  25. Explainable deep learning models in medical image analysis

    Amitojdeep Singh, Sourya Sengupta, Vasudevan Lakshminarayanan

    cs.CVcs.LGeess.IVarXiv:2005.13799v12020
  26. VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

    Lyuke Wang, Zhuo Li, Guangxu Zhu

    cs.CVcs.AIarXiv:2608.24063v12026
  27. Zero-Shot Learning via Semantic Similarity Embedding

    Ziming Zhang, Venkatesh Saligrama

    cs.CVstat.MLarXiv:1509.04767v22015
  28. Kimera: an Open-Source Library for Real-Time Metric-Semantic Localization and Mapping

    Antoni Rosinol, Marcus Abate, Yun Chang +1

    cs.ROcs.CVarXiv:1910.02490v32019
  29. Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation

    Feng Li, Hao Zhang, Huaizhe xu +4

    cs.CVarXiv:2206.02777v32022
  30. Categorical Depth Distribution Network for Monocular 3D Object Detection

    Cody Reading, Ali Harakeh, Julia Chae +1

    cs.CVarXiv:2103.01100v22021
  31. Learning to Estimate 3D Human Pose and Shape from a Single Color Image

    Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou +1

    cs.CVarXiv:1805.04092v12018
  32. PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models

    Sachit Menon, Alexandru Damian, Shijia Hu +2

    cs.CVcs.LGeess.IVarXiv:2003.03808v32020
  33. NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

    Holger Caesar, Juraj Kabzan, Kok Seang Tan +6

    cs.CVarXiv:2106.11810v42021
  34. Unbiased Teacher for Semi-Supervised Object Detection

    Yen-Cheng Liu, Chih-Yao Ma, Zijian He +6

    cs.CVcs.LGarXiv:2102.09480v12021
  35. Expectation-Maximization Attention Networks for Semantic Segmentation

    Xia Li, Zhisheng Zhong, Jianlong Wu +3

    cs.CVarXiv:1907.13426v22019
  36. A Baseline for Few-Shot Image Classification

    Guneet S. Dhillon, Pratik Chaudhari, Avinash Ravichandran +1

    cs.LGcs.CVstat.MLarXiv:1909.02729v52019
  37. SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

    Zefan Tian, Yuteng Ye, Yiheng Zhang +5

    cs.CVarXiv:2608.23930v12026
  38. PoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers

    Xumin Yu, Yongming Rao, Ziyi Wang +3

    cs.CVcs.AIcs.LGarXiv:2108.08839v12021
  39. TDAN: Temporally Deformable Alignment Network for Video Super-Resolution

    Yapeng Tian, Yulun Zhang, Yun Fu +1

    cs.CVarXiv:1812.02898v12018
  40. Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

    Xin Wang, Qiuyuan Huang, Asli Celikyilmaz +5

    cs.CVcs.AIcs.CLarXiv:1811.10092v22018
  41. Constrained Convolutional Neural Networks for Weakly Supervised Segmentation

    Deepak Pathak, Philipp Krähenbühl, Trevor Darrell

    cs.CVcs.LGarXiv:1506.03648v22015
  42. HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

    Tianrui Guan, Fuxiao Liu, Xiyang Wu +9

    cs.CVcs.CLarXiv:2310.14566v52023
  43. A Fourier-based Framework for Domain Generalization

    Qinwei Xu, Ruipeng Zhang, Ya Zhang +2

    cs.CVarXiv:2105.11120v12021
  44. Generating 3D faces using Convolutional Mesh Autoencoders

    Anurag Ranjan, Timo Bolkart, Soubhik Sanyal +1

    cs.CVarXiv:1807.10267v32018
  45. Interacted Planes Reveal 3D Line Mapping

    Zeran Ke, Bin Tan, Gui-Song Xia +2

    cs.CVarXiv:2602.01296v12026
  46. Data-Free Quantization Through Weight Equalization and Bias Correction

    Markus Nagel, Mart van Baalen, Tijmen Blankevoort +1

    cs.LGcs.CVstat.MLarXiv:1906.04721v32019
  47. On the Relationship between Self-Attention and Convolutional Layers

    Jean-Baptiste Cordonnier, Andreas Loukas, Martin Jaggi

    cs.LGcs.CLcs.CVarXiv:1911.03584v22019
  48. Baking Neural Radiance Fields for Real-Time View Synthesis

    Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall +2

    cs.CVcs.GRarXiv:2103.14645v12021
  49. Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation

    Hongliang Yan, Yukang Ding, Peihua Li +3

    cs.CVarXiv:1705.00609v12017
  50. Joint Distribution Alignment for Universal Domain Adaptation

    Shizhe Li, Hongshan Pu, Mengying Xie +2

    cs.LGcs.CVarXiv:2608.24429v12026
  51. Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

    Qihang Sun, Zhongxiao Liu, Bailiang Jian +6

    eess.IVcs.CVarXiv:2608.24486v12026
  52. Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks

    Deng-Ping Fan, Zheng Lin, Jia-Xing Zhao +5

    cs.CVarXiv:1907.06781v22019
  53. Correlation Congruence for Knowledge Distillation

    Baoyun Peng, Xiao Jin, Jiaheng Liu +5

    cs.CVarXiv:1904.01802v12019
  54. Bayesian Loss for Crowd Count Estimation with Point Supervision

    Zhiheng Ma, Xing Wei, Xiaopeng Hong +1

    cs.CVarXiv:1908.03684v12019
  55. SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery

    Yezhen Cong, Samar Khanna, Chenlin Meng +6

    cs.CVcs.AIarXiv:2207.08051v32022
  56. Detecting and Recognizing Human-Object Interactions

    Georgia Gkioxari, Ross Girshick, Piotr Dollár +1

    cs.CVarXiv:1704.07333v32017
  57. EXPANSE: A Deep Continual / Progressive Learning System for Deep Transfer Learning

    Mohammadreza Iman, John A. Miller, Khaled Rasheed +2

    cs.LGcs.CVarXiv:2205.10356v22022
  58. When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

    Zhengxiang Wang, Owen Rambow

    cs.AIcs.CVarXiv:2608.23978v12026
  59. A Fourier Perspective on Model Robustness in Computer Vision

    Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens +2

    cs.LGcs.CVstat.MLarXiv:1906.08988v32019
  60. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

    Anas Awadalla, Irena Gao, Josh Gardner +13

    cs.CVcs.AIcs.LGarXiv:2308.01390v22023