Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
14,461 to 14,520 of 18,817
Decomposing Motion and Content for Natural Video Sequence Prediction
Ruben Villegas, Jimei Yang, Seunghoon Hong +2
cs.CVarXiv:1706.08033v22017Learning Depth from Monocular Videos using Direct Methods
Chaoyang Wang, Jose Miguel Buenaposada, Rui Zhu +1
cs.CVarXiv:1712.00175v12017Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector
Qi Fan, Wei Zhuo, Chi-Keung Tang +1
cs.CVarXiv:1908.01998v42019GAIA-1: A Generative World Model for Autonomous Driving
Anthony Hu, Lloyd Russell, Hudson Yeo +5
cs.CVcs.AIcs.ROarXiv:2309.17080v12023Med-Flamingo: a Multimodal Medical Few-shot Learner
Michael Moor, Qian Huang, Shirley Wu +6
cs.CVcs.AIarXiv:2307.15189v12023GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images
Jun Gao, Tianchang Shen, Zian Wang +6
cs.CVarXiv:2209.11163v12022U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation
Junho Kim, Minjae Kim, Hyeonwoo Kang +1
cs.CVeess.IVarXiv:1907.10830v42019Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation
Anurag Ranjan, Varun Jampani, Lukas Balles +4
cs.CVarXiv:1805.09806v32018Video Instance Segmentation
Linjie Yang, Yuchen Fan, Ning Xu
cs.CVarXiv:1905.04804v42019Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
Sixiang Chen, Jiaming Liu, Jixian Wu +7
cs.ROcs.CVarXiv:2608.24885v12026Scaling Laws for Autoregressive Generative Modeling
Tom Henighan, Jared Kaplan, Mor Katz +16
cs.LGcs.CLcs.CVarXiv:2010.14701v22020KLTNet: Learning Sparse Feature Tracking for Robust and Accurate Monocular Visual-Inertial Odometry
Renbiao Jin, Danping Zou, Wenxian Yu
cs.CVarXiv:2608.24544v12026Going Deeper into Action Recognition: A Survey
Samitha Herath, Mehrtash Harandi, Fatih Porikli
cs.CVarXiv:1605.04988v22016Speaker-Follower Models for Vision-and-Language Navigation
Daniel Fried, Ronghang Hu, Volkan Cirik +7
cs.CVcs.CLarXiv:1806.02724v22018Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Minghui Liao, Pengyuan Lyu, Minghang He +3
cs.CVarXiv:1908.08207v12019OCNet: Object Context Network for Scene Parsing
Yuhui Yuan, Lang Huang, Jianyuan Guo +3
cs.CVarXiv:1809.00916v42018Vision GNN: An Image is Worth Graph of Nodes
Kai Han, Yunhe Wang, Jianyuan Guo +2
cs.CVarXiv:2206.00272v32022DeepViT: Towards Deeper Vision Transformer
Daquan Zhou, Bingyi Kang, Xiaojie Jin +5
cs.CVarXiv:2103.11886v42021DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object Detection
Haibao Yu, Yizhen Luo, Mao Shu +8
cs.CVcs.AIarXiv:2204.05575v12022SLIP: Self-supervision meets Language-Image Pre-training
Norman Mu, Alexander Kirillov, David Wagner +1
cs.CVarXiv:2112.12750v12021Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation
Xinning Yao, Jingjing Wang, Jinghua Yue +3
cs.CVarXiv:2608.24541v12026Vision Language Model Fusion for Explainable Face Recognition
Ana Estrada-Real, Lydia Alapatt, Christoph Busch +1
cs.CVarXiv:2608.24430v12026All are Worth Words: A ViT Backbone for Diffusion Models
Fan Bao, Shen Nie, Kaiwen Xue +4
cs.CVcs.AIcs.LGarXiv:2209.12152v42022GANimation: Anatomically-aware Facial Animation from a Single Image
Albert Pumarola, Antonio Agudo, Aleix M. Martinez +2
cs.CVarXiv:1807.09251v22018Explainable deep learning models in medical image analysis
Amitojdeep Singh, Sourya Sengupta, Vasudevan Lakshminarayanan
cs.CVcs.LGeess.IVarXiv:2005.13799v12020VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
Lyuke Wang, Zhuo Li, Guangxu Zhu
cs.CVcs.AIarXiv:2608.24063v12026Zero-Shot Learning via Semantic Similarity Embedding
Ziming Zhang, Venkatesh Saligrama
cs.CVstat.MLarXiv:1509.04767v22015Kimera: an Open-Source Library for Real-Time Metric-Semantic Localization and Mapping
Antoni Rosinol, Marcus Abate, Yun Chang +1
cs.ROcs.CVarXiv:1910.02490v32019Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
Feng Li, Hao Zhang, Huaizhe xu +4
cs.CVarXiv:2206.02777v32022Categorical Depth Distribution Network for Monocular 3D Object Detection
Cody Reading, Ali Harakeh, Julia Chae +1
cs.CVarXiv:2103.01100v22021Learning to Estimate 3D Human Pose and Shape from a Single Color Image
Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou +1
cs.CVarXiv:1805.04092v12018PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models
Sachit Menon, Alexandru Damian, Shijia Hu +2
cs.CVcs.LGeess.IVarXiv:2003.03808v32020NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
Holger Caesar, Juraj Kabzan, Kok Seang Tan +6
cs.CVarXiv:2106.11810v42021Unbiased Teacher for Semi-Supervised Object Detection
Yen-Cheng Liu, Chih-Yao Ma, Zijian He +6
cs.CVcs.LGarXiv:2102.09480v12021Expectation-Maximization Attention Networks for Semantic Segmentation
Xia Li, Zhisheng Zhong, Jianlong Wu +3
cs.CVarXiv:1907.13426v22019A Baseline for Few-Shot Image Classification
Guneet S. Dhillon, Pratik Chaudhari, Avinash Ravichandran +1
cs.LGcs.CVstat.MLarXiv:1909.02729v52019SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image
Zefan Tian, Yuteng Ye, Yiheng Zhang +5
cs.CVarXiv:2608.23930v12026PoinTr: Diverse Point Cloud Completion with Geometry-Aware Transformers
Xumin Yu, Yongming Rao, Ziyi Wang +3
cs.CVcs.AIcs.LGarXiv:2108.08839v12021TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
Yapeng Tian, Yulun Zhang, Yun Fu +1
cs.CVarXiv:1812.02898v12018Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz +5
cs.CVcs.AIcs.CLarXiv:1811.10092v22018Constrained Convolutional Neural Networks for Weakly Supervised Segmentation
Deepak Pathak, Philipp Krähenbühl, Trevor Darrell
cs.CVcs.LGarXiv:1506.03648v22015HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Tianrui Guan, Fuxiao Liu, Xiyang Wu +9
cs.CVcs.CLarXiv:2310.14566v52023A Fourier-based Framework for Domain Generalization
Qinwei Xu, Ruipeng Zhang, Ya Zhang +2
cs.CVarXiv:2105.11120v12021Generating 3D faces using Convolutional Mesh Autoencoders
Anurag Ranjan, Timo Bolkart, Soubhik Sanyal +1
cs.CVarXiv:1807.10267v32018Interacted Planes Reveal 3D Line Mapping
Zeran Ke, Bin Tan, Gui-Song Xia +2
cs.CVarXiv:2602.01296v12026Data-Free Quantization Through Weight Equalization and Bias Correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort +1
cs.LGcs.CVstat.MLarXiv:1906.04721v32019On the Relationship between Self-Attention and Convolutional Layers
Jean-Baptiste Cordonnier, Andreas Loukas, Martin Jaggi
cs.LGcs.CLcs.CVarXiv:1911.03584v22019Baking Neural Radiance Fields for Real-Time View Synthesis
Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall +2
cs.CVcs.GRarXiv:2103.14645v12021Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation
Hongliang Yan, Yukang Ding, Peihua Li +3
cs.CVarXiv:1705.00609v12017Joint Distribution Alignment for Universal Domain Adaptation
Shizhe Li, Hongshan Pu, Mengying Xie +2
cs.LGcs.CVarXiv:2608.24429v12026Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation
Qihang Sun, Zhongxiao Liu, Bailiang Jian +6
eess.IVcs.CVarXiv:2608.24486v12026Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
Deng-Ping Fan, Zheng Lin, Jia-Xing Zhao +5
cs.CVarXiv:1907.06781v22019Correlation Congruence for Knowledge Distillation
Baoyun Peng, Xiao Jin, Jiaheng Liu +5
cs.CVarXiv:1904.01802v12019Bayesian Loss for Crowd Count Estimation with Point Supervision
Zhiheng Ma, Xing Wei, Xiaopeng Hong +1
cs.CVarXiv:1908.03684v12019SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery
Yezhen Cong, Samar Khanna, Chenlin Meng +6
cs.CVcs.AIarXiv:2207.08051v32022Detecting and Recognizing Human-Object Interactions
Georgia Gkioxari, Ross Girshick, Piotr Dollár +1
cs.CVarXiv:1704.07333v32017EXPANSE: A Deep Continual / Progressive Learning System for Deep Transfer Learning
Mohammadreza Iman, John A. Miller, Khaled Rasheed +2
cs.LGcs.CVarXiv:2205.10356v22022When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
Zhengxiang Wang, Owen Rambow
cs.AIcs.CVarXiv:2608.23978v12026A Fourier Perspective on Model Robustness in Computer Vision
Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens +2
cs.LGcs.CVstat.MLarXiv:1906.08988v32019OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Anas Awadalla, Irena Gao, Josh Gardner +13
cs.CVcs.AIcs.LGarXiv:2308.01390v22023