Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,121 to 3,180 of 18,839
Lite Vision Transformer with Enhanced Self-Attention
Chenglin Yang, Yilin Wang, Jianming Zhang +4
cs.CVarXiv:2112.10809v12021Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
Zhenhailong Wang, Manling Li, Ruochen Xu +10
cs.CVcs.AIarXiv:2205.10747v42022RigNet: Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan, Kun Wang, Xiang Li +3
cs.CVarXiv:2107.13802v52021Multiregion Bilinear Convolutional Neural Networks for Person Re-Identification
Evgeniya Ustinova, Yaroslav Ganin, Victor Lempitsky
cs.CVarXiv:1512.05300v52015Source-Free Domain Adaptation via Distribution Estimation
Ning Ding, Yixing Xu, Yehui Tang +3
cs.CVarXiv:2204.11257v12022Severity Assessment of Coronavirus Disease 2019 (COVID-19) Using Quantitative Features from Chest CT Images
Zhenyu Tang, Wei Zhao, Xingzhi Xie +4
eess.IVcs.CVarXiv:2003.11988v12020SLOAM: Semantic Lidar Odometry and Mapping for Forest Inventory
Steven W. Chen, Guilherme V. Nardari, Elijah S. Lee +4
cs.ROcs.CVcs.LGarXiv:1912.12726v12019The Medical Segmentation Decathlon
Michela Antonelli, Annika Reinke, Spyridon Bakas +56
eess.IVcs.CVcs.LGarXiv:2106.05735v12021Task-Driven Convolutional Recurrent Models of the Visual System
Aran Nayebi, Daniel Bear, Jonas Kubilius +5
q-bio.NCcs.AIcs.CVarXiv:1807.00053v22018Pay Attention to What You Read: Non-recurrent Handwritten Text-Line Recognition
Lei Kang, Pau Riba, Marçal Rusiñol +2
cs.CVarXiv:2005.13044v12020LookOut: Diverse Multi-Future Prediction and Planning for Self-Driving
Alexander Cui, Sergio Casas, Abbas Sadat +2
cs.ROcs.AIcs.CVarXiv:2101.06547v32021Do Feature Attribution Methods Correctly Attribute Features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro +1
cs.LGcs.CVarXiv:2104.14403v22021Image Denoising: The Deep Learning Revolution and Beyond -- A Survey Paper --
Michael Elad, Bahjat Kawar, Gregory Vaksman
eess.IVcs.CVarXiv:2301.03362v12023AutoHR: A Strong End-to-end Baseline for Remote Heart Rate Measurement with Neural Searching
Zitong Yu, Xiaobai Li, Xuesong Niu +2
cs.CVarXiv:2004.12292v12020CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
Rundi Wu, Ruiqi Gao, Ben Poole +4
cs.CVarXiv:2411.18613v22024Simultaneous Corn and Soybean Yield Prediction from Remote Sensing Data Using Deep Transfer Learning
Saeed Khaki, Hieu Pham, Lizhi Wang
cs.CVcs.LGeess.IVarXiv:2012.03129v32020Risk Stratification of Lung Nodules Using 3D CNN-Based Multi-task Learning
Sarfaraz Hussein, Kunlin Cao, Qi Song +1
cs.CVcs.LGarXiv:1704.08797v12017A Comprehensive Review for Breast Histopathology Image Analysis Using Classical and Deep Neural Networks
Xiaomin Zhou, Chen Li, Md Mamunur Rahaman +6
eess.IVcs.CVarXiv:2003.12255v22020Solo-learn: A Library of Self-supervised Methods for Visual Representation Learning
Victor G. Turrisi da Costa, Enrico Fini, Moin Nabi +2
cs.CVarXiv:2108.01775v42021iTAML: An Incremental Task-Agnostic Meta-learning Approach
Jathushan Rajasegaran, Salman Khan, Munawar Hayat +2
cs.LGcs.CVstat.MLarXiv:2003.11652v12020Event-Based Motion Segmentation by Motion Compensation
Timo Stoffregen, Guillermo Gallego, Tom Drummond +2
cs.CVarXiv:1904.01293v42019SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding
Hezhen Hu, Weichao Zhao, Wengang Zhou +1
cs.CVarXiv:2305.04868v12023Rethinking the Trigger of Backdoor Attack
Yiming Li, Tongqing Zhai, Baoyuan Wu +3
cs.CRcs.CVcs.LGarXiv:2004.04692v32020Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification
Yulin Wang, Kangchen Lv, Rui Huang +3
cs.CVcs.AIcs.LGarXiv:2010.05300v12020HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval
Song Liu, Haoqi Fan, Shengsheng Qian +3
cs.CVcs.AIarXiv:2103.15049v22021In or Out? Fixing ImageNet Out-of-Distribution Detection Evaluation
Julian Bitterwolf, Maximilian Müller, Matthias Hein
cs.LGcs.CVarXiv:2306.00826v12023On Robustness and Transferability of Convolutional Neural Networks
Josip Djolonga, Jessica Yung, Michael Tschannen +11
cs.CVcs.LGarXiv:2007.08558v22020Neural SDE: Stabilizing Neural ODE Networks with Stochastic Noise
Xuanqing Liu, Tesi Xiao, Si Si +3
cs.LGcs.AIcs.CVarXiv:1906.02355v12019ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Emotional Reaction Intensity Estimation Challenges
Dimitrios Kollias, Panagiotis Tzirakis, Alice Baird +2
cs.CVcs.LGarXiv:2303.01498v32023The Algorithmic Automation Problem: Prediction, Triage, and Human Effort
Maithra Raghu, Katy Blumer, Greg Corrado +3
cs.CVcs.AIcs.LGarXiv:1903.12220v12019Going Deeper through the Gleason Scoring Scale: An Automatic end-to-end System for Histology Prostate Grading and Cribriform Pattern Detection
Julio Silva-Rodríguez, Adrián Colomer, María A. Sales +2
eess.IVcs.CVarXiv:2105.10490v12021Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
Tan Wang, Xing Xu, Yang Yang +3
cs.CVarXiv:1908.04011v22019Deep Binary Reconstruction for Cross-modal Hashing
Xuelong Li, Di Hu, Feiping Nie
cs.CVcs.MMarXiv:1708.05127v22017ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and Removal
Bin Ding, Chengjiang Long, Ling Zhang +1
cs.CVarXiv:1908.01323v12019SMART Frame Selection for Action Recognition
Shreyank N Gowda, Marcus Rohrbach, Laura Sevilla-Lara
cs.CVarXiv:2012.10671v12020Memory Bounded Deep Convolutional Networks
Maxwell D. Collins, Pushmeet Kohli
cs.CVarXiv:1412.1442v12014Weakly Supervised Dense Event Captioning in Videos
Xuguang Duan, Wenbing Huang, Chuang Gan +3
cs.CVarXiv:1812.03849v12018Excessive Invariance Causes Adversarial Vulnerability
Jörn-Henrik Jacobsen, Jens Behrmann, Richard Zemel +1
cs.LGcs.AIcs.CVarXiv:1811.00401v42018Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention
Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh +2
cs.CVarXiv:1908.07236v22019FENeRF: Face Editing in Neural Radiance Fields
Jingxiang Sun, Xuan Wang, Yong Zhang +4
cs.CVarXiv:2111.15490v22021Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations
Wei-Jin Huang, Yue-Yi Zhang, Yi-Lin Wei +5
cs.ROcs.AIcs.CVarXiv:2601.09518v12026ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries
Junru Gu, Chenxu Hu, Tianyuan Zhang +4
cs.CVcs.ROarXiv:2208.01582v32022MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
Xinchen Yan, Akash Rastogi, Ruben Villegas +5
cs.LGcs.AIcs.CVarXiv:1808.04545v12018Learning towards Minimum Hyperspherical Energy
Weiyang Liu, Rongmei Lin, Zhen Liu +4
cs.LGcs.CVstat.MLarXiv:1805.09298v92018N-Gram in Swin Transformers for Efficient Lightweight Image Super-Resolution
Haram Choi, Jeongmin Lee, Jihoon Yang
cs.CVarXiv:2211.11436v32022Guaranteed Tensor Recovery Fused Low-rankness and Smoothness
Hailin Wang, Jiangjun Peng, Wenjin Qin +2
cs.LGcs.AIcs.CVarXiv:2302.02155v12023Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation
Tianfei Zhou, Meijie Zhang, Fang Zhao +1
cs.CVarXiv:2203.09653v22022Robust Online Matrix Factorization for Dynamic Background Subtraction
Hongwei Yong, Deyu Meng, Wangmeng Zuo +1
cs.CVarXiv:1705.10000v12017Liquid Structural State-Space Models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang +3
cs.LGcs.AIcs.CLarXiv:2209.12951v12022The Pros and Cons: Rank-aware Temporal Attention for Skill Determination in Long Videos
Hazel Doughty, Walterio Mayol-Cuevas, Dima Damen
cs.CVarXiv:1812.05538v22018Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
Yangchen Zeng, Zhenyu Yu, Dongming Jiang +5
cs.CVarXiv:2604.15065v12026MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang +18
cs.CVcs.AIcs.CLarXiv:2406.09411v22024Tora: Trajectory-oriented Diffusion Transformer for Video Generation
Zhenghao Zhang, Junchao Liao, Menghao Li +5
cs.CVarXiv:2407.21705v42024Visual Coreference Resolution in Visual Dialog using Neural Module Networks
Satwik Kottur, José M. F. Moura, Devi Parikh +2
cs.CVcs.AIcs.CLarXiv:1809.01816v12018MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features
Shakti N. Wadekar, Abhishek Chaurasia
cs.CVcs.AIcs.LGarXiv:2209.15159v22022Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking
Sarthak Sharma, Junaid Ahmed Ansari, J. Krishna Murthy +1
cs.ROcs.CVarXiv:1802.09298v22018Asynchronous Temporal Fields for Action Recognition
Gunnar A. Sigurdsson, Santosh Divvala, Ali Farhadi +1
cs.CVarXiv:1612.06371v22016On Attention Models for Human Activity Recognition
Vishvak S Murahari, Thomas Ploetz
cs.CVcs.AIcs.LGarXiv:1805.07648v12018AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala
cs.CVcs.CLarXiv:2103.16002v12021ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
Siyuan Qiao, Yukun Zhu, Hartwig Adam +2
cs.CVarXiv:2012.05258v12020