Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,121 to 3,180 of 18,839

  1. Lite Vision Transformer with Enhanced Self-Attention

    Chenglin Yang, Yilin Wang, Jianming Zhang +4

    cs.CVarXiv:2112.10809v12021
  2. Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

    Zhenhailong Wang, Manling Li, Ruochen Xu +10

    cs.CVcs.AIarXiv:2205.10747v42022
  3. RigNet: Repetitive Image Guided Network for Depth Completion

    Zhiqiang Yan, Kun Wang, Xiang Li +3

    cs.CVarXiv:2107.13802v52021
  4. Multiregion Bilinear Convolutional Neural Networks for Person Re-Identification

    Evgeniya Ustinova, Yaroslav Ganin, Victor Lempitsky

    cs.CVarXiv:1512.05300v52015
  5. Source-Free Domain Adaptation via Distribution Estimation

    Ning Ding, Yixing Xu, Yehui Tang +3

    cs.CVarXiv:2204.11257v12022
  6. Severity Assessment of Coronavirus Disease 2019 (COVID-19) Using Quantitative Features from Chest CT Images

    Zhenyu Tang, Wei Zhao, Xingzhi Xie +4

    eess.IVcs.CVarXiv:2003.11988v12020
  7. SLOAM: Semantic Lidar Odometry and Mapping for Forest Inventory

    Steven W. Chen, Guilherme V. Nardari, Elijah S. Lee +4

    cs.ROcs.CVcs.LGarXiv:1912.12726v12019
  8. The Medical Segmentation Decathlon

    Michela Antonelli, Annika Reinke, Spyridon Bakas +56

    eess.IVcs.CVcs.LGarXiv:2106.05735v12021
  9. Task-Driven Convolutional Recurrent Models of the Visual System

    Aran Nayebi, Daniel Bear, Jonas Kubilius +5

    q-bio.NCcs.AIcs.CVarXiv:1807.00053v22018
  10. Pay Attention to What You Read: Non-recurrent Handwritten Text-Line Recognition

    Lei Kang, Pau Riba, Marçal Rusiñol +2

    cs.CVarXiv:2005.13044v12020
  11. LookOut: Diverse Multi-Future Prediction and Planning for Self-Driving

    Alexander Cui, Sergio Casas, Abbas Sadat +2

    cs.ROcs.AIcs.CVarXiv:2101.06547v32021
  12. Do Feature Attribution Methods Correctly Attribute Features?

    Yilun Zhou, Serena Booth, Marco Tulio Ribeiro +1

    cs.LGcs.CVarXiv:2104.14403v22021
  13. Image Denoising: The Deep Learning Revolution and Beyond -- A Survey Paper --

    Michael Elad, Bahjat Kawar, Gregory Vaksman

    eess.IVcs.CVarXiv:2301.03362v12023
  14. AutoHR: A Strong End-to-end Baseline for Remote Heart Rate Measurement with Neural Searching

    Zitong Yu, Xiaobai Li, Xuesong Niu +2

    cs.CVarXiv:2004.12292v12020
  15. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    Rundi Wu, Ruiqi Gao, Ben Poole +4

    cs.CVarXiv:2411.18613v22024
  16. Simultaneous Corn and Soybean Yield Prediction from Remote Sensing Data Using Deep Transfer Learning

    Saeed Khaki, Hieu Pham, Lizhi Wang

    cs.CVcs.LGeess.IVarXiv:2012.03129v32020
  17. Risk Stratification of Lung Nodules Using 3D CNN-Based Multi-task Learning

    Sarfaraz Hussein, Kunlin Cao, Qi Song +1

    cs.CVcs.LGarXiv:1704.08797v12017
  18. A Comprehensive Review for Breast Histopathology Image Analysis Using Classical and Deep Neural Networks

    Xiaomin Zhou, Chen Li, Md Mamunur Rahaman +6

    eess.IVcs.CVarXiv:2003.12255v22020
  19. Solo-learn: A Library of Self-supervised Methods for Visual Representation Learning

    Victor G. Turrisi da Costa, Enrico Fini, Moin Nabi +2

    cs.CVarXiv:2108.01775v42021
  20. iTAML: An Incremental Task-Agnostic Meta-learning Approach

    Jathushan Rajasegaran, Salman Khan, Munawar Hayat +2

    cs.LGcs.CVstat.MLarXiv:2003.11652v12020
  21. Event-Based Motion Segmentation by Motion Compensation

    Timo Stoffregen, Guillermo Gallego, Tom Drummond +2

    cs.CVarXiv:1904.01293v42019
  22. SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding

    Hezhen Hu, Weichao Zhao, Wengang Zhou +1

    cs.CVarXiv:2305.04868v12023
  23. Rethinking the Trigger of Backdoor Attack

    Yiming Li, Tongqing Zhai, Baoyuan Wu +3

    cs.CRcs.CVcs.LGarXiv:2004.04692v32020
  24. Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification

    Yulin Wang, Kangchen Lv, Rui Huang +3

    cs.CVcs.AIcs.LGarXiv:2010.05300v12020
  25. HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval

    Song Liu, Haoqi Fan, Shengsheng Qian +3

    cs.CVcs.AIarXiv:2103.15049v22021
  26. In or Out? Fixing ImageNet Out-of-Distribution Detection Evaluation

    Julian Bitterwolf, Maximilian Müller, Matthias Hein

    cs.LGcs.CVarXiv:2306.00826v12023
  27. On Robustness and Transferability of Convolutional Neural Networks

    Josip Djolonga, Jessica Yung, Michael Tschannen +11

    cs.CVcs.LGarXiv:2007.08558v22020
  28. Neural SDE: Stabilizing Neural ODE Networks with Stochastic Noise

    Xuanqing Liu, Tesi Xiao, Si Si +3

    cs.LGcs.AIcs.CVarXiv:1906.02355v12019
  29. ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Emotional Reaction Intensity Estimation Challenges

    Dimitrios Kollias, Panagiotis Tzirakis, Alice Baird +2

    cs.CVcs.LGarXiv:2303.01498v32023
  30. The Algorithmic Automation Problem: Prediction, Triage, and Human Effort

    Maithra Raghu, Katy Blumer, Greg Corrado +3

    cs.CVcs.AIcs.LGarXiv:1903.12220v12019
  31. Going Deeper through the Gleason Scoring Scale: An Automatic end-to-end System for Histology Prostate Grading and Cribriform Pattern Detection

    Julio Silva-Rodríguez, Adrián Colomer, María A. Sales +2

    eess.IVcs.CVarXiv:2105.10490v12021
  32. Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking

    Tan Wang, Xing Xu, Yang Yang +3

    cs.CVarXiv:1908.04011v22019
  33. Deep Binary Reconstruction for Cross-modal Hashing

    Xuelong Li, Di Hu, Feiping Nie

    cs.CVcs.MMarXiv:1708.05127v22017
  34. ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and Removal

    Bin Ding, Chengjiang Long, Ling Zhang +1

    cs.CVarXiv:1908.01323v12019
  35. SMART Frame Selection for Action Recognition

    Shreyank N Gowda, Marcus Rohrbach, Laura Sevilla-Lara

    cs.CVarXiv:2012.10671v12020
  36. Memory Bounded Deep Convolutional Networks

    Maxwell D. Collins, Pushmeet Kohli

    cs.CVarXiv:1412.1442v12014
  37. Weakly Supervised Dense Event Captioning in Videos

    Xuguang Duan, Wenbing Huang, Chuang Gan +3

    cs.CVarXiv:1812.03849v12018
  38. Excessive Invariance Causes Adversarial Vulnerability

    Jörn-Henrik Jacobsen, Jens Behrmann, Richard Zemel +1

    cs.LGcs.AIcs.CVarXiv:1811.00401v42018
  39. Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention

    Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh +2

    cs.CVarXiv:1908.07236v22019
  40. FENeRF: Face Editing in Neural Radiance Fields

    Jingxiang Sun, Xuan Wang, Yong Zhang +4

    cs.CVarXiv:2111.15490v22021
  41. Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations

    Wei-Jin Huang, Yue-Yi Zhang, Yi-Lin Wei +5

    cs.ROcs.AIcs.CVarXiv:2601.09518v12026
  42. ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries

    Junru Gu, Chenxu Hu, Tianyuan Zhang +4

    cs.CVcs.ROarXiv:2208.01582v32022
  43. MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

    Xinchen Yan, Akash Rastogi, Ruben Villegas +5

    cs.LGcs.AIcs.CVarXiv:1808.04545v12018
  44. Learning towards Minimum Hyperspherical Energy

    Weiyang Liu, Rongmei Lin, Zhen Liu +4

    cs.LGcs.CVstat.MLarXiv:1805.09298v92018
  45. N-Gram in Swin Transformers for Efficient Lightweight Image Super-Resolution

    Haram Choi, Jeongmin Lee, Jihoon Yang

    cs.CVarXiv:2211.11436v32022
  46. Guaranteed Tensor Recovery Fused Low-rankness and Smoothness

    Hailin Wang, Jiangjun Peng, Wenjin Qin +2

    cs.LGcs.AIcs.CVarXiv:2302.02155v12023
  47. Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

    Tianfei Zhou, Meijie Zhang, Fang Zhao +1

    cs.CVarXiv:2203.09653v22022
  48. Robust Online Matrix Factorization for Dynamic Background Subtraction

    Hongwei Yong, Deyu Meng, Wangmeng Zuo +1

    cs.CVarXiv:1705.10000v12017
  49. Liquid Structural State-Space Models

    Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang +3

    cs.LGcs.AIcs.CLarXiv:2209.12951v12022
  50. The Pros and Cons: Rank-aware Temporal Attention for Skill Determination in Long Videos

    Hazel Doughty, Walterio Mayol-Cuevas, Dima Damen

    cs.CVarXiv:1812.05538v22018
  51. Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection

    Yangchen Zeng, Zhenyu Yu, Dongming Jiang +5

    cs.CVarXiv:2604.15065v12026
  52. MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

    Fei Wang, Xingyu Fu, James Y. Huang +18

    cs.CVcs.AIcs.CLarXiv:2406.09411v22024
  53. Tora: Trajectory-oriented Diffusion Transformer for Video Generation

    Zhenghao Zhang, Junchao Liao, Menghao Li +5

    cs.CVarXiv:2407.21705v42024
  54. Visual Coreference Resolution in Visual Dialog using Neural Module Networks

    Satwik Kottur, José M. F. Moura, Devi Parikh +2

    cs.CVcs.AIcs.CLarXiv:1809.01816v12018
  55. MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features

    Shakti N. Wadekar, Abhishek Chaurasia

    cs.CVcs.AIcs.LGarXiv:2209.15159v22022
  56. Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking

    Sarthak Sharma, Junaid Ahmed Ansari, J. Krishna Murthy +1

    cs.ROcs.CVarXiv:1802.09298v22018
  57. Asynchronous Temporal Fields for Action Recognition

    Gunnar A. Sigurdsson, Santosh Divvala, Ali Farhadi +1

    cs.CVarXiv:1612.06371v22016
  58. On Attention Models for Human Activity Recognition

    Vishvak S Murahari, Thomas Ploetz

    cs.CVcs.AIcs.LGarXiv:1805.07648v12018
  59. AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning

    Madeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh Agrawala

    cs.CVcs.CLarXiv:2103.16002v12021
  60. ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation

    Siyuan Qiao, Yukun Zhu, Hartwig Adam +2

    cs.CVarXiv:2012.05258v12020