Computer Vision and Pattern Recognition

Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

12,661 to 12,720 of 18,867

  1. Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

    Rohit Patel, Dieuwke Hupkes, Sloan Strader

    cs.CVcs.AIcs.MMarXiv:2608.26317v12026
  2. Consistent Video Depth Estimation

    Xuan Luo, Jia-Bin Huang, Richard Szeliski +2

    cs.CVarXiv:2004.15021v22020
  3. RON: Reverse Connection with Objectness Prior Networks for Object Detection

    Tao Kong, Fuchun Sun, Anbang Yao +3

    cs.CVarXiv:1707.01691v12017
  4. Clustered Object Detection in Aerial Images

    Fan Yang, Heng Fan, Peng Chu +2

    cs.CVarXiv:1904.08008v32019
  5. Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

    Chao Li, Cheng Deng, Ning Li +3

    cs.CVarXiv:1804.01223v12018
  6. EDGE: Editable Dance Generation From Music

    Jonathan Tseng, Rodrigo Castellon, C. Karen Liu

    cs.SDcs.CVcs.GRarXiv:2211.10658v22022
  7. A Survey of Deep Learning Techniques for Weed Detection from Images

    A S M Mahmudul Hasan, Ferdous Sohel, Dean Diepeveen +2

    cs.CVcs.LGarXiv:2103.01415v12021
  8. Global Second-order Pooling Convolutional Networks

    Zilin Gao, Jiangtao Xie, Qilong Wang +1

    cs.CVarXiv:1811.12006v22018
  9. Skeleton-based Action Recognition with Convolutional Neural Networks

    Chao Li, Qiaoyong Zhong, Di Xie +1

    cs.CVarXiv:1704.07595v12017
  10. Understanding and Robustifying Differentiable Architecture Search

    Arber Zela, Thomas Elsken, Tonmoy Saikia +3

    cs.LGcs.AIcs.CVarXiv:1909.09656v22019
  11. Data Augmentation Can Improve Robustness

    Sylvestre-Alvise Rebuffi, Sven Gowal, Dan A. Calian +3

    cs.CVcs.LGstat.MLarXiv:2111.05328v12021
  12. Translating and Segmenting Multimodal Medical Volumes with Cycle- and Shape-Consistency Generative Adversarial Network

    Zizhao Zhang, Lin Yang, Yefeng Zheng

    cs.CVarXiv:1802.09655v22018
  13. Who Remains, What Changes: Identity Anchored Composed Gait Retrieval

    Jingchen Fei, Zengbin Wang, Yukun Liu +3

    cs.CVarXiv:2608.26632v12026
  14. Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation

    YuXuan Liu, Abhishek Gupta, Pieter Abbeel +1

    cs.LGcs.AIcs.CVarXiv:1707.03374v22017
  15. A General Optimization-based Framework for Local Odometry Estimation with Multiple Sensors

    Tong Qin, Jie Pan, Shaozu Cao +1

    cs.CVarXiv:1901.03638v12019
  16. Peeking into the Future: Predicting Future Person Activities and Locations in Videos

    Junwei Liang, Lu Jiang, Juan Carlos Niebles +2

    cs.CVarXiv:1902.03748v32019
  17. Zero-Shot Object Detection

    Ankan Bansal, Karan Sikka, Gaurav Sharma +2

    cs.CVarXiv:1804.04340v22018
  18. Visual Classification via Description from Large Language Models

    Sachit Menon, Carl Vondrick

    cs.CVcs.LGarXiv:2210.07183v22022
  19. Effective Whole-body Pose Estimation with Two-stages Distillation

    Zhendong Yang, Ailing Zeng, Chun Yuan +1

    cs.CVarXiv:2307.15880v22023
  20. Learning Convolutional Networks for Content-weighted Image Compression

    Mu Li, Wangmeng Zuo, Shuhang Gu +2

    cs.CVarXiv:1703.10553v22017
  21. SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion

    Vikram Voleti, Chun-Han Yao, Mark Boss +6

    cs.CVarXiv:2403.12008v12024
  22. Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals Measurement

    Xin Liu, Josh Fromm, Shwetak Patel +1

    eess.SPcs.CVeess.IVarXiv:2006.03790v22020
  23. LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

    Yushe Cao, Shikun Feng, Ruxiang Duan +4

    cs.CVcs.AIarXiv:2608.26714v12026
  24. Image and Video Compression with Neural Networks: A Review

    Siwei Ma, Xinfeng Zhang, Chuanmin Jia +3

    cs.CVarXiv:1904.03567v22019
  25. KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations

    Chenchen Ge, Hanwen Shen, Bowen Jing +6

    cs.CVcs.AIarXiv:2608.27365v12026
  26. Is Robustness the Cost of Accuracy? -- A Comprehensive Study on the Robustness of 18 Deep Image Classification Models

    Dong Su, Huan Zhang, Hongge Chen +3

    cs.CVarXiv:1808.01688v22018
  27. ConceptFusion: Open-set Multimodal 3D Mapping

    Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu +14

    cs.CVcs.AIcs.ROarXiv:2302.07241v32023
  28. Contextual Action Recognition with R*CNN

    Georgia Gkioxari, Ross Girshick, Jitendra Malik

    cs.CVarXiv:1505.01197v32015
  29. RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network

    Bing Han, Gopalakrishnan Srinivasan, Kaushik Roy

    cs.NEcs.CVcs.LGarXiv:2003.01811v22020
  30. Learning Spatiotemporal Features with 3D Convolutional Networks

    Du Tran, Lubomir Bourdev, Rob Fergus +2

    cs.CVarXiv:1412.0767v42014
  31. Orthographic Feature Transform for Monocular 3D Object Detection

    Thomas Roddick, Alex Kendall, Roberto Cipolla

    cs.CVarXiv:1811.08188v12018
  32. Hyperbolic Image Embeddings

    Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova +2

    cs.CVcs.LGarXiv:1904.02239v22019
  33. InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation

    Xingchao Liu, Xiwen Zhang, Jianzhu Ma +2

    cs.LGcs.CVarXiv:2309.06380v22023
  34. DeepCache: Accelerating Diffusion Models for Free

    Xinyin Ma, Gongfan Fang, Xinchao Wang

    cs.CVcs.AIarXiv:2312.00858v22023
  35. ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

    Fei Yu, Jiji Tang, Weichong Yin +4

    cs.CVcs.CLarXiv:2006.16934v32020
  36. Track to Detect and Segment: An Online Multi-Object Tracker

    Jialian Wu, Jiale Cao, Liangchen Song +3

    cs.CVarXiv:2103.08808v12021
  37. MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

    Ajay Mandlekar, Soroush Nasiriany, Bowen Wen +5

    cs.ROcs.AIcs.CVarXiv:2310.17596v12023
  38. Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline

    Penghao Wu, Xiaosong Jia, Li Chen +3

    cs.CVcs.AIcs.ROarXiv:2206.08129v22022
  39. The Contextual Loss for Image Transformation with Non-Aligned Data

    Roey Mechrez, Itamar Talmi, Lihi Zelnik-Manor

    cs.CVcs.LGarXiv:1803.02077v42018
  40. Automated 2D and 3D Segmentation of AMD and DME Lesions in OCT

    Lucia Sundberg, Zhihao Zhao, M. Ali Nasseri

    cs.CVarXiv:2608.27095v12026
  41. Towards a Visual Privacy Advisor: Understanding and Predicting Privacy Risks in Images

    Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz

    cs.CVcs.CRcs.CYarXiv:1703.10660v22017
  42. Self-Supervised Representation Learning: Introduction, Advances and Challenges

    Linus Ericsson, Henry Gouk, Chen Change Loy +1

    cs.LGcs.CVstat.MLarXiv:2110.09327v12021
  43. MultiMAE: Multi-modal Multi-task Masked Autoencoders

    Roman Bachmann, David Mizrahi, Andrei Atanov +1

    cs.CVcs.LGarXiv:2204.01678v12022
  44. TANet: Robust 3D Object Detection from Point Clouds with Triple Attention

    Zhe Liu, Xin Zhao, Tengteng Huang +3

    cs.CVarXiv:1912.05163v12019
  45. Shallow Attention Network for Polyp Segmentation

    Jun Wei, Yiwen Hu, Ruimao Zhang +3

    cs.CVarXiv:2108.00882v12021
  46. Few-Shot Object Detection and Viewpoint Estimation for Objects in the Wild

    Yang Xiao, Vincent Lepetit, Renaud Marlet

    cs.CVarXiv:2007.12107v22020
  47. Adversarial Neuron Pruning Purifies Backdoored Deep Models

    Dongxian Wu, Yisen Wang

    cs.LGcs.CRcs.CVarXiv:2110.14430v12021
  48. Parameter Efficient Continual Learning for Sparse Event-Based Transformers

    Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur

    cs.CVarXiv:2608.26720v12026
  49. Towards Robust Blind Face Restoration with Codebook Lookup Transformer

    Shangchen Zhou, Kelvin C. K. Chan, Chongyi Li +1

    cs.CVarXiv:2206.11253v22022
  50. Classifying and Segmenting Microscopy Images Using Convolutional Multiple Instance Learning

    Oren Z. Kraus, Lei Jimmy Ba, Brendan Frey

    cs.CVq-bio.SCstat.MLarXiv:1511.05286v12015
  51. Exploiting Local Features from Deep Networks for Image Retrieval

    Joe Yue-Hei Ng, Fan Yang, Larry S. Davis

    cs.CVarXiv:1504.05133v22015
  52. You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection

    Yuxin Fang, Bencheng Liao, Xinggang Wang +5

    cs.CVcs.AIcs.LGarXiv:2106.00666v32021
  53. PointLLM: Empowering Large Language Models to Understand Point Clouds

    Runsen Xu, Xiaolong Wang, Tai Wang +3

    cs.CVcs.AIcs.CLarXiv:2308.16911v32023
  54. 6-DoF Object Pose from Semantic Keypoints

    Georgios Pavlakos, Xiaowei Zhou, Aaron Chan +2

    cs.CVcs.ROarXiv:1703.04670v12017
  55. Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

    Haodong Li, Shaoteng Liu, Tianyu Wang +7

    cs.CVarXiv:2608.09926v12026
  56. Lung Infection Quantification of COVID-19 in CT Images with Deep Learning

    Fei Shan, Yaozong Gao, Jun Wang +6

    cs.CVeess.IVq-bio.QMarXiv:2003.04655v32020
  57. Predicting Cardiovascular Risk Factors from Retinal Fundus Photographs using Deep Learning

    Ryan Poplin, Avinash V. Varadarajan, Katy Blumer +5

    cs.CVarXiv:1708.09843v22017
  58. Image reconstruction by domain transform manifold learning

    Bo Zhu, Jeremiah Z. Liu, Bruce R. Rosen +1

    cs.CVarXiv:1704.08841v12017
  59. AiATrack: Attention in Attention for Transformer Visual Tracking

    Shenyuan Gao, Chunluan Zhou, Chao Ma +2

    cs.CVarXiv:2207.09603v22022
  60. High-Quality Self-Supervised Deep Image Denoising

    Samuli Laine, Tero Karras, Jaakko Lehtinen +1

    cs.LGcs.CVcs.NEarXiv:1901.10277v32019