Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
12,661 to 12,720 of 18,867
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
Rohit Patel, Dieuwke Hupkes, Sloan Strader
cs.CVcs.AIcs.MMarXiv:2608.26317v12026Consistent Video Depth Estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski +2
cs.CVarXiv:2004.15021v22020RON: Reverse Connection with Objectness Prior Networks for Object Detection
Tao Kong, Fuchun Sun, Anbang Yao +3
cs.CVarXiv:1707.01691v12017Clustered Object Detection in Aerial Images
Fan Yang, Heng Fan, Peng Chu +2
cs.CVarXiv:1904.08008v32019Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval
Chao Li, Cheng Deng, Ning Li +3
cs.CVarXiv:1804.01223v12018EDGE: Editable Dance Generation From Music
Jonathan Tseng, Rodrigo Castellon, C. Karen Liu
cs.SDcs.CVcs.GRarXiv:2211.10658v22022A Survey of Deep Learning Techniques for Weed Detection from Images
A S M Mahmudul Hasan, Ferdous Sohel, Dean Diepeveen +2
cs.CVcs.LGarXiv:2103.01415v12021Global Second-order Pooling Convolutional Networks
Zilin Gao, Jiangtao Xie, Qilong Wang +1
cs.CVarXiv:1811.12006v22018Skeleton-based Action Recognition with Convolutional Neural Networks
Chao Li, Qiaoyong Zhong, Di Xie +1
cs.CVarXiv:1704.07595v12017Understanding and Robustifying Differentiable Architecture Search
Arber Zela, Thomas Elsken, Tonmoy Saikia +3
cs.LGcs.AIcs.CVarXiv:1909.09656v22019Data Augmentation Can Improve Robustness
Sylvestre-Alvise Rebuffi, Sven Gowal, Dan A. Calian +3
cs.CVcs.LGstat.MLarXiv:2111.05328v12021Translating and Segmenting Multimodal Medical Volumes with Cycle- and Shape-Consistency Generative Adversarial Network
Zizhao Zhang, Lin Yang, Yefeng Zheng
cs.CVarXiv:1802.09655v22018Who Remains, What Changes: Identity Anchored Composed Gait Retrieval
Jingchen Fei, Zengbin Wang, Yukun Liu +3
cs.CVarXiv:2608.26632v12026Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel +1
cs.LGcs.AIcs.CVarXiv:1707.03374v22017A General Optimization-based Framework for Local Odometry Estimation with Multiple Sensors
Tong Qin, Jie Pan, Shaozu Cao +1
cs.CVarXiv:1901.03638v12019Peeking into the Future: Predicting Future Person Activities and Locations in Videos
Junwei Liang, Lu Jiang, Juan Carlos Niebles +2
cs.CVarXiv:1902.03748v32019Zero-Shot Object Detection
Ankan Bansal, Karan Sikka, Gaurav Sharma +2
cs.CVarXiv:1804.04340v22018Visual Classification via Description from Large Language Models
Sachit Menon, Carl Vondrick
cs.CVcs.LGarXiv:2210.07183v22022Effective Whole-body Pose Estimation with Two-stages Distillation
Zhendong Yang, Ailing Zeng, Chun Yuan +1
cs.CVarXiv:2307.15880v22023Learning Convolutional Networks for Content-weighted Image Compression
Mu Li, Wangmeng Zuo, Shuhang Gu +2
cs.CVarXiv:1703.10553v22017SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
Vikram Voleti, Chun-Han Yao, Mark Boss +6
cs.CVarXiv:2403.12008v12024Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals Measurement
Xin Liu, Josh Fromm, Shwetak Patel +1
eess.SPcs.CVeess.IVarXiv:2006.03790v22020LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
Yushe Cao, Shikun Feng, Ruxiang Duan +4
cs.CVcs.AIarXiv:2608.26714v12026Image and Video Compression with Neural Networks: A Review
Siwei Ma, Xinfeng Zhang, Chuanmin Jia +3
cs.CVarXiv:1904.03567v22019KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations
Chenchen Ge, Hanwen Shen, Bowen Jing +6
cs.CVcs.AIarXiv:2608.27365v12026Is Robustness the Cost of Accuracy? -- A Comprehensive Study on the Robustness of 18 Deep Image Classification Models
Dong Su, Huan Zhang, Hongge Chen +3
cs.CVarXiv:1808.01688v22018ConceptFusion: Open-set Multimodal 3D Mapping
Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu +14
cs.CVcs.AIcs.ROarXiv:2302.07241v32023Contextual Action Recognition with R*CNN
Georgia Gkioxari, Ross Girshick, Jitendra Malik
cs.CVarXiv:1505.01197v32015RMP-SNN: Residual Membrane Potential Neuron for Enabling Deeper High-Accuracy and Low-Latency Spiking Neural Network
Bing Han, Gopalakrishnan Srinivasan, Kaushik Roy
cs.NEcs.CVcs.LGarXiv:2003.01811v22020Learning Spatiotemporal Features with 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus +2
cs.CVarXiv:1412.0767v42014Orthographic Feature Transform for Monocular 3D Object Detection
Thomas Roddick, Alex Kendall, Roberto Cipolla
cs.CVarXiv:1811.08188v12018Hyperbolic Image Embeddings
Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova +2
cs.CVcs.LGarXiv:1904.02239v22019InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
Xingchao Liu, Xiwen Zhang, Jianzhu Ma +2
cs.LGcs.CVarXiv:2309.06380v22023DeepCache: Accelerating Diffusion Models for Free
Xinyin Ma, Gongfan Fang, Xinchao Wang
cs.CVcs.AIarXiv:2312.00858v22023ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Fei Yu, Jiji Tang, Weichong Yin +4
cs.CVcs.CLarXiv:2006.16934v32020Track to Detect and Segment: An Online Multi-Object Tracker
Jialian Wu, Jiale Cao, Liangchen Song +3
cs.CVarXiv:2103.08808v12021MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
Ajay Mandlekar, Soroush Nasiriany, Bowen Wen +5
cs.ROcs.AIcs.CVarXiv:2310.17596v12023Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
Penghao Wu, Xiaosong Jia, Li Chen +3
cs.CVcs.AIcs.ROarXiv:2206.08129v22022The Contextual Loss for Image Transformation with Non-Aligned Data
Roey Mechrez, Itamar Talmi, Lihi Zelnik-Manor
cs.CVcs.LGarXiv:1803.02077v42018Automated 2D and 3D Segmentation of AMD and DME Lesions in OCT
Lucia Sundberg, Zhihao Zhao, M. Ali Nasseri
cs.CVarXiv:2608.27095v12026Towards a Visual Privacy Advisor: Understanding and Predicting Privacy Risks in Images
Tribhuvanesh Orekondy, Bernt Schiele, Mario Fritz
cs.CVcs.CRcs.CYarXiv:1703.10660v22017Self-Supervised Representation Learning: Introduction, Advances and Challenges
Linus Ericsson, Henry Gouk, Chen Change Loy +1
cs.LGcs.CVstat.MLarXiv:2110.09327v12021MultiMAE: Multi-modal Multi-task Masked Autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov +1
cs.CVcs.LGarXiv:2204.01678v12022TANet: Robust 3D Object Detection from Point Clouds with Triple Attention
Zhe Liu, Xin Zhao, Tengteng Huang +3
cs.CVarXiv:1912.05163v12019Shallow Attention Network for Polyp Segmentation
Jun Wei, Yiwen Hu, Ruimao Zhang +3
cs.CVarXiv:2108.00882v12021Few-Shot Object Detection and Viewpoint Estimation for Objects in the Wild
Yang Xiao, Vincent Lepetit, Renaud Marlet
cs.CVarXiv:2007.12107v22020Adversarial Neuron Pruning Purifies Backdoored Deep Models
Dongxian Wu, Yisen Wang
cs.LGcs.CRcs.CVarXiv:2110.14430v12021Parameter Efficient Continual Learning for Sparse Event-Based Transformers
Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur
cs.CVarXiv:2608.26720v12026Towards Robust Blind Face Restoration with Codebook Lookup Transformer
Shangchen Zhou, Kelvin C. K. Chan, Chongyi Li +1
cs.CVarXiv:2206.11253v22022Classifying and Segmenting Microscopy Images Using Convolutional Multiple Instance Learning
Oren Z. Kraus, Lei Jimmy Ba, Brendan Frey
cs.CVq-bio.SCstat.MLarXiv:1511.05286v12015Exploiting Local Features from Deep Networks for Image Retrieval
Joe Yue-Hei Ng, Fan Yang, Larry S. Davis
cs.CVarXiv:1504.05133v22015You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
Yuxin Fang, Bencheng Liao, Xinggang Wang +5
cs.CVcs.AIcs.LGarXiv:2106.00666v32021PointLLM: Empowering Large Language Models to Understand Point Clouds
Runsen Xu, Xiaolong Wang, Tai Wang +3
cs.CVcs.AIcs.CLarXiv:2308.16911v320236-DoF Object Pose from Semantic Keypoints
Georgios Pavlakos, Xiaowei Zhou, Aaron Chan +2
cs.CVcs.ROarXiv:1703.04670v12017Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
Haodong Li, Shaoteng Liu, Tianyu Wang +7
cs.CVarXiv:2608.09926v12026Lung Infection Quantification of COVID-19 in CT Images with Deep Learning
Fei Shan, Yaozong Gao, Jun Wang +6
cs.CVeess.IVq-bio.QMarXiv:2003.04655v32020Predicting Cardiovascular Risk Factors from Retinal Fundus Photographs using Deep Learning
Ryan Poplin, Avinash V. Varadarajan, Katy Blumer +5
cs.CVarXiv:1708.09843v22017Image reconstruction by domain transform manifold learning
Bo Zhu, Jeremiah Z. Liu, Bruce R. Rosen +1
cs.CVarXiv:1704.08841v12017AiATrack: Attention in Attention for Transformer Visual Tracking
Shenyuan Gao, Chunluan Zhou, Chao Ma +2
cs.CVarXiv:2207.09603v22022High-Quality Self-Supervised Deep Image Denoising
Samuli Laine, Tero Karras, Jaakko Lehtinen +1
cs.LGcs.CVcs.NEarXiv:1901.10277v32019