Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,621 to 1,680 of 18,782
Space-time Mixing Attention for Video Transformer
Adrian Bulat, Juan-Manuel Perez-Rua, Swathikiran Sudhakaran +2
cs.CVcs.AIcs.LGarXiv:2106.05968v22021TINYCD: A (Not So) Deep Learning Model For Change Detection
Andrea Codegoni, Gabriele Lombardi, Alessandro Ferrari
cs.CVcs.LGeess.IVarXiv:2207.13159v22022Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
Bo Yang, Sen Wang, Andrew Markham +1
cs.CVcs.AIcs.LGarXiv:1808.00758v22018Robot Navigation in Crowds by Graph Convolutional Networks with Attention Learned from Human Gaze
Yuying Chen, Congcong Liu, Ming Liu +1
cs.ROcs.AIcs.CVarXiv:1909.10400v12019VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Jiazheng Xu, Yu Huang, Jiale Cheng +19
cs.CVarXiv:2412.21059v42024DeepDRR -- A Catalyst for Machine Learning in Fluoroscopy-guided Procedures
Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee +4
physics.med-phcs.CVarXiv:1803.08606v12018StopThePop: Sorted Gaussian Splatting for View-Consistent Real-time Rendering
Lukas Radl, Michael Steiner, Mathias Parger +3
cs.GRcs.CVarXiv:2402.00525v32024Training CNNs with Low-Rank Filters for Efficient Image Classification
Yani Ioannou, Duncan Robertson, Jamie Shotton +2
cs.CVcs.LGcs.NEarXiv:1511.06744v32015Face Recognition Using Deep Multi-Pose Representations
Wael AbdAlmageed, Yue Wua, Stephen Rawlsa +9
cs.CVarXiv:1603.07388v12016PreDiff: Precipitation Nowcasting with Latent Diffusion Models
Zhihan Gao, Xingjian Shi, Boran Han +6
cs.LGcs.AIcs.CVarXiv:2307.10422v22023CubiCasa5K: A Dataset and an Improved Multi-Task Model for Floorplan Image Analysis
Ahti Kalervo, Juha Ylioinas, Markus Häikiö +2
cs.CVarXiv:1904.01920v12019TEACHTEXT: CrossModal Generalized Distillation for Text-Video Retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu +4
cs.CVarXiv:2104.08271v22021A multilevel thresholding algorithm using Electromagnetism Optimization
Diego Oliva, Erik Cuevas, Gonzalo Pajares +2
cs.CVarXiv:1406.6336v12014Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge Distillation
Bingchen Zhao, Kai Han
cs.CVarXiv:2107.03358v22021MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving
Jiale Li, Hang Dai, Hao Han +1
cs.CVarXiv:2303.08600v12023Beyond Physical Connections: Tree Models in Human Pose Estimation
Fang Wang, Yi Li
cs.CVarXiv:1305.2269v12013Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction
Zhizhong Han, Xiyang Wang, Yu-Shen Liu +1
cs.CVarXiv:1907.12704v12019Planar Prior Assisted PatchMatch Multi-View Stereo
Qingshan Xu, Wenbing Tao
cs.CVarXiv:1912.11744v12019Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta +1
cs.ROcs.CVarXiv:2405.01527v22024A Closer Look at the Explainability of Contrastive Language-Image Pre-training
Yi Li, Hualiang Wang, Yiqun Duan +2
cs.CVarXiv:2304.05653v22023Total Denoising: Unsupervised Learning of 3D Point Cloud Cleaning
Pedro Hermosilla, Tobias Ritschel, Timo Ropinski
cs.CVcs.GRarXiv:1904.07615v22019Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator
Jeff Zhang, Tianyu Gu, Kanad Basu +1
cs.LGcs.ARcs.CVarXiv:1802.04657v22018VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
Xiang Li, Jian Ding, Mohamed Elhoseiny
cs.CVarXiv:2406.12384v22024Out-of-Distribution Detection for Generalized Zero-Shot Action Recognition
Devraj Mandal, Sanath Narayan, Saikumar Dwivedi +4
cs.CVarXiv:1904.08703v22019Improved Techniques for Training Adaptive Deep Networks
Hao Li, Hong Zhang, Xiaojuan Qi +2
cs.CVarXiv:1908.06294v12019Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model
Yuming Du, Robin Kips, Albert Pumarola +3
cs.CVarXiv:2304.08577v12023Low-Resolution Face Recognition
Zhiyi Cheng, Xiatian Zhu, Shaogang Gong
cs.CVarXiv:1811.08965v22018GlyphControl: Glyph Conditional Control for Visual Text Generation
Yukang Yang, Dongnan Gui, Yuhui Yuan +4
cs.CVarXiv:2305.18259v22023CLIP-Count: Towards Text-Guided Zero-Shot Object Counting
Ruixiang Jiang, Lingbo Liu, Changwen Chen
cs.CVcs.AIarXiv:2305.07304v22023Filmy Cloud Removal on Satellite Imagery with Multispectral Conditional Generative Adversarial Nets
Kenji Enomoto, Ken Sakurada, Weimin Wang +4
cs.CVarXiv:1710.04835v12017CovidAID: COVID-19 Detection Using Chest X-Ray
Arpan Mangal, Surya Kalia, Harish Rajgopal +4
eess.IVcs.CVcs.LGarXiv:2004.09803v12020Joint-task Self-supervised Learning for Temporal Correspondence
Xueting Li, Sifei Liu, Shalini De Mello +3
cs.CVarXiv:1909.11895v12019Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Di Hu, Rui Qian, Minyue Jiang +5
cs.CVcs.LGcs.MMarXiv:2010.05466v12020FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute Learning
Chenxu Zhang, Yifan Zhao, Yifei Huang +4
cs.CVarXiv:2108.07938v12021GRiT: A Generative Region-to-text Transformer for Object Understanding
Jialian Wu, Jianfeng Wang, Zhengyuan Yang +4
cs.CVarXiv:2212.00280v12022View-Structured Conformal Prediction for 3D Gaussian Splatting
Junzheng Chu, Bin Pan, Zhenwei Shi
cs.LGcs.CVarXiv:2609.10307v12026Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
Phuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis +4
cs.CVarXiv:2312.10671v32023Learning A Single Network for Scale-Arbitrary Super-Resolution
Longguang Wang, Yingqian Wang, Zaiping Lin +3
cs.CVarXiv:2004.03791v22020GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Drew A. Hudson, Christopher D. Manning
cs.CLcs.AIcs.CVarXiv:1902.09506v32019Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
Fangyun Wei, Xiao Sun, Hongyang Li +2
cs.CVarXiv:2007.02846v42020Towards Flexible Blind JPEG Artifacts Removal
Jiaxi Jiang, Kai Zhang, Radu Timofte
eess.IVcs.CVarXiv:2109.14573v12021SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition
Mingxin Huang, Yuliang Liu, Zhenghao Peng +6
cs.CVarXiv:2203.10209v12022Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
Xuan Ju, Ailing Zeng, Yuxuan Bian +2
cs.CVarXiv:2310.01506v22023Text-Only Training for Image Captioning using Noise-Injected CLIP
David Nukrai, Ron Mokady, Amir Globerson
cs.CVcs.AIcs.LGarXiv:2211.00575v12022Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food
Quin Thames, Arjun Karpur, Wade Norris +4
cs.CVcs.LGarXiv:2103.03375v22021Feature Learning from Incomplete EEG with Denoising Autoencoder
Junhua Li, Zbigniew Struzik, Liqing Zhang +1
cs.CVq-bio.NCarXiv:1410.0818v12014Uncertainty-Informed Deep Learning Models Enable High-Confidence Predictions for Digital Histopathology
James M Dolezal, Andrew Srisuwananukorn, Dmitry Karpeyev +13
q-bio.QMcs.CVeess.IVarXiv:2204.04516v12022Curriculum Model Adaptation with Synthetic and Real Data for Semantic Foggy Scene Understanding
Dengxin Dai, Christos Sakaridis, Simon Hecker +1
cs.CVarXiv:1901.01415v22019Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers
Junhyeong Cho, Kim Youwang, Tae-Hyun Oh
cs.CVcs.AIcs.LGarXiv:2207.13820v12022MMM: Generative Masked Motion Model
Ekkasit Pinyoanuntapong, Pu Wang, Minwoo Lee +1
cs.CVcs.AIcs.LGarXiv:2312.03596v22023DigiFace-1M: 1 Million Digital Face Images for Face Recognition
Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis +5
cs.CVarXiv:2210.02579v12022Hashing on Nonlinear Manifolds
Fumin Shen, Chunhua Shen, Qinfeng Shi +3
cs.CVarXiv:1412.0826v12014Manipulation by Feel: Touch-Based Control with Deep Predictive Models
Stephen Tian, Frederik Ebert, Dinesh Jayaraman +4
cs.ROcs.AIcs.CVarXiv:1903.04128v12019Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
Yixin Zhu, Tao Gao, Lifeng Fan +9
cs.AIcs.CVcs.LGarXiv:2004.09044v12020Photorealistic Image Synthesis for Object Instance Detection
Tomas Hodan, Vibhav Vineet, Ran Gal +6
cs.CVcs.AIcs.ROarXiv:1902.03334v12019Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition
Peng Wang, Yuanzhouhan Cao, Chunhua Shen +2
cs.CVarXiv:1503.01224v22015Controlling Vision-Language Models for Multi-Task Image Restoration
Ziwei Luo, Fredrik K. Gustafsson, Zheng Zhao +2
cs.CVarXiv:2310.01018v22023Active Learning for Deep Detection Neural Networks
Hamed H. Aghdam, Abel Gonzalez-Garcia, Joost van de Weijer +1
cs.CVcs.LGarXiv:1911.09168v12019Unsupervised Domain Adaptation via Disentangled Representations: Application to Cross-Modality Liver Segmentation
Junlin Yang, Nicha C. Dvornek, Fan Zhang +3
eess.IVcs.CVarXiv:1907.13590v22019Newtonian Image Understanding: Unfolding the Dynamics of Objects in Static Images
Roozbeh Mottaghi, Hessam Bagherinezhad, Mohammad Rastegari +1
cs.CVarXiv:1511.04048v12015