Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,521 to 2,580 of 18,795
Point Cloud Oversegmentation with Graph-Structured Deep Metric Learning
Loic Landrieu, Mohamed Boussaha
cs.CVcs.LGarXiv:1904.02113v12019SALD: Sign Agnostic Learning with Derivatives
Matan Atzmon, Yaron Lipman
cs.CVcs.GRcs.LGarXiv:2006.05400v22020Accurate Image Restoration with Attention Retractable Transformer
Jiale Zhang, Yulun Zhang, Jinjin Gu +3
cs.CVarXiv:2210.01427v42022Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
Masaki Saito, Shunta Saito, Masanori Koyama +1
cs.CVarXiv:1811.09245v22018LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose Estimation
Keunhong Park, Arsalan Mousavian, Yu Xiang +1
cs.CVcs.GRcs.ROarXiv:1912.00416v32019Context-aware Feature Generation for Zero-shot Semantic Segmentation
Zhangxuan Gu, Siyuan Zhou, Li Niu +2
cs.CVarXiv:2008.06893v12020Pixtral 12B
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +39
cs.CVcs.CLarXiv:2410.07073v22024Selective Sensor Fusion for Neural Visual-Inertial Odometry
Changhao Chen, Stefano Rosa, Yishu Miao +4
cs.CVcs.AIcs.LGarXiv:1903.01534v12019Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
Jiaxi Gu, Xiaojun Meng, Guansong Lu +9
cs.CVcs.LGarXiv:2202.06767v42022Multi-Image Matching via Fast Alternating Minimization
Xiaowei Zhou, Menglong Zhu, Kostas Daniilidis
cs.CVarXiv:1505.04845v22015Document Image Binarization with Fully Convolutional Neural Networks
Chris Tensmeyer, Tony Martinez
cs.CVarXiv:1708.03276v12017Perceptual Image Quality Assessment with Transformers
Manri Cheon, Sung-Jun Yoon, Byungyeon Kang +1
cs.CVeess.IVarXiv:2104.14730v22021Panoptic-PolarNet: Proposal-free LiDAR Point Cloud Panoptic Segmentation
Zixiang Zhou, Yang Zhang, Hassan Foroosh
cs.CVarXiv:2103.14962v12021Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
Chenglei Si, Yanzhe Zhang, Ryan Li +3
cs.CLcs.CVcs.CYarXiv:2403.03163v32024Learned Image Downscaling for Upscaling using Content Adaptive Resampler
Wanjie Sun, Zhenzhong Chen
cs.CVcs.MMeess.IVarXiv:1907.12904v22019Few-shot Scene-adaptive Anomaly Detection
Yiwei Lu, Frank Yu, Mahesh Kumar Krishna Reddy +1
cs.CVcs.LGarXiv:2007.07843v12020Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth
Doyeon Kim, Woonghyun Ka, Pyungwhan Ahn +3
cs.CVarXiv:2201.07436v32022Image Synthesis From Reconfigurable Layout and Style
Wei Sun, Tianfu Wu
cs.CVstat.MLarXiv:1908.07500v12019Dynamic Spatial Propagation Network for Depth Completion
Yuankai Lin, Tao Cheng, Qi Zhong +2
cs.CVarXiv:2202.09769v12022NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media
Grace Luo, Trevor Darrell, Anna Rohrbach
cs.CVcs.CLarXiv:2104.05893v22021Single Image Reflection Removal through Cascaded Refinement
Chao Li, Yixiao Yang, Kun He +2
cs.CVarXiv:1911.06634v22019Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
Lin Wu, Yang Wang, Junbin Gao +1
cs.CVarXiv:1706.03160v22017Unsupervised Deep Learning for Structured Shape Matching
Jean-Michel Roufosse, Abhishek Sharma, Maks Ovsjanikov
cs.GRcs.CVarXiv:1812.03794v32018SFNet: Learning Object-aware Semantic Correspondence
Junghyup Lee, Dohyung Kim, Jean Ponce +1
cs.CVarXiv:1904.01810v22019Action Unit Detection with Region Adaptation, Multi-labeling Learning and Optimal Temporal Fusing
Wei Li, Farnaz Abitahi, Zhigang Zhu
cs.CVarXiv:1704.03067v12017Learning Open-World Object Proposals without Learning to Classify
Dahun Kim, Tsung-Yi Lin, Anelia Angelova +2
cs.CVarXiv:2108.06753v12021Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
Ning Zhang, Manohar Paluri, Yaniv Taigman +2
cs.CVarXiv:1501.05703v22015MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Ho Kei Cheng, Masato Ishii, Akio Hayakawa +3
cs.CVcs.LGcs.SDarXiv:2412.15322v22024Understanding Zero-Shot Adversarial Robustness for Large-Scale Models
Chengzhi Mao, Scott Geng, Junfeng Yang +2
cs.CVarXiv:2212.07016v22022Explainable Deep Learning Methods in Medical Image Classification: A Survey
Cristiano Patrício, João C. Neves, Luís F. Teixeira
eess.IVcs.AIcs.CVarXiv:2205.04766v32022Spectral DiffuserCam: lensless snapshot hyperspectral imaging with a spectral filter array
Kristina Monakhova, Kyrollos Yanny, Neerja Aggarwal +1
eess.IVcs.CVphysics.opticsarXiv:2006.08565v22020Vehicle Instance Segmentation from Aerial Image and Video Using a Multi-Task Learning Residual Fully Convolutional Network
Lichao Mou, Xiao Xiang Zhu
cs.CVarXiv:1805.10485v12018Estimating and Exploiting the Aleatoric Uncertainty in Surface Normal Estimation
Gwangbin Bae, Ignas Budvytis, Roberto Cipolla
cs.CVarXiv:2109.09881v12021Instance Localization for Self-supervised Detection Pretraining
Ceyuan Yang, Zhirong Wu, Bolei Zhou +1
cs.CVarXiv:2102.08318v22021AttentionNet: Aggregating Weak Directions for Accurate Object Detection
Donggeun Yoo, Sunggyun Park, Joon-Young Lee +2
cs.CVcs.LGarXiv:1506.07704v22015Deep Learning and Conditional Random Fields-based Depth Estimation and Topographical Reconstruction from Conventional Endoscopy
Faisal Mahmood, Nicholas J. Durr
cs.CVarXiv:1710.11216v32017Data augmentation instead of explicit regularization
Alex Hernández-García, Peter König
cs.CVarXiv:1806.03852v52018Robot Learning in Homes: Improving Generalization and Reducing Dataset Bias
Abhinav Gupta, Adithyavairavan Murali, Dhiraj Gandhi +1
cs.ROcs.AIcs.CVarXiv:1807.07049v12018CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations
Mohammadreza Zolfaghari, Yi Zhu, Peter Gehler +1
cs.CVcs.AIcs.LGarXiv:2109.14910v12021Mitosis domain generalization in histopathology images -- The MIDOG challenge
Marc Aubreville, Nikolas Stathonikos, Christof A. Bertram +32
eess.IVcs.CVphysics.med-pharXiv:2204.03742v12022SpectFormer: Frequency and Attention is what you need in a Vision Transformer
Badri N. Patro, Vinay P. Namboodiri, Vijay Srinivas Agneeswaran
cs.CVcs.AIcs.CLarXiv:2304.06446v22023A Survey of Stealth Malware: Attacks, Mitigation Measures, and Steps Toward Autonomous Open World Solutions
Ethan M. Rudd, Andras Rozsa, Manuel Günther +1
cs.CRcs.CVarXiv:1603.06028v22016Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev +2
cs.CVarXiv:2103.16553v12021Feature Pyramid Network for Multi-Class Land Segmentation
Selim S. Seferbekov, Vladimir I. Iglovikov, Alexander V. Buslaev +1
cs.CVarXiv:1806.03510v22018Neural Compatibility Modeling with Attentive Knowledge Distillation
Xuemeng Song, Fuli Feng, Xianjing Han +3
cs.CVcs.MMarXiv:1805.00313v12018Imposing Hard Constraints on Deep Networks: Promises and Limitations
Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua
cs.CVarXiv:1706.02025v12017Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2
cs.AIcs.CLcs.CVarXiv:2609.08025v12026Towards the Detection of Diffusion Model Deepfakes
Jonas Ricker, Simon Damm, Thorsten Holz +1
cs.CVarXiv:2210.14571v42022ManipulaTHOR: A Framework for Visual Object Manipulation
Kiana Ehsani, Winson Han, Alvaro Herrasti +5
cs.CVcs.AIcs.LGarXiv:2104.11213v12021RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
Diwas Lamsal, Juha Carlon, Reinhard Claeys +8
cs.AIcs.CVarXiv:2609.08090v12026Unsupervised Change Detection in Multi-temporal VHR Images Based on Deep Kernel PCA Convolutional Mapping Network
Chen Wu, Hongruixuan Chen, Bo Do +1
eess.IVcs.CVarXiv:1912.08628v12019Review of Deep Learning
Rong Zhang, Weiping Li, Tong Mo
cs.LGcs.CVcs.NEarXiv:1804.01653v22018Diffusion-SDF: Text-to-Shape via Voxelized Diffusion
Muheng Li, Yueqi Duan, Jie Zhou +1
cs.CVcs.AIcs.GRarXiv:2212.03293v22022Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao +3
cs.CVarXiv:1812.02347v32018Interpretable and Accurate Fine-grained Recognition via Region Grouping
Zixuan Huang, Yin Li
cs.CVcs.AIcs.LGarXiv:2005.10411v12020BEVBert: Multimodal Map Pre-training for Language-guided Navigation
Dong An, Yuankai Qi, Yangguang Li +4
cs.CVcs.AIcs.CLarXiv:2212.04385v22022What is a salient object? A dataset and a baseline model for salient object detection
Ali Borji
cs.CVarXiv:1412.5027v12014Uncovering convolutional neural network decisions for diagnosing multiple sclerosis on conventional MRI using layer-wise relevance propagation
Fabian Eitel, Emily Soehler, Judith Bellmann-Strobl +10
cs.CVarXiv:1904.08771v12019OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Raghav Kapoor, Yash Parag Butala, Melisa Russak +4
cs.AIcs.CLcs.CVarXiv:2402.17553v32024Towards Transferable Adversarial Attacks on Vision Transformers
Zhipeng Wei, Jingjing Chen, Micah Goldblum +3
cs.CVcs.AIarXiv:2109.04176v32021