Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,881 to 2,940 of 18,815
Fashion Forward: Forecasting Visual Style in Fashion
Ziad Al-Halah, Rainer Stiefelhagen, Kristen Grauman
cs.CVarXiv:1705.06394v32017A Theoretical Explanation for Perplexing Behaviors of Backpropagation-based Visualizations
Weili Nie, Yang Zhang, Ankit Patel
cs.CVcs.AIarXiv:1805.07039v42018Channel Splitting Network for Single MR Image Super-Resolution
Xiaole Zhao, Yulun Zhang, Tao Zhang +1
cs.CVarXiv:1810.06453v32018COVID_MTNet: COVID-19 Detection with Multi-Task Deep Learning Approaches
Md Zahangir Alom, M M Shaifur Rahman, Mst Shamima Nasrin +2
eess.IVcs.CVcs.LGarXiv:2004.03747v32020Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
Fengyu Yang, Chao Feng, Ziyang Chen +8
cs.CVcs.ROarXiv:2401.18084v12024Discrimination-aware Network Pruning for Deep Model Compression
Jing Liu, Bohan Zhuang, Zhuangwei Zhuang +4
cs.CVarXiv:2001.01050v22020Experiments of Federated Learning for COVID-19 Chest X-ray Images
Boyi Liu, Bingjie Yan, Yize Zhou +2
eess.IVcs.CVcs.LGarXiv:2007.05592v12020Contrastive Boundary Learning for Point Cloud Segmentation
Liyao Tang, Yibing Zhan, Zhe Chen +2
cs.CVarXiv:2203.05272v22022Semi-Supervised Graph Classification: A Hierarchical Graph Perspective
Jia Li, Yu Rong, Hong Cheng +3
cs.CVcs.LGarXiv:1904.05003v12019TextMesh: Generation of Realistic 3D Meshes From Text Prompts
Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni +2
cs.CVarXiv:2304.12439v12023SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Cheng Yin, Wang Xu, Junpeng Yang +8
cs.CVcs.LGcs.ROarXiv:2609.05533v12026Exponential Moving Average Normalization for Self-supervised and Semi-supervised Learning
Zhaowei Cai, Avinash Ravichandran, Subhransu Maji +3
cs.LGcs.AIcs.CVarXiv:2101.08482v22021FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation
Xiaoyu Shi, Zhaoyang Huang, Dasong Li +6
cs.CVarXiv:2303.01237v12023Mean Deviation Similarity Index: Efficient and Reliable Full-Reference Image Quality Evaluator
Hossein Ziaei Nafchi, Atena Shahkolaei, Rachid Hedjam +1
cs.CVarXiv:1608.07433v42016Scaling the Scattering Transform: Deep Hybrid Networks
Edouard Oyallon, Eugene Belilovsky, Sergey Zagoruyko
cs.CVcs.LGarXiv:1703.08961v22017Deep Model Intellectual Property Protection via Deep Watermarking
Jie Zhang, Dongdong Chen, Jing Liao +4
cs.CVcs.CRarXiv:2103.04980v12021Deep Attentive Tracking via Reciprocative Learning
Shi Pu, Yibing Song, Chao Ma +2
cs.CVarXiv:1810.03851v22018TAP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models
Mehedi Hasan, Ashfak Yeafi, Md Khairul Islam
cs.CVcs.AIarXiv:2609.04071v12026Object Detection for Graphical User Interface: Old Fashioned or Deep Learning or a Combination?
Jieshan Chen, Mulong Xie, Zhenchang Xing +4
cs.CVcs.HCcs.LGarXiv:2008.05132v22020End-to-end Prostate Cancer Detection in bpMRI via 3D CNNs: Effects of Attention Mechanisms, Clinical Priori and Decoupled False Positive Reduction
Anindo Saha, Matin Hosseinzadeh, Henkjan Huisman
eess.IVcs.CVarXiv:2101.03244v102021Long-tailed Visual Recognition via Gaussian Clouded Logit Adjustment
Mengke Li, Yiu-ming Cheung, Yang Lu
cs.CVarXiv:2305.11733v12023Complementary Patch for Weakly Supervised Semantic Segmentation
Fei Zhang, Chaochen Gu, Chenyue Zhang +1
cs.CVarXiv:2108.03852v12021RRNet: Relational Reasoning Network with Parallel Multi-scale Attention for Salient Object Detection in Optical Remote Sensing Images
Runmin Cong, Yumo Zhang, Leyuan Fang +3
cs.CVarXiv:2110.14223v22021Transformation-Equivariant 3D Object Detection for Autonomous Driving
Hai Wu, Chenglu Wen, Wei Li +3
cs.CVarXiv:2211.11962v32022KoDF: A Large-scale Korean DeepFake Detection Dataset
Patrick Kwon, Jaeseong You, Gyuhyeon Nam +2
cs.CVcs.LGarXiv:2103.10094v22021Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
Sambhav R. Jain, Albert Gural, Michael Wu +1
cs.CVcs.AIcs.LGarXiv:1903.08066v32019A Unified Framework for Multi-View Multi-Class Object Pose Estimation
Chi Li, Jin Bai, Gregory D. Hager
cs.CVarXiv:1803.08103v22018Causal Attention for Unbiased Visual Recognition
Tan Wang, Chang Zhou, Qianru Sun +1
cs.CVarXiv:2108.08782v12021RoomNet: End-to-End Room Layout Estimation
Chen-Yu Lee, Vijay Badrinarayanan, Tomasz Malisiewicz +1
cs.CVarXiv:1703.06241v22017The Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations
Emanuele Cardinale, Marco Proietti, Alessandro Cacciatore +3
cs.CVcs.AIarXiv:2609.04009v12026Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition
Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju +3
cs.CVcs.AIcs.ROarXiv:2609.03995v12026Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning
Samarth Tripathi, Sarthak Tripathi, Homayoon Beigi
cs.AIcs.CVcs.HCarXiv:1804.05788v32018MLIC: Multi-Reference Entropy Model for Learned Image Compression
Wei Jiang, Jiayu Yang, Yongqi Zhai +3
eess.IVcs.CVarXiv:2211.07273v102022Implicit Style-Content Separation using B-LoRA
Yarden Frenkel, Yael Vinker, Ariel Shamir +1
cs.CVarXiv:2403.14572v22024Learning to Generate Text-grounded Mask for Open-world Semantic Segmentation from Only Image-Text Pairs
Junbum Cha, Jonghwan Mun, Byungseok Roh
cs.CVarXiv:2212.00785v22022RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting
Tomas Guija-Valiente, Blanca Rodriguez-Gonzalez, Norberto Malpica +1
cs.CVcs.AIcs.LGarXiv:2609.03956v12026Translating Images into Maps
Avishkar Saha, Oscar Mendez Maldonado, Chris Russell +1
cs.CVarXiv:2110.00966v22021Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement
Jiaxiang Tang, Hang Zhou, Xiaokang Chen +4
cs.CVarXiv:2303.02091v22023Maximum-Entropy Fine-Grained Classification
Abhimanyu Dubey, Otkrist Gupta, Ramesh Raskar +1
cs.CVcs.LGarXiv:1809.05934v22018Towards real-time unsupervised monocular depth estimation on CPU
Matteo Poggi, Filippo Aleotti, Fabio Tosi +1
cs.CVcs.ROarXiv:1806.11430v32018GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs
Junqing Du, Fernando Ropero, Erkin Turkoz +2
cs.CVcs.AIcs.ROarXiv:2609.03892v12026Reducing Information Bottleneck for Weakly Supervised Semantic Segmentation
Jungbeom Lee, Jooyoung Choi, Jisoo Mok +1
cs.CVcs.LGarXiv:2110.06530v12021Boundary and Entropy-driven Adversarial Learning for Fundus Image Segmentation
Shujun Wang, Lequan Yu, Kang Li +3
cs.CVarXiv:1906.11143v22019The impact of phase information for few-shot fine-grained image classification
Ruiling Liu, Linyue Zhang, Wenyi Zeng +5
cs.CVcs.AIarXiv:2609.03829v12026Sim2real transfer learning for 3D human pose estimation: motion to the rescue
Carl Doersch, Andrew Zisserman
cs.CVarXiv:1907.02499v22019Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography
Louis Chen, Torbjörn E. M. Nordling
cs.CVcs.AIeess.IVarXiv:2609.03663v12026EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu +6
cs.CVcs.AIcs.LGarXiv:2209.13064v12022Robust Camera Location Estimation by Convex Programming
Onur Ozyesil, Amit Singer
cs.CVarXiv:1412.0165v22014Gaze Estimation using Transformer
Yihua Cheng, Feng Lu
cs.CVarXiv:2105.14424v12021A Vision-based Social Distancing and Critical Density Detection System for COVID-19
Dongfang Yang, Ekim Yurtsever, Vishnu Renganathan +2
eess.IVcs.CVarXiv:2007.03578v22020Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving
Zhenxun Yuan, Xiao Song, Lei Bai +3
cs.CVarXiv:2011.13628v12020On Low-Resolution Face Recognition in the Wild: Comparisons and New Techniques
Pei Li, Loreto Prieto, Domingo Mery +1
cs.CVarXiv:1805.11529v22018Deep Model-Based 6D Pose Refinement in RGB
Fabian Manhardt, Wadim Kehl, Nassir Navab +1
cs.CVarXiv:1810.03065v12018Multichannel Compressive Sensing MRI Using Noiselet Encoding
Kamlesh Pawar, Gary F. Egan, Jingxin Zhang
physics.med-phcs.CVarXiv:1407.5536v22014Automatic Face Reenactment
Pablo Garrido, Levi Valgaerts, Ole Rehmsen +3
cs.CVcs.GRarXiv:1602.02651v12016Make Pixels Dance: High-Dynamic Video Generation
Yan Zeng, Guoqiang Wei, Jiani Zheng +4
cs.CVarXiv:2311.10982v12023CPTR: Full Transformer Network for Image Captioning
Wei Liu, Sihan Chen, Longteng Guo +2
cs.CVarXiv:2101.10804v320213D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information
Sungheon Park, Jihye Hwang, Nojun Kwak
cs.CVarXiv:1608.03075v22016Convolutional Neural Network-Based Image Representation for Visual Loop Closure Detection
Yi Hou, Hong Zhang, Shilin Zhou
cs.ROcs.CVarXiv:1504.05241v12015Generative Models from the perspective of Continual Learning
Timothée Lesort, Hugo Caselles-Dupré, Michael Garcia-Ortiz +2
cs.LGcs.AIcs.CVarXiv:1812.09111v12018