Computer Vision and Pattern Recognition
Papers filed under cs.CV on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,181 to 9,240 of 18,830
Explainable Medical Imaging AI Needs Human-Centered Design: Guidelines and Evidence from a Systematic Review
Haomin Chen, Catalina Gomez, Chien-Ming Huang +1
cs.HCcs.CVcs.LGarXiv:2112.12596v42021Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis
Kai Zhang, Yawei Li, Jingyun Liang +6
cs.CVcs.GReess.IVarXiv:2203.13278v42022Parametric Multimodal User Memory: Storing What Captions Cannot Carry
Bojie Li, Noah Shi
cs.CLcs.AIcs.CVarXiv:2608.28609v12026Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks
Hyeonseob Nam, Hyo-Eun Kim
cs.CVarXiv:1805.07925v32018DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
Jiashu Zhu, Yanhao Zheng, Ruitian Tian +7
cs.CVcs.SDarXiv:2608.31106v12026Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition
Koray Kavukcuoglu, Marc'Aurelio Ranzato, Yann LeCun
cs.CVcs.LGarXiv:1010.3467v12010Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training
Haichao Zhang, Jianyu Wang
cs.CVcs.CRcs.LGarXiv:1907.10764v42019Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision
Kuan Fang, Yuke Zhu, Animesh Garg +4
cs.ROcs.CVcs.LGarXiv:1806.09266v12018Image Deformation Meta-Networks for One-Shot Learning
Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3
cs.CVarXiv:1905.11641v22019Shield: Fast, Practical Defense and Vaccination for Deep Learning using JPEG Compression
Nilaksh Das, Madhuri Shanbhogue, Shang-Tse Chen +5
cs.CVcs.AIcs.CRarXiv:1802.06816v12018Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models
Guillermo Ortiz-Jimenez, Alessandro Favero, Pascal Frossard
cs.LGcs.CVarXiv:2305.12827v32023Unsupervised Domain Adaptation through Self-Supervision
Yu Sun, Eric Tzeng, Trevor Darrell +1
cs.LGcs.CVstat.MLarXiv:1909.11825v22019HyperSeg: Patch-wise Hypernetwork for Real-time Semantic Segmentation
Yuval Nirkin, Lior Wolf, Tal Hassner
cs.CVarXiv:2012.11582v22020AdvSim: Generating Safety-Critical Scenarios for Self-Driving Vehicles
Jingkang Wang, Ava Pun, James Tu +5
cs.ROcs.AIcs.CVarXiv:2101.06549v42021Multispectral and Hyperspectral Image Fusion by MS/HS Fusion Net
Qi Xie, Minghao Zhou, Qian Zhao +3
cs.CVarXiv:1901.03281v120193D Gaussian Splatting as Markov Chain Monte Carlo
Shakiba Kheradmand, Daniel Rebain, Gopal Sharma +6
cs.CVarXiv:2404.09591v32024FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model Extraction
Samiul Alam, Luyang Liu, Ming Yan +1
cs.LGcs.CRcs.CVarXiv:2212.01548v22022BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu +6
cs.CVarXiv:2608.31113v12026Adversarial Unlearning of Backdoors via Implicit Hypergradient
Yi Zeng, Si Chen, Won Park +3
cs.LGcs.CRcs.CVarXiv:2110.03735v42021DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms
Hua Qi, Qing Guo, Felix Juefei-Xu +5
cs.CVarXiv:2006.07634v22020ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
Baoguang Shi, Cong Yao, Minghui Liao +6
cs.CVarXiv:1708.09585v32017Unsupervised Person Re-identification by Deep Learning Tracklet Association
Minxian Li, Xiatian Zhu, Shaogang Gong
cs.CVarXiv:1809.02874v12018Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Zhenxin Li, Kailin Li, Shihao Wang +9
cs.CVarXiv:2406.06978v42024Cross View Fusion for 3D Human Pose Estimation
Haibo Qiu, Chunyu Wang, Jingdong Wang +2
cs.CVarXiv:1909.01203v12019Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection Generalization
Shichao Dong, Jin Wang, Renhe Ji +3
cs.CVarXiv:2210.14457v22022Practical Full Resolution Learned Lossless Image Compression
Fabian Mentzer, Eirikur Agustsson, Michael Tschannen +2
eess.IVcs.CVcs.LGarXiv:1811.12817v32018BioCLIP: A Vision Foundation Model for the Tree of Life
Samuel Stevens, Jiaman Wu, Matthew J Thompson +9
cs.CVcs.CLcs.LGarXiv:2311.18803v32023PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7
cs.CLcs.CVarXiv:2608.30241v12026Masked Discrimination for Self-Supervised Learning on Point Clouds
Haotian Liu, Mu Cai, Yong Jae Lee
cs.CVarXiv:2203.11183v22022PCAN: 3D Attention Map Learning Using Contextual Information for Point Cloud Based Retrieval
Wenxiao Zhang, Chunxia Xiao
cs.CVcs.ROarXiv:1904.09793v12019MNIST-C: A Robustness Benchmark for Computer Vision
Norman Mu, Justin Gilmer
cs.CVcs.LGarXiv:1906.02337v12019Social Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition
Timur Bagautdinov, Alexandre Alahi, François Fleuret +2
cs.CVarXiv:1611.09078v12016Dress Code: High-Resolution Multi-Category Virtual Try-On
Davide Morelli, Matteo Fincato, Marcella Cornia +3
cs.CVcs.AIcs.GRarXiv:2204.08532v22022KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA
Kenneth Marino, Xinlei Chen, Devi Parikh +2
cs.CVcs.CLarXiv:2012.11014v12020Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Minghan Qin, Yuang Wang, Xiuyu Yang +6
cs.CVcs.AIarXiv:2608.30821v12026Accurate Optical Flow via Direct Cost Volume Processing
Jia Xu, René Ranftl, Vladlen Koltun
cs.CVarXiv:1704.07325v12017Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction
Guanglei Yang, Hao Tang, Mingli Ding +2
cs.CVarXiv:2103.12091v22021Box-driven Class-wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation
Chunfeng Song, Yan Huang, Wanli Ouyang +1
cs.CVarXiv:1904.11693v12019RegNet: Multimodal Sensor Registration Using Deep Neural Networks
Nick Schneider, Florian Piewak, Christoph Stiller +1
cs.CVcs.AIcs.LGarXiv:1707.03167v12017LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu, Hao Cheng, Haotian Liu +10
cs.CVcs.AIcs.CLarXiv:2311.05437v12023Video Captioning via Hierarchical Reinforcement Learning
Xin Wang, Wenhu Chen, Jiawei Wu +2
cs.CVcs.AIcs.CLarXiv:1711.11135v32017Learning Human-Object Interaction Detection using Interaction Points
Tiancai Wang, Tong Yang, Martin Danelljan +3
cs.CVarXiv:2003.14023v12020BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled Images
Thu Nguyen-Phuoc, Christian Richardt, Long Mai +2
cs.CVarXiv:2002.08988v42020PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight Pruning
Wei Niu, Xiaolong Ma, Sheng Lin +5
cs.LGcs.CVcs.DCarXiv:2001.00138v42020Low-light Image Enhancement via Breaking Down the Darkness
Qiming Hu, Xiaojie Guo
cs.CVarXiv:2111.15557v12021iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes
Bokui Shen, Fei Xia, Chengshu Li +12
cs.AIcs.CVcs.ROarXiv:2012.02924v62020Learning image representations tied to ego-motion
Dinesh Jayaraman, Kristen Grauman
cs.CVcs.AIstat.MLarXiv:1505.02206v22015An Experimental-based Review of Image Enhancement and Image Restoration Methods for Underwater Imaging
Yan Wang, Wei Song, Giancarlo Fortino +3
eess.IVcs.CVcs.MMarXiv:1907.03246v12019Pillar-based Object Detection for Autonomous Driving
Yue Wang, Alireza Fathi, Abhijit Kundu +4
cs.CVcs.LGcs.ROarXiv:2007.10323v22020Variational Relational Point Completion Network
Liang Pan, Xinyi Chen, Zhongang Cai +4
cs.CVcs.LGarXiv:2104.10154v12021Ear Recognition: More Than a Survey
Žiga Emeršič, Vitomir Štruc, Peter Peer
cs.CVarXiv:1611.06203v22016Seeing Voices and Hearing Faces: Cross-modal biometric matching
Arsha Nagrani, Samuel Albanie, Andrew Zisserman
cs.CVarXiv:1804.00326v22018The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation
Junjie Huang, Zheng Zhu, Feng Guo +2
cs.CVarXiv:1911.07524v22019Trustworthy clinical AI solutions: a unified review of uncertainty quantification in deep learning models for medical image analysis
Benjamin Lambert, Florence Forbes, Alan Tucholka +3
eess.IVcs.AIcs.CVarXiv:2210.03736v12022Self-trained Deep Ordinal Regression for End-to-End Video Anomaly Detection
Guansong Pang, Cheng Yan, Chunhua Shen +2
cs.CVarXiv:2003.06780v12020Single-Stage Multi-Person Pose Machines
Xuecheng Nie, Jianfeng Zhang, Shuicheng Yan +1
cs.CVarXiv:1908.09220v12019Differential Treatment for Stuff and Things: A Simple Unsupervised Domain Adaptation Method for Semantic Segmentation
Zhonghao Wang, Mo Yu, Yunchao Wei +5
cs.CVcs.LGeess.IVarXiv:2003.08040v32020Are GAN generated images easy to detect? A critical analysis of the state-of-the-art
Diego Gragnaniello, Davide Cozzolino, Francesco Marra +2
cs.CVcs.AIarXiv:2104.02617v12021AdaCos: Adaptively Scaling Cosine Logits for Effectively Learning Deep Face Representations
Xiao Zhang, Rui Zhao, Yu Qiao +2
cs.CVarXiv:1905.00292v22019Automatic Breast Ultrasound Image Segmentation: A Survey
Min Xian, Yingtao Zhang, H. D. Cheng +3
cs.CVcs.LGarXiv:1704.01472v22017