Sound
Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
61 to 120 of 1,027
Audio Deepfake Detection: A Survey
Jiangyan Yi, Chenglong Wang, Jianhua Tao +3
cs.SDeess.ASarXiv:2308.14970v12023High-Quality, Low-Delay Music Coding in the Opus Codec
Jean-Marc Valin, Gregory Maxwell, Timothy B. Terriberry +1
cs.MMcs.SDarXiv:1602.04845v12016Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam
Marc Delcroix, Tsubasa Ochiai, Katerina Zmolikova +4
eess.AScs.CLcs.SDarXiv:2001.08378v12020Xiaomi-CocktailASR-1 Technical Report
Yiru Zhang, Hang Su, Lichun Fan +10
cs.SDcs.CLeess.ASarXiv:2609.11274v12026A Review of Machine Learning Methods Applied to Structural Dynamics and Vibroacoustic
Barbara Cunha, Christophe Droz, Abdelmalek Zine +2
cs.LGcs.SDeess.ASarXiv:2204.06362v22022A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
Shulei Ji, Jing Luo, Xinyu Yang
cs.SDcs.LGeess.ASarXiv:2011.06801v12020ASR is all you need: cross-modal distillation for lip reading
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
cs.CVcs.SDeess.ASarXiv:1911.12747v22019Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech
Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2
cs.CLcs.SDarXiv:2609.11545v12026SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ
Huy Hoang Le, Long-Bao Nguyen, Minh Tri Dao
cs.CLcs.SDarXiv:2609.11355v12026Performance-Efficiency Trade-offs in Unsupervised Pre-training for Speech Recognition
Felix Wu, Kwangyoun Kim, Jing Pan +3
cs.CLcs.LGcs.SDarXiv:2109.06870v12021Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation
Maria Frangiadaki, Dimitrios Damianos, Kosmas Kritsis +1
cs.CLcs.SDarXiv:2609.11302v12026Large Margin Softmax Loss for Speaker Verification
Yi Liu, Liang He, Jia Liu
cs.SDeess.ASarXiv:1904.03479v12019LLM-Anchored Paralinguistic Enrichment for Alzheimer's Disease Detection
Xiao Wei, Yuqin Lin, Yaru Cao +6
cs.CLcs.SDarXiv:2609.10896v12026X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
Haojun Zhang, Yi Zou, Min Chen +7
cs.SDcs.AIarXiv:2609.11412v12026A spelling correction model for end-to-end speech recognition
Jinxi Guo, Tara N. Sainath, Ron J. Weiss
eess.AScs.AIcs.CLarXiv:1902.07178v12019Understanding Optical Music Recognition
Jorge Calvo-Zaragoza, Jan Hajič, Alexander Pacha
cs.CVcs.AIcs.IRarXiv:1908.03608v32019DPT-FSNet: Dual-path Transformer Based Full-band and Sub-band Fusion Network for Speech Enhancement
Feng Dang, Hangting Chen, Pengyuan Zhang
cs.SDeess.ASarXiv:2104.13002v22021CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
Linhao Dong, Bo Xu
cs.CLcs.LGcs.NEarXiv:1905.11235v42019EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation
Hsiao-Tzu Hung, Joann Ching, Seungheon Doh +3
cs.SDcs.MMeess.ASarXiv:2108.01374v12021Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis
Hong Nguyen, Sean Foley, Christina Hagedorn +4
cs.SDcs.CVarXiv:2609.09757v12026Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis
Yu Wang, Xinsheng Wang, Pengcheng Zhu +6
cs.SDcs.DBeess.ASarXiv:2201.07429v22022Voice2Series: Reprogramming Acoustic Models for Time Series Classification
Chao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu Chen
cs.LGcs.AIcs.NEarXiv:2106.09296v32021Bird detection in audio: a survey and a challenge
Dan Stowell, Mike Wood, Yannis Stylianou +1
cs.SDarXiv:1608.03417v12016Proactive Detection of Voice Cloning with Localized Watermarking
Robin San Roman, Pierre Fernandez, Alexandre Défossez +3
cs.SDcs.AIcs.CRarXiv:2401.17264v22024TimeCues Studio: A Workspace for Music Annotation and Algorithm Prototyping
Sapir Caduri, Yoav Goldberg
cs.SDcs.HCcs.LGarXiv:2609.10338v12026Orukeet: Multilingual ASR with Frozen Gabor Kernels
Nathan Roll, Irene Yi, Büşra Marşan +6
cs.SDcs.LGeess.ASarXiv:2609.10054v12026Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations
Soumyadeep Roy
cs.SDcs.LGarXiv:2609.10051v12026WavLLM: Towards Robust and Adaptive Speech Large Language Model
Shujie Hu, Long Zhou, Shujie Liu +9
cs.CLcs.AIcs.SDarXiv:2404.00656v32024Empirical Study of Drone Sound Detection in Real-Life Environment with Deep Neural Networks
Sungho Jeon, Jong-Woo Shin, Young-Jun Lee +3
cs.SDcs.LGarXiv:1701.05779v12017Sparsity-based Algorithm for Detecting Faults in Rotating Machines
Wangpeng He, Yin Ding, Yanyang Zi +1
cs.SDarXiv:1511.00067v12015The VoicePrivacy 2020 Challenge: Results and findings
Natalia Tomashenko, Xin Wang, Emmanuel Vincent +11
cs.CLcs.SDeess.ASarXiv:2109.00648v42021Wavenet based low rate speech coding
W. Bastiaan Kleijn, Felicia S. C. Lim, Alejandro Luebs +4
eess.AScs.SDeess.SParXiv:1712.01120v12017Robust and fine-grained prosody control of end-to-end speech synthesis
Younggun Lee, Taesu Kim
cs.CLcs.LGcs.SDarXiv:1811.02122v22018Spirit LM: Interleaved Spoken and Written Language Model
Tu Anh Nguyen, Benjamin Muller, Bokai Yu +13
cs.CLcs.SDeess.ASarXiv:2402.05755v22024Lung Sound Classification Using Co-tuning and Stochastic Normalization
Truc Nguyen, Franz Pernkopf
eess.AScs.LGcs.SDarXiv:2108.01991v12021Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Di Hu, Rui Qian, Minyue Jiang +5
cs.CVcs.LGcs.MMarXiv:2010.05466v12020Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
Alexei Baevski, Arun Babu, Wei-Ning Hsu +1
cs.LGcs.CLcs.SDarXiv:2212.07525v22022Clova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020
Hee Soo Heo, Bong-Jin Lee, Jaesung Huh +1
eess.AScs.SDarXiv:2009.14153v12020AudioBench: A Universal Benchmark for Audio Large Language Models
Bin Wang, Xunlong Zou, Geyu Lin +6
cs.SDcs.CLeess.ASarXiv:2406.16020v52024Sound Event Detection Using Spatial Features and Convolutional Recurrent Neural Network
Sharath Adavanne, Pasi Pertilä, Tuomas Virtanen
cs.SDcs.LGarXiv:1706.02291v12017MT3: Multi-Task Multitrack Music Transcription
Josh Gardner, Ian Simon, Ethan Manilow +2
cs.SDcs.LGeess.ASarXiv:2111.03017v42021Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference
Yaohan Guan, Yen-Ju Lu, Yuzhe Wang +5
cs.MAcs.CLcs.SDarXiv:2609.09628v12026Do speech foundation models really learn words?
Robin Huo, Ewan Dunbar
cs.CLcs.SDarXiv:2609.10434v12026Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis +1
cs.SDcs.CLcs.LGarXiv:2609.10022v12026FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement
Jun Chen, Zilin Wang, Deyi Tuo +3
cs.SDcs.AIeess.ASarXiv:2203.12188v22022Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal +3
cs.SDcs.AIcs.CLarXiv:2404.09956v42024Transfer Learning from Adult to Children for Speech Recognition: Evaluation, Analysis and Recommendations
Prashanth Gurunath Shivakumar, Panayiotis Georgiou
eess.AScs.CLcs.SDarXiv:1805.03322v12018The Implementation of Low-cost Urban Acoustic Monitoring Devices
Charlie Mydlarz, Justin Salamon, Juan Pablo Bello
cs.SDarXiv:1605.08450v12016Iterative Pseudo-Labeling for Speech Recognition
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn +3
cs.CLcs.SDeess.ASarXiv:2005.09267v22020StreamAlign: Streaming Text-Aligned Speech Tokenization
Kang-wook Kim, Jinyoung Park, Jinsoo Kim +3
cs.CLcs.SDeess.ASarXiv:2609.09719v12026FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
Shengkui Zhao, Bin Ma, Karn N. Watcharasupat +1
cs.SDeess.ASarXiv:2206.07293v32022Chatty Maps: Constructing sound maps of urban areas from social media data
Luca Maria Aiello, Rossano Schifanella, Daniele Quercia +1
cs.SIcs.CYcs.SDarXiv:1603.07813v12016HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
Jiaqi Su, Zeyu Jin, Adam Finkelstein
eess.AScs.LGcs.SDarXiv:2006.05694v22020NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments
Niramay M. Patel, Bibek Behera, Raksha Sharma
cs.SDcs.AIcs.CLarXiv:2609.10058v12026From Audio to Semantics: Approaches to end-to-end spoken language understanding
Parisa Haghani, Arun Narayanan, Michiel Bacchiani +6
eess.AScs.CLcs.SDarXiv:1809.09190v12018Contrastive Learning of Musical Representations
Janne Spijkervet, John Ashley Burgoyne
cs.SDcs.LGeess.ASarXiv:2103.09410v22021Deep Spoken Keyword Spotting: An Overview
Iván López-Espejo, Zheng-Hua Tan, John Hansen +1
cs.SDcs.HCcs.LGarXiv:2111.10592v12021Self-supervised Moving Vehicle Tracking with Stereo Sound
Chuang Gan, Hang Zhao, Peihao Chen +2
cs.CVcs.LGcs.SDarXiv:1910.11760v12019Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models
Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1
cs.SDcs.AIarXiv:2609.09263v12026AENet: Learning Deep Audio Features for Video Analysis
Naoya Takahashi, Michael Gygli, Luc Van Gool
cs.MMcs.CVcs.SDarXiv:1701.00599v22017