Sound
Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
721 to 780 of 1,027
An Unsupervised Autoregressive Model for Speech Representation Learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang +1
cs.CLcs.LGcs.SDarXiv:1904.03240v22019DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
Chandan K A Reddy, Vishak Gopal, Ross Cutler
eess.AScs.SDarXiv:2110.01763v42021Summaries:한국어A multi-device dataset for urban acoustic scene classification
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen
eess.AScs.SDarXiv:1807.09840v22018Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
Yao Qin, Nicholas Carlini, Ian Goodfellow +2
eess.AScs.LGcs.SDarXiv:1903.10346v22019Exploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory Sound Data
Chloë Brown, Jagmohan Chauhan, Andreas Grammenos +6
cs.SDcs.LGeess.ASarXiv:2006.05919v32020A Review of Speaker Diarization: Recent Advances with Deep Learning
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis +3
eess.AScs.CLcs.SDarXiv:2101.09624v42021Diffsound: Discrete Diffusion Model for Text-to-sound Generation
Dongchao Yang, Jianwei Yu, Helin Wang +4
cs.SDcs.AIeess.ASarXiv:2207.09983v22022MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection
Harsh Purohit, Ryo Tanabe, Kenji Ichige +4
cs.SDcs.LGeess.ASarXiv:1909.09347v12019HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
Ke Chen, Xingjian Du, Bilei Zhu +3
cs.SDcs.AIcs.IRarXiv:2202.00874v12022Single-Channel Multi-Speaker Separation using Deep Clustering
Yusuf Isik, Jonathan Le Roux, Zhuo Chen +2
cs.LGcs.SDstat.MLarXiv:1607.02173v12016Computational bioacoustics with deep learning: a review and roadmap
Dan Stowell
cs.SDeess.ASq-bio.QMarXiv:2112.06725v12021Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators
Jonas Länzlinger, Katharina O. E. Müller, Burkhard Stiller +1
cs.CLcs.SDeess.ASarXiv:2608.26148v12026Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores
Eichi Uehara
cs.CLcs.LGcs.SDarXiv:2608.26137v12026Machine learning in acoustics: theory and applications
Michael J. Bianco, Peter Gerstoft, James Traer +4
eess.SPcs.LGcs.SDarXiv:1905.04418v42019From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents
Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
cs.CLcs.AIcs.HCarXiv:2608.26163v12026Recent Advances in End-to-End Automatic Speech Recognition
Jinyu Li
eess.AScs.AIcs.CLarXiv:2111.01690v22021Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson +1
cs.SDcs.AIcs.LGarXiv:1502.04149v42015Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
Hang Zhou, Yasheng Sun, Wayne Wu +3
cs.CVcs.LGcs.MMarXiv:2104.11116v12021AudioGen: Textually Guided Audio Generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak +6
cs.SDcs.CLcs.LGarXiv:2209.15352v22022The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
Chandan K. A. Reddy, Vishak Gopal, Ross Cutler +10
eess.AScs.LGcs.SDarXiv:2005.13981v32020pyannote.audio: neural building blocks for speaker diarization
Hervé Bredin, Ruiqing Yin, Juan Manuel Coria +7
eess.AScs.SDarXiv:1911.01255v12019A Survey on Neural Speech Synthesis
Xu Tan, Tao Qin, Frank Soong +1
eess.AScs.CLcs.LGarXiv:2106.15561v32021A Wavenet for Speech Denoising
Dario Rethage, Jordi Pons, Xavier Serra
cs.SDarXiv:1706.07162v32017AudioPaLM: A Large Language Model That Can Speak and Listen
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27
cs.CLcs.AIcs.SDarXiv:2306.12925v12023Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications
Chanameth Boonpramuk, Winn Voravuthikunchai, Songpol Bunyang
cs.SDcs.AIarXiv:2608.24916v12026In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun +7
eess.AScs.SDarXiv:2003.11982v22020LPCNet: Improving Neural Speech Synthesis Through Linear Prediction
Jean-Marc Valin, Jan Skoglund
eess.AScs.LGcs.SDarXiv:1810.11846v22018Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings
Leonardo Pepino, Pablo Riera, Luciana Ferrer
cs.SDcs.LGeess.ASarXiv:2104.03502v12021BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Sang-gil Lee, Wei Ping, Boris Ginsburg +2
cs.SDcs.CLcs.LGarXiv:2206.04658v22022Masked Autoencoders that Listen
Po-Yao Huang, Hu Xu, Juncheng Li +5
cs.SDcs.AIcs.LGarXiv:2207.06405v32022Deepfakes Generation and Detection: State-of-the-art, open challenges, countermeasures, and way forward
Momina Masood, Marriam Nawaz, Khalid Mahmood Malik +2
cs.CRcs.LGcs.SDarXiv:2103.00484v22021CREPE: A Convolutional Representation for Pitch Estimation
Jong Wook Kim, Justin Salamon, Peter Li +1
eess.AScs.LGcs.SDarXiv:1802.06182v12018SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification
Giries Abu Ayoub, Loay Mualem, Simon Korman
cs.SDcs.AIarXiv:2608.25054v12026Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Tianle Wang, Xinyi Tong, Liangke Zhao +7
cs.SDcs.AIarXiv:2608.25621v12026Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia +1
eess.AScs.CVcs.SDarXiv:2201.02184v22022DDSP: Differentiable Digital Signal Processing
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu +1
cs.LGcs.SDeess.ASarXiv:2001.04643v12020MidiNet: A Convolutional Generative Adversarial Network for Symbolic-domain Music Generation
Li-Chia Yang, Szu-Yu Chou, Yi-Hsuan Yang
cs.SDcs.AIarXiv:1703.10847v22017Hello Edge: Keyword Spotting on Microcontrollers
Yundong Zhang, Naveen Suda, Liangzhen Lai +1
cs.SDcs.CLcs.LGarXiv:1711.07128v32017Dawn of the transformer era in speech emotion recognition: closing the valence gap
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf +4
eess.AScs.LGcs.SDarXiv:2203.07378v42022WHAM!: Extending Speech Separation to Noisy Environments
Gordon Wichern, Joe Antognini, Michael Flynn +5
cs.SDcs.CLcs.LGarXiv:1907.01160v12019CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Zhihao Du, Qian Chen, Shiliang Zhang +9
cs.SDcs.AIeess.ASarXiv:2407.05407v22024AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra +1
cs.SDarXiv:2608.25244v12026Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Rongjie Huang, Jiawei Huang, Dongchao Yang +7
cs.SDcs.LGcs.MMarXiv:2301.12661v12023Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths
Osamu Hoshuyama
eess.AScs.SDarXiv:2608.25413v12026Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
Yigitcan Özer, Xin Wang, Zhe Zhang +1
cs.SDcs.AIarXiv:2608.25289v12026Self-Supervised Speech Representation Learning: A Review
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt +9
cs.CLcs.SDeess.ASarXiv:2205.10643v32022NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
Gabriel Mittag, Babak Naderi, Assmaa Chehadi +1
eess.AScs.AIcs.LGarXiv:2104.09494v12021A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
Yigitcan Özer, Zhe Zhang, Wanying Ge +2
cs.SDcs.AIarXiv:2608.25285v12026AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
Haohe Liu, Yi Yuan, Xubo Liu +7
cs.SDcs.AIcs.MMarXiv:2308.05734v32023F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Yushen Chen, Zhikang Niu, Ziyang Ma +5
eess.AScs.SDarXiv:2410.06885v32024Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data
Rene Glitza, Luca Becker, Rainer Martin
cs.LGcs.DCcs.SDarXiv:2608.25794v12026Convolutional Recurrent Neural Networks for Music Classification
Keunwoo Choi, George Fazekas, Mark Sandler +1
cs.NEcs.LGcs.MMarXiv:1609.04243v32016AI4COVID-19: AI Enabled Preliminary Diagnosis for COVID-19 from Cough Samples via an App
Ali Imran, Iryna Posokhova, Haneya N. Qureshi +6
eess.AScs.LGcs.SDarXiv:2004.01275v62020Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
Qian Zhang, Han Lu, Hasim Sak +4
eess.AScs.CLcs.SDarXiv:2002.02562v22020What Do Audio-Visual Synchronization Metrics Actually Measure?
Jai Kumar Sharma, Peeyush Tapadiya
cs.CVcs.MMcs.SDarXiv:2608.25157v12026ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection
Junichi Yamagishi, Xin Wang, Massimiliano Todisco +8
eess.AScs.CRcs.LGarXiv:2109.00537v12021Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +8
cs.SDcs.AIcs.CLarXiv:2608.24958v12026A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
Adam Roberts, Jesse Engel, Colin Raffel +2
cs.LGcs.SDeess.ASarXiv:1803.05428v52018AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Kaizhi Qian, Yang Zhang, Shiyu Chang +2
eess.AScs.AIcs.LGarXiv:1905.05879v22019Objects that Sound
Relja Arandjelović, Andrew Zisserman
cs.CVcs.LGcs.MMarXiv:1712.06651v22017