Sound

Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

721 to 780 of 1,027

  1. An Unsupervised Autoregressive Model for Speech Representation Learning

    Yu-An Chung, Wei-Ning Hsu, Hao Tang +1

    cs.CLcs.LGcs.SDarXiv:1904.03240v22019
  2. DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors

    Chandan K A Reddy, Vishak Gopal, Ross Cutler

    eess.AScs.SDarXiv:2110.01763v42021
    Summaries:한국어
  3. A multi-device dataset for urban acoustic scene classification

    Annamaria Mesaros, Toni Heittola, Tuomas Virtanen

    eess.AScs.SDarXiv:1807.09840v22018
  4. Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition

    Yao Qin, Nicholas Carlini, Ian Goodfellow +2

    eess.AScs.LGcs.SDarXiv:1903.10346v22019
  5. Exploring Automatic Diagnosis of COVID-19 from Crowdsourced Respiratory Sound Data

    Chloë Brown, Jagmohan Chauhan, Andreas Grammenos +6

    cs.SDcs.LGeess.ASarXiv:2006.05919v32020
  6. A Review of Speaker Diarization: Recent Advances with Deep Learning

    Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis +3

    eess.AScs.CLcs.SDarXiv:2101.09624v42021
  7. Diffsound: Discrete Diffusion Model for Text-to-sound Generation

    Dongchao Yang, Jianwei Yu, Helin Wang +4

    cs.SDcs.AIeess.ASarXiv:2207.09983v22022
  8. MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection

    Harsh Purohit, Ryo Tanabe, Kenji Ichige +4

    cs.SDcs.LGeess.ASarXiv:1909.09347v12019
  9. HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

    Ke Chen, Xingjian Du, Bilei Zhu +3

    cs.SDcs.AIcs.IRarXiv:2202.00874v12022
  10. Single-Channel Multi-Speaker Separation using Deep Clustering

    Yusuf Isik, Jonathan Le Roux, Zhuo Chen +2

    cs.LGcs.SDstat.MLarXiv:1607.02173v12016
  11. Computational bioacoustics with deep learning: a review and roadmap

    Dan Stowell

    cs.SDeess.ASq-bio.QMarXiv:2112.06725v12021
  12. Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

    Jonas Länzlinger, Katharina O. E. Müller, Burkhard Stiller +1

    cs.CLcs.SDeess.ASarXiv:2608.26148v12026
  13. Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

    Eichi Uehara

    cs.CLcs.LGcs.SDarXiv:2608.26137v12026
  14. Machine learning in acoustics: theory and applications

    Michael J. Bianco, Peter Gerstoft, James Traer +4

    eess.SPcs.LGcs.SDarXiv:1905.04418v42019
  15. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

    Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee

    cs.CLcs.AIcs.HCarXiv:2608.26163v12026
  16. Recent Advances in End-to-End Automatic Speech Recognition

    Jinyu Li

    eess.AScs.AIcs.CLarXiv:2111.01690v22021
  17. Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation

    Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson +1

    cs.SDcs.AIcs.LGarXiv:1502.04149v42015
  18. Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

    Hang Zhou, Yasheng Sun, Wayne Wu +3

    cs.CVcs.LGcs.MMarXiv:2104.11116v12021
  19. AudioGen: Textually Guided Audio Generation

    Felix Kreuk, Gabriel Synnaeve, Adam Polyak +6

    cs.SDcs.CLcs.LGarXiv:2209.15352v22022
  20. The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results

    Chandan K. A. Reddy, Vishak Gopal, Ross Cutler +10

    eess.AScs.LGcs.SDarXiv:2005.13981v32020
  21. pyannote.audio: neural building blocks for speaker diarization

    Hervé Bredin, Ruiqing Yin, Juan Manuel Coria +7

    eess.AScs.SDarXiv:1911.01255v12019
  22. A Survey on Neural Speech Synthesis

    Xu Tan, Tao Qin, Frank Soong +1

    eess.AScs.CLcs.LGarXiv:2106.15561v32021
  23. A Wavenet for Speech Denoising

    Dario Rethage, Jordi Pons, Xavier Serra

    cs.SDarXiv:1706.07162v32017
  24. AudioPaLM: A Large Language Model That Can Speak and Listen

    Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27

    cs.CLcs.AIcs.SDarXiv:2306.12925v12023
  25. Domain-Adaptive ASR for Telephony AI Agents: Fine-tuning Canary Flash Models for Enterprise Contact Center Applications

    Chanameth Boonpramuk, Winn Voravuthikunchai, Songpol Bunyang

    cs.SDcs.AIarXiv:2608.24916v12026
  26. In defence of metric learning for speaker recognition

    Joon Son Chung, Jaesung Huh, Seongkyu Mun +7

    eess.AScs.SDarXiv:2003.11982v22020
  27. LPCNet: Improving Neural Speech Synthesis Through Linear Prediction

    Jean-Marc Valin, Jan Skoglund

    eess.AScs.LGcs.SDarXiv:1810.11846v22018
  28. Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

    Leonardo Pepino, Pablo Riera, Luciana Ferrer

    cs.SDcs.LGeess.ASarXiv:2104.03502v12021
  29. BigVGAN: A Universal Neural Vocoder with Large-Scale Training

    Sang-gil Lee, Wei Ping, Boris Ginsburg +2

    cs.SDcs.CLcs.LGarXiv:2206.04658v22022
  30. Masked Autoencoders that Listen

    Po-Yao Huang, Hu Xu, Juncheng Li +5

    cs.SDcs.AIcs.LGarXiv:2207.06405v32022
  31. Deepfakes Generation and Detection: State-of-the-art, open challenges, countermeasures, and way forward

    Momina Masood, Marriam Nawaz, Khalid Mahmood Malik +2

    cs.CRcs.LGcs.SDarXiv:2103.00484v22021
  32. CREPE: A Convolutional Representation for Pitch Estimation

    Jong Wook Kim, Justin Salamon, Peter Li +1

    eess.AScs.LGcs.SDarXiv:1802.06182v12018
  33. SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification

    Giries Abu Ayoub, Loay Mualem, Simon Korman

    cs.SDcs.AIarXiv:2608.25054v12026
  34. Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding

    Tianle Wang, Xinyi Tong, Liangke Zhao +7

    cs.SDcs.AIarXiv:2608.25621v12026
  35. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

    Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia +1

    eess.AScs.CVcs.SDarXiv:2201.02184v22022
  36. DDSP: Differentiable Digital Signal Processing

    Jesse Engel, Lamtharn Hantrakul, Chenjie Gu +1

    cs.LGcs.SDeess.ASarXiv:2001.04643v12020
  37. MidiNet: A Convolutional Generative Adversarial Network for Symbolic-domain Music Generation

    Li-Chia Yang, Szu-Yu Chou, Yi-Hsuan Yang

    cs.SDcs.AIarXiv:1703.10847v22017
  38. Hello Edge: Keyword Spotting on Microcontrollers

    Yundong Zhang, Naveen Suda, Liangzhen Lai +1

    cs.SDcs.CLcs.LGarXiv:1711.07128v32017
  39. Dawn of the transformer era in speech emotion recognition: closing the valence gap

    Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf +4

    eess.AScs.LGcs.SDarXiv:2203.07378v42022
  40. WHAM!: Extending Speech Separation to Noisy Environments

    Gordon Wichern, Joe Antognini, Michael Flynn +5

    cs.SDcs.CLcs.LGarXiv:1907.01160v12019
  41. CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

    Zhihao Du, Qian Chen, Shiliang Zhang +9

    cs.SDcs.AIeess.ASarXiv:2407.05407v22024
  42. AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP

    Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra +1

    cs.SDarXiv:2608.25244v12026
  43. Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

    Rongjie Huang, Jiawei Huang, Dongchao Yang +7

    cs.SDcs.LGcs.MMarXiv:2301.12661v12023
  44. Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths

    Osamu Hoshuyama

    eess.AScs.SDarXiv:2608.25413v12026
  45. Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

    Yigitcan Özer, Xin Wang, Zhe Zhang +1

    cs.SDcs.AIarXiv:2608.25289v12026
  46. Self-Supervised Speech Representation Learning: A Review

    Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt +9

    cs.CLcs.SDeess.ASarXiv:2205.10643v32022
  47. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

    Gabriel Mittag, Babak Naderi, Assmaa Chehadi +1

    eess.AScs.AIcs.LGarXiv:2104.09494v12021
  48. A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

    Yigitcan Özer, Zhe Zhang, Wanying Ge +2

    cs.SDcs.AIarXiv:2608.25285v12026
  49. AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

    Haohe Liu, Yi Yuan, Xubo Liu +7

    cs.SDcs.AIcs.MMarXiv:2308.05734v32023
  50. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

    Yushen Chen, Zhikang Niu, Ziyang Ma +5

    eess.AScs.SDarXiv:2410.06885v32024
  51. Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data

    Rene Glitza, Luca Becker, Rainer Martin

    cs.LGcs.DCcs.SDarXiv:2608.25794v12026
  52. Convolutional Recurrent Neural Networks for Music Classification

    Keunwoo Choi, George Fazekas, Mark Sandler +1

    cs.NEcs.LGcs.MMarXiv:1609.04243v32016
  53. AI4COVID-19: AI Enabled Preliminary Diagnosis for COVID-19 from Cough Samples via an App

    Ali Imran, Iryna Posokhova, Haneya N. Qureshi +6

    eess.AScs.LGcs.SDarXiv:2004.01275v62020
  54. Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss

    Qian Zhang, Han Lu, Hasim Sak +4

    eess.AScs.CLcs.SDarXiv:2002.02562v22020
  55. What Do Audio-Visual Synchronization Metrics Actually Measure?

    Jai Kumar Sharma, Peeyush Tapadiya

    cs.CVcs.MMcs.SDarXiv:2608.25157v12026
  56. ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection

    Junichi Yamagishi, Xin Wang, Massimiliano Todisco +8

    eess.AScs.CRcs.LGarXiv:2109.00537v12021
  57. Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

    Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +8

    cs.SDcs.AIcs.CLarXiv:2608.24958v12026
  58. A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music

    Adam Roberts, Jesse Engel, Colin Raffel +2

    cs.LGcs.SDeess.ASarXiv:1803.05428v52018
  59. AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss

    Kaizhi Qian, Yang Zhang, Shiyu Chang +2

    eess.AScs.AIcs.LGarXiv:1905.05879v22019
  60. Objects that Sound

    Relja Arandjelović, Andrew Zisserman

    cs.CVcs.LGcs.MMarXiv:1712.06651v22017