Sound

Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

781 to 840 of 1,027

  1. Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

    Sophia Riaz, Haoze Zheng, Amos Roche +5

    cs.SDcs.AIarXiv:2608.22273v12026
  2. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

    Matthew Le, Apoorv Vyas, Bowen Shi +8

    eess.AScs.CLcs.LGarXiv:2306.15687v22023
  3. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

    Zhifei Xie, Jiaqi Lang, Ze An +7

    eess.AScs.AIcs.IRarXiv:2608.26005v12026
  4. W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

    Yu-An Chung, Yu Zhang, Wei Han +4

    cs.LGcs.SDeess.ASarXiv:2108.06209v22021
  5. Separating Voice from Age in COPD Screening

    George P. Kafentzis, Nikoletta Arvaniti

    eess.AScs.LGcs.SDarXiv:2608.21599v12026
  6. AudioCLIP: Extending CLIP to Image, Text and Audio

    Andrey Guzhov, Federico Raue, Jörn Hees +1

    cs.SDcs.CVeess.ASarXiv:2106.13043v12021
  7. MusPyExpress: Extending MusPy with Enhanced Expression Text Support

    Phillip Long, Hao-Wen Dong, Julian McAuley +1

    cs.SDcs.LGeess.ASarXiv:2608.21678v12026
  8. Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

    Wentao Jiang, Youchen Xie, Haidi Fan +4

    cs.HCcs.CVcs.SDarXiv:2608.24909v12026
  9. DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressors

    Chandan K A Reddy, Vishak Gopal, Ross Cutler

    cs.SDcs.LGeess.ASarXiv:2010.15258v22020
  10. The Sound of Pixels

    Hang Zhao, Chuang Gan, Andrew Rouditchenko +3

    cs.CVcs.SDeess.ASarXiv:1804.03160v42018
  11. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

    Max Bain, Jaesung Huh, Tengda Han +1

    cs.SDeess.ASarXiv:2303.00747v22023
  12. Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks

    Sharath Adavanne, Archontis Politis, Joonas Nikunen +1

    cs.SDeess.ASarXiv:1807.00129v32018
  13. Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

    Curtis Hawthorne, Andriy Stasyuk, Adam Roberts +6

    cs.SDcs.LGeess.ASarXiv:1810.12247v52018
  14. EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

    Jiawen Wang, Xiaoxue Gao, Zi Haur Pang +1

    cs.SDcs.AIarXiv:2608.23758v12026
  15. Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026

    Jinxing Zhou, Suiyi Zhao, Yanghao Zhou +1

    cs.MMcs.CVcs.SDarXiv:2608.22337v12026
  16. Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron

    RJ Skerry-Ryan, Eric Battenberg, Ying Xiao +6

    cs.CLcs.LGcs.SDarXiv:1803.09047v12018
  17. End-to-End Text-Dependent Speaker Verification

    Georg Heigold, Ignacio Moreno, Samy Bengio +1

    cs.LGcs.SDarXiv:1509.08062v12015
  18. Clotho: An Audio Captioning Dataset

    Konstantinos Drossos, Samuel Lipping, Tuomas Virtanen

    cs.SDcs.CLcs.LGarXiv:1910.09387v12019
  19. Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection

    Emre Çakır, Giambattista Parascandolo, Toni Heittola +2

    cs.LGcs.SDarXiv:1702.06286v12017
  20. GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

    Guoguo Chen, Shuzhou Chai, Guanbo Wang +18

    cs.SDcs.CLeess.ASarXiv:2106.06909v12021
  21. SampleRNN: An Unconditional End-to-End Neural Audio Generation Model

    Soroush Mehri, Kundan Kumar, Ishaan Gulrajani +5

    cs.SDcs.AIarXiv:1612.07837v22016
  22. Attentive Statistics Pooling for Deep Speaker Embedding

    Koji Okabe, Takafumi Koshinaka, Koichi Shinoda

    eess.AScs.SDarXiv:1803.10963v22018
  23. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

    Jaehyeon Kim, Sungwon Kim, Jungil Kong +1

    eess.AScs.SDarXiv:2005.11129v22020
  24. FMA: A Dataset For Music Analysis

    Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst +1

    cs.SDcs.IRarXiv:1612.01840v32016
  25. FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

    Junjie Li, Xuelong Geng, Kun Xie +13

    cs.CLcs.SDarXiv:2608.24168v12026
  26. MuseGAN: Multi-track Sequential Generative Adversarial Networks for Symbolic Music Generation and Accompaniment

    Hao-Wen Dong, Wen-Yi Hsiao, Li-Chia Yang +1

    eess.AScs.AIcs.LGarXiv:1709.06298v22017
  27. Real Time Speech Enhancement in the Waveform Domain

    Alexandre Defossez, Gabriel Synnaeve, Yossi Adi

    eess.AScs.LGcs.SDarXiv:2006.12847v32020
  28. YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

    Edresson Casanova, Julian Weber, Christopher Shulby +3

    cs.SDcs.CLeess.ASarXiv:2112.02418v42021
  29. Moshi: a speech-text foundation model for real-time dialogue

    Alexandre Défossez, Laurent Mazaré, Manu Orsini +5

    eess.AScs.AIcs.CLarXiv:2410.00037v22024
  30. SALMONN: Towards Generic Hearing Abilities for Large Language Models

    Changli Tang, Wenyi Yu, Guangzhi Sun +6

    cs.SDcs.CLeess.ASarXiv:2310.13289v22023
  31. Deep Voice: Real-time Neural Text-to-Speech

    Sercan O. Arik, Mike Chrzanowski, Adam Coates +9

    cs.CLcs.LGcs.NEarXiv:1702.07825v22017
  32. Pisets: A Robust Speech Recognition System for Lectures and Interviews

    Ivan Bondarenko, Daniil Grebenkin, Oleg Sedukhin +3

    cs.CLcs.SDeess.ASarXiv:2601.18415v12026
  33. UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

    Takaaki Saeki, Detai Xin, Wataru Nakata +3

    cs.SDeess.ASarXiv:2204.02152v22022
  34. Adaptive Evidence Weighting for Audio-Spatiotemporal Fusion

    Oscar Ovanger, Levi Harris, Timothy H. Keitt

    cs.SDcs.AIarXiv:2602.03817v12026
  35. NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control

    Yufan Wen, Zhaocheng Liu, YeGuo Hua +4

    cs.SDcs.AIeess.ASarXiv:2602.09070v22026
  36. Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization

    Mandip Goswami

    eess.AScs.SDeess.SParXiv:2602.12299v12026
  37. Preliminary sonification of ENSO using traditional Javanese gamelan scales

    Sandy Hardian Susanto Herho, Rusmawan Suwarman, Nurjanna Joko Trilaksono +2

    physics.soc-phcs.SDphysics.ao-pharXiv:2602.14560v22026
  38. BEATs: Audio Pre-Training with Acoustic Tokenizers

    Sanyuan Chen, Yu Wu, Chengyi Wang +4

    eess.AScs.AIcs.CLarXiv:2212.09058v12022
  39. MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

    Zhongxi Wang, Yueqian Lin, Jingyang Zhang +2

    cs.LGcs.CLcs.CVarXiv:2603.02482v12026
  40. Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models

    Kyudan Jung, Jihwan Kim, Soyoon Kim +3

    cs.SDcs.AIeess.ASarXiv:2603.25750v22026
  41. SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection

    Kyudan Jung, Jihwan Kim, Minwoo Lee +4

    cs.SDcs.AIarXiv:2603.20686v12026
  42. Adversarial Audio Synthesis

    Chris Donahue, Julian McAuley, Miller Puckette

    cs.SDcs.LGarXiv:1802.04208v32018
  43. RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation

    Mandip Goswami

    eess.AScs.CLcs.SDarXiv:2601.19949v12026
  44. Scaling Speech Technology to 1,000+ Languages

    Vineel Pratap, Andros Tjandra, Bowen Shi +13

    cs.CLcs.SDeess.ASarXiv:2305.13516v12023
  45. Segment Length Matters: A Study of Segment Lengths on Audio Fingerprinting Performance

    Ziling Gong, Yunyan Ouyang, Iram Kamdar +5

    cs.SDcs.AIcs.IRarXiv:2601.17690v12026
  46. Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation

    Daniel Stoller, Sebastian Ewert, Simon Dixon

    cs.SDeess.ASstat.MLarXiv:1806.03185v12018
  47. Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

    Aoke Zhang, Bo Wang, Xihong Wu +2

    cs.SDcs.AIarXiv:2608.22420v12026
  48. The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines

    Jon Barker, Shinji Watanabe, Emmanuel Vincent +1

    cs.SDcs.AIeess.ASarXiv:1803.10609v12018
  49. JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

    Zhan Liu, Changli Tang, Yuxin Wang +7

    cs.CVcs.AIcs.SDarXiv:2602.18527v32026
  50. MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

    Yize Li, Ningyuan Yang, Sile Yin +6

    cs.SDcs.LGeess.ASarXiv:2608.22236v12026
  51. Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription

    Nicolas Boulanger-Lewandowski, Yoshua Bengio, Pascal Vincent

    cs.LGcs.SDstat.MLarXiv:1206.6392v12012
  52. Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

    Masoud Jalayer, Changyi Li, Yu Xiao

    cs.CVcs.AIcs.SDarXiv:2608.22359v12026
  53. TasNet: time-domain audio separation network for real-time, single-channel speech separation

    Yi Luo, Nima Mesgarani

    cs.SDcs.LGcs.MMarXiv:1711.00541v22017
  54. MusicLM: Generating Music From Text

    Andrea Agostinelli, Timo I. Denk, Zalán Borsos +10

    cs.SDcs.LGeess.ASarXiv:2301.11325v12023
  55. V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

    Yan-Bo Lin, Jonah Casebeer, Long Mai +3

    cs.CVcs.AIcs.LGarXiv:2603.11042v22026
  56. Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition

    Umberto Cappellazzo, Stavros Petridis, Maja Pantic

    eess.AScs.CVcs.SDarXiv:2603.12046v22026
  57. VoXtream2: Full-stream TTS with dynamic speaking rate control

    Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

    eess.AScs.CLcs.HCarXiv:2603.13518v12026
  58. PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

    Mohammad Javad Ranjbar Kalahroodi, Mohammad Amini, Parmis Bathayan +2

    cs.CLcs.SDarXiv:2603.14456v12026
  59. ReactMotion: Generating Reactive Listener Motions from Speaker Utterance

    Cheng Luo, Bizhu Wu, Bing Li +5

    cs.CVcs.AIcs.HCarXiv:2603.15083v12026
  60. FSD50K: An Open Dataset of Human-Labeled Sound Events

    Eduardo Fonseca, Xavier Favory, Jordi Pons +2

    cs.SDcs.LGeess.ASarXiv:2010.00475v22020