Sound

Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

121 to 180 of 1,024

  1. HEAR: Holistic Evaluation of Audio Representations

    Joseph Turian, Jordie Shier, Humair Raj Khan +20

    cs.SDcs.AIcs.LGarXiv:2203.03022v32022
  2. From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection

    Mengzhe Geng, Yujia Lu, Patrick Littell +2

    cs.SDcs.CLeess.ASarXiv:2609.08899v22026
  3. TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

    Fritz Cremer, Jonathan Cremer

    cs.SDcs.CLcs.LGarXiv:2609.08703v12026
  4. CN-Celeb: multi-genre speaker recognition

    Lantian Li, Ruiqi Liu, Jiawen Kang +6

    cs.SDeess.ASarXiv:2012.12468v22020
  5. ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

    Richard Yucheng He, Baodong Cao, Chen Xu +2

    cs.CLcs.SDeess.ASarXiv:2609.08147v12026
  6. A neural attention model for speech command recognition

    Douglas Coimbra de Andrade, Sabato Leo, Martin Loesener Da Silva Viana +1

    eess.AScs.SDarXiv:1808.08929v12018
  7. Foley Music: Learning to Generate Music from Videos

    Chuang Gan, Deng Huang, Peihao Chen +2

    cs.CVcs.LGcs.SDarXiv:2007.10984v12020
  8. Sample Efficient Adaptive Text-to-Speech

    Yutian Chen, Yannis Assael, Brendan Shillingford +11

    cs.LGcs.SDstat.MLarXiv:1809.10460v32018
  9. FunASR: A Fundamental End-to-End Speech Recognition Toolkit

    Zhifu Gao, Zerui Li, Jiaming Wang +8

    cs.SDcs.CLeess.ASarXiv:2305.11013v12023
  10. Two-Pass End-to-End Speech Recognition

    Tara N. Sainath, Ruoming Pang, David Rybach +9

    cs.CLcs.SDeess.ASarXiv:1908.10992v12019
  11. X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR

    Zhiwei Lin, Kaiqi Fu, Rime Wen +5

    cs.SDcs.AIarXiv:2609.08672v12026
  12. Sudo rm -rf: Efficient Networks for Universal Audio Source Separation

    Efthymios Tzinis, Zhepei Wang, Paris Smaragdis

    eess.AScs.CLcs.LGarXiv:2007.06833v12020
  13. On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement

    Morten Kolbæk, Zheng-Hua Tan, Søren Holdt Jensen +1

    cs.SDcs.LGeess.ASarXiv:1909.01019v22019
  14. SpecAugment on Large Scale Datasets

    Daniel S. Park, Yu Zhang, Chung-Cheng Chiu +5

    eess.AScs.CLcs.LGarXiv:1912.05533v12019
  15. Generative Spoken Dialogue Language Modeling

    Tu Anh Nguyen, Eugene Kharitonov, Jade Copet +8

    cs.CLcs.LGcs.SDarXiv:2203.16502v22022
  16. Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

    Bella Godiva, Yeonju Kim, Yong Man Ro

    cs.SDcs.AIcs.HCarXiv:2609.08390v12026
  17. Learning Hidden Unit Contributions for Unsupervised Acoustic Model Adaptation

    Pawel Swietojanski, Jinyu Li, Steve Renals

    cs.CLcs.LGcs.SDarXiv:1601.02828v22016
  18. Audio Super Resolution using Neural Networks

    Volodymyr Kuleshov, S. Zayd Enam, Stefano Ermon

    cs.SDcs.LGarXiv:1708.00853v12017
  19. Transformer-Transducer: End-to-End Speech Recognition with Self-Attention

    Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar +6

    eess.AScs.CLcs.SDarXiv:1910.12977v12019
  20. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict

    Yosuke Higuchi, Shinji Watanabe, Nanxin Chen +2

    eess.AScs.SDarXiv:2005.08700v22020
  21. Multi-scale Multi-band DenseNets for Audio Source Separation

    Naoya Takahashi, Yuki Mitsufuji

    cs.SDcs.CLcs.MMarXiv:1706.09588v12017
  22. Audio Surveillance: a Systematic Review

    Marco Crocco, Marco Cristani, Andrea Trucco +1

    cs.SDcs.CVcs.MMarXiv:1409.7787v12014
  23. Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens

    Rafael Valle, Jason Li, Ryan Prenger +1

    cs.SDcs.LGeess.ASarXiv:1910.11997v12019
  24. MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

    Nicolas M. Müller, Piotr Kawa, Wei Herng Choong +6

    cs.SDeess.ASarXiv:2401.09512v112024
  25. The Third DIHARD Diarization Challenge

    Neville Ryant, Prachi Singh, Venkat Krishnamohan +6

    eess.AScs.SDarXiv:2012.01477v32020
  26. Zipformer: A faster and better encoder for automatic speech recognition

    Zengwei Yao, Liyong Guo, Xiaoyu Yang +6

    eess.AScs.LGcs.SDarXiv:2310.11230v42023
  27. Recent Progresses in Deep Learning based Acoustic Models (Updated)

    Dong Yu, Jinyu Li

    eess.AScs.CLcs.SDarXiv:1804.09298v22018
  28. MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

    Ke Chen, Yusong Wu, Haohe Liu +3

    cs.SDcs.AIcs.LGarXiv:2308.01546v12023
  29. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    Ho Kei Cheng, Masato Ishii, Akio Hayakawa +3

    cs.CVcs.LGcs.SDarXiv:2412.15322v22024
  30. What Did I Just Say? Self-Listening for Full-Duplex Speech Models

    Xuanning Zhou, Junyi Ao, Xiaotong Liu +3

    cs.SDarXiv:2609.05592v12026
  31. Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation

    Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu +4

    eess.AScs.CLcs.SDarXiv:2005.08575v52020
  32. Counterpoint by Convolution

    Cheng-Zhi Anna Huang, Tim Cooijmans, Adam Roberts +2

    cs.LGcs.SDeess.ASarXiv:1903.07227v12019
  33. RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification

    Jee-weon Jung, Hee-Soo Heo, Ju-ho Kim +2

    eess.AScs.LGcs.SDarXiv:1904.08104v22019
  34. Margin Matters: Towards More Discriminative Deep Neural Network Embeddings for Speaker Recognition

    Xu Xiang, Shuai Wang, Houjun Huang +2

    eess.AScs.CLcs.SDarXiv:1906.07317v12019
  35. Intermediate Loss Regularization for CTC-based Speech Recognition

    Jaesong Lee, Shinji Watanabe

    eess.AScs.CLcs.SDarXiv:2102.03216v12021
  36. Dense CNN with Self-Attention for Time-Domain Speech Enhancement

    Ashutosh Pandey, DeLiang Wang

    eess.AScs.SDarXiv:2009.01941v22020
  37. RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

    Ngo Truong Dinh, Tung-Lam Bui, Chi-Trung Duong +3

    cs.AIcs.LGcs.MMarXiv:2609.07409v12026
  38. The LOCATA Challenge: Acoustic Source Localization and Tracking

    Christine Evers, Heinrich Loellmann, Heinrich Mellmann +4

    eess.AScs.SDeess.SParXiv:1909.01008v32019
  39. Look, Listen, and Act: Towards Audio-Visual Embodied Navigation

    Chuang Gan, Yiwei Zhang, Jiajun Wu +2

    cs.CVcs.LGcs.ROarXiv:1912.11684v22019
  40. ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    Peng Wang, Shijie Wang, Junyang Lin +5

    cs.CVcs.CLcs.SDarXiv:2305.11172v12023
  41. Speaker Anonymization Using X-vector and Neural Waveform Models

    Fuming Fang, Xin Wang, Junichi Yamagishi +4

    eess.AScs.CLcs.LGarXiv:1905.13561v12019
  42. Serialized Output Training for End-to-End Overlapped Speech Recognition

    Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang +2

    cs.CLcs.SDeess.ASarXiv:2003.12687v22020
  43. Beyond .WAV: Design and Software Verification of VocalCap, a Traceable Browser-Based Audio Capture System for Vocal Biomarker Research

    Augusto Camargo

    cs.SDcs.LGarXiv:2609.03320v12026
  44. The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use

    Bob L. Sturm

    cs.SDarXiv:1306.1461v22013
  45. ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers

    Kaizhi Qian, Yang Zhang, Heting Gao +5

    cs.SDcs.AIeess.ASarXiv:2204.09224v22022
  46. Topic Modeling Based Multi-modal Depression Detection

    Yuan Gong, Christian Poellabauer

    cs.CLcs.IRcs.LGarXiv:1803.10384v12018
  47. SoundStorm: Efficient Parallel Audio Generation

    Zalán Borsos, Matt Sharifi, Damien Vincent +3

    cs.SDcs.LGeess.ASarXiv:2305.09636v12023
  48. A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion

    Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaïdi +3

    eess.AScs.SDarXiv:2111.02392v22021
  49. AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

    Ziyang Ma, Zhikang Niu, Wenming Tu +30

    cs.SDcs.CLcs.MMarXiv:2609.08936v12026
  50. Omni Interaction Agent Technical Report

    Orantqing, Shengpeng Ji, Junlong Tong +20

    eess.AScs.AIcs.LGarXiv:2609.08977v12026
  51. General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

    Eduardo Fonseca, Manoj Plakal, Frederic Font +4

    cs.SDcs.LGeess.ASarXiv:1807.09902v32018
  52. AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

    Yihui Fu, Luyao Cheng, Shubo Lv +10

    cs.SDeess.ASarXiv:2104.03603v42021
  53. Speech Denoising with Deep Feature Losses

    Francois G. Germain, Qifeng Chen, Vladlen Koltun

    eess.AScs.SDarXiv:1806.10522v22018
  54. Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition

    Chao-Han Huck Yang, Jun Qi, Samuel Yen-Chi Chen +4

    cs.SDcs.LGcs.NEarXiv:2010.13309v22020
  55. Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

    Yoto Fujita, Simon Leglaive, Laurent Girin

    cs.SDcs.AIarXiv:2609.03940v12026
  56. Speech Emotion Recognition using Self-Supervised Features

    Edmilson Morais, Ron Hoory, Weizhong Zhu +3

    cs.SDcs.AIcs.LGarXiv:2202.03896v12022
  57. Deep Clustering and Conventional Networks for Music Separation: Stronger Together

    Yi Luo, Zhuo Chen, John R. Hershey +2

    stat.MLcs.LGcs.SDarXiv:1611.06265v22016
  58. Broadcasted Residual Learning for Efficient Keyword Spotting

    Byeonggeun Kim, Simyung Chang, Jinkyu Lee +1

    cs.SDcs.LGeess.ASarXiv:2106.04140v42021
  59. Test-time adaptation for speech enhancement with an autoregressive speech prior

    Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda +1

    cs.SDcs.AIarXiv:2609.03622v12026
  60. A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition

    Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu +3

    eess.AScs.AIcs.CLarXiv:1807.10857v22018