Sound
Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
901 to 960 of 1,026
Woosh: A Sound Effects Foundation Model
Gaëtan Hadjeres, Marc Ferras, Khaled Koutini +7
cs.SDcs.AIcs.LGarXiv:2604.01929v32026Do Audio-Visual Large Language Models Really See and Hear?
Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3
cs.AIcs.SDarXiv:2604.02605v12026A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
Tianle Chen, Deepti Ghadiyaram
cs.CVcs.SDarXiv:2604.03995v12026Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +15
cs.SDcs.AIcs.CLarXiv:2604.10905v12026Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
Zeyue Tian, Binxin Yang, Zhaoyang Liu +8
cs.SDcs.AIcs.CVarXiv:2604.10708v22026BMdataset: A Musicologically Curated LilyPond Dataset
Matteo Spanio, Ilay Guler, Antonio Rodà
cs.SDcs.CLcs.IRarXiv:2604.10628v22026Hierarchical Codec Diffusion for Video-to-Speech Generation
Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4
cs.SDcs.CVarXiv:2604.15923v12026ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
Heewon Oh
cs.SDeess.ASarXiv:2604.16254v22026VoxMind: An End-to-End Agentic Spoken Dialogue System
Tianle Liang, Yifu Chen, Shengpeng Ji +7
cs.SDarXiv:2604.15710v12026Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
Jaechul Roh, Amir Houmansadr
cs.CRcs.SDarXiv:2604.16659v12026MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
Szu-Chi Chen, I-Ning Tsai, Yi-Cheng Lin +2
cs.CLcs.AIcs.SDarXiv:2604.17435v12026Tadabur: A Large-Scale Quran Audio Dataset
Faisal Alherran
cs.SDcs.AIarXiv:2604.18932v12026AudioWorldSim: Realistic Binaural Audio Datasets For World Models
Luis Vitor Zerkowski, Luiz Velho
cs.SDcs.LGarXiv:2608.21075v12026Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
Kawshik Kumar Paul, Md. Nafiul Alam Fuji
cs.CLcs.SDeess.ASarXiv:2608.20346v12026A Factorial Ablation of a Speech-to-SFT Pipeline: Differential Effects on Data Quality and Downstream Transfer
Wonsup Shin, Jingu Kim
cs.SDcs.CLarXiv:2608.20394v12026Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing
L. Choy, A. S. Khan, S. Patrizi +3
cs.CLcs.SDarXiv:2608.20396v12026TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
Vladimir Bataev, Lilit Grigoryan, Andrei Andrusenko +3
eess.AScs.AIcs.CLarXiv:2608.21343v12026DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization
Naiyuan Li, Li Dong, Diqun Yan
cs.SDcs.AIarXiv:2608.21176v12026Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding
Haoyue Liu, Zhichao Wang, Ye Chen +2
cs.SDcs.AIarXiv:2608.20769v12026XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Arun Babu, Changhan Wang, Andros Tjandra +10
cs.CLcs.SDeess.ASarXiv:2111.09296v32021MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere +6
eess.AScs.CLcs.LGarXiv:1910.06711v32019Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
Yusong Wu, Ke Chen, Tianyu Zhang +4
cs.SDeess.ASarXiv:2211.06687v42022WaveGlow: A Flow-based Generative Network for Speech Synthesis
Ryan Prenger, Rafael Valle, Bryan Catanzaro
cs.SDcs.AIcs.LGarXiv:1811.00002v12018State-of-the-art Speech Recognition With Sequence-to-Sequence Models
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu +11
cs.CLcs.SDeess.ASarXiv:1712.01769v62017A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri +1
cs.CVcs.LGcs.SDarXiv:2008.10010v12020SUPERB: Speech processing Universal PERformance Benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang +17
cs.CLcs.SDeess.ASarXiv:2105.01051v42021The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
Venkata Pushpak Teja Menta
cs.CLcs.SDarXiv:2605.03073v12026Liberating LLM Capabilities in Full-Duplex Speech Models
Luoyuan Zhang, Bokai Xu, Junbo Cui +4
cs.CLcs.AIcs.SDarXiv:2606.07547v12026APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
Jaavid Aktar Husain, Dorien Herremans
cs.SDcs.AIcs.LGarXiv:2605.03395v22026High Fidelity Neural Audio Compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve +1
eess.AScs.AIcs.SDarXiv:2210.13438v12022SEGAN: Speech Enhancement Generative Adversarial Network
Santiago Pascual, Antonio Bonafonte, Joan Serrà
cs.LGcs.NEcs.SDarXiv:1703.09452v32017Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu +10
cs.CLcs.SDeess.ASarXiv:2301.02111v12023Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Jaehyeon Kim, Jungil Kong, Juhee Son
cs.SDeess.ASarXiv:2106.06103v12021LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Viet Dang, Rob Clark +5
cs.SDeess.ASarXiv:1904.02882v12019AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, James Glass
cs.SDcs.AIarXiv:2104.01778v32021Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification
Justin Salamon, Juan Pablo Bello
cs.SDcs.CVcs.LGarXiv:1608.04363v22016Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock +3
cs.CVcs.AIcs.LGarXiv:2103.03206v22021PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal +3
cs.SDeess.ASarXiv:1912.10211v52019SDR - half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan +1
cs.SDeess.ASarXiv:1811.02508v12018MUSAN: A Music, Speech, and Noise Corpus
David Snyder, Guoguo Chen, Daniel Povey
cs.SDarXiv:1510.08484v12015MERIT: Learning Disentangled Music Representations for Audio Similarity
Abhinaba Roy, Junyi Liang, Dorien Herremans
cs.SDarXiv:2605.27346v12026Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
Yi Wang
cs.LGcs.AIcs.SDarXiv:2608.18025v12026DEMON: Diffusion Engine for Musical Orchestrated Noise
Ryan Fosdick
cs.SDarXiv:2605.28657v12026ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
Tiantian Feng, Anfeng Xu, Xuan Shi +10
cs.SDarXiv:2605.29257v12026FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Yi Ren, Chenxu Hu, Xu Tan +4
eess.AScs.CLcs.LGarXiv:2006.04558v82020Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Hang Zhang, Xin Li, Lidong Bing
cs.CLcs.CVcs.SDarXiv:2306.02858v42023DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang +2
eess.AScs.CLcs.LGarXiv:2009.09761v32020ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Brecht Desplanques, Jenthe Thienpondt, Kris Demuynck
eess.AScs.SDarXiv:2005.07143v32020Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton +11
cs.CLcs.LGcs.SDarXiv:1703.10135v22017Audio Interaction Model
Zhifei Xie, Zihang Liu, Ze An +8
cs.SDcs.AIcs.CLarXiv:2606.05121v12026Where Flow Matching Leaks: Characterising Membership Signals Along the Interpolation Path
Thomas Sesmat, Gabriel Meseguer-Brocal, Geoffroy Peeters
cs.LGcs.AIcs.SDarXiv:2606.07271v32026How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling
Jinju Lee
cs.SDcs.LGarXiv:2606.07334v42026dots.tts Technical Report
Shi Lian, Changtao Li, Bohan Li +6
cs.SDcs.AIeess.ASarXiv:2606.07080v22026VoxCeleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, Andrew Zisserman
cs.SDarXiv:1706.08612v22017VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, Andrew Zisserman
cs.SDcs.CVeess.ASarXiv:1806.05622v22018HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, Jaekyoung Bae
cs.SDcs.LGeess.ASarXiv:2010.05646v22020CNN Architectures for Large-Scale Audio Classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis +10
cs.SDcs.LGstat.MLarXiv:1609.09430v22016Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
Yash Vishe, Eric Xue, Xunyi Jiang +4
cs.SDcs.AIarXiv:2512.14629v22025Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Umberto Cappellazzo, Xubo Liu, Stavros Petridis +1
eess.AScs.AIcs.SDarXiv:2608.19863v12026Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Ryota Komatsu, Kota Kawakita, Takuma Okamoto +1
cs.CLcs.AIcs.SDarXiv:2607.04064v12026