Sound

Papers filed under cs.SD on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,021 to 1,027 of 1,027

  1. Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

    Fangxu Yu, Tao Feng, Dehai Min +6

    cs.SDcs.CLarXiv:2608.02831v12026
  2. Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP

    Błażej Kotowski, Frederic Font

    cs.SDcs.HCcs.LGarXiv:2608.13724v12026
  3. Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

    Ilia Semenkov, Daria Kleeva, Ivan Dakhtin +2

    cs.LGcs.SDq-bio.NCarXiv:2608.01481v12026
  4. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    Yu Zhang, Ruiqi Li, Changhao Pan +3

    eess.AScs.SDarXiv:2608.02023v22026
  5. From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

    Yuanhe Zhang, Weiliu Wang, Jie Ren +7

    cs.SDcs.AIarXiv:2608.09158v12026
  6. AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

    Yuqing Wen, Yukai Huang, Qianqian Xie +6

    cs.MMcs.CVcs.SDarXiv:2607.24821v12026
  7. KVAE: Family of Tokenizers for Multimodal Generative Models

    Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11

    cs.CVcs.LGcs.SDarXiv:2608.05798v12026