Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection Bin Chen, Hong Jia, Jisheng Bai, Ting Dang Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation Jiahe Lei, Jiaxing Yu, Kejun Zhang, Lei Wang Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | E10 / R9 (90%) | - |
| RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction Ambuj Mehrish, Sebastiano Vascon Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents Botian Jiang, Kexin Huang, Min Liang, Shuang Chen Published: 2026-08-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-09 | cs.SD | ai-safety, cssd, preprint | E8 / R6 (93%) | - |
| PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue Junfeng Ma, Libo Wang, Shibo Wang, Zicheng Zhang Published: 2026-08-07Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint | 2026-08-07 | cs.SD | ai-safety, cssd, preprint | E8 / R7 (94%) | 15 |
| Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Alexandre D'Hooge, Eoin Cummins, Yaolong Ju, Zhongyi Huang Published: 2026-08-06Area: cs.SDCitations: 1 Tags: ai-safety, cssd, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, preprint | E8 / R8 (91%) | 1 |
| Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Alexandre D'Hooge, Eoin Cummins, Yaolong Ju, Zhongyi Huang Published: 2026-08-06Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, preprint | E7 / R7 (92%) | 15 |
| MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering Jakub Po膰wiardowski, Mateusz Modrzejewski Published: 2026-08-06Area: cs.SDCitations: - Tags: ai-safety, cssd, interpretability, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, interpretability, preprint | E11 / R13 (91%) | - |
| AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation Jinting Wang, Li Liu, Shan Yang, Shengyu Li Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-08-05 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | - | - |
| HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models Ding Luo, Jin Zeng, Ruikang Zhang, Tian Jin Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-05 | cs.SD | ai-safety, cssd, preprint | E10 / R10 (94%) | - |
| Masked diffusion enables coherent beat tracking Filip Korzeniowski, Francesco Foscarin, Richard Vogl Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-05 | cs.SD | ai-safety, cssd, preprint | - | - |
| AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities Adam Dubrowski, Bill Kapralos, KC Collins, Priyamvada Tripathi Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| Equivariant Music Transformer Simon Dixon, Zixun Guo Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion Alon Ziv, Harel Pogoda, Yossi Adi Published: 2026-08-04Area: cs.SDCitations: 18 Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | E13 / R11 (89%) | 18 |
| Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping Gus Xia, Jingwei Zhao, Ye Wang, Ziyu Wang Published: 2026-08-04Area: cs.SDCitations: 45 Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | E15 / R10 (93%) | 45 |
| Multi-Task Multi-Frame Visual Piano Transcription Alexander Lerch, Hoyeol Sohn, Juhan Nam, Yonghyun Kim Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification Ahmad ElShiekh, Ahmed Rashad, Ghassan Al-Sumaidaee, Sajjad Abdoli Published: 2026-08-03Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-03 | cs.SD | ai-safety, cssd, preprint | - | - |
| Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior Ehsan Yaghoubi, Florian Haselbeck Published: 2026-08-03Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-03 | cs.SD | ai-safety, cssd, preprint | - | - |
| dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model Bohan Li, Colin Zhang, Da Zheng, Hankun Wang Published: 2026-08-02Area: cs.SDCitations: 45 Tags: ai-safety, cssd, preprint | 2026-08-02 | cs.SD | ai-safety, cssd, preprint | E11 / R8 (92%) | 45 |
| UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations Chenhao Lin, Nan Nan, Xiaohong Guan, Ziyue Kang Published: 2026-08-01Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-01 | cs.SD | ai-safety, cssd, preprint | E10 / R8 (91%) | - |
| DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs Zhenhua Tan, Zhuomin Zhu, Ziwei Cheng Published: 2026-07-31Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-31 | cs.SD | ai-safety, cssd, preprint | E11 / R9 (94%) | - |
| Teffic-Audio: Tell Fact from Fiction Jindong Wang, Kunyu Feng, Li Wang, Wan Lin Published: 2026-07-30Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-30 | cs.SD | ai-safety, cssd, preprint | E10 / R9 (93%) | - |
| TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models Abhishek Mukherji, Aryan Vijay Bhosale, Dinesh Manocha, Harshit Rajgarhia Published: 2026-07-30Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-30 | cs.SD | ai-safety, cssd, preprint | E18 / R17 (91%) | - |
| VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition EngSiong Chng, Tianrui Wang, Xinyu Yang, Yukun Chen Published: 2026-07-30Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-30 | cs.SD | ai-safety, cssd, preprint | E32 / R17 (87%) | - |
| Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection Gencheng Liu, Haotian Mo, Jie Liu, Keqi Yang Published: 2026-07-29Area: cs.SDCitations: 1 Tags: ai-safety, cssd, preprint | 2026-07-29 | cs.SD | ai-safety, cssd, preprint | E7 / R8 (93%) | 1 |
| MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Junbo Li, Jun Fang, Lin Li, Qingyang Hong Published: 2026-07-29Area: cs.SDCitations: 28 Tags: ai-safety, cssd, preprint | 2026-07-29 | cs.SD | ai-safety, cssd, preprint | E14 / R12 (87%) | 28 |
| MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Junbo Li, Jun Fang, Lin Li, Qingyang Hong Published: 2026-07-29Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-29 | cs.SD | ai-safety, cssd, preprint | - | - |
| MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation Chih-Pin Tan, Fang-Duo Tsai, Hsuan-Yu Yeh, Ting-Yi Hu Published: 2026-07-29Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-29 | cs.SD | ai-safety, cssd, preprint | E10 / R8 (91%) | - |
| Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage Ernest Fokou茅, Vivek Senthil, Zhiqiang Tao Published: 2026-07-27Area: cs.SDCitations: 5 Tags: ai-safety, cssd, preprint | 2026-07-27 | cs.SD | ai-safety, cssd, preprint | E7 / R5 (90%) | 5 |
| Probing Speaker Identity Sensitivity in Audio Deepfake Detectors Arun Ross, Daniyal Kabir Dar Published: 2026-07-23Area: cs.SDCitations: 34 Tags: ai-safety, cssd, preprint | 2026-07-23 | cs.SD | ai-safety, cssd, preprint | E9 / R8 (93%) | 34 |