Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs Zhenhua Tan, Zhuomin Zhu, Ziwei Cheng Published: 2026-07-31Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-07-31 | cs.SD | ai-safety, cssd, preprint | E11 / R9 (94%) | - |
| UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations Chenhao Lin, Nan Nan, Xiaohong Guan, Ziyue Kang Published: 2026-08-01Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-01 | cs.SD | ai-safety, cssd, preprint | E10 / R8 (91%) | - |
| dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model Bohan Li, Colin Zhang, Da Zheng, Hankun Wang Published: 2026-08-02Area: cs.SDCitations: 45 Tags: ai-safety, cssd, preprint | 2026-08-02 | cs.SD | ai-safety, cssd, preprint | E11 / R8 (92%) | 45 |
| Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification Ahmad ElShiekh, Ahmed Rashad, Ghassan Al-Sumaidaee, Sajjad Abdoli Published: 2026-08-03Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-03 | cs.SD | ai-safety, cssd, preprint | - | - |
| Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior Ehsan Yaghoubi, Florian Haselbeck Published: 2026-08-03Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-03 | cs.SD | ai-safety, cssd, preprint | - | - |
| AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities Adam Dubrowski, Bill Kapralos, KC Collins, Priyamvada Tripathi Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| Equivariant Music Transformer Simon Dixon, Zixun Guo Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion Alon Ziv, Harel Pogoda, Yossi Adi Published: 2026-08-04Area: cs.SDCitations: 18 Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | E13 / R11 (89%) | 18 |
| Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping Gus Xia, Jingwei Zhao, Ye Wang, Ziyu Wang Published: 2026-08-04Area: cs.SDCitations: 45 Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | E15 / R10 (93%) | 45 |
| Multi-Task Multi-Frame Visual Piano Transcription Alexander Lerch, Hoyeol Sohn, Juhan Nam, Yonghyun Kim Published: 2026-08-04Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-04 | cs.SD | ai-safety, cssd, preprint | - | - |
| AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation Jinting Wang, Li Liu, Shan Yang, Shengyu Li Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-08-05 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | - | - |
| HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models Ding Luo, Jin Zeng, Ruikang Zhang, Tian Jin Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-05 | cs.SD | ai-safety, cssd, preprint | E10 / R10 (94%) | - |
| Masked diffusion enables coherent beat tracking Filip Korzeniowski, Francesco Foscarin, Richard Vogl Published: 2026-08-05Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-05 | cs.SD | ai-safety, cssd, preprint | - | - |
| Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Alexandre D'Hooge, Eoin Cummins, Yaolong Ju, Zhongyi Huang Published: 2026-08-06Area: cs.SDCitations: 1 Tags: ai-safety, cssd, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, preprint | E8 / R8 (91%) | 1 |
| Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Alexandre D'Hooge, Eoin Cummins, Yaolong Ju, Zhongyi Huang Published: 2026-08-06Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, preprint | E7 / R7 (92%) | 15 |
| MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering Jakub Po膰wiardowski, Mateusz Modrzejewski Published: 2026-08-06Area: cs.SDCitations: - Tags: ai-safety, cssd, interpretability, preprint | 2026-08-06 | cs.SD | ai-safety, cssd, interpretability, preprint | E11 / R13 (91%) | - |
| PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue Junfeng Ma, Libo Wang, Shibo Wang, Zicheng Zhang Published: 2026-08-07Area: cs.SDCitations: 15 Tags: ai-safety, cssd, preprint | 2026-08-07 | cs.SD | ai-safety, cssd, preprint | E8 / R7 (94%) | 15 |
| CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents Botian Jiang, Kexin Huang, Min Liang, Shuang Chen Published: 2026-08-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-09 | cs.SD | ai-safety, cssd, preprint | E8 / R6 (93%) | - |
| Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions Anke Koelzer, Fanny Riols, Hoang H Nguyen, Jash Shah Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation Haitao Qian, Qianxiao Fang, Wanyi Ning, Wei Zhou Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs Chen Li, Haoran Gao, Jie Ren, Kun Wang Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection Bin Chen, Hong Jia, Jisheng Bai, Ting Dang Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation Jiahe Lei, Jiaxing Yu, Kejun Zhang, Lei Wang Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | E10 / R9 (90%) | - |
| RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction Ambuj Mehrish, Sebastiano Vascon Published: 2026-08-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-10 | cs.SD | ai-safety, cssd, preprint | - | - |
| DuplexWorld: Can voice agents help you get through the day? Abhishek Mukherji, Akhil Pothanapalli, Aryan Vijay Bhosale, Asif Shaik Published: 2026-08-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-11 | cs.SD | ai-safety, cssd, preprint | E14 / R13 (94%) | - |
| Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models Dongxiao Liu, Ji Zhang, Kunlan Xiang, Mingxuan Li Published: 2026-08-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-11 | cs.SD | ai-safety, cssd, preprint | E9 / R7 (93%) | - |
| Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition Gaopeng Xu, Haitao Yao, Yinfeng Xia, Zheng Xue Published: 2026-08-11Area: cs.SDCitations: 20 Tags: ai-safety, cssd, preprint | 2026-08-11 | cs.SD | ai-safety, cssd, preprint | E10 / R7 (93%) | 20 |
| HybridSB-MoE: Dual-Domain Schr枚dinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement Aswini Sivakumar, Jie Hu, Yao Qiang, Zhengyi Lu Published: 2026-08-13Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-13 | cs.SD | ai-safety, cssd, preprint | E8 / R7 (92%) | - |
| Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels Andrii Shevtsov, Vadym Vilhurin, Volodymyr Sydorskyi Published: 2026-08-14Area: cs.SDCitations: 39 Tags: ai-safety, cssd, preprint | 2026-08-14 | cs.SD | ai-safety, cssd, preprint | E10 / R9 (94%) | 39 |
| Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task Alexandru-Stefan Morosanu, Laura Erhan, Stefan-Daniel Achirei, Valerian Cecan Published: 2026-08-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-08-14 | cs.SD | ai-safety, cssd, preprint | E8 / R8 (88%) | - |