Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model Guiping Zhong, Haiyun Li, Hui Lu, Huimeng Wang Published: 2026-04-21Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-21 | cs.SD | ai-safety, cssd, preprint | E10 / R5 (97%) | - |
| ATIR: Towards Audio-Text Interleaved Contextual Retrieval Chenghao Zhang, Tong Zhao, Yutao Zhu, Zhicheng Dou Published: 2026-04-22Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-22 | cs.SD | ai-safety, cssd, preprint | E7 / R5 (97%) | - |
| Enhancing Speaker Verification with Whispered Speech via Post-Processing Magdalena Go艂臋biowska, Piotr Syga Published: 2026-04-22Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-22 | cs.SD | ai-safety, cssd, preprint | E9 / R4 (97%) | - |
| ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence Fanhong Meng, Haoran Luo, Luu Anh Tuan, Menghe Ma Published: 2026-04-22Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-22 | cs.SD | ai-safety, cssd, preprint | E11 / R4 (97%) | - |
| All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation Chen-An Li, Chih-Kai Yang, Hung-yi Lee, Ke-Han Lu Published: 2026-04-27Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-04-27 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E10 / R4 (100%) | - |
| RAS: a Reliability Oriented Metric for Automatic Speech Recognition Bohan Li, Hankun Wang, Jing Peng, Kai Yu Published: 2026-04-27Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-27 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (97%) | - |
| Speech Enhancement Based on Drifting Models Bastiaan Kleijn, Diego Caviedes-Nozal, Liang Xu, Longfei Felix Yan Published: 2026-04-27Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-27 | cs.SD | ai-safety, cssd, preprint | E10 / R5 (99%) | - |
| AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction Amir M. Rahmani, Henry Peng Zou, Hoang H Nguyen, Honghui Xu Published: 2026-06-14Area: cs.SDCitations: - Tags: ai-safety, alignment-training, cssd, preprint | 2026-06-14 | cs.SD | ai-safety, alignment-training, cssd, preprint | E10 / R4 (98%) | - |
| Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech Eng Siong Chng, Haoyang Li, Hardik B. Sailor, Hexin Liu Published: 2026-06-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-14 | cs.SD | ai-safety, cssd, preprint | E8 / R4 (97%) | - |
| NVMOS: Non-Verbal Vocalization Quality Assessment in Speech Jialong Mai, Jinxin Ji, Wencui Liu, Xiangmin Xu Published: 2026-06-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-14 | cs.SD | ai-safety, cssd, preprint | E7 / R4 (96%) | - |
| ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion Byungchan Hwang, Hak Gu Kim, Hyung Kyu Kim Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (98%) | - |
| ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition Fuliang Weng, Shu Shang, Yaqian Zhou, Zeqian Hu Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (99%) | - |
| Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection Chunhong Yuan, Hugen Lv, Xiangyu Li, Zhuodong Liu Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E15 / R7 (97%) | - |
| MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild Gabriel Skantze, Haotian Qi Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E9 / R4 (98%) | - |
| Probing Low Frame Rate Degradation in Neural Audio Codecs Alex Gichamba, Moise Busogi Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E12 / R4 (97%) | - |
| TuneJury: An Open Metric for Improving Music Generation Preference Alignment Chris Donahue, Haiwen Xia, Junghyun Koo, Junwon Lee Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, alignment-training, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, alignment-training, cssd, preprint | E9 / R5 (99%) | - |
| Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control Dong-Min Byun, Joon-Seung Choi, Seong-Whan Lee Published: 2026-06-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-15 | cs.SD | ai-safety, cssd, preprint | E8 / R5 (98%) | - |
| A Neuromorphic Trigger for Efficient Audio Event Detection Benjamin Hatton, Luca Peres, Oliver Rhodes Published: 2026-06-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-16 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (98%) | - |
| Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD) Ryan Quinn, Sinclair Gurny Published: 2026-06-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-16 | cs.SD | ai-safety, cssd, preprint | E7 / R4 (97%) | - |
| L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification Deok-Hyeon Cho, Hyung-Seok Oh, Seong-Whan Lee, Seung-Bin Kim Published: 2026-06-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-16 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (97%) | - |
| MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data Jason Li, Paarth Neekhara, Roy Fejgin, Ryan Langman Published: 2026-06-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-16 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (97%) | - |
| Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation Ioannis Prokopiou, Maximos Kaliakatsos-Papakostas, Pantelis Vikatos, Themos Stafylakis Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E7 / R5 (98%) | - |
| Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification Ryan Quinn, Sinclair Gurny Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E8 / R4 (97%) | - |
| FlowFake: Liquid Networks for Audio Deepfake Detection Dinesh Kumar Vishwakarma, Divyansh Sharma, Shivaay Dhondiyal Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E10 / R5 (98%) | - |
| NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization Junqiao Fan, Lihua Xie, Shenghai Yuan, Yizhuo Yang Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (97%) | - |
| PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets Donald S. Williamson, Junyi Fan Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E13 / R5 (97%) | - |
| QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement Hideaki Tamori, Makoto Sakai, Shogo Yamauchi, Tohru Nitta Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E9 / R5 (97%) | - |
| Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors Benjamin Brazowski, Daniel Segal, Eitan Richardson, Michael Finkelson Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E8 / R4 (97%) | - |
| RIVET: Robust Idempotent Voice Attribute Editing Bhiksha Raj, Bhuvan Koduru, Dareen Alharthi, Rita Singh Published: 2026-06-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-17 | cs.SD | ai-safety, cssd, preprint | E7 / R5 (97%) | - |
| Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow Dongyu Wang, Jean-Yves Guillemaut, Liting Gao, Shubin Zhang Published: 2026-06-18Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-06-18 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (97%) | - |