Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models Martin Takáč, Salem Lahlou, Yanda Li, Yuhan Liu Published: 2026-04-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-16 | cs.SD | ai-safety, cssd, preprint | E9 / R4 (98%) | - |
| Comparison of window shapes and lengths in short-time feature extraction for classification of heart sound signals Abeer FathAllah Brery, Mahmoud Fakhry Published: 2026-04-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-15 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Ian McLoughlin, Jun Liu, Lirong Dai, Nan Jiang Published: 2026-04-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-15 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Audio Source Separation in Reverberant Environments using $β$-divergence based Nonnegative Factorization Mahmoud Fakhry, Maurizio Omologo, Piergiorgio Svaizer Published: 2026-04-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-14 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Elastic Net Regularization and Gabor Dictionary for Classification of Heart Sound Signals using Deep Learning Ascensión Gallardo-Antolín, Mahmoud Fakhry Published: 2026-04-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-14 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing Wei Xue, Xi Chen, Yike Guo Published: 2026-04-13Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-13 | cs.SD | ai-safety, cssd, preprint | E7 / R3 (96%) | - |
| Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music Arushi Goel, Aya Aljafari, Bryan Catanzaro, Chao-Han Huck Yang Published: 2026-04-13Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-13 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (96%) | - |
| Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing Binxin Yang, Chen Li, Hubery Yin, Jiexuan Zhang Published: 2026-04-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-12 | cs.SD | ai-safety, cssd, preprint | E5 / R4 (96%) | - |
| Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music Ajay Pundhir, Shivam Chauhan Published: 2026-04-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-12 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (96%) | - |
| MeloTune: On-Device Arousal Learning and Peer-to-Peer Mood Coupling for Proactive Music Curation Hongwei Xu Published: 2026-04-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-12 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (97%) | - |
| MeloTune: On-Device Arousal Learning and Peer-to-Peer Mood Coupling for Proactive Music Curation Hongwei Xu Published: 2026-04-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-12 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (96%) | - |
| VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories Qian Zhang, Xiongkuo Min, Yixuan Gao, Yuqin Cao Published: 2026-04-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-12 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (96%) | - |
| AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models Bo Li, Chen Fang, Mintong Kang Published: 2026-04-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-10 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (97%) | - |
| DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech Sebastian Stober, Suhita Ghosh, Yamini Sinha Published: 2026-04-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-10 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (96%) | - |
| GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking Di Wu, Guocong Quan, Hengyuan Na, Miao Hu Published: 2026-04-10Area: cs.SDCitations: - Tags: adversarial-robustness, ai-safety, cssd, preprint | 2026-04-10 | cs.SD | adversarial-robustness, ai-safety, cssd, preprint | E5 / R3 (94%) | - |
| Noise-Aware In-Context Learning for Hallucination Mitigation in ALLMs Khalid Zaman, Masashi Unoki, Qixuan Huang Published: 2026-04-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-10 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan Guangtao Zhai, Haonan Cheng, Hengyan Huang, Jian Liu Published: 2026-04-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-04-09 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Selective Attention System (SAS): Device-Addressed Speech Detection for Real-Time On-Device Voice AI Bonny Banerjee, Daniyal Anjum, David Joohun Kim, Omar Abbasi Published: 2026-04-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-09 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (96%) | - |
| Towards Real-Time Human-AI Musical Co-Performance: Accompaniment Generation with Latent Diffusion Models and MAX/MSP Shlomo Dubnov, Tornike Karchkhadze Published: 2026-04-08Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-08 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (98%) | - |
| Anchored Cyclic Generation: A Novel Paradigm for Long-Sequence Symbolic Music Generation Boyu Cao, Dehan Li, Haoyu Gu, Lekai Qian Published: 2026-04-07Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-07 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (95%) | - |
| A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech Evangelos Kanoulas, Hongyi Zhu, Jia-Hong Huang, Prayag Tiwari Published: 2026-04-07Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-07 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (97%) | - |
| Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Jihua Zhu, Wenyu Wang, Xin Gao, Yiquan Zhou Published: 2026-04-07Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-07 | cs.SD | ai-safety, cssd, preprint | E5 / R4 (95%) | - |
| Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization Adam Rothschild, Ahmed Hassoon, Andrew Perrault, David Grünert Published: 2026-04-07Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-07 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction Jia Li, Yinfeng Yu Published: 2026-04-06Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-06 | cs.SD | ai-safety, cssd, preprint | E7 / R4 (97%) | - |
| YMIR: A new Benchmark Dataset and Model for Arabic Yemeni Music Genre Classification Using Convolutional Neural Networks Abdulrahman A. AlKannad, Eiad Almekhlafi, Moeen AL-Makhlafi, Nawaf Q. Othman Ahmed Mohammed Published: 2026-04-06Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-06 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (99%) | - |
| Split and Conquer Partial Deepfake Speech Haim Permuter, Inbal Rimon, Oren Gal Published: 2026-04-03Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-03 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Acoustic and perceptual differences between standard and accented Chinese speech and their voice clones Chengzhe Sun, Phil Rose, Siwei Lyu, Tianle Yang Published: 2026-04-02Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-02 | cs.SD | ai-safety, cssd, preprint | E7 / R3 (95%) | - |
| Woosh: A Sound Effects Foundation Model Alexandre Bittar, Benno Weck, Gaëtan Hadjeres, Hakim Missoum Published: 2026-04-02Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-02 | cs.SD | ai-safety, cssd, preprint | E7 / R4 (97%) | - |
| Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors Anton Firc, Kamil Malinka, Lukáš Sekanina, Martin Perešíni Published: 2026-04-01Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-01 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| TRACE: Training-Free Partial Audio Deepfake Detection via Embedding Trajectory Analysis of Speech Foundation Models Awais Khan, Khalid Malik, Kutub Uddin, Muhammad Umar Farooq Published: 2026-04-01Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-04-01 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (98%) | - |