Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| EDMFormer: Genre-Specific Self-Supervised Learning for Music Structure Segmentation Joel Song Bae, Krish Patel, Oscar Chung, Sahal Sajeer Published: 2026-03-08Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-08 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (96%) | - |
| Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech Sai Praneeth Karimireddy, Shrikanth Narayanan, Thanapat Trachu, Thanathai Lertpetchpun Published: 2026-03-08Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-08 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (95%) | - |
| VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription Puneeth N Ail, Sugandha Sharma, Sumit Ranjan, Ubaid Abbas Published: 2026-03-08Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-08 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (97%) | - |
| Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio Chris Donahue, Phillip Long, Zachary Novack Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction Abhirup Ghosh, Hong Jia, Jean Honorio, Jiaheng Dong Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E7 / R3 (95%) | - |
| Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models Lucas Rakotoarivony Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E6 / R4 (96%) | - |
| Fish Audio S2 Technical Report Dawei Han, Jiahua Liu, Qingzheng Wang, Ruoyi Zhang Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E5 / R4 (94%) | - |
| Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis Aishwarya Fursule, Anderson R. Avila, Shruti Kshirsagar Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E7 / R3 (96%) | - |
| VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs Hezhao Zhang, Huang-Cheng Chou, Shrikanth Narayanan, Thomas Hain Published: 2026-03-09Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-09 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models Chih-Kai Yang, Hung-Wei Chen, Hung-yi Lee, Ke-Han Lu Published: 2026-03-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-10 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (99%) | - |
| Physics-Informed Neural Engine Sound Modeling with Differentiable Pulse-Train Synthesis Lonce Wyse, Robin Doerfler Published: 2026-03-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-10 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (95%) | - |
| SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases Angelina Wang, Laya Iyer, Sanmi Koyejo Published: 2026-03-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-10 | cs.SD | ai-safety, cssd, preprint | E7 / R5 (99%) | - |
| TimberAgent: Gram-Guided Retrieval for Executable Music Effect Control Fang Liu, Shengli Zhang, Shihao He, Taotao Wang Published: 2026-03-10Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-10 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow Duojia Li, Haizhou Li, Lin Li, Qingyang Hong Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-11 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Probabilistic Verification of Voice Anti-Spoofing Models Alexandr Kozodaev, Dmitrii Korzh, Evgeny Kushnir, Mikhail Pautov Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-11 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (95%) | - |
| Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation Jesus Villalba-Lopez, Laureano Moro-Velazquez, Najim Dehak, Thomas Thebaud Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint, safety-evaluation | 2026-03-11 | cs.SD | ai-safety, cssd, preprint, safety-evaluation | E5 / R3 (96%) | - |
| Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning Artem Dvirniak, Artem Iudin, Dmitrii Korzh, Dmitrii Tarasov Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-11 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS Aditya Choudhary, Anupam Purwar Published: 2026-03-11Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-11 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2 Suvendu Sekhar Mohanty Published: 2026-03-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-12 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (94%) | - |
| Toward Complex-Valued Neural Networks for Waveform Generation Deok-Hyeon Cho, Hyung-Seok Oh, Seong-Whan Lee, Seung-Bin Kim Published: 2026-03-12Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-12 | cs.SD | ai-safety, cssd, preprint | E5 / R4 (96%) | - |
| LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement Chih-Ning Chen, Fan-Gang Zeng, Hsin-Min Wang, Jen-Cheng Hou Published: 2026-03-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-14 | cs.SD | ai-safety, cssd, preprint | E5 / R4 (97%) | - |
| What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection 脡va Sz茅kely, Harm Lameris, Joakim Gustafson, Shree Harsha Bokkahalli Satish Published: 2026-03-14Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-14 | cs.SD | ai-safety, cssd, preprint | E6 / R3 (93%) | - |
| Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models An-Yu Cheng, Chia-Chien Chen, Chih-Kai Yang, Hung-yi Lee Published: 2026-03-15Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-15 | cs.SD | ai-safety, cssd, preprint | E10 / R3 (93%) | - |
| Music Genre Classification: A Comparative Analysis of Classical Machine Learning and Deep Learning Approaches Abhishek Karna, OmPrakash Dhakl, Sachin Prajuli Published: 2026-03-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-16 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (99%) | - |
| NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation Dekun Chen, Huan Liao, Qinke Ni, Yuxiang Wang Published: 2026-03-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-16 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (94%) | - |
| VorTEX: Various overlap ratio for Target speech EXtraction Bugeun Kim, Jihwan Seol, Ro-hoon Oh Published: 2026-03-16Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-16 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| Diffusion Models for Joint Audio-Video Generation Alejandro Paredes La Torre Published: 2026-03-17Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-17 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (95%) | - |
| MOSS-TTS Technical Report Botian Jiang, Cheng Chang, Dong Hong, Hanfu Chen Published: 2026-03-18Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-18 | cs.SD | ai-safety, cssd, preprint | E4 / R3 (98%) | - |
| MOSS-TTS Technical Report Botian Jiang, Cheng Chang, Dong Hong, Hanfu Chen Published: 2026-03-18Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-18 | cs.SD | ai-safety, cssd, preprint | E5 / R3 (96%) | - |
| Voice Privacy from an Attribute-based Perspective Cristian Tejedor Garc铆a, Martha Larson, Mehtab Ur Rahman Published: 2026-03-19Area: cs.SDCitations: - Tags: ai-safety, cssd, preprint | 2026-03-19 | cs.SD | ai-safety, cssd, preprint | E25 / R29 (96%) | - |