Paper deep dive
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler, Yuzhe Wang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velázquez
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 7/5/2026, 4:21:41 AM
Summary
The paper introduces DyadEE, a new dataset for emotional entrainment detection in dyadic speech, and TRACE, a window-level framework that models interaction as ordered sequences of acoustic embeddings. TRACE uses emotion-fine-tuned Whisper representations and incorporates conversational context and social relationship information to improve detection accuracy. Experimental results show that TRACE achieves 97.01% accuracy by leveraging both contextual and relational signals, outperforming baselines like Emotion MLP and DyadFormer.
Entities (7)
Relation Signals (4)
DyadEE → builton → Seamless Interaction corpus
confidence 100% · Built on the Seamless Interaction corpus [15], our dataset includes...
TRACE → evaluatedon → DyadEE
confidence 100% · Experimental results on DyadEE show that... TRACE achieving the best accuracy of 97.01%
TRACE → outperforms → DyadFormer
confidence 100% · TRACE achieves the best accuracy of 97.01%... DyadFormer... achieves competitive speech-only accuracy (92.20%)
TRACE → uses → Whisper
confidence 100% · derived from emotion fine-tuned Whisper representations
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by social relationships and conversational context, influencing affective coordination over time. We introduce DyadEE, a dataset for emotional entrainment detection in dyadic speech interactions, containing both emotionally entrained conversations and synthetic interactions where entrainment is disrupted through partner swapping and emotion resynthesis. We further propose TRACE, a window-level framework that models dyadic interaction as ordered sequences of acoustic embeddings derived from emotion fine-tuned Whisper representations, treating each sample as an interaction trace rather than pooled utterances. Experimental results on DyadEE show that incorporating conversational context and relationship information improves emotional entrainment detection, with TRACE achieving the best accuracy of 97.01%.
Tags
Links
- Source: https://arxiv.org/abs/2606.30543v1
- Canonical: https://arxiv.org/abs/2606.30543v1
Trouble viewing inline? Open PDF directly →
Full Text
28,983 characters extracted from source content.
Expand or collapse full text
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech Sathvik Manikantan Napa Ugandhar 1,*,† , Hao Zhang 1,*,† , Alison Gunzler 1 , Yuzhe Wang 2 Thomas Thebaud 2 , Georgi Tinchev 3 , Venkatesh Ravichandran 4 , Laureano Moro-Vel ́ azquez 2,† 1 Department of Computer Science, Johns Hopkins University, USA 2 Department of Electrical and Computer Engineering, Johns Hopkins University, USA 3 Amazon Research, UK 4 Amazon, USA snapaug1@jh.edu, hzhan276@jh.edu, laureano@jhu.edu Abstract—With the proliferation of speech AI agents, under- standing emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by social relationships and conversational context, influencing affective coordination over time. We introduce DyadEE, a dataset for emotional entrainment detection in dyadic speech interac- tions, containing both emotionally entrained conversations and synthetic interactions where entrainment is disrupted through partner swapping and emotion resynthesis. We further propose TRACE, a window-level framework that models dyadic inter- action as ordered sequences of acoustic embeddings derived from emotion fine-tuned Whisper representations, treating each sample as an interaction trace rather than pooled utterances. Experimental results on DyadEE show that incorporating conver- sational context and relationship information improves emotional entrainment detection, with TRACE achieving the best accuracy of 97.01%. Index Terms—emotional entrainment detection, dyadic speech, computational paralinguistics I. INTRODUCTION Entrainment is the natural tendency of two participants in a conversation to align with each other over the course of an interaction. Dyadic speech interaction manifests across multiple dimensions, including speech rate, lexical choice, prosody, and affective adaptation [1], [2]. Among these, emo- tional entrainment, the mutual adaptation of affective states between speakers during interaction, is strongly shaped by the relationship between speakers and the conversational context [3]. In natural conversations, emotional alignment is not lim- ited to mirroring: partners converge affectively in affiliative contexts, while complementary or regulatory responses are more natural in conflict, negotiation, or support-seeking sce- narios [4]. Emotional entrainment is therefore best understood as a relationship-conditioned and context-conditioned process rather than a context-free single global similarity score [5]. Understanding emotional entrainment has direct practical value for conversational speech AI. In speech-to-speech agent * Equal contribution. † Corresponding author. systems, a response style that is appropriate in peer interaction may be inappropriate or even unsafe in a clinician-patient setting. Models deployed in emotionally sensitive domains such as companionship, mental health support, or professional assistance must therefore adapt their prosodic and affective behavior to both their social role and the evolving conversa- tional context [6], [7]. Evaluations that ignore these factors can produce misleading conclusions, rewarding globally pleasant responses that violate role-specific norms and masking failures that emerge only under particular interaction settings [8]. This creates a clear need for benchmarks that assess emotional en- trainment under explicit contextual and relational constraints. Prior work has explored related aspects of affective coor- dination in conversation, including empathetic alignment in NLP, emotional synchrony in dyadic interaction, and acoustic– prosodic entrainment in speech [9], [10]. These lines of work suggest that alignment is shaped not only by lexi- cal content, but also by temporal interaction structure and broader conversational conditions [11], [12]. At the modeling level, prior approaches range from lightweight speech emotion recognition baselines such as Emotion MLP [13] to long- range dyadic interaction models such as DyadFormer [14], but neither is designed to evaluate emotional entrainment under explicit conversational context and social relationship constraints. However, to the best of our knowledge, prior work has rarely evaluated emotional entrainment in dyadic speech under the joint constraints of conversational context and social relationship. This gap motivates our study. We outline emotional entrainment as interaction plausibility: whether a dyad exhibits coherent interaction dynamics and affective flow rather than mere acoustic realism. We formulate emotional entrainment detection as a binary classification task over two-speaker audio streams, distinguishing coherent entrainment from disrupted interaction. Built on the Seamless Interaction corpus [15], our dataset includes relationship la- bels, emotionally informative contexts, disrupted dyads from partner swapping and emotion-contradicting TTS resynthesis, arXiv:2606.30543v1 [cs.CL] 29 Jun 2026 TABLE I DYADEE DATASET COMPOSITION ACROSS EMOTIONALLY ENTRAINED (E) AND NON-EMOTIONALLY ENTRAINED (NE) DYADIC SPEECH, WITH VARIOUS CONTROLLED PERTURBATIONS Data TypeCountDuration (h) E+Original Dyad2,125— E+Voice Conversion2,125— NE: Emotion Resynthesis2,125— NE: Emotion Resynthesis+Voice Conversion2,125— Total8,500— and voice-converted entrained dyads for artifact-robust train- ing. We model each dyad as a temporal sequence of window- level embeddings from an emotion-finetuned Whisper encoder, and show that relationship and conversational context are both crucial for capturing emotional entrainment. Our contributions are: (i) TRACE, a discriminator that models dyadic interaction as temporal sequences of window- level acoustic embeddings conditioned on relationship and context signals and (i) DyadEE, a relationship- and context- conditioned dataset for emotional entrainment detection, with complementary strategies for disrupting entrainment and voice conversion augmentation to reduce acoustic shortcuts; We re- lease both the dataset and codebase to support future research. 1 I. DYADIC EMOTIONAL ENTRAINMENT DATASET We introduce DyadEE, a dataset for emotional entrainment in dyadic speech interaction. Figure 1 (left) summarizes its construction from entrained dyads, including augmented en- trained variants obtained via voice conversion and denoising, and disrupted dyads created through same-context partner swapping, different-context partner swapping, and emotion resynthesis. Each dyad is annotated with conversational con- text and relationship labels, which serve as conditioning vari- ables. Table I reports the data types, number of dyads, and total duration. A. Natural Dyads Selection We build DyadEE on top of the Seamless Interaction dataset [15], a large-scale collection of naturalistic dyadic conversa- tions recorded under diverse speaker prompts designed to elicit rich social and emotional interactions. The dataset provides relationship metadata and specific contextual prompts that make it well-suited for studying emotional entrainment in conversational speech. In our curated dataset, the relationship distribution includes 6 categories in total, including friends (13,979), coworkers (1,388), family-generic (2,558), romantic partners (2,522), classmates (848), and siblings (301). To focus on emotion- ally informative interactions, using OpenAI GPT-5 [16], we abstract and manually curate the original fine-grained speaker prompts into 14 general context categories: challenges and support (2,414), communication and feelings (2,480), deliv- ering bad news (1,360), difficult work relationships (455), 1 https://github.com/anonymoususer276/TRACE isolation and support needs (1,695), leadership and decision- making (1,200), mistakes and apologies (1,441), overcoming challenges (2,919), past compromise (538), regret and apolo- gies (1,616), relationship withdrawal (1,040), self-defense events (2,149), trust and decision-making (1,076), and unre- solved disagreements (1,213). These contexts are utilized to curate emotionally entrained dyads and to construct context- controlled disrupted interactions. B. Dyad Expansion Strategies We expand DyadEE through controlled construction of non-entrained dyads and augmentation of entrained dyads. Emotion-contradicting Resynthesis: We create emotion- inconsistent variants by resynthesizing one interlocutor while keeping the other unchanged. This one-sided manipulation yields more subtle negatives, as the dyad still contains one real human channel while emotional alignment is disrupted. For each conversation, we first run the acoustics-only par- alinguistic classifier VoxProfile [17] on the original audio to estimate the speaker’s emotion. We then preserve the original transcript for the selected speaker and resynthesize it with an expressive TTS system, using either EmotiVoice [18] (34.83 %) or CapSpeech [19] (65.17%), in a deliberately contradic- tory emotional style (e.g., happy→sad, calm→angry), thereby breaking emotional entrainment within the interaction. Data Augmentation of Entrained Dyads: To reduce re- liance on low-level acoustic artifacts, we augment emotionally entrained dyads with voice conversion using FreeVC [20] and speech denoising using MossFormer2 [21]. Voice conversion is applied to one randomly selected speaker in each dyad, while the other speaker is left unchanged. This one-sided augmentation preserves a more partially original interaction while introducing speaker-level acoustic variation without altering the underlying conversational structure. Denoising is applied to recover cleaner speech from noisy recordings while retaining the original interaction pattern. The two aug- mentations are applied independently and may overlap. By introducing acoustic variation without changing the interaction itself, these augmented examples encourage the model to attend to emotional entrainment rather than speaker identity, channel conditions, or synthetic artifacts. I. TRACE FRAMEWORK We propose TRACE, the first emotional entrainment classi- fier that leverages a temporally ordered sequence of alternat- ing speaker windows and explicitly conditions prediction on acoustic, contextual, and relationship information to determine whether a dyadic interaction is emotionally entrained or non- entrained. Figure 1 illustrates the TRACE framework (right). A. Dyadic Representation and Input Features Temporal Dyadic Representation: Given a dyad with two speakers A and B recorded in separate channels, we represent the interaction as a temporal sequence of alternating fixed t-length speech windows A n and B n , obtained from each channel: A 1 → B 1 → A 2 → B 2 → · → A N → B N , DyadEE Dataset Original Dyads Voice Conversion Denoising Swap Partners Same Context Swap Partners Diff Context Emotion Resynthesis Multi-head Self Attention Feedforward Network Bidirectionl Attention RoPE RMSNorm SwiGLU ×6 TRACE Framework Emotion Fine-tuned Whisper Embeddings LLaMA-style Blocks C l a s s i f i c a t i o n H e a d Cat E n t r a i n e d N o n - e n t r a i n e d Emotion Entrained Non-Emotion Entrained B 1 B 2 . . . Temporal- ordered Audio Sequence Speaker A Speech Windows A 1 A 2 . . . A1 B1 A2 B2 . . . Speaker B H i g h - l e v e l C o n t e x t R e l a t i o n s h i p C a t e g o r y R o m a n t i c P a r t n e r U n r e s o l v e d D i s a g r e e m e n t Speaker A Speaker B Speech Windows One-hot Categorical Embedding SBERT Embeddings Fig. 1. TRACE pipeline. Left: DyadEE dataset construction with natural dyads (original conversations and their voice-converted and denoised variants) and disrupted dyads generated via partner swapping and emotion-contradicting resynthesis, each annotated with conversational context and relationship category. Right: TRACE models dyadic interaction using window-level embeddings from an emotion-fine-tuned Whisper encoder for emotional entrainment classification. where N is the total number of windows for this dyad. This representation preserves local interaction flow between interlocutors. It also captures prosodic and affective variation at a finer granularity than utterance-level pooling, remaining more stable and tractable than turn-level modeling. Acoustic Emotion Embeddings: For each speech window, we extract a dense acoustic embedding from the final hidden layer of an emotion-finetuned whisper-large-v3 encoder [22] provided by the VoxProfile suite [17] 2 . The encoder is trained for valence, arousal, and dominance prediction, enabling it to capture affect-relevant prosodic and acoustic cues such as intonation, intensity, and spectral variation over short tem- poral segments. The resulting window-level embeddings are arranged according to the temporal dyadic sequence and serve as the primary acoustic input to TRACE. This allows the model to track how local affective patterns evolve across alternating speaker windows rather than relying on a single utterance-level summary. Context and Relationship Features: In addition to acoustic cues, TRACE conditions on two dyad-level signals: con- versational context and speaker relationship. We derive the conversational context prompt from speaker prompts and en- code them as fixed-dimensional semantic embeddings using a sentence-transformers SBERT [23], which captures high- level information about the interaction setting (e.g., apology, disagreement, support-seeking). We encode the relationship category (e.g., friends, coworkers, romantic partners) as a one- hot vector over the set of relationship labels, followed by a learnable embedding layer to obtain a dense categorical representation. Unlike the window-level acoustic sequence, these two signals remain fixed for a given dyad and provide global conditioning information about the social and situa- tional constraints under which emotional entrainment should 2 Pretrained model accessed on January 8th 2026 be interpreted. B. Emotional Entrainment Modeling Emotional Interaction Modeling: Given the temporal sequence of window-level acoustic emotion embeddings, TRACE models dyadic interaction with a lightweight stack of 6 LLaMA-style blocks [24]. Each block adopts RoPE positional encoding, RMSNorm, and feed-forward layers with SwiGLU activation, while replacing the original causal at- tention with bidirectional self-attention to capture dependen- cies across both preceding and following speaker windows. This design is better suited to dyadic sequence classification, where emotional entrainment depends on interaction patterns distributed across the full conversation rather than left-to- right generation alone. Through these stacked blocks, TRACE transforms local acoustic emotion cues into higher-level rep- resentations that encode interaction dynamics and emotional coordination across the dyad. Feature Integration and Classification: To incorporate higher-level social and situational information for improved emotional entrainment detection, we concatenate the sequence- level acoustic representation with the context embedding and the relationship embedding. The resulting feature vector is then fed into a multilayer perceptron (MLP) classification head, which produces a binary prediction (entrained vs. non- entrained) together with a calibrated probability. By fusing these signals, TRACE evaluates emotional entrainment not only from interaction-driven acoustic dynamics, but also under the contextual and relational constraints that shape appropriate emotional coordination. IV. EXPERIMENTS Our experiments include 2 parts: a performance comparison of our system with existing baselines and a feature ablation. First, we compare TRACE with two baselines: Emotion MLP [13], a lightweight speech emotion recognition model with a multilayer perceptron classifier, and DyadFormer [14], a Transformer for long-range dyadic interaction modeling. For fair comparison, both baselines are adapted to use the same input feature settings as TRACE, except for our temporal dyadic representation. We split DyadEE into training and test sets at an approximate 6:4 ratio, yielding 12,896 training dyads and 8,700 test dyads. We report accuracy, ROC-AUC, and F1 score. Then, our ablation study evaluates 4 input settings: Speech only (Sp), Speech + Context (Sp+Ctx), Speech + Rela- tionship (Sp+Rel), and Speech + Context + Relationship (Sp+Ctx+Rel). We further analyze model behavior across different relationship and context categories. All models are trained and evaluated under the same ap- proximate 6:4 train-test split and optimization settings. We make sure that the same participant speakers do not occur in both train and test split to avoid data leakage. We use a single compute node with two A100 80GB GPUs and optimize all models with cross-entropy loss for binary emotional entrain- ment classification. A. Evaluation Data B. Comparative Results Table I compares the performance of different models across various feature combinations using accuracy, ROC- AUC, and macro F1 metrics on a test set. The results demonstrate that incorporating conversational context and re- lationship information yields consistent and substantial im- provements in entrainment detection. TRACE improves mono- tonically as conditioning signals are added: from 85.56% speech-only to 94.49% with context (+8.93 p), 96.11% with relationship (+10.55 p), and 97.01% under joint conditioning (+11.45 p), the largest absolute gain observed across any model and feature configuration. DyadFormer, a transformer- based sequence model, achieves competitive speech-only ac- curacy (92.20%) but fails to benefit from the additional condi- tioning signals, declining to 91.68% under full conditioning (−0.52 p). Emotion MLP, which discards temporal order entirely through mean pooling, serves as a diagnostic lower bound: its insensitivity to context (78.37% under Sp+Ctx, below its speech-only baseline of 82.13%) further validates that contextual and relationship features are not recoverable from global acoustic statistics alone, but require interaction- level temporal structure to be useful. C. Conditional Embeddings Ablation Relationship Conditioning: Table I shows the accuracy gain over speech-only baseline across different relationship categories. It is shown that these gains are not uniform across social categories. Coworkers benefit most from relationship conditioning (+0.97 p) and achieve the largest total gain un- der full conditioning (+2.08 p), consistent with role-defined interactions where affective norms are structurally constrained. Romantic Partner benefits predominantly from relationship TABLE I ACCURACY (%), ROC-AUC, AND MACRO F1 ON THE TEST SET. SP = SPEECH ONLY, CTX = CONTEXT, REL = RELATIONSHIP. FeaturesModelAcc. ROC-AUCF1 Sp Emotion MLP [13] 82.130.8520.631 DyadFormer [14]92.200.9260.866 TRACE85.560.9070.843 Sp+Ctx Emotion MLP [13] 78.370.8820.536 DyadFormer [14]93.580.9280.877 TRACE94.490.9870.981 Sp+Rel Emotion MLP [13] 84.480.9310.657 DyadFormer [14]95.600.9430.922 TRACE96.110.9910.965 Sp+Ctx+Rel Emotion MLP [13] 82.590.8460.745 DyadFormer [14]91.680.9440.905 TRACE97.010.9960.972 TABLE I ACCURACY GAIN (PERCENTAGE POINTS, P) OVER THE SPEECH-ONLY BASELINE BY RELATIONSHIP CATEGORY FOR EACH FEATURE CONFIGURATION. “−” DENOTES NEGLIGIBLE CHANGE. Relationship+Ctx (p) +Rel (p) +Ctx+Rel (p) Friends+0.39 −1.02 Coworkers+0.16+0.97+2.08 Classmates+0.80+0.61+0.83 Family+0.89+0.34+0.56 Siblings−+0.50+0.60 Romantic Partner −0.08+1.04+0.92 conditioning (+1.04 p), reflecting that naturalness in close relationships is more strongly governed by social closeness than situational framing. Family shows the strongest response to context alone (+0.89 p) with modest additional gains from relationship (+0.34 p), suggesting that family interactions are more sensitive to situational framing than to explicit social role. Friends show negligible gain from relationship alone and slightly degrade under full conditioning (−1.02 p), suggesting that high affective variability within friendship interactions makes social-situational signals harder to leverage jointly. Taken together, these patterns confirm that the relationship is the stronger of the two conditioning signals, and that its benefit is most pronounced in interactions with well-defined social norms. Context Conditioning. Table IV reports the accuracy gain over the speech-only baseline per conversational context, for each feature configuration. As shown in Table I, incorporating conversational context yields a gain of+8.93 p over the speech-only baseline (85.56% to 94.49%), demonstrating that situational framing carries meaningful signal for entrainment detection beyond acoustics alone. The per-context breakdown reveals where this gain is concentrated: Communication & Feelings (+0.84 p), Trust & Decision-Making (+0.79 p), Isolation & Support Needs (+0.91 p), and Unresolved Dis- agreements (+1.46 p) show consistent positive contributions, reflecting scenarios where situational framing shapes affective expectations in a recoverable way. Role-asymmetric contexts TABLE IV ACCURACY GAIN (PERCENTAGE POINTS, P) OVER THE SPEECH-ONLY BASELINE BY CONTEXT CATEGORY FOR EACH FEATURE CONFIGURATION. “−” DENOTES NEGLIGIBLE CHANGE. ContextSp+Ctx Sp+Rel Sp+Ctx+Rel Challenges & Support−1.06 − −2.12 Communication & Feelings+0.84 +0.84+1.68 Delivering Bad News+0.54 +2.69 −3.23 Difficult Work Relationships − −+1.94 Isolation & Support Needs+0.91 +0.91+0.91 Leadership & Decision-Making −2.17 −1.09+2.17 Mistakes & Apologies−+1.82 −3.64 Overcoming Challenges−1.92 − Past Compromise− − −2.60 Regret & Apologies− −4.32+1.44 Relationship Withdrawal− − −0.92 Self-Defense Event− −1.68+1.68 Trust & Decision-Making+0.79 +0.79+1.58 Unresolved Disagreements+1.46 +2.19+0.73 such as Leadership & Decision-Making (−2.17 p) degrade under context alone but recover strongly under joint condition- ing (+2.17 p), while Delivering Bad News benefits most from relationship conditioning (+2.69 p), indicating that social role is the primary naturalness driver in that scenario. V. CONCLUSION AND FUTURE WORK In this work, we introduced DyadEE, a dataset for emo- tional entrainment detection in dyadic speech interaction, and TRACE, a strong temporal framework that combines window-level acoustic emotion representations with contextual and relationship information. Our results show that modeling emotional entrainment does not depend on the speech signal alone, as both conversational context and social relationship are important for modeling appropriate affective coordination. Although our approach detecting emotional entrainment has yielded outstanding performance, the non-entrained condition is instantiated via a limited set of synthetic transformations, which may not fully reflect the diversity of non-entrained conversational speech. Future work should expand DyadEE with more naturalistic and diverse non-entrained speech, and model emotional entrainment beyond binary classification, for example as a graded or time-varying phenomenon. It would also be valuable to extend DyadEE to multimodal settings, broader social contexts, and cross-cultural or multilingual sce- narios, as well as to explore how entrainment-aware modeling can improve conversational speech agents. LIMITATIONS VI. GENERATIVE AI USE DISCLOSURE The authors used Claude (Anthropic) to assist with gram- mar checking and proofreading the manuscript. These tools were not used to generate any scientific content, results, or conclusions. All authors reviewed and take full responsibility for the final content of this paper. REFERENCES [1] C. J. Wynn, T. S. Barrett, and S. A. Borrie, “Rhythm perception, speak- ing rate entrainment, and conversational quality: A mediated model,” Journal of Speech, Language, and Hearing Research, vol. 65, no. 6, p. 2187–2203, 2022. [2] J. Kruyt, D. de Jong, A. D’Ausilio, and ˇ S. Be ˇ nu ˇ s, “Measuring prosodic entrainment in conversation: A review and comparison of different methods,” Journal of Speech, Language, and Hearing Research, vol. 66, no. 11, p. 4280–4314, 2023. [3] J. Kejriwal, “Relationship between speech entrainment and emotion,” in 2022 10th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW).IEEE, 2022, p. 1–4. [4] K. S. Zee and N. Bolger, “Physiological coregulation during social support discussions.” Emotion, vol. 23, no. 3, p. 825, 2023. [5] ˇ S. Be ˇ nu ˇ s, “Social aspects of entrainment in spoken interaction,” Cogni- tive Computation, vol. 6, no. 4, p. 802–813, 2014. [6] J. Phang, M. Lampe, L. Ahmad, S. Agarwal, C. M. Fang, A. R. Liu, V. Danry, E. Lee, S. W. Chan, P. Pataranutaporn et al., “Investigating affective use and emotional well-being on chatgpt,” arXiv preprint arXiv:2504.03888, 2025. [7] A. Thakkar, A. Gupta, and A. De Sousa, “Artificial intelligence in positive mental health: a narrative review,” Frontiers in digital health, vol. 6, p. 1280235, 2024. [8] Y. Sun and H. Ding, “Unpacking the gender-role interaction of prosodic entrainment in chinese long-and-short turn-taking: evidence from per- ceptual and acoustic similarities,” Humanities and Social Sciences Communications, vol. 11, no. 1, p. 1618, 2024. [9] R. Levitan and J. Hirschberg, “Measuring acoustic-prosodic entrainment with respect to multiple levels and dimensions,” in Interspeech 2011, 2011, p. 3081–3084. [10] J. Yang and D. Jurgens, “Modeling empathetic alignment in conversa- tion,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, p. 3127–3148. [11] M. McNeill and R. Levitan, “Autoregressive cross-interlocutor attention scores meaningfully capture conversational dynamics.” ISCA, 2024. [12] V. R. D. M. Herbuela and Y. Nagai, “Spatiotemporal emotional syn- chrony in dyadic interactions: The role of speech conditions in facial and vocal affective alignment,” arXiv preprint arXiv:2505.13455, 2025. [13] J. Joy, A. Kannan, S. Ram, and S. Rama, “Speech emotion recognition using neural network and mlp classifier,” Ijesc, vol. 2020, p. 25 170– 25 172, 2020. [14] D. Curto, A. Clap ́ es, J. Selva, S. Smeureanu, J. Junior, J. CS, D. Gallardo-Pujol, G. Guilera, D. Leiva, T. B. Moeslund et al., “Dyad- former: A multi-modal transformer for long-range modeling of dyadic interactions,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, p. 2177–2188. [15] V. Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Cheng et al., “Seamless inter- action: Dyadic audiovisual motion modeling and large-scale dataset,” arXiv preprint arXiv:2506.22554, 2025. [16] A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram et al., “Openai gpt-5 system card,” arXiv preprint arXiv:2601.03267, 2025. [17] T. Feng, J. Lee, A. Xu, Y. Lee, T. Lertpetchpun, X. Shi, H. Wang, T. Thebaud, L. Moro-Velazquez, D. Byrd et al., “Vox-profile: A speech foundation model benchmark for characterizing diverse speaker and speech traits,” arXiv preprint arXiv:2505.14648, 2025. [18] NetEase Youdao, “Emotivoice: a multi-voice and prompt-controlled tts engine,” 2024, gitHub repository, commit. Accessed 2026-02-25. [19] H. Wang, J. Hai, D. Chong, K. Thakkar, T. Feng, D. Yang, J. Lee, T. Thebaud, L. M. Velazquez, J. Villalba et al., “Capspeech: En- abling downstream applications in style-captioned text-to-speech,” arXiv preprint arXiv:2506.02863, 2025. [20] J. Li, W. Tu, and L. Xiao, “Freevc: Towards high-quality text-free one-shot voice conversion,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, p. 1–5. [21] S. Zhao, Y. Ma, C. Ni, C. Zhang, H. Wang, T. H. Nguyen, K. Zhou, J. Q. Yip, D. Ng, and B. Ma, “Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2024, p. 10 356–10 360. [22] V. Timmel, C. Paonessa, M. Vogel, D. Perruchoud, and R. Kakooee, “Fine-tuning whisper on low-resource languages for real-world applica- tions,” in Proceedings of the 10th edition of the Swiss Text Analytics Conference, 2025, p. 57–65. [23] A. M. Korga, S. Wefers, K. Hanken, R. B. Tareaf, B. Steemers, and H. Avvad, “Does size matter? examining sentence similarity perfor- mance in large language models,” in 2025 International Conference on Information Networking (ICOIN). IEEE, 2025, p. 595–600. [24] Q. Fang, Y. Zhou, S. Guo, S. Zhang, and Y. Feng, “Llama-omni 2: Llm- based real-time spoken chatbot with autoregressive streaming speech synthesis,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, p. 18 617–18 629.