Paper deep dive
Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language
Yoshiki Tanaka, Ryuichi Uehara, Koji Inoue, Michimasa Inaba
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/13/2026, 12:31:07 AM
Summary
The paper introduces 'Emotion Transcription in Conversation' (ETC), a novel task that generates natural language descriptions of emotional states in dialogue. The authors constructed a Japanese dataset of 1,002 dialogues with self-reported emotional transcriptions and emotion category labels. They benchmarked LLMs (GPT-4.1 and Llama-3.1) and proposed a fine-grained evaluation method based on atomic unit decomposition to assess content faithfulness.
Entities (5)
Relation Signals (3)
ETC Dataset â supportstask â Emotion Transcription in Conversation
confidence 100% · To realize the proposed ETC task, our initial undertaking involves the meticulous collection of a suitable dataset
GPT-4.1 â evaluatedon â ETC Dataset
confidence 95% · We investigated the performance of two LLMs: GPT-4.1... on our dataset
Llama-3.1-Swallow â evaluatedon â ETC Dataset
confidence 95% · Additionally, we performed supervised fine-tuning on Llama-3.1 with our dataset.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Emotion Recognition in Conversation (ERC) is critical for enabling natural human-machine interactions. However, existing methods predominantly employ categorical or dimensional emotion annotations, which often fail to adequately represent complex, subtle, or culturally specific emotional nuances. To overcome this limitation, we propose a novel task named Emotion Transcription in Conversation (ETC). This task focuses on generating natural language descriptions that accurately reflect speakers' emotional states within conversational contexts. To address the ETC, we constructed a Japanese dataset comprising text-based dialogues annotated with participants' self-reported emotional states, described in natural language. The dataset also includes emotion category labels for each transcription, enabling quantitative analysis and its application to ERC. We benchmarked baseline models, finding that while fine-tuning on our dataset enhances model performance, current models still struggle to infer implicit emotional states. The ETC task will encourage further research into more expressive emotion understanding in dialogue. The dataset is publicly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.07138v1
- Canonical: https://arxiv.org/abs/2603.07138v1
Trouble viewing inline? Open PDF directly â
Full Text
62,756 characters extracted from source content.
Expand or collapse full text
Emotion Transcription in Conversation: A Benchmark for Capturing Subtle and Complex Emotional States through Natural Language Yoshiki Tanaka 1 , Ryuichi Uehara 1 , Koji Inoue 2 , Michimasa Inaba 1 1 The University of Electro-Communications, Tokyo, Japan 2 Kyoto University, Kyoto, Japan y-tanaka,r-uehara,m-inaba@uec.ac.jp inoue.koji.3x@kyoto-u.ac.jp Abstract Emotion Recognition in Conversation (ERC) is critical for enabling natural human-machine interactions. However, existing methods predominantly employ categorical or dimensional emotion annotations, which often fail to ade- quately represent complex, subtle, or culturally specific emotional nuances. To overcome this limitation, we propose a novel task named Emotion Transcription in Conversation (ETC). This task focuses on generating natural language descriptions that accurately reflect speakersâ emotional states within conversational contexts. To address the ETC, we constructed a Japanese dataset comprising text-based dialogues annotated with participantsâ self-reported emotional states, described in natural language. The dataset also includes emotion category labels for each transcription, enabling quantitative analysis and its application to ERC. We benchmarked baseline models, finding that while fine-tuning on our dataset enhances model performance, current models still struggle to infer implicit emotional states. The ETC task will encourage further research into more expressive emotion understanding in dialogue. The dataset is publicly available at https://github.com/UEC-InabaLab/ETCDataset. Keywords:Emotion Transcription in Conversation, Emotion Recognition in Conversation, Corpus Construc- tion, Dialogue 1.Introduction Emotion Recognition in Conversation (ERC) is a critical technology for enabling natural and seam- less human-machine interactions, and its signifi- cance has grown substantially in recent years ( Po- ria et al.,2019;Pereira et al.,2025). Accu- rately comprehending user emotions is indispens- able for developing more empathetic and human- like conversational systems. Research in this do- main has advanced dramatically over the past decade, with the advent of deep learning and large- scale language models propelling recognition ac- curacy toward practical application levels ( Fu et al., 2025;Pereira et al.,2025). Numerous benchmark datasets have been established (Busso et al., 2008;Li et al.,2017;Poria et al.,2019), and emo- tion classification algorithms have evolved from traditional machine learning to sophisticated deep learning architectures. However, a prevalent limi- tation persists: dialogue data predominantly used in existing research often relies on specific sce- narios or includes posed emotional expressions, which do not always reflect spontaneous, every- day conversations ( Busso et al.,2008;Poria et al., 2019). Consequently, this presents a consider- able challenge for conversational systems and robots designed for diverse real-world applica- tions. Addressing this gap, we propose a novel task for emotion recognition: âEmotion Transcription in Conversation (ETC).â Unlike conventional ERC ap- I feel like you can't really build relationships without trust, but you get betrayed sometimes. It's complicated, isn't it? I worried this deep question might make them uncomfortable, but I wanted to ask to get to know them better. I deeply empathized with the complexity of relationships and wanted to express that we need to approach them sincerely and build trust. Relationships are really hard, aren't they? What you think is helpful can bother someone, (...) There's no right answer âyou just figure it out as you go. I totally agree with you. I've also had an experience where people misunderstood my intentions and unfairly blamed me. I was glad they responded as if it were their own experience. I empathized with their experience and felt sympathy for how unfair one-sided blame is. That sounds like a very painful experience. When someone blames you unfairly, they won't listen even when you reach out, making it harder to build trust. Thank you. Yeah, it was my boss, a man in his 50s. Our thinking was just fundamentally different, so I distanced myself from him. I was glad to receive their kind words about my painful experience. UtteranceEmotion Transcription ... Figure 1:Example of a dialogue with emotion transcriptions from our dataset (translated from Japanese). Each transcription is also annotated with multi-label emotion categories. proaches, ETC aims to generate natural language descriptions of a speakerâs emotional state for each utterance within a conversation. By capturing emotions through natural language, we anticipate the ability to represent richer and more detailed af- fective information that eludes traditional categor- ical or numerical frameworks. This includes sub- arXiv:2603.07138v1 [cs.CL] 7 Mar 2026 tle emotional shifts, blended emotions, and cul- turally specific expressions of feeling, which has been attempted by soft label of emotional cate- gories so far (Wen et al.,2023). For instance, com- plex states such as âa sense of exasperation mixed with concern for the other personâ or âdisappoint- ment at an unmet expectation while trying to main- tain composureâ are diïŹicult to articulate within ex- isting paradigms. Describing emotions in natural language opens a new avenue for machines to understand and articulate the inherent richness of human affective states. Although there has been a limited number of attempts to represent emo- tional states of dialogue participants in natural lan- guage (Shinoda et al.,2025), those were artificially generated by LLMs. This formulation is particularly exigent in the current era of rapidly advancing LLM capabilities, offering the potential for machines to more accurately grasp the depths of human emo- tion and thereby foster genuinely empathetic inter- actions. To realize the proposed ETC task, our initial un- dertaking involves the meticulous collection of a suitable dataset and the establishment of a robust benchmark. Specifically, we leverage crowdsourc- ing to collect text-based dialogues in Japanese that simulate a variety of everyday scenarios. Cru- cially, during this process, dialogue participants will be prompted to describe their own emotional state at each turn of the conversation using free- form natural language, as depicted in Figure1. Furthermore, to enable quantitative analysis and its application to ERC, we collect emotion category labels corresponding to each transcription. This methodology yields a large-scale dataset compris- ing pairs of naturalistic emotional expressions and their corresponding linguistic descriptions. Sub- sequently, using this curated dataset, we develop and evaluate models designed to generate these emotional state descriptions from dialogue context. This empirical validation ascertains the feasibility of the ETC task and the eïŹicacy of our proposed approach. The primary contributions of this research are threefold: (1) We introduce âEmotion Transcrip- tion in Conversation (ETC)â as a new task in con- versational emotion recognition, enabling a more nuanced and detailed understanding of emotions beyond what is possible with existing categorical or dimensional approaches. This represents a paradigm shift in how machines can interpret and represent affective states. (2) We construct and release a novel dialogue dataset annotated with speakersâ own natural language descriptions of their emotions. This dataset will serve as a valu- able resource for the research community, partic- ularly for studies aiming to achieve more human- like emotional understanding in AI. (3) We de- velop baseline models for the ETC task using this dataset and empirically demonstrate their effec- tiveness. In doing so, this research charts a new direction for emotion recognition technology and contributes to the advancement of emotional in- telligence in future conversational systems and robotics, paving the way for more sophisticated and empathetic human-AI collaboration. 2.Related Work ERC has garnered significant attention within Nat- ural Language Processing (NLP), driven by the exponential increase in recognition models with benchmark datasets. Recognition ModelsPre-trained language models (PLMs), such as BERT (Devlin et al., 2019) and RoBERTa (Liu et al.,2019), have become foundational in text-based ERC due to their powerful contextual understanding. Many approaches fine-tune these PLMs on specific ERC datasets, often augmenting them with additional layers to integrate conversational features like speaker identities and dialogue history (Qin et al., 2023;Yu et al.,2024). Graph Neural Networks (GNNs) have also gained prominence in modeling intricate conversational dynamics, particularly speaker interactions and utterance dependencies. Notable examples include DialogueGCN (Ghosal et al.,2019), which represents conversations as graphs with utterances as nodes and interrelations as edges. Additionally, models like ESIHGNN have introduced event-state interactions within heterogeneous GNN frameworks to enhance ERC performance ( Zha et al.,2024). Large Language Models (LLMs) have further advanced ERC through diverse prompting strategies, such as zero-shot, few-shot, and chain-of-thought prompt- ing, as well as targeted LoRA fine-tuning ( Feng et al.,2024). Recent studies like LaERC-S specif- ically explore enhancing LLM-based ERC through explicit integration of speaker characteristics (Fu et al. ,2025). Datasets and Emotion CategoriesSeveral benchmark datasets have been developed to ad- vance ERC research, often employing discrete emotion categories. MELD (Multimodal Emo- tionLines Dataset), derived from the âFriendsâ TV series, captures multi-party interactions an- notated with Ekmanâs six basic emotions and neutral, making it widely used for studying emo- tional shifts in conversational contexts ( Poria et al. ,2019). The IEMOCAP (Interactive Emo- tional Dyadic Motion Capture Database) dataset provides dyadic interactions with both scripted and improvised dialogues, supplemented by mul- timodal data, and features custom emotional la- bels such as anger, happiness/excitement, neu- tral, and sadness (Busso et al.,2008). MSP- Conversation (Martinez-Lucas et al.,2020) and MSP-Podcast (Busso et al.,2025) are audio- based corpora sourced from podcast recordings. MSP-Conversation provides time-continuous an- notations of emotional dimensions (arousal, va- lence, and dominance) over entire conversations, while MSP-Podcast offers segment-level anno- tations with categorical emotion labels. MSP- Podcast also utilizes typed descriptions as supple- mentary annotations from annotators when they select the âotherâ emotion category, allowing an- notators to specify emotions beyond predefined categories (Chou et al.,2022). DailyDialog offers cleaner, human-generated dyadic dialogues anno- tated with Ekmanâs basic emotions and includes communicative intent labels (Li et al.,2017). Emo- tionLines, also based on âFriends,â is a text-only predecessor of MELD (Hsu et al.,2018). Similarly, EmoryNLP, although also drawn from âFriends,â provides distinct data splits and annotations (Zahiri and Choi,2018). EmoContext focuses on short, three-turn dialogues with simplified emotion cate- gories (happiness, sadness, anger, others) ( Chat- terjee et al.,2019). CAPE, a Chinese appraisal- based emotional dataset, specifically targets emo- tional generation in LLMs (Liu et al.,2025). While some attempts have been made to represent di- alogue participantsâ emotional states explicitly in natural language, similar to our ETC task, these ef- forts primarily relied on artificially generated data from LLMs ( Shinoda et al.,2025). 3.Data Collection 3.1.Procedure We used CrowdWorks 1 , a Japanese crowdsourc- ing platform, for this data collection. To character- ize our participant pool, we measured the Big Five personality traits of all crowdworkers using a 10- item questionnaire based on the Japanese version of the Ten-Item Personality Inventory (TIPI-J) ( Os- hio et al.,2012). To effectively elicit rich and natu- rally occurring emotional expressions, we adopted a speaker-listener dialogue setting inspired by the EmpatheticDialogue paradigm ( Rashkin et al., 2019). Crowdworkers were assigned one of two roles: a âSpeakerâ or a âListener.â The core task for the Speaker was to recount a personal episode related to a specific emotion. To guide this, Speak- ers were provided with one of 32 emotion labels, which were Japanese translations of the emotions used in the EmpatheticDialogue corpus ( Rashkin 1 https://crowdworks.jp Dialogues1,002 Utterances / Emotion Transcripts10,020 Avg. Length per Utterance42.72 by Speaker44.64 by Listener40.80 Avg. Length per Emotion Transcripts 28.89 by Speaker28.92 by Listener28.86 Table 1:Statistics of the Dataset et al.,2019). They were instructed to convey, through dialogue, a concrete personal experience where they felt the designated emotion. Listeners, in turn, were tasked with actively engaging with the Speakerâs narrative. Dialogues were structured to begin with an utterance from the Speaker, followed by alternating turns between the Speaker and Lis- tener. Each dialogue concluded after 5 turns, re- sulting in a total of 10 utterances. Immediately after inputting each of their utterances, both the Speaker and the Listener were required to provide a free-form textual description of their own inter- nal emotional state at that specific moment of ut- tering the text. We defined this emotional state description as âa verbalization of the internal emo- tional state or intention a participant held at the time of their utterance.â Crowdworkers were ex- plicitly guided to describe, in as much detail as possible, the psychological context behind their ut- terance. Further details on the data collection are provided in AppendixA. 3.2.Data Collection Results We collected a total of 1,002 dialogues. An exam- ple of collected data is presented in Figure1. A total of 199 unique crowdworkers participated in the data collection. The distribution of personality traits among participants is provided in Appendix B for reference. The median number of dialogues per worker was 6.0, with the most active worker contributing to 38 dialogues. We aimed for an even distribution of dialogues across the 32 emotion la- bels; the number of dialogues per emotion ranged from a minimum of 30 to a maximum of 32. Table1shows key statistics of the collected dataset. On average, Speakersâ utterances were slightly longer (in characters) than those of Listen- ers, while the lengths of emotional transcriptions were comparable across both roles. 3.3.Annotation for Emotion Categorization While the emotion transcriptions we collected con- tain fine-grained details of the speakersâ emotional states, they are challenging to analyze quantita- tively. To address this, we additionally annotated each emotion transcription in our dataset with cor- responding emotion labels. This annotation not only enables a quantitative analysis of the emo- tions expressed in the transcriptions but also al- lows our dataset to be applied to traditional ERC tasks. We adopted a set of seven emotion categories, comprising the six basic emotions from Ekmanâs six universal emotions (Ekman et al.,1987)joy, sadness, fear, anger, surprise, and disgust (, ,,,,)supplemented with a âneutralâ () category. This set is widely used in established ERC datasets (Hsu et al.,2018;Poria et al.,2019;Li et al.,2017). To account for the possibility that a single emo- tion transcription may express multiple emotions, we employed a multi-label annotation approach. Specifically, annotators assigned all relevant la- bels to each emotion transcription. âNeutralâ was used only for transcriptions that did not express any specific emotion. Each emotion transcription was annotated by three different annotators on CrowdWorks. An- notators were provided with the full dialogue con- text and the corresponding emotion transcription, and made a binary answer on whether each of the seven emotion categories was present in the tran- scription. Following previous studies (Busso et al., 2008;Hsu et al.,2018), we used majority voting to aggregate the results; an emotion category was assigned as the final label if two or more of the three annotators agreed on it. If none of the emo- tion categories were agreed upon, the transcrip- tion was labeled as âNeutralâ only. 4.Experiments To evaluate the utility of our dataset for the ETC task, we trained baseline models and assessed their performance. We divided our dataset into training, validation, and test sets in an 8:1:1 ratio. 4.1.Task The ETC task requires models to generate natural language descriptions reflecting a speakerâs emo- tional state. To evaluate this ability, we formulate the ETC as follows. Each dialogueDin our dataset consists of N= 10utterances and is represented as a se- quence of utterance-speaker pairs, denoted as D= ((u 1 , s 1 ),(u 2 , s 2 ), . . . ,(u N , s N )), whereu i denotes thei-th utterance ands i represents its speaker. In the ETC task, the modelMis given the dialogue contextC n up to then-th utterance, denoted asC n = ((u 1 , s 1 ),(u 2 , s 2 ), . . . ,(u n , s n )), wherencan be any utterance in the dialogue (i.e., 1â€nâ€N). The objective of this task is to predict the emotion transcriptione n for an utter- anceu n made by speakers n based onC n . i.e., e n =M(C n ). Here, the emotion transcriptione n is a natural language description of the emotional state of speakers n at the time of utteringu n . 4.2.Models We investigated the performance of two LLMs: GPT-4.1 2 , the latest advanced chat model devel- oped by OpenAI, andLlama-3.1-Swallow 3 (Fu- jii et al.,2024;Okazaki et al.,2024) , an open- source LLM known for its strong proficiency in Japanese. Both LLMs were evaluated using zero- shot and 4-shot prompting. The prompt template we designed for this experiment is presented in Ap- pendix C. Additionally, we performed supervised fine-tuning on Llama-3.1 with our dataset. For the fine-tuned Llama-3.1, we report the average per- formance across models trained with five random seeds to ensure robustness. Further implementa- tion details are provided in AppendixD. 4.3.Evaluation Metrics 4.3.1.Traditional Automatic Evaluation Metrics To evaluate the quality of the emotion transcrip- tions, we use two traditional automatic evalua- tion metrics: BLEU (Papineni et al.,2002) and ROUGE (Lin,2004), and a semantic similarity met- ric: BERTScore (Zhang et al.,2020). 4.3.2.Fine-grained and Interpretable Evaluation of Content Faithfulness Recent studies have proposed using LLMs as au- tomatic evaluators to approximate human prefer- ences more closely ( Zheng et al.,2023;Liu et al., 2023). Following this trend, we initially explored a LLM-based evaluation metric to assess the se- mantic alignment between the generated emotion transcriptions and the ground truth. However, relying on a single, holistic score presents significant challenges for our task. The emotional states are often complex, containing multiple components, such as âfeeling angry that my feelings are not understood, but also sad.â For such complex expressions, a single score strug- gles to accurately penalize the omission of some emotional components or the hallucination of oth- ers. To address this gap, we introduce a fine- grained evaluation approach of content faithful- ness. This method, inspired by FActScore ( Min et al.,2023) is based on two steps. 2 We used gpt-4.1-2025-04-14. 3 We usedLlama-3.1-Swallow-8B-Instruct-v0.3. ModelsSetting B-1 B-2 B-3 B-4 R-1 R-2 R-L BS Prec. Rec. F1 GPT-4.1 zero-shot 16.89 7.99 3.93 2.09 23.61 4.87 17.93 57.66 14.7342.2713.99 4-shot 26.89 13.34 7.36 4.12 28.065.7822.98 59.67 20.5927.4013.78 Llama-3.1 zero-shot 17.40 7.20 3.55 1.84 20.25 3.08 16.26 55.24 9.18 17.98 5.77 4-shot 29.4114.868.464.8427.51 5.58 23.2259.9514.83 14.18 7.84 fine-tuning36.07 23.01 15.59 9.98 31.95 8.79 28.54 62.64 28.5019.7114.29 Table 2:Automatic evaluation results for ETC task. Reported metrics include BLEU (B), ROUGE (R), BERTScore (BS), Precision (Prec.), Recall (Rec.), and F1-score (F1). Best values are inbold, and the second-best values in underlined. ModelSetting # Units (SD) GPT-4.1 zero-shot 3.39 (1.14) 4-shot1.96 (0.76) Llama-3.1 zero-shot 1.99 (1.04) 4-shot1.58 (0.58) fine-tuning 1.39 (0.60) Reference â1.35 (0.64) Table 3:Average number of atomic units decom- posed from the predicted transcription and from the ground truth (Reference). Decomposition into Atomic Units. First, the source emotion transcription is decomposed into atomic units, each conveying a single piece of in- formation about the emotional state. For example, feeling angry that my feelings are not understood, but also sad.â is split into two units:I am angry that my feelings are not understoodâ and âI am sad that my feelings are not understood.â Note that some input texts may not contain any emotional descrip- tion. We used Gemini-2.5-Flash 4 for this step. Assessing Support for Atomic Units. Next, we classify whether each atomic unit from the source transcription is supported by the target emotion transcription into three categories:Supported, Not Supported, andNeutral. For example, the unit âI am angry that my feelings are not under- stoodâ is supported by the target âfeeling angry that my feelings are not understood, but also sad.â and is thus classified asSupported. In contrast, the unit âI am happy that my feelings are understoodâ is not supported by the target and is classified as Not Supported. TheNeutralapplies when a core component matches, but the validity of a sec- ondary component is uncertain. For instance, rela- tive to the target âI am angry,â the unit âI am angry that my feelings are not understoodâ is classified asNeutralbecause the core emotion (âangryâ) matches, but the reason (âthat my feelings are not understoodâ) cannot be verified. This classification step was performed using Gemini-2.5-Flash. The 4 https://ai.google.dev/gemini-api/ docs/models prompts used in the above two steps are presented in AppendixE. Following these two steps, we compute two scores for each predicted emotion transcription. Precision. We decomposed the predicted emo- tion transcription into atomic units in the first step and calculated the proportion of these units that areSupportedby the ground truth transcription. Recall. We decomposed the ground truth emotion transcription into atomic units and calculated the proportion of these units that areSupportedby the predicted emotion transcription. For both metrics, units classified asNeutral are not counted as correct in the calculation. We exclude cases where the ground-truth transcrip- tion does not describe any emotional state in our evaluation, as the above scores are undefined in such cases. If the generated transcription con- tains no emotional description, but the ground truth does, we set both scores to 0. Additionally, we cal- culate theF1-score, the harmonic mean of Preci- sion and Recall, as an overall measure of perfor- mance. 4.4.Results Table2presents the results of automatic evalu- ation for the ETC task. For widely used, tradi- tional metricsBLEU, ROUGE, and BERTScore the fine-tuned Llama-3.1 model outperformed the other models by a significant margin. In terms of Precision, the Llama-3.1 model fine- tuned on our dataset achieves the highest score of 28.50%. For this metric, the 4-shot setting con- sistently outperforms the zero-shot setting, align- ing with the trend observed in traditional automatic evaluation metrics. Recall, however, exhibits an opposite trend. The zero-shot setting yields higher scores than the 4-shot setting for both models, with the zero-shot GPT-4.1 achieving the highest Recall of 42.27%significantly outperforming all other models. Finally, for the F1-score, the fine- tuned Llama-3.1 model secures the best result at 14.29%. Nevertheless, GPT-4.1 remains highly competitive, with its zero-shot setting achieving a close score of 13.99%. Context Speaker A: The other day, I was driving at night when a cyclist with no lights suddenly shot out in front of me... It was terrifying. ( ) Speaker B: That sounds so scary. An experience like that could take years off your life. ( ) Speaker A: I thought my heart was going to stop. I was relieved we didnât crash, but afterward, I just got so angry. ( ) Output Ground-Truth: I was happy that the dialogue partner understood my feelings, and I wanted to share my feelings even more. ( ) GPT-4.1 (zero-shot): Iâm feeling a mix of fear from the truly dangerous situation I was in, relief that it didnât turn into an accident, and anger that is now welling up toward the cyclist who shot out without lights. Iâm speaking with the hope that the dialogue partner understands my shock and anger from that moment. ( ) GPT-4.1 (4-shot): I wanted the dialogue partner to understand my feelings of fear and anger. ( ) Llama-3.1 (zero-shot): I was really scared, and I also feel angry. ( ) Llama-3.1 (4-shot): I want the dialogue partner to understand my fear and anger. ( ) Llama-3.1 (fine-tuned): I was happy for their empathy and wanted them to hear more details. ( ) Table 4:Case study of emotion transcription generation. (translated from Japanese) The outputs are the emotion transcriptions for Speaker Aâs final utterance. The fine-grained evaluation of content faithful- ness reveals the distinct characteristics of each model. We present the average number of atomic units decomposed from the generated transcrip- tion and from the ground truth in Table3. The fine- tuned Llama-3.1 model generates transcriptions with an average atomic unit of 1.39, which is clos- est to the ground truth of 1.35 among all models. In contrast, models in the zero-shot setting, par- ticularly the GPT variants, tend to generate higher numbers of atomic units per transcription. The trends observed in the information content explain the Precision-Recall trade-off. Transcrip- tions with a higher number of atomic units, such as those from the zero-shot models, are more likely to cover the atomic units present in the ground truth, thereby enhancing their Recall scores. Con- versely, this verbosity increases the likelihood of including redundant information, which negatively impacts their Precision. In contrast, the fine-tuned Llama-3.1 modelâs output quantity (1.39 units) not only aligns closely with the ground truth (1.35) but also achieves the highest Precision. This suggests that fine-tuning helped the model accurately iden- tify the speakerâs emotional state. Moreover, the ETC task remains challenging, as evidenced by the overall low scores across all eval- uation metrics. Even the best-performing model, the fine-tuned Llama-3.1, achieves only 14.29% in F1-score, indicating significant room for improve- ment in accurately capturing speakersâ emotional states in dialogues. Additionally, lower F1 scores than both Precision and Recall for all models suggest a notable imbalance between these two aspects, highlighting the diïŹiculty of generating emotion transcriptions that are both comprehen- sive and precise. Bridging this gap positions our dataset as a valuable and challenging benchmark for the community, advancing future research on understanding emotional states. 5.Case Study To illustrate the challenges faced by current ETC models, we present a case study in Table 4. In this example, Speaker Aâs final utterance explicitly mentions feelings of relief and anger related to a past event. The ground-truth transcription, how- ever, reveals that the speakerâs actual emotional state at that moment was happiness, stemming from the listenerâs empathetic response. This gap between the narrated emotions and the speakerâs actual emotional state poses a significant chal- lenge for the ETC task. Faced with this challenge, most models failed to capture the speakerâs true emotional state. As shown in Table4, the outputs from GPT-4.1 and the zero-shot and 4-shot Llama-3.1 models fo- cused only on the negative emotions explicitly mentioned in the narrative. For example, the generation from GPT-4.1 in the zero-shot setting, which is the most lengthy, not only identified the emotions present in the utterance but also elabo- rated on them to express a desire for the listenerâs understanding. However, it still failed to predict the emergent feeling of happiness. In contrast, the Llama-3.1 model, fine-tuned on our dataset, suc- cessfully inferred the speakerâs happiness derived from the empathetic interaction with the listener, suggesting that our dataset provides valuable train- ing signals for capturing nuanced emotional states in dialogues. 6.Corpus Analysis The annotated emotion labels enable a structured analysis of our dataset, providing insights into the characteristics of the emotion transcriptions. In this section, we present a quantitative analysis of the annotated emotion labels, focusing on three aspects: the distribution and agreement of the la- bels (Section 6.1), the co-occurrence of emotion labels within utterances (Section6.2), and the tex- tual differences between utterances and emotion transcriptions (Section6.3). 6.1.Label Distribution and Agreement We first analyze the distribution of emotion labels assigned to the emotion transcriptions, a process detailed in Section 3.3. Table5presents their distribution, showing the prevalence of each la- bel for Speakers, Listeners, and all participants. In this analysis, prevalence is calculated as the proportion of transcriptions assigned a given la- bel through majority voting. The table also reports the overall inter-annotator agreement, measured by Fleissâ kappa ( Fleiss,1971). Overall, approximately 50% of the emotion tran- scriptions were labeled as âNeutral,â indicating that nearly half of the transcriptions expressed one or more emotions other than neutral. Among the emotion labels, âJoyâ was the most frequently ex- pressed. When analyzed by role, Speakersâ tran- scriptions have a higher proportion of emotion labels assigned compared to those of Listeners. This difference is likely to be attributed to the task instructions for each role. Speakers were explicitly instructed to talk about the topic related to a spe- cific emotion, a prompt that naturally encouraged more emotional expression and recall, compared to the Listeners. Interestingly, âSurpriseâ was the only emotion label expressed more frequently in the Listenerâs transcriptions than in the Speakerâs Emotion Spk. (%) Lsn. (%) All (%)Îș Joy29.920.025.0 0.603 Sadness10.88.89.8 0.359 Fear7.15.36.2 0.476 Anger3.72.83.2 0.529 Surprise2.44.73.5 0.560 Disgust6.33.34.8 0.233 Neutral43.457.250.3 0.400 Overallâ 0.533 Table 5:Distribution of emotion labels in emotion transcriptions by dialogue role and inter-annotator agreement (Fleissâ kappa). Columns denote di- alogue role: speaker (Spk.), listener (Lsn.), and both roles (All). Figure 2:Co-occurrence matrix of emotion labels. Each cell indicates the number of transcriptions an- notated with both corresponding labels. (4.7% vs. 2.4%). Surprise is an emotion often experienced as a reaction to an event or new in- formation. For example, our dataset contains sev- eral Listener transcriptions expressing surprise in response to the Speakerâs narrative, such as âIt was somewhat unexpected, so I tried to imagine the situation myself.â and âI was surprised that something like this could really happen.â, thereby illustrating this trend. The overall inter-annotator agreement of Fleissâ kappa is 0.533, indicating moderate agree- ment. Examining emotion labels individually, âjoyâ showed the highest agreement among annota- tors, followed by âangerâ (Îș= 0.529) and âsur- prise.â (Îș= 0.560) Conversely, âdisgustâ exhibited the lowest agreement (Îș= 0.233). 6.2.Co-occurrence of Emotion Labels In this section, we analyze the co-occurrence of emotion labels in the emotion transcriptions to un- derstand the complexity of emotional expression. Figure 2illustrates the co-occurrence relationships of emotion labels within the transcriptions. We found that approximately 5.6% of all non-neutral transcriptions contained multiple emotion labels. A closer look at the co-occurrence matrix in Fig- ure 2reveals several patterns. First, emotions with Emotion UtteranceEmotion transcription Joy(give),(excited),(moved), (eat),(enjoy) (happy),(excited), (positive),(moved),(excited) Sadness(sad),(lonely),(painful), (disappointed),(enough) (sad),(regretful),(lonely), (painful),(sympathy) Fear(news),(terrifying), (incident),(worried),(earthquake) (fear),(worried),(terrifying), (imagine),(worry) Anger(irritated),(angry),(stand), (bad),(anger) (anger),(irritation),(strong), (irritated),(angry) Surprise(what),(surprised), (about),(amazing),(times) (surprise),(unexpected), (impressed),(surprised),(imagine) Disgust(dislike),(disgust),(bad), (angry),(averse) (dislike),(disgust),(imagine), (harbor),(sorry) Neutral(little),(give),(trust), (eat),(point) (write),(agree),(answer), (specific),(explain) Table 6:Top-5 frequent words by category for utterances and emotion transcriptions. similar valence frequently co-occur, such as âDis- gustâ with âAngerâ and âSadness.â Second, inter- estingly, emotions with opposing natures, like âJoyâ and âSadness,â also co-occur frequently. This co- occurrence is often not contradictory, as the two emotions can stem from different aspects of the situation. For example, transcriptions from the dataset included statements such as, âI am recall- ing both pleasant and bitter memories.â and âlis- tener is empathizing with me, but I also still feel sad when I recall the shopping scene.â Further- more, âDisgustâ tends to co-occur broadly with a wide range of other emotions. This broad associa- tion may contribute to its ambiguity, potentially ex- plaining why it received the lowest inter-annotator agreement. These results suggest the complexity of human emotion, which often involves multiple, sometimes conflicting feelings simultaneously. 6.3.Frequent Words for Emotion Categories In this section, to analyze the differences in lin- guistic features between utterances and emotion transcriptions, we extract and compare their re- spective characteristic words for each emotion la- bel. This analysis is based on the method con- ducted by Ide and Kawahara(2022). Specifically, we first use Juman++ (Tolmachev et al.,2018) to identify words appearing in an utterance or emotion transcription. Next, to filter out common words that appear across all emotion labels, we re- move any word with an IDF value less than half of the maximum IDF. Finally, we extract the highest- frequency words as characteristic words. Table 6shows the Top-5 most frequent words for each emotion label. A common trend in both utter- ances and emotion transcriptions is the extraction of words that express the corresponding emotion category. For example, for joy, words like âexcitedâ are frequent, and for sadness, âsadâ is the most fre- quent. A key difference emerged between the two text types. In utterances, the characteristic words tended to include not only emotion words but also words describing events and actions that trigger emotions. For instance, for joy, the word âeatâ is extracted, while for fear, words like âincidentâ and âearthquakeâ are prominent. In contrast, the words that are frequent in emotion transcriptions tend to express the subtle nuances of the emotion. For example, in the category of joy, words like âhappyâ and âexcitedâ appeared, while for fear, âworriedâ and âterrifyingâ are extracted. This shows a linguis- tic difference in expression even within the same emotion category. For the neutral category has a trend of general words, rather than emotion words. Notably, for the emotion transcriptions, the words often used to de- scribe an utteranceâs intent were extracted as char- acteristic words, âagree,â âanswer,â and âexplain.â 7.Conclusion and Future Work We introduced Emotion Transcription in Conver- sation (ETC), a novel task for generating verbal- izations of a speakerâs emotional states, captur- ing fine-grained emotional nuances. To support this task, we constructed a new dataset of 1k Japanese dialogues annotated with self-reported emotion transcriptions and benchmarked baseline models. Our experimental results indicate the util- ity of our dataset for the ETC task, while also high- lighting the challenges. A key challenge is the po- tential gap between explicitly expressed emotions and the actual emotional state; even fine-tuned models failed to capture these nuances (see Ap- pendixFfor failure cases). This suggests the need for models with enhanced capabilities to infer nu- anced emotional states. Future work can address this challenge from several directions. One is to explore sophisti- cated prompting (e.g., chain-of-thought) and ad- vanced fine-tuning objectives, such as RLHF or contrastive learning, to distinguish subtle expres- sions. Furthermore, speaker modeling could en- hance model capabilities. Building on prior ERC research, this approach could involve incorporat- ing speaker embeddings (Hu et al.,2021) or lever- aging personality traits we have collected, which previous studies have linked to emotional states in dialogue (Wang et al.,2024;Shinoda et al.,2025). Additionally, designing human evaluation for the ETC task is a critical future direction. Assess- ing emotion transcriptions requires capturing sub- tle nuances in emotional expressions. While LLM- based evaluators have shown promise in this re- gard, their reliability and potential biases war- rant careful examination. Comparing LLM-based and human evaluations would provide valuable in- sights into the validity of automatic evaluation for this task. Limitations This study has the following limitations and chal- lenges: First, the ETC dataset was constructed using crowdsourcing with Japanese-speaking par- ticipants, which may introduce cultural and linguis- tic biases. To confirm the generalizability of this dataset, similar data collection efforts across di- verse languages and cultural contexts are required. Second, our dataset is limited to the text modal- ity, which may lack nonverbal cues essential for understanding humansâ emotional states. There- fore, constructing multimodal datasets that incor- porate audio and visual information is a promis- ing direction for enriching emotion understanding in real-world conversations. Third, the collected dialogues are relatively brief, comprising only 5 turns (10 utterances). This short context might be insuïŹicient to adequately capture longer con- versational dynamics and the complexity of emo- tional states evolving over time. Additionally, our dataset comprises 1k dialogues, and it remains un- clear whether this scale is suïŹicient for training ro- bust ETC models. However, the crowdsourcing- based collection protocol employed in this study provides a feasible and promising approach for scaling up ETC datasets in the future. Lastly, addi- tional consideration is necessary regarding poten- tial use cases of emotional transcriptions (ET). Fu- ture research should explore how recognized ET can be effectively integrated into conversational systems to inform interactions and enhance emo- tional intelligence. Ethical Considerations This study involves the use of data that cap- tures internal emotional states of individuals ex- pressed through natural language, which neces- sitates careful ethical consideration regarding the sensitivity of emotional content and the potential risks associated with AI applications. First, we col- lected the dialogue data and corresponding emo- tion transcriptions for each utterance through a Japanese crowdsourcing platform. Prior to the data collection, the participants were provided with clear explanations of the studyâs purpose, the na- ture of the data being collected. The participants were asked to consent to the data use before they engaged in the dialogue. The collected data contains no personally identifiable information and was stored in an anonymized format. Also, the proposed Emotion Transcription in Conversation (ETC) task aims to generate natural language de- scriptions of a speakerâs internal emotional state based on conversational context. This technol- ogy holds potential for applications in dialogue sys- tems, educational support, and counseling. How- ever, if misused, it may raise serious ethical con- cerns, such as excessive inference of emotions, reinforcement of cultural biases, or surveillance and manipulation of emotional states. To address these risks, we evaluated the alignment between model-generated outputs and human annotations, and highlighted the current limitations and areas for improvement. Nevertheless, continued atten- tion must be paid to the ethical implications of how such trained models are deployed and used in real- world contexts. Acknowledgments This work was supported by JSPS KAKENHI Grant Number 25H01382. 8.Bibliographical References Huang-Cheng Chou, Wei-Cheng Lin, Chi-Chun Lee, and Carlos Busso. 2022.Exploiting anno- tatorstyped description of emotion perception to maximize utilization of ratings for speech emo- tion recognition . InICASSP2022-2022IEEE InternationalConferenceonAcoustics,Speech andSignalProcessing(ICASSP), pages 7717â 7721. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. InNorthAmericanChapterof theAssociationforComputationalLinguistics: HumanLanguageTechnologies(NAACL-HLT), pages 4171â4186. P. Ekman, W. V. Friesen, M. J. OSullivan, A. K. Chan, I. Diacoyanni-Tarlatzis, K. G. Heider, R. Krause, W. A. LeCompte, T. K. Pitcairn, P. E. Ricci-Bitti, K. R. Scherer, M. Tomita, and A. Tzavaras. 1987.Universals and cultural dif- ferences in the judgments of facial expressions of emotion.volume 53, pages 712â717. Shutong Feng, Guangzhi Sun, Nurul Lubis, Wen Wu, Chao Zhang, and Milica GasiÄ. 2024. Af- fect recognition in conversations using large lan- guage models. InAnnualMeetingoftheSpecial InterestGrouponDiscourseandDialogue (SIGDIAL), pages 259â273. J. L. Fleiss. 1971.Measuring nominal scale agreement among many raters.Psychological Bulletin, 76:378â382. Yumeng Fu, Junjie Wu, Zhongjie Wang, Meishan Zhang, Lili Shan, Yulin Wu, and Bingquan Liu. 2025. LaERC-S: Improving LLM-based emo- tion recognition in conversation with speaker characteristics. In InternationalConference onComputationalLinguistics(COLING), pages 6748â6761. Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Hi- roki Iida, Masanari Ohi, Kakeru Hattori, Hirai Shota, Sakae Mizuki, Rio Yokota, and Naoaki Okazaki. 2024. Continual pre-training for cross- lingual LLM adaptation: Enhancing japanese language capabilities. InFirstConferenceon LanguageModeling. Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Erik Cambria. 2019. DialogueGCN: A graph convolutional neural network for emotion recognition in conversa- tion. InProceedingsofthe2019Conference onEmpiricalMethodsinNaturalLanguage Processingandthe9thInternationalJoint ConferenceonNaturalLanguageProcessing (EMNLP-IJCNLP), pages 154â164, Hong Kong, China. Association for Computational Linguis- tics. Jingwen Hu, Yuchen Liu, Jinming Zhao, and Qin Jin. 2021. MMGCN: Multimodal fu- sion via deep graph convolution network for emotion recognition in conversation . In Proceedingsofthe59thAnnualMeetingof theAssociationforComputationalLinguistics andthe11thInternationalJointConferenceon NaturalLanguageProcessing(Volume1:Long Papers), pages 5666â5675, Online. Association for Computational Linguistics. Chin-Yew Lin. 2004.ROUGE: A package for automatic evaluation of summaries. InText SummarizationBranchesOut, pages 74â81, Barcelona, Spain. Association for Computa- tional Linguistics. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023.G- eval: NLG evaluation using gpt-4 with better hu- man alignment. InProceedingsofthe2023 ConferenceonEmpiricalMethodsinNatural LanguageProcessing, pages 2511â2522, Sin- gapore. Association for Computational Linguis- tics. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoy- anov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXivpreprint arXiv:1907.11692. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evalua- tion of factual precision in long form text gener- ation . InProceedingsofthe2023Conference onEmpiricalMethodsinNaturalLanguage Processing, pages 12076â12100, Singapore. Association for Computational Linguistics. Atsushi Oshio, ABE Shingo, and Pino Cutrone. 2012. Development, reliability, and validity of the japanese version of ten item personality in- ventory (tipi-j).JapaneseJournalofPersonality, 21(1). Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002.Bleu: a method for au- tomatic evaluation of machine translation. In Proceedingsofthe40thAnnualMeetingon AssociationforComputationalLinguistics, ACL â02, page 311Ăąïżż318, USA. Association for Com- putational Linguistics. PatrĂcia Pereira, Helena Moniz, and Joao Paulo Carvalho. 2025. Deep emotion recognition in textual conversations: A survey. Artificial IntelligenceReview, 58(1):1â37. Soujanya Poria, Navonil Majumder, Rada Mihal- cea, and Eduard Hovy. 2019. Emotion recog- nition in conversation: Research challenges, datasets, and recent advances. IEEEaccess, 7:100943â100953. Xincheng Qin, Zhen Wu, Teng Zhang, Yuxuan Li, Jiaxing Luan, Bo Wang, Lidong Wang, and Jian Cui. 2023. BERT-ERC: Fine-tuning BERT is enough for emotion recognition in conversation. InAAAIConferenceonArtificialIntelligence (AAAI), volume 37, pages 13492â13500. Arseny Tolmachev, Daisuke Kawahara, and Sadao Kurohashi. 2018.Juman++: A mor- phological analysis toolkit for scriptio con- tinua. InProceedingsofthe2018Conference onEmpiricalMethodsinNaturalLanguage Processing:SystemDemonstrations, pages 54â59, Brussels, Belgium. Association for Com- putational Linguistics. Yan Wang, Bo Wang, Yachao Zhao, Dongming Zhao, Xiaojia Jin, Jijun Zhang, Ruifang He, and Yuexian Hou. 2024.Emotion recognition in conversation via dynamic personality. In Proceedingsofthe2024JointInternational ConferenceonComputationalLinguistics, LanguageResourcesandEvaluation (LREC-COLING2024), pages 5711â5722, Torino, Italia. ELRA and ICCL. Jintao Wen, Geng Tu, Rui Li, Dazhi Jiang, and Wenhua Zhu. 2023. Learning more from mixed emotions: a label refinement method for emo- tion recognition in conversations. Transactions oftheAssociationforComputationalLinguistics, 11:1485â1499. Fangxu Yu, Junjie Guo, Zhen Wu, and Xinyu Dai. 2024. Emotion-anchored contrastive learn- ing framework for emotion recognition in con- versation. In NorthAmericanChapterof theAssociationforComputationalLinguistics: HumanLanguageTechnologies(NAACL-HLT), pages 4521â4534. Xupeng Zha, Huan Zhao, and Zixing Zhang. 2024. ESIHGNN: Event-state interactions infused het- erogeneous graph neural network for conver- sational emotion recognition. In International ConferenceonAcoustics,SpeechandSignal Processing(ICASSP), pages 11136â11140. Tianyi Zhang, Varsha Kishore, Felix Wu, Kil- ian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. InInternationalConferenceonLearning Representations. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena . InThirty-seventh ConferenceonNeuralInformationProcessing SystemsDatasetsandBenchmarksTrack. 9.Language Resource References C. Busso, R. Lotfian, K. Sridhar, A.N. Salman, W.-C. Lin, L. Goncalves, S. Parthasarathy, A. Reddy Naini, S.-G. Leem, L. Martinez- Lucas, H.-C. Chou, and P. Mote. 2025. The MSP-Podcast corpus.ArXive-prints (arXiv:2509.09791), pages 1â20. Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008. IEMOCAP: Interactive emotional dyadic motion capture database.LanguageResourcesandEvaluation, 42(4):335â359. Ankush Chatterjee, Kedhar Nath Narahari, Meghana Joshi, and Puneet Agrawal. 2019. Semeval-2019 task 3: Emocontext contextual emotion detection in text. In Proceedingsof the13thinternationalworkshoponsemantic evaluation, pages 39â48. Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang, and Lun-Wei Ku. 2018. EmotionLines: An emotion corpus of multi-party conversations. InInternationalConferenceon LanguageResourcesandEvaluation(LREC). Tatsuya Ide and Daisuke Kawahara. 2022.Build- ing a dialogue corpus annotated with expressed and experienced emotions. InProceedingsof the60thAnnualMeetingoftheAssociationfor ComputationalLinguistics:StudentResearch Workshop, pages 21â30, Dublin, Ireland. Asso- ciation for Computational Linguistics. Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017.Dailydialog: A manually labelled multi-turn dialogue dataset. InInternationalJointConferenceonNatural LanguageProcessing(IJCNLP), pages 986â 995, Taipei, Taiwan. Asian Federation of Natural Language Processing. June M. Liu, He Cao, Renliang Sun, Rui Wang, Yu Li, and Jiaxing Zhang. 2025. CAPE: A Chinese dataset for appraisal-based emo- tional generation in large language models. In NorthAmericanChapteroftheAssociationfor ComputationalLinguistics:HumanLanguage Technologies(NAACL-HLT):Findings, pages 6291â6309. L. Martinez-Lucas, Mohammed Abdelwahab, and Carlos Busso. 2020. The MSP-conversation cor- pus. InInterspeech2020, pages 1823â1827, Shanghai, China. Naoaki Okazaki, Kakeru Hattori, Hirai Shota, Hi- roki Iida, Masanari Ohi, Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Rio Yokota, and Sakae Mizuki. 2024.Building a large japanese web corpus for large language models. InFirst ConferenceonLanguageModeling. Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Rishi Naik, Erik Cambria, and Alexander Hoffmann. 2019.MELD: A multi- modal multi-party dataset for emotion recogni- tion in conversations. InAnnualMeetingofthe AssociationforComputationalLinguistics(ACL), pages 527â536, Florence, Italy. Association for Computational Linguistics. Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019.Towards em- pathetic open-domain conversation models: A new benchmark and dataset. InProceedingsof the57thAnnualMeetingoftheAssociationfor ComputationalLinguistics, pages 5370â5381, Florence, Italy. Association for Computational Linguistics. Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno, Keita Suzuki, Ryo Ma- sumura, Hiroaki Sugiyama, and Kuniko Saito. 2025. ToMATO: Verbalizing the mental states of role-playing LLMs for benchmarking theory of mind. AAAIConferenceonArtificialIntelligence, 39(2):1520â1528. Seyyed M. Zahiri and Jinho D. Choi. 2018. Emotion detection on tv show transcripts with sequence-based convolutional neural networks. InAAAIWorkshops. A.Data Collection Details We conducted the collection of dialogue and cor- responding emotion transcriptions using a custom- built chat interface hosted on Google Sheets. The interface consists of multiple dialogue slots (as shown in Figure3), each associated with a prede- fined emotion label. Each dialogue slot in the interface is composed of several key areas. Figure3-A shows metadata for the dialogue, including a predefined emotion la- bel and a unique dialogue slot ID. This area also includes two cells for entering CrowdWorks user- namesone for the Speaker and one for the Lis- tener. In these cells, participants were instructed to enter their username to join the conversation. They then wait to be paired with a dialogue partner or start the conversation if someone has already joined. Figure 3-B allows participants to exchange coordination messages during the task. This area is designed for task-related communication that does not directly affect the dialogue itself and fa- cilitates the accurate and smooth progression of tasks. For instance, participants used it to con- firm matching by entering âNice to meet youâ, in- form dialogue partners of their readiness with âIâm readyâ, or notify other workers of a brief absence while waiting for matching with âIâl be back by 4:00 PMâ. Figure 3-C is the main body of the interaction, where participants input their utterances across five turns each. After a participant (e.g., the Speaker) enters their utterance into a cell in this area, the other participant (e.g., the Listener) re- sponds by entering their utterance in the next cell. Subsequently, while the Listener is typing their re- sponse, the Speaker writes an emotion transcrip- tion for their previous utterance by inserting a note into the same cell. These steps are repeated al- ternately for five turns. During the dialogue ses- sion, we required participants to write at least five characters in length for both the utterance and the corresponding emotion transcription. Figure 3-D consists of checkboxes for tracking task progress. The left column is for utterance in- put, and the right column is for emotion transcrip- tion. At each turn, participants check the corre- sponding box after completing each action. The di- alogue session is considered complete only when all checkboxes have been marked by both partici- pants. B.Distribution of Participantsâ Personality Traits Figure4illustrates the distribution of Big Five per- sonality traits among the 199 participants, mea- sured using the Japanese version of the Ten-Item Personality Inventory (TIPI-J) (Oshio et al.,2012). Scores for each dimension range from 2 to 14, with higher values indicating stronger tendencies. C.Prompt Template for ETC Task The prompt template for the ETC task, trans- lated from Japanese, is shown in Figure5. The model is expected to generate only the emotion transcription, in accordance with the definition. [TARGET_SPEAKER]is a placeholder for the tar- get speaker whose emotion transcription is to be generated, and[DIALOGUE]is a placeholder for the dialogue history. D.Implementation Details We executed supervised fine-tuning of Llama-3.1- Swallow on two NVIDIA A100 80GB GPUs, using a batch size of8and training for2epochs. We applied 4-bit quantization to enhance memory eïŹi- ciency and used thePagedAdamW8bitoptimizer during training. Validation was conducted every 200steps. For hyperparameter tuning, we em- ployed a grid search with learning rates tested in 1e-05, 5e-05and warm-up steps in300, 700. As a result, we selected a learning rate of 1e-05and300warm-up steps. During the inference, to ensure reproducibility, we setdo_sample=Falsefor Llama-3.1-Swallow andtemperatureto0.0for GPT-4.1. E.Prompt Template for Fine-grained Evaluation The prompt templates for the fine-grained eval- uation of content faithfulness, translated from Japanese, are shown in Figure 6and Figure7. The decomposition prompt (Figure6) instructs the model to output a list of atomic units of in- formation contained in the emotion transcription. [TRANSCRIPTION]is a placeholder for the emo- tion transcription to be decomposed. The entail- ment prompt (Figure7) asks the model to deter- mine whether each unit of information is entailed by the target transcription.[SOURCE_TEXT]and [TARGET_TEXT]are placeholders for the source and target emotion transcriptions, respectively. F.Failure Case Analysis for the ETC Task Table7presents a cherry-picked example of where the models failed to predict the speakerâs emotion transcription. In this case, the ground- truth transcription expresses the speakerâs delight at the partnerâs willingness to delve into details. A B C D Figure 3:The dialogue slot used for conducting conversations and entering emotion transcriptions. Figure 4:Distribution of Big Five personality traits among the 199 participants. However, every model instead focuses on the speakerâs shock, illustrating a tendency to prior- itize immediately preceding negative cues over more nuanced, interaction-driven emotions. This suggests that ETC models struggle to incorporate the partnerâs supportive attitude when inferring a speakerâs emotional state. # Task Definition You will be given a multi-turn dialogue between two speakers, Speaker A and Speaker B. Your task is to infer the emotion transcription of the speaker who spoke the last utterance (either Speaker A or Speaker B) at the time it was uttered. # Definition of Emotion Transcription Emotion transcription is defined as a verbalization of the internal emotional state or intention that a participant in the dialogue held at the time of their utterance. It is a linguistic expression of the psychological background behind the utterance, including âwhat they were thinking,â âwhy they chose those words,â or âhow they were feelingâ when they spoke. # Output Rules 1. Predict the emotion transcription corresponding to the last utterance of the specified target speaker. 2. Generate the emotion transcription in a tone as if the speaker themselves were describing it. 3. Do not enclose the emotion transcription in quotation marks or brackets. 4. Do not include any extra explanations after the emotion transcription, nor include the emotion transcriptions of other speakers. 5. When referring to the dialogue partner, use the term âmy dialogue partnerâ; when referring to yourself, use âI.â Target speaker for emotion transcription prediction: [TARGET_SPEAKER] Dialogue: ----- [DIALOGUE] ----- [TARGET_SPEAKER]âs emotion transcription: Figure 5:The prompt template for the ETC task (Translated from Japanese). The model is expected to generate only the emotion transcription. Your task is to decompose a given emotion transcription into semantically independent minimal emotional units. # Definition of Emotion Transcription Emotion transcription is defined as a verbalization of the internal emotional state or intention that a participant in the dialogue held at the time of their utterance. It is a linguistic expression of the psychological background behind the utterance, including âwhat they were thinking,â âwhy they chose those words,â or âhow they were feelingâ when they spoke. # Guidelines - Extract only internal states aligned with the definition above: emotions (e.g., âfelt sad,â âempathized,â âwas gratefulâ) and intentions (e.g., âtried to ...,â âwanted to ...,â âfelt that ...â). - Do not include descriptions of objective facts or actions (e.g., âlikes/dislikes ...,â âperformed ...,â âis explaining ...â). - However, if the transcription contains a dialogue participant's âfeelings,â âthoughts,â or ârecollectionsâ regarding those facts or actions, those descriptions must be extracted. - Ensure that each atomic unit is semantically complete on its own. In particular, the use of demonstratives such as âthat,â âthis,â or âitâ is prohibited. - If no emotional element is present, output "[No emotional elements]." (8 examples) Decompose the following emotion transcription into semantically independent minimal emotional units: [TRANSCRIPTION] Figure 6:The prompt template for the decomposition step in the fine-grained evaluation (Translated from Japanese) You are given two emotion transcriptions: a âpremiseâ and a âhypothesis.â Your task is to determine whether the hypothesis is supported by the premise by **comparing only their surface-level content**. # Mandatory Rules - **Inferring implicit causal relationships or value judgments is prohibited.** - Verify whether the **subject (whose emotion or intention) and the object** of the emotional states match between the premise and the hypothesis. - Any emotional states not **explicitly** stated in the premise must not be considered as implied. - Ignore differences in tense in your judgment. # Definition of Emotion Transcription Emotion transcription is defined as a verbalization of the internal emotional state or intention that a participant in the dialogue held at the time of their utterance. It is a linguistic expression of the psychological background behind the utterance, including âwhat they were thinking,â âwhy they chose those words,â or âhow they were feelingâ when they spoke. # Guidelines **Supported** - The hypothesis can be **definitively determined to be entailed** solely from the content of the premise. **Not Supported** - The emotional element in the hypothesis is **not present in the premise and cannot be definitively determined**, or it **clearly contradicts** the premise. **Neutral** - Only the **core emotional element matches**, while the **validity of secondary elements is uncertain**. # Output Format Return only a pure JSON object strictly following the format below. ```json âreasoningâ: â<Briefly describe the rationale for your judgment here>â, âanswerâ: â<One of: Supported / Neutral / Not Supported>â ``` (9 examples) Example 10 Premise: [SOURCE_TEXT] Hypothesis: [TARGET_TEXT] Figure 7:The prompt template for the entailment step in the fine-grained evaluation (Translated from Japanese). Context Speaker A: The other day, my favorite baseball player suffered a serious injury and will be out for three months. (3) Speaker B: Three months? Are you talking about Okamoto? (3 ) Speaker A: Yes. That news was really shocking. () Speaker B: I thought so. He got hurt in that collision with Nakano from the Hanshin Tigers, and I know how devastating that is for a Giants fan. ( ) Speaker A: Right. At first they reported it as a one-month muscle injury, but after a detailed examina- tion it turned out to be a ligament tear requiring three months of recovery, which made it even more shocking. (1 3) Output Ground-Truth: I feel very happy that my dialogue partner could dive so deeply and discuss the details with me. () GPT-4.1 (zero-shot): I feel even more disappointed and shocked because the injury was worse than initially reported, and Iâm explaining the details so they understand how bad it really is. ( ) GPT-4.1 (4-shot): I wanted my partner to understand how shocked I was. ( ) Llama-3.1 (zero-shot): That incident was really shocking. () Llama-3.1 (4-shot): I feel relieved that my partner understands my feelings and Iâm seeking empa- thy. () Llama-3.1 (fine-tuned): I answered their question, recalled what happened, and told them how much more shocked I felt. ( ) Table 7:Failure case study of emotion transcription generation (translated from Japanese). The outputs are the emotion transcriptions for Speaker Aâs final utterance.