Paper deep dive
Emergent Communication for Co-constructed Emotion Between Embodied Agents via Collective Predictive Coding
Zehang Zhang, Nguyen Le Hoang, Tadahiro Taniguchi, Takato Horii
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/8/2026, 12:47:53 PM
Summary
This study investigates how shared emotion categories emerge between embodied agents through emergent communication, grounded in the Collective Predictive Coding (CPC) framework and the Metropolis-Hastings Naming Game (MHNG). Using an Inter-GMM+MVAE model with visual, auditory, and interoceptive inputs, the research demonstrates that communicative interaction significantly improves the alignment and clarity of learned emotion categories at the symbolic layer. Notably, robust alignment persists even when agents have divergent interoceptive dynamics, providing computational support for the co-constructionist view that physiological heterogeneity is constitutive of, rather than an obstacle to, shared emotional meaning.
Entities (7)
Relation Signals (6)
MHNG â enables â Emergent Communication
confidence 95% · modeling emergent communication between two embodied agents using the Metropolis-Hastings Naming Game (MHNG)
MHNG-based communication â improves â Alignment of emotion categories
confidence 95% · MHNG-based communication significantly improves the alignment, clarity, and inter-agent agreement of the learned emotion categories
Divergent interoceptive dynamics â doesnotprevent â Robust categorical alignment
confidence 90% · even when the two agents have systematically divergent interoceptive dynamics, communication still produces robust categorical alignment
CPC â extends â Predictive Coding
confidence 90% · extends predictive coding from the individual to the social level
Inter-GMM+MVAE â implements â CPC framework
confidence 90% · apply Inter-GMM+MVAE, a computational model of CPC, to model the co-construction of emotion
Interoceptive heterogeneity â isconstitutiveof â Shared emotional meaning
confidence 85% · interoceptive heterogeneity is constitutive of, rather than an obstacle to, shared emotional meaning
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:According to the theory of constructed emotion, the brain actively forms emotion categories by integrating multimodal bodily signals, and constructs emotional experiences by using these categories to predict and interpret sensory inputs. While research has advanced in modeling individual emotion construction, the social process of co-construction-how a shared understanding of emotions emerges between individuals-remains computationally underexplored. This study investigates this process by modeling emergent communication between two embodied agents using the Metropolis-Hastings Naming Game (MHNG), grounded in the Collective Predictive Coding (CPC) framework. Our experiments, using visual, auditory, and simulated interoceptive inputs, yield two main findings. First, MHNG-based communication significantly improves the alignment, clarity, and inter-agent agreement of the learned emotion categories compared to non-communicative and non-selective baselines, with the alignment effect concentrated at the symbolic layer rather than the perceptual latent representation. Second, even when the two agents have systematically divergent interoceptive dynamics, communication still produces robust categorical alignment, with distinct, category-specific reshaping patterns of each agent's emotion categories-consistent with the constructed-emotion view that interoceptive heterogeneity is constitutive of, rather than an obstacle to, shared emotional meaning. These findings provide computational support for the co-constructionist view of emotion and extend the CPC framework from physical to socially-grounded domains.
Tags
Links
- Source: https://arxiv.org/abs/2605.09522v1
- Canonical: https://arxiv.org/abs/2605.09522v1
Trouble viewing inline? Open PDF directly â
Full Text
69,445 characters extracted from source content.
Expand or collapse full text
IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS1 Emergent Communication for Co-constructed Emotion Between Embodied Agents via Collective Predictive Coding Zehang Zhang, Student Member, IEEE, Nguyen Le Hoang, Member, IEEE, Tadahiro Taniguchi, Member, IEEE, and Takato Horii, Member, IEEE, AbstractâAccording to the theory of constructed emotion, the brain actively forms emotion categories by integrating mul- timodal bodily signals, and constructs emotional experiences by using these categories to predict and interpret sensory inputs. While research has advanced in modeling individual emotion construction, the social process of co-constructionâhow a shared understanding of emotions emerges between individu- alsâremains computationally underexplored. This study investi- gates this process by modeling emergent communication between two embodied agents using the Metropolis-Hastings Naming Game (MHNG), grounded in the Collective Predictive Coding (CPC) framework. Our experiments, using visual, auditory, and simulated interoceptive inputs, yield two main findings. First, MHNG-based communication significantly improves the align- ment, clarity, and inter-agent agreement of the learned emotion categories compared to non-communicative and non-selective baselines, with the alignment effect concentrated at the symbolic layer rather than the perceptual latent representation. Second, even when the two agents have systematically divergent interocep- tive dynamics, communication still produces robust categorical alignment, with distinct, category-specific reshaping patterns of each agentâs emotion categoriesâconsistent with the constructed- emotion view that interoceptive heterogeneity is constitutive of, rather than an obstacle to, shared emotional meaning. These findings provide computational support for the co-constructionist view of emotion and extend the CPC framework from physical to socially-grounded domains. Index TermsâTheory of constructed emotion, predictive cod- ing, symbol emergence, emergent communication, Metropolis- Hastings, naming game, deep generative model, machine learn- ing. I. INTRODUCTION H OW do we acquire concepts of emotions from our own sensory experiences? And how do we come to be able to share and understand those emotions with others? Emotions are understood as internal psychological states that profoundly affect individual cognition, experience, and behavior, constituting complex responses to environmental stimuli [1]. Of particular note in the multifaceted nature of emotions is their physical embodiment: emotions encompass This work was supported by JSPS KAKENHI Grant Number JP23H04834 and 23H04835. Z. Zhang is with the Graduate School of Engineering Science, The University of Osaka, Toyonaka 560-8531, Japan. N. L. Hoang and T. Taniguchi are with the Graduate School of Informatics, Kyoto University, Kyoto 606-8501, Japan. T. Horii is with the Graduate School of Engineering Science, The University of Osaka, Toyonaka 560-8531, Japan. (e-mail: takato@sys.es.osaka-u.ac.jp) Manuscript received Month D, Y; revised Month D, Y. not only subjective feelings but also physiological reactions, cognitive evaluations, and behavioral tendencies [2]. Recent findings emphasize the role of interoceptionâsensory signals originating within the bodyâas a crucial determinant in the experience and formation of emotions [3]. These interocep- tions are thought to eventually become integrated into a basic state known as core affect, which represents feelings of valence (pleasureâdispleasure) and arousal (highâlow activation) [4]. Unlike specific emotion categories, core affect acts as the continuous physiological basis from which emotions are con- structed. Beyond internal mechanisms, emotions are also modulated by external social and cultural contexts â cross-cultural studies have demonstrated that the expression, perception, and regulation of emotions are significantly influenced by cultural frameworks [5], [6]. Yet despite emotionâs critical role in cognition and social life, it remains insufficiently understood how emotion categories are constructed and how they come to be shared among individuals [7], [8], [9]. Among theories of emotion, Ekmanâs basic emotion theory posits that a small set of emotionsâsuch as joy, sadness, anger, fear, disgust, and surpriseâare biologically hardwired, each underpinned by a distinct neural circuit, and universally expressed across cultures [10]. However, this essentialist view has faced substantial criticism. Meta-analyses of neuroimaging studies have failed to identify consistent, emotion-specific brain signatures [11], [12], and cross-cultural research has revealed significant variation in how emotions are recognized and categorized [13], [14], challenging the assumption of universality. In contrast, the theory of constructed emotion [7] challenges this essentialist perspective. It argues that humans do not possess biologically hardwired reactions for specific emotions. Instead, the theory distinguishes between emotion categories and emotional experiences. Emotion categories are concep- tual knowledge structures formed through the integration of interoception and exteroception, shaped by personal beliefs and memories. Emotional experiences, in turn, are the specific instances constructed when the brain applies these categories to predict and interpret continuous sensory inputs. In Barrettâs framework, the formation of emotion categories is a prerequi- site for these experiences; without the corresponding category, an individual would perceive only raw bodily sensations rather than specific emotions. Specifically, this construction unfolds in two conceptual stages. First, the brain acquires emotion arXiv:2605.09522v1 [cs.MA] 10 May 2026 IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS2 Agent A Agent B Happy Surprise Object or Event Exteroception Interoception Core Affect Latent space Exteroception Interoception Core Affect Latent space â Observations of Agent Aâ Observations of Agent B âĄRepresentation Learning âĄRepresentation Learning âąA speak to B âŁB speak to A â Observationsof Agent Aâ Observations of Agent B âĄRepresentation Learning âĄRepresentation Learning âąA speak to B âŁB speak to A â Observations of Agent Aâ Observations of Agent B âĄRepresentation Learning âĄRepresentation Learning âąA speak to B âŁB speak to A Fig. 1: The process of two agents (Agent A and Agent B) forming and sharing emotion categories after observing the same object through communication. Each agent receives the similar external stimulus (A joint attention object) and responds to it through exteroception and interoception (e.g., heartbeat, visceral state, etc.).Interoception is integrated within the agent to generate core affect, which is further integrated with exteroception to promote the formation of emotion categories. Subsequently, the agents communicate symbolically (language interaction) based on emotion category inference. Through this emotion communication mechanism, the two agents dynamically update their respective emotion categories and achieve the co-construction of emotion. categories (e.g., âangerâ, âjoyâ) learned from statistical reg- ularities in past interoceptive and exteroceptive experiences. Second, specific emotion are experienced by applying these categories to current sensory inputs to resolve ambiguity. For instance, an accelerated heartbeat (interoception) is merely physiological noise until the brain categorizes it. If the brain predicts this state as âanxietyâ in a stressful context, it is experienced as distress; if categorized as âexcitementâ in a positive context, the same physiological state is perceived as thrill. Thus, the emotion category shapes the subjective quality of the experience. This constructionist view begs for a computational explanation of how such category formation and predictive experience might be implemented in the brain. The theory of predictive coding, proposed by Friston [15], offers a unifying computational framework that provides a natural explanation for these processes. In this framework, the brain continuously constructs internal models that predict incoming sensory stimuli and strives to minimize the dis- crepancy between predictions and actual inputs, a discrepancy termed prediction error. Importantly, predictive coding extends beyond exteroceptive signals to interoceptive signals, as argued by Seth et al. [16], who proposed that emotional experiences arise from minimizing prediction errors related to internal bodily states. Thus, predictive coding serves as a theoretical foundation that complements and reinforces the mechanisms proposed by the theory of constructed emotion, framing emotional experiences as emergent outcomes of hierarchical predictive processes. Building on this approach, prior research has explored the internal formation of emotion categories based on em- bodied sensory inputs [17], [18], but did not address how shared structures of emotion categories could emerge between individualsâa key aspect in understanding the social and cultural co-construction of emotion. In contrast, Gendron and Barrett [8] proposed a framework in which emotion perception is a dynamic process of conceptual synchrony between two individuals, mediated in part by emotion words. While the- oretically influential, their proposal has not been instantiated computationally. Addressing this gap requires a computational framework capable of explaining how shared meaning emerges between individuals through interaction. The theory of symbol emer- gence [19] provides such a foundation: shared symbols are not predefined but arise through a bottom-up process in IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS3 which agents form internal categories from their own sensory experiences and gradually align these categories through com- municative interaction [20]. Taniguchi et al. further formalized this process through CPC [21], which extends predictive cod- ing from the individual to the social level â agents not only minimize prediction errors within their own sensory models but also coordinate their internal representations through the exchange of symbolic signs, thereby achieving inter-agent alignment. The co-construction of emotion maps naturally onto this framework: emotion categories, formed from each individualâs unique bodily and sensory experience, can be viewed as internal representations that become socially aligned through communicative exchange. This study thus situates the co-construction of emotion within the CPC framework and models it as a process of CPC, as conceptually illustrated in Fig. 1. Specifically, we employ the Inter-Gaussian Mixture Model+MultivariateVariationalAutoencoder(Inter- GMM+MVAE) [22], which integrates a multimodal deep generative model with the MHNG. We use this framework to simulate emergent communicationâa bottom-up process where shared symbol systems evolve through local interactions without external supervisionâbetween two agents processing visual, auditory, and interoceptive inputs. Our primary goal is to investigate how this communication influences the formation and alignment of emotion categories. To do so, we compare the structure and coherence of categories acquired in scenarios with and without MHNG-based communication. Crucially, the theory of constructed emotion posits that physiological variations are ubiquitous across individuals; the same category (e.g., âangerâ) may be constructed from different interoception depending on the person [7]. This raises a fundamental question: does shared emotional understanding require agents to have identical bodies, or can it be achieved despite physiological differences? Gendron and Barrett [8] suggest that emotion perception relies on âconceptual synchronyâ rather than physiological mirroring. To computationally validate this hypothesis, we conduct experiments introducing agents with divergent interoceptive dynamics. By manipulating the similarity of interoception in the two agents, we investigate whether emergent communication can bridge the âinteroceptive gapâ and enable the co-construction of shared emotion categories even in the absence of physiological isomorphism. To probe the crucial role of embodied reaction, we conduct additional experiments manipulating the agentsâ interoceptive inputs. We specifically examine conditions where embodimentâ specifically interoceptive reactivityâdiffers between the two agents, in addition to conditions with a complete absence of interoception. The main contributions of this study are as follows: âą We apply Inter-GMM+MVAE, a computational model of CPC, to model the co-construction of emotion and em- pirically confirm the effect of emergent communication on emotional alignment. âą We investigate the effect of interoceptive differences on the co-construction of emotion by introducing agents with systematically varied embodiments. Our experi- ments demonstrate that MHNG-based communication not only supports the formation of similar emotion category structures between agents, but also absorbs differences in bodily reactions, enabling robust co-construction of emotion even when the two agents have divergent in- teroceptive dynamics. I. RELATED STUDY A. Emotion model Several studies have explored how emotion categories emerge from embodied sensory experience, spanning both theoretical frameworks and computational models. Horii et al. [17] proposed a developmental model of emotion perception in which an infant agent learns to recognize a caregiverâs emotions by integrating visual, auditory, and tactile signals through hierarchically structured restricted Boltzmann ma- chines, showing that tactile dominance and perceptual im- provement jointly facilitate emotion differentiation. Hieida et al. [18] developed a computational model of emotion that integrates visual stimuli with simulated interoceptive signals, framing emotion formation as a generative process driven by the interaction between internal and external appraisal. Seth [23] proposed the interoceptive inference model, which applies predictive coding to interoception, arguing that emotional ex- periences arise from the brainâs active inference on the causes of internal bodily signals rather than from passive bottom- up processing. While these studies successfully examined the internal formation of emotion categories within a single agent, none addressed how shared emotion categories could emerge between individuals â the social co-construction of emotion process that is the focus of this study. B. Co-construction of Concepts Gendron and Barrett [8] proposed that emotion perception is not the passive detection of discrete signals but a dynamic process of conceptual synchrony between two individuals. Drawing on predictive coding and grounded cognition, they argued that both the perceiver and the target continuously generate and refine predictions about each otherâs internal states through multimodal sensory cues such as facial ex- pressions, vocal changes, and bodily movements. Language plays a central role in this process: emotion words serve as efficient activators of conceptual knowledge and as explicit bids for mutual understanding, enabling what they termed the co-construction of emotion. However, their framework remained at the theoretical level, without proposing specific computational models. To computationally model such co-construction processes, the theory of symbol emergence systems (SES) [19] and its formalization through CPC [21] provide a promising foun- dation. In SESs, agents form internal representations through physical interactions with the environment and simultaneously organize shared external representations (symbols) through semiotic communication with other agents. Taniguchi formal- ized this dynamics through the CPC hypothesis, which extends predictive coding from a single brain to a multi-agent system: symbol emergence is cast as decentralized Bayesian inference, IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS4 where agents collectively infer shared representations that maximize the predictability of their distributed sensory obser- vations. The MHNG [20], [24] provides a concrete algorithmic realization of this process â agents exchange signs without explicit feedback, and the acceptance or rejection of proposed signs follows a Metropolis-Hastings criterion, guaranteeing convergence to the posterior distribution over shared symbols. Several computational models have instantiated this frame- work with increasing complexity. Taniguchi et al. [24] pro- posed the Inter-GMM+VAE, the first deep generative model for emergent communication based on the MH naming game, in which two agents observing the same objects from different viewpoints cooperatively form internal representations via VAEs, learn categories via GMMs, and share signs without ex- plicit feedback. Hagiwara et al. [25] proposed the Inter-MDM, which extended the framework to multimodal agents equipped with visual, auditory, and haptic modalities, demonstrating that emergent communication enables agents to form shared categories even when some sensory modalities are missing. Hoang et al. [22] further advanced the approach with the Inter- GMM+MVAE, integrating multimodal VAEs with the MH naming game and systematically comparing fusion strategies (PoE, MoE, MoPoE), showing that PoE consistently produced the most effective latent spaces for emergent communication. Beyond object categorization, Sakurai et al. [26] demon- strated the generality of the CPC framework by applying it to collaborative music generation through MH-MuG, where agents with distinct musical knowledge collectively produce stylistically fused compositions via decentralized Bayesian inference. While these models have been applied to object categorization and music generation, no prior work has applied the CPC framework to the domain of emotion â the focus of the present study. I. PRELIMINARIES We adopt the Inter-GMM+MVAE framework proposed by Hoang et al. [22], which integrates MVAE with a MHNG to support symbol emergence between agents. In this section, we first introduce the Product-of-Experts MVAE (PoE-MVAE), which serves as the generative model for integrating intero- ceptive and exteroceptive information. Second, we describe the generation of core affect, which serves as the continuous physiological basis for the agentâs emotional experience and acts as a crucial input to the model. Finally, we detail the MHNG, the mechanism that enables agents to communicate and align emotion categories without explicit feedback. A. Multimodal Variational Autoencoder(MVAE) The theory of constructed emotion posits that emotion cat- egories are formed through the integration of multiple sensory streams â including interoception, visual perception, and au- ditory input â rather than from any single modality alone [7]. Computationally modeling this process therefore requires a generative framework capable of fusing heterogeneous sensory inputs into a unified latent representation. The MVAE provides such a framework: it learns a shared latent space from multiple modalities, enabling both the integration of complementary information and the reconstruction of individual modalities from this shared representation [27]. In this study, we adopt the PoE-MVAE [27] to integrate multiple modalities into a shared latent representation. PoE achieves this by taking the product of posterior distributions from modality-specific encoders, resulting in a consensus representation that emphasizes the overlapping beliefs across modalities. Formally, given modality-specific posterior approximations q Ï m (z | x m ) for each modality m â 1,...,M, the joint latent posterior is defined as: q PoE (z | X)â M Y m=1 q Ï m (z | x m ),(1) where z denotes the shared latent variable, x m represents the observation for the m-th modality, and X =x 1 ,...,x M is the set of all multimodal observations.This fusion mechanism makes PoE especially robust to noisy or missing modalities, while being able to generate a compact and consistent latent space. Compared with other multimodal fusion strategies, such as Mixture-of-Experts (MoE) [28] and Mixture-of-Product-of- Experts (MoPoE) [29], PoE consistently demonstrates superior performance in both clustering metrics and latent space qual- ity, particularly in emergent communication tasks using the MHNG [22]. Critically for emotion modeling, PoE effectively resolves the ambiguity often present in individual sensory channels. For example, even if the posterior probability of an emotion cat- egory is ambiguous based on a single modality (e.g., a subtle facial expression), PoE integrates complementary signals from other modalities (e.g., voice or interoception) to produce a sharper, more confident joint posterior. In contrast, averaging- based methods like MoE tend to dilute such signals, resulting in flatter or multi-peaked distributions. This ability to yield a concentrated and unimodal latent representation not only handles multimodal ambiguity but also aligns closely with the Gaussian assumptions made by GMM, thereby facilitating more stable and effective parameter estimation and clustering performance. B. Metropolis-Hastings Naming Game (MHNG) The MHNG [20], [24] is a probabilistic communication game that enables two agents to develop shared symbolic representations which does not rely on explicit feedback. In each interaction, both agents jointly attend to the same object, forming internal representations through their own sensations. One agent, designated as a speaker, samples a sign from a posterior distribution conditioned on its internal representation inferred from sensory signals and sends it to a listener. The listener evaluates the received sign based on its own inferred internal state, applying a MH acceptance criterion: r = min 1, P (z Li d |ÎŒ Li , Î Li ,w Sp d ) P (z Li d |ÎŒ Li , Î Li ,w Li d ) ! ,(2) where z Li d is the listenerâs latent representation, and w Sp d and w Li d denote the speakerâs and listenerâs signs, respectively. If the listener accepts the received sign according to this prob- ability, it updates its internal model parameters accordingly. IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS5 Algorithm 1: MHNG 1 Procedure MHNG(z Sp ,ÎŒ Sp , Î Sp ,z Li ,ÎŒ Li , Î Li ,w Li d ) 2 w Sp d ⌠P (w Sp d | z Sp d ,ÎŒ Sp , Î Sp ) 3 r ⌠min 1, P (z Li d |ÎŒ Li ,Î Li ,w Sp d ) P (z Li d |ÎŒ Li ,Î Li ,w Li d ) 4 u⌠Unif(0, 1) 5if u†r then 6w d = w Sp d 7else 8w d = w Li d 9end 10 end The roles of speaker and listener are then alternated, and the process is iteratively repeated across data points. The entire MHNG procedure is summarized in Algorithm 1. Algorithm 1 details the computational steps of a single com- municative interaction. First, the speaker generates a proposal sign w Sp d by sampling from its distribution conditioned on its latent state z Sp d . Second, the listener evaluates this proposal against its own current sign w Li d by calculating the acceptance ratio r using equation 2. This ratio effectively measures whether the speakerâs sign explains the listenerâs internal state (z Li d ) better than the listenerâs existing sign. Finally, the listener accepts the new sign w Sp d with probability min(1,r), thereby dynamically aligning its symbolic representation with the speakerâs without requiring direct access to the speakerâs internal state. C. Model of emotional communication:Inter-GMM+MVAE The Inter-GMM+MVAE [22] combines MVAE and MHNG to model emergent communication based on multimodal infor- mation. The two agents first learn the multimodal information of the joint-attention objects through MVAE to obtain the latent representation, and then modulate the sign learned by their respective GMMs through MHNG. The definition of Inter-GMM+MVAE can refer to Table I and Fig. 2. Its generative process is defined as follows: w d ⌠Cat(Ï)d = 1,...,D(3) ÎŒ â k , Î â k âŒN (ÎŒ â k |m, (αΠâ k ) â1 )W(Î â k |Μ,ÎČ) k = 1,...,K(4) z â d âŒN (z â d |ÎŒ â w d , (Î â w d ) â1 )d = 1,...,D(5) o â â,d ⌠P Ξ â â (o â â,d |z â d )d = 1,...,D(6) Algorithm 2 outlines the iterative learning dynamics where percep- tion, communication, and learning are interleaved. In each epoch, agents first infer their individual latent states z d from multimodal ob- servations. Subsequently, they engage in the naming game (MHNG) to update the shared signs w d . Crucially, these communicated signs then serve as pseudo-labels for the parameter update step: each agent optimizes its GMM parameters (ÎŒ, Î) and VAE decoder parameters (Ξ) to maximize the likelihood of the observations given the agreed- upon signs. By alternating the roles of speaker and listener, both agents mutually adapt their internal representations to align with the socially constructed symbol system. Algorithm 2: Iterative Co-construction of Emotion Categories via Inter-GMM+MVAE 1 Initialize all parameters 2 for t = 1 to T do 3for d = 1 to D do 4z A d ⌠P (z A d | Ξ A ,o A d ,w A d ,ÎŒ A , Î A ) 5z B d ⌠P (z B d | Ξ B ,o B d ,w B d ,ÎŒ B , Î B ) 6end 7// Agent A speaks to Agent B 8for d = 1 to D do 9w B d â MHNG(z A ,ÎŒ A , Î A ,z B ,ÎŒ B , Î B ,w B d ) 10end 11// Learning by Agent B 12 ÎŒ B , Î B ⌠P (ÎŒ B , Î B | w B ,z B ,α B ,m B ,ÎČ B ,v B ) 13 Ξ B ⌠P (Ξ B | o B ,z B ) 14// Agent B speaks to Agent A 15for d = 1 to D do 16w A d â MHNG(z B ,ÎŒ B , Î B ,z A ,ÎŒ A , Î A ,w A d ) 17end 18// Learning by Agent A 19 ÎŒ A , Î A ⌠P (ÎŒ A , Î A | w A ,z A ,α A ,m A ,ÎČ A ,v A ) 20 Ξ A ⌠P (Ξ A | o A ,z A ) 21 end Fig. 2: Graphical model of Inter-GMM+MVAE used in this research. Each agent ââA,B (left: blue, right: red) has its own GMM and MVAE. The interaction process between the two agents proceeds as follows: (1) Both agents are exposed to the same emotional stimulus E d (e.g., watching a movie together) and each agent ex- presses individual reactions (such as facial expressions, vocalizations, and changes in interoception). (2) Each agent observes their own reactions as observation data o â â,d and infers the latent variable z â d using their GMM+MVAE models. (3) One agent takes the role of the speaker and generates a IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS6 TABLE I: Definition of variables used in the generative process of Inter-GMM+MVAE (Eqs. (3)â(6)) and the inference algorithms (Algorithms 1 and 2). Variables with superscript â are agent-specific (ââA,B) SymbolDefinition ââA,BAgent identifier (A and B). ââi,v,aModality index: interoception (i), vision (v), and audio (a). KNumber of signs (categories). DNumber of data points. w d Discrete variable drawn from a categorical distribution; represents the assigned sign for the d-th data point. z â d Latent variable inferred by each agent based on multimodal observations. o â,d Observed multimodal sensory data corresponding to each data point. Ξ â Variational parameters associated with the VAE for each modality. ÎŒ â k ,Î â k Mean vector and precision matrix parameters of the kth multivariate normal distribution within the GMM. a,m,ÎČ,Μ,ÏHyperparameters for the distributions ÎŒ,Î, and w. sign w Sp d inferred from z Sp d , which is then sent to the another agent, i.e., the listener agent. (4) (4) The listener agent computes the acceptance proba- bility r defined in Eq. (2) and probabilistically accepts the speakerâs sign. If accepted, the listener updates the parameters of its GMM+MVAE model according to the new sign; otherwise, it retains its previous sign. (5) The roles of speaker and listener are then switched, and steps (3) and (4) are repeated. (6) Steps (1) through (5) are iteratively repeated. IV. EXPERIMENT SETTING We examine the effect of the presence or absence of MHNG, i.e., the interaction between two agents through the exchange of signs, on the formation of each agentâs emotion category. Furthermore, by giving Agent B core affect that is different from that of Agent A, we consider the effect that differences in core affect have on the co-construction of emotions. Specifically, we use the Adjusted Rand Index (ARI) [30] in two complementary ways. First, we measure ARI between the agent-emerged signs and the stimulus-level emotion category provided by RAVDESS (i.e., the emotion the actor was in- structed to portray). We emphasise that this is not a supervised classification targetâthe model receives no label during learn- ingâbut rather a reference label that records which stimuli were intended to evoke the same emotion, against which we can assess whether the unsupervised co-construction process recovers psychologically meaningful clusters. We use Cohenâs Kappa coefficient [31] as a complementary, chance-corrected measure of inter-agent agreement, and visualise the structure of the formed categories using heatmaps of recall against the RAVDESS reference labels. Furthermore, we visualize the latent variable z d using PCA and t-distributed Stochastic Neighbor Embedding (t-SNE) [32], and evaluate the structure of emotion categories by calculating the similarity of the latent space z of the two agents using TopSim [33]. We also evaluate the differentiation of emotion categories using the Davies- Bouldin Score (DBS) [34]. Based on our theoretical framework, we formally propose the following hypotheses: H1: Emergentcommunicationfacilitatestheco- construction of emotion. We predict that the MHNG scenario will result in: (a) a stronger correspondence between learned categories and the RAVDESS reference categories, as measured by a higher ARI; (b) higher inter-agent agreement on emotional symbols (i.e., higher Kappa); and (c) more clearly differentiated individual emotion categories (i.e., lower DBS), compared to the No Communication scenario. H2: Interoceptive similarity enhances the quality of co- constructed categories. We predict that agents with similar core affect that communicate via MHNG will develop: (a) more accurate emotion categories (higher ARI); (b) more distinct emotion categories (lower DBS); and (c) more structurally similar latent representations (higher TopSim), compared to agents with divergent core affect models. A. Dataset:Agentâs body reactions as observations of the model When considering the emergence of emotion symbols, we use information based on the agentâs own body reactions as the observation of the model. In other words, the agentâs body reactions, such as facial expressions, vocal expressions, and interoception, when viewing a certain emotion evocative stimulus E d are treated as the input of the MVAE of each agent. We used the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) [35] as the facial and vocal expressions corresponding to this emotion evocative stimulus. RAVDESS is a video data set of speech and song in eight emotion categories (Neutral, Calm, Happy, Sad, Angry, Fear- ful, Disgust, Surprised) collected from 12 male and 12 female actors. In the experiment, we used only the speech video data, and obtained Facial Action Units (FAUs) at each time as facial expressions and Mel-frequency Cepstral Coefficients (MFCCs) as vocal expressions of agent. We used OpenFace [36] for facial expressions. Using this method, the strengths of 35 FAUs were obtained for each frame of the video. In addition, 20- dimensional MFCC, âMFCC, and âMFCC were obtained from the video as vocal expressions. The frame length for obtaining MFCC was 2048 samples, and calculations were performed every 512 samples. B. Generation of core affect We model core affect as a stochastic process to capture the inherent temporal fluctuations and physiological noise characteristic of bodily sensations. Specifically, we assume IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS7 1.00.50.00.51.0 Valence 1.0 0.5 0.0 0.5 1.0 Arousal Disgust Angry Fearful Sad Happy Surprised Calm Neutral (a) Original core affect 1.00.50.00.51.0 Valence 1.0 0.5 0.0 0.5 1.0 Arousal Disgust Angry Fearful Sad Happy Surprised Calm Neutral (b) Happy inverse core affect 1.00.50.00.51.0 Valence 1.0 0.5 0.0 0.5 1.0 Arousal Disgust Angry Fearful Sad Happy Surprised Calm Neutral (c) Low valence focus core affect 1.00.50.00.51.0 Valence 1.0 0.5 0.0 0.5 1.0 Arousal Disgust Angry Fearful Sad Happy Surprised Calm Neutral (d) Low arousal focus core affect Fig. 3: We used only Original core affect for Agent A, and each of these four core affect for Agent B in the experiment.Among them, Happy inverse core affect changes the mean of Happy from (0.9, 0.5) to (-0.9, -0.5), low valence focus core affect changes ÎŒ v , Ξ v , Ï v to one-fourth of Original core affect, and low arousal focus core affect changes ÎŒ a , Ξ a , Ï a to one-fourth of Original core affect. 1.00.50.00.51.0 Valence 1.0 0.5 0.0 0.5 1.0 Arousal Neutral(0.0, 0.0) Calm(0.8, -0.5) Happy(0.9, 0.5) Sad(-0.7, -0.5) Angry(-0.6, 0.6) Fearful(-0.8, 0.7) Disgust(-0.9, 0.2) Surprised(0.0, 0.8) Means Valence-Arousal Coordinates of Emotions Neutral Calm Happy Sad Angry Fearful Disgust Surprised Fig. 4: Mean valenceâarousal coordinates for each emotion category. Colors are consistent with emotion labels that core affect is governed by homeostatic regulationâthe tendency of the bodyâs internal state to return to a baseline after being perturbed by stimuli. To mathematically represent this dynamics, we employ the Ornstein-Uhlenbeck (OU) pro- cess, which is a mean-reverting stochastic process defined as follows: dX t = Ξ(ÎŒâ X t )dt + Ï dW t ,(7) where: âą X t : the value of the stochastic process at time t; ⹠Ξ > 0: the mean reversion rate, which determines how quickly the process reverts to the mean ÎŒ; âą ÎŒ: the long-term mean toward which the process tends; âą Ï: the volatility coefficient, controlling the magnitude of random fluctuations; âą dt: an infinitesimal time increment; âą dW t : an infinitesimal increment of the Wiener process at time t. In this process, the variable X t starts from a specific initial value and stochastically fluctuates, tending to revert towards the long-term mean ÎŒ. We referred to Russellâs Circumplex Model [4], which reduces core affect to two dimensions: va- lence and arousal, and set a target mean for each emotion. The target mean for each emotion is shown in Fig. 4. We set the target mean of the OU process for each sample in the dataset to the mean vector corresponding to its labeled emotion. To simulate the ecological dynamics of emotional experience, we assume a sequential process where an agent transitions from a preceding emotional state to a new state triggered by a specific event. To capture these transition dynamics, each data sample is replicated seven times, with each replica initialized using the mean of one of the seven other emotion categories (different from the target). This initialization strategy effectively models the trajectory of core affect as it shifts from various prior states toward the new attractor defined by the current stimulus, thereby mitigating bias from any single fixed initial condition. Figs. 5a and 5b illustrate examples where the target means correspond to Neutral and Happy, respectively. The specific values for the mean reversion rate Ξ and volatility Ï assigned to each emotion category are detailed in Table I in the Appendix. 1) Asymmetric interoceptive profiles: A central tenet of constructed emotion theory is that interoceptive variability across individuals is the rule, not the exception [7]. Empirical work in affective neuroscience supports this: interoceptive accuracy modulates the intensity and quality of subjective emotional experience [37], and individuals high in alexithymia show atypical interoceptive sensitivity along the valence and arousal dimensions [38], [39]. To probe whether MHNG- mediated communication can bridge such heterogeneity, we construct three asymmetric core-affect profiles in addition to the Original baseline (Fig. 3): âą Happy-inverse: the Happiness attractor is inverted from (0.9, 0.5) to (â0.9,â0.5), simulating an agent for whom positively valenced events elicit a negatively valenced bodily response. âą Low-valence-focus: ÎŒ v , Ξ v , Ï v are reduced to one-quarter of the Original values, simulating an agent whose intero- ceptive readout along the valence axis is attenuated. âą Low-arousal-focus: ÎŒ a , Ξ a , Ï a are reduced to one- IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS8 quarter, simulating an agent with attenuated arousal sen- sitivity. In all asymmetric experiments, Agent A retains the Original profile while Agent B adopts one of the three asymmetric profiles. We emphasise that these profiles are not intended as models of pathological interoception; rather, they probe the broader space of interoceptive variation that is normal in any human population. â1.0â0.50.00.51.0 Valence â1.0 â0.5 0.0 0.5 1.0 Arousal Neutral Calm Happy Sad Angry Fearful Disgust Surprised Angry Happy Sad Disgust Surprised Fearful Calm (a) Target emotion: Neutral â1.0â0.50.00.51.0 Valence â1.0 â0.5 0.0 0.5 1.0 Arousal Neutral Calm Happy Sad Angry Fearful Disgust Surprised Neutral Disgust Calm Surprised Angry Fearful Sad (b) Target emotion: Happy Fig. 5: Visualization of core affect trajectories generated by the Ornstein-Uhlenbeck stochastic process. The plots illustrate the stochastic nature of bodily reactions, showing temporal transitions from various initial states converging toward the mean vector of the target emotion ((a) Neutral and (b) Happy) in the Valence-Arousal space. C. Experimental setup To investigate the impact of communication within the framework of the co-construction of emotion, we investigate the latent variable z d of the GMM+MVAE for both agents in a model that adopts the MHNG framework for the Inter- GMM+MVAE (hereafter referred to as the MHNG scenario), a model in which the acceptance probability r in MHNG is always 0 (hereafter referred to as the No communication scenario), and a model in which the acceptance probability r is always 1 (hereafter referred to as the All-acceptance scenario). Here, the computational model for the No communication condition can be considered as each agent being equipped with its own GMM+MVAE. In this experiment, the data of 12 males from RAVDESS was used as the observation information for agent A, and the data of 12 females was used as the observation in- formation for agent B. The number of input dimensions for each mode of MVAE is 35Ă109[frame] for o vd , 3815 dimensions, 60Ă345[frame] for o ad , 20700 dimensions, and 2Ă345[frames] for o id , 690 dimensions. For RAVDESS video data with shorter video length, the data is padded with the last frame. The PoE is used for MVAE, and the number of dimensions of the latent variables in GMM+MVAE is set to 9. Experiments were performed 10 times for each scenario, and the results were evaluated using the ARI, DBS, heat map, Kappa coefficient, TopSim, and visualizations of the latent variables using PCA and t-SNE. V. EXPERIMENT RESULTS In this section, we showed the experimental results, orga- nized around our primary research questions. We first analyze the general effect of emergent communication on the formation of emotion categories. We then examine the impact of com- munication on the underlying latent space structure. Finally, we investigate how divergence in core affect between agents affects the co-construction process. A. Effect of Emergent Communication on Category Formation To analysis the effect of emergent communication, we compared the three main experimental scenarios. The quanti- tative results, presented in Table I, show a clear performance hierarchy. The MHNG scenario consistently outperformed the base- lines across all key metrics. It achieved the highest ARI, indicating the strongest correspondence between the learned signs and the RAVDESS reference labels. Similarly, it yielded the highest Cohenâs Kappa coefficient, signifying the strongest inter-agent agreement on the emerged emotional symbols. For the Davies-Bouldin Score (DBS), where lower values indicate better-defined and more separated clusters, the MHNG sce- nario produced the lowest (best) scores, as shown in Table I. In contrast, the No Communication scenario performed moderately, serving as a baseline for independent learning. Notably, the All-acceptance scenario, which forces agents to accept all proposed symbols, consistently resulted in the poorest performance across ARI and DBS. These findings suggest that emergent communication, when mediated by an intelligent and selective mechanism like MHNG, is highly effective for the co-construction of emotion categories. It not only helps agents to differentiate their own emotional concepts more clearly (lower DBS) but also enables IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS9 TABLE I: Clustering performance and quality metrics (mean ± std) under different emotional conditions and communication strategies. The top table displays the Adjusted Rand Index (ARI) and Cohenâs Kappa, where higher values are better (â). The bottom table shows the Davies-Bouldin Score (DBS), where lower values are better (â), and Topic Similarity (TopSim), where higher values are better (â). For all metrics, bold indicates the best result and underline indicates the second-best. Condition ARI (â)Kappa (â) No Com.MHNGAll Acc.No Com.MHNGAll Acc. Agent AAgent BAgent AAgent BAgent AAgent B Vision+Audio / Vision+Audio0.14±0.03 0.17±0.040.19±0.04 0.20±0.04 0.17±0.070.15±0.07 -0.00±0.04 0.38±0.07 0.34±0.11 Original / Original0.28±0.080.21±0.050.41±0.08 0.35±0.08 0.12±0.05 0.12±0.050.01±0.060.51±0.09 0.22±0.08 Original / Happy inverse0.26±0.070.23±0.050.43±0.08 0.41±0.07 0.09±0.02 0.08±0.02 -0.01±0.03 0.49±0.07 0.16±0.04 Original / Low valence focus0.30±0.08 0.16±0.040.30±0.08 0.26±0.07 0.06±0.010.07±0.02 -0.01±0.03 0.39±0.07 0.14±0.02 Original / Low arousal focus0.28±0.070.09±0.010.34±0.07 0.20±0.04 0.02±0.01 0.03±0.01 -0.00±0.02 0.39±0.04 0.06±0.02 Condition DBS (â)TopSim (â) No Com.MHNGAll Acc.No Com.MHNGAll Acc. Agent AAgent BAgent AAgent BAgent AAgent B Vision+Audio / Vision+Audio4.03±0.683.75±0.933.25±0.50 3.43±0.815.91±2.347.37±4.320.20±0.050.14±0.07 0.24±0.11 Original / Original5.16±0.60 5.44±1.114.41±0.85 4.47±0.8010.86±2.9011.07±2.73 0.18±0.06 0.22±0.060.25±0.08 Original / Happy inverse4.28±0.635.68±1.183.65±0.70 4.38±1.0415.71±6.6314.05±4.76 0.12±0.06 0.16±0.07 0.16±0.08 Original / Low valence focus4.69±0.716.34±1.324.39±0.84 5.44±1.0113.86±3.5716.51±6.33 0.13±0.12 0.15±0.060.25±0.04 Original / Low arousal focus4.58±0.747.34±1.694.03±0.71 6.03±1.08 32.99±10.27 29.66±6.67 0.10±0.06 0.14±0.03 0.12±0.05 AgentAAgentB All AcceptanceNo Communication MHNG Kappa: 0.13 ARI A: 0.05 ARI B: 0.07 PCA t-SNE AgentAAgentB Kappa: 0.43 ARI A: 0.35 ARI B: 0.26 PCA t-SNE AgentAAgentB Kappa: 0.02 ARI A: 0.38 ARI B: 0.08 PCA t-SNE Fig. 6: Result of PCA and t-SNE them to converge on a shared, accurate symbolic system (higher ARI and Kappa). The failure of the All-acceptance scenario highlights that mere information exchange is insuffi- cient; the ability to selectively reject incongruent symbols is crucial for robust social learning. B. Analysis of Latent Space Structure We next investigated whether communication fundamentally reshaped the agentsâ overall internal representations or primar- ily aligned their symbolic labels. Quantitative analysis of the latent space using TopSim (Table I) revealed no significant difference in structural similarity between the MHNG and No Communication scenarios. The PCA plots in Figure 6 visually corroborate this finding, revealing no significant differences in the global data structure across the three scenarios. However, a more fine-grained visualization using t-SNE (Figure 6) suggests a subtle but important effect. In the MHNG scenario, some clusters (e.g., the purple and gray clusters) appear more compact and distinctly formed between the two agents compared to the No Communication scenario. Taken together, the primary influence of MHNG is not on fundamentally reshaping the agentsâ latent spaces (z d ), which remain heavily grounded in their own sensory experience. IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS10 neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 56%2%12%6%5%1%9%9% 8%56%15%4%3%4%5%6% 9%25%27%7%9%4%11%8% 8%4%7%61%6%6%4%5% 15%3%5%1%57%1%10%8% 3%4%4%3%3%62%14%7% 9%5%13%5%4%12%38%13% 15%2%16%2%2%10%13%41% No Communication - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 48%4%5%0%5%2%13%22% 11%43%14%10%3%2%13%4% 9%16%37%7%6%5%9%11% 7%3%10%53%5%9%7%7% 9%7%10%5%51%4%8%7% 11%8%8%2%7%45%7%12% 9%9%10%6%11%12%29%15% 13%5%5%6%4%17%17%32% No Communication - Agent B neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 72%2%2%0%0%1%5%18% 1%76%15%2%0%1%3%2% 5%20%46%7%3%2%11%5% 7%5%3%62%3%5%10%4% 1%5%6%2%74%3%3%5% 3%1%5%2%1%69%13%8% 9%2%11%5%1%7%41%24% 24%1%7%2%1%10%12%43% MHNG - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 52%4%3%1%0%2%10%28% 7%59%19%5%4%1%2%3% 6%17%53%4%6%1%9%4% 6%1%6%61%7%6%8%5% 7%6%4%2%63%9%3%6% 5%3%4%1%4%63%14%6% 12%10%7%2%2%8%37%20% 18%3%3%1%1%11%20%42% MHNG - Agent B neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 69%7%4%3%6%3%5%3% 6%37%23%5%10%5%9%5% 9%18%27%13%10%8%11%4% 6%6%8%47%10%10%10%4% 7%9%13%6%31%9%16%9% 9%6%7%11%9%38%11%8% 12%10%10%9%16%12%16%15% 12%5%7%6%7%9%11%42% All Acceptance - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 66%11%4%2%6%2%5%3% 7%37%21%5%9%6%9%6% 11%18%27%12%11%7%11%4% 6%6%8%47%10%12%7%4% 7%10%13%7%33%8%14%9% 9%7%7%10%10%39%11%7% 12%12%11%7%15%11%16%16% 13%5%7%6%8%8%12%41% All Acceptance - Agent B (a) Heat map of Original/Original neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 50%5%9%4%4%5%12%11% 12%51%14%1%5%3%5%7% 9%19%38%3%4%6%11%11% 5%9%10%61%2%5%4%5% 9%2%10%1%63%2%7%5% 6%2%7%3%1%61%11%9% 12%4%12%0%2%13%39%19% 17%3%8%2%3%9%21%36% No Communication - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 25%12%13%9%7%9%8%17% 10%32%14%11%9%6%14%4% 7%14%39%10%8%7%11%4% 12%10%6%24%16%9%11%13% 10%8%12%8%31%14%11%6% 9%6%5%15%10%27%9%19% 9%8%12%12%9%9%28%11% 18%5%4%11%6%17%8%31% No Communication - Agent B neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 53%6%3%3%1%1%9%24% 4%64%15%7%2%3%3%3% 6%17%48%10%2%4%8%5% 10%5%10%52%2%5%13%3% 5%3%2%2%68%2%14%4% 7%2%2%1%3%64%14%5% 17%10%7%7%4%5%33%17% 18%3%8%6%2%4%11%48% MHNG - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 32%13%3%8%6%6%11%20% 11%42%20%9%6%6%4%2% 6%13%53%7%8%4%6%2% 11%7%6%38%8%7%11%12% 7%9%3%6%53%7%11%4% 11%9%4%7%4%43%13%8% 8%13%12%10%5%13%31%9% 13%6%7%9%2%8%15%40% MHNG - Agent B neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels neutral calm happy sad angry fearful disgust surprised True Labels 30%10%10%7%8%10%9%16% 10%18%14%12%13%11%13%8% 8%13%17%15%14%12%13%8% 8%12%12%24%12%11%12%8% 9%12%14%11%19%12%14%9% 11%11%12%13%11%20%11%12% 10%12%13%14%12%13%15%11% 17%8%10%11%8%12%11%24% All Acceptance - Agent A neutralcalmhappysadangryfearfuldisgustsurprised Predicted Labels 40%10%8%6%8%8%7%14% 10%21%15%10%14%11%12%8% 11%16%17%13%13%10%12%8% 9%10%11%28%13%10%12%7% 10%12%14%10%20%9%14%12% 9%10%11%12%11%22%12%12% 12%12%13%11%13%11%16%12% 16%9%10%9%9%13%10%24% All Acceptance - Agent B (b) Heat map of Original/Low arousal focus Fig. 7: The heat map uses recall to evaluate how well the model can recognize data with the same labels. Because the agents exchange only a single, low-bandwidth discrete sign (w d ) for each emotional stimulusânot the rich, high-dimensional sequence dataâit is logical that the main impact would be on the symbolic mapping rather than a fundamental reorganization of the underlying latent space. Consequently, this process establishes a shared symbolic layer upon individual, subjective experiences, facilitating mutual understanding without requiring identical internal states. C. Effect of Interoceptive Divergence We next explored the role of interoceptive similarity in the co-construction process by introducing systematic divergence in the agentsâ core affect. As a baseline, the introduction of any core affect yielded substantially better ARI and Kappa scores than experiments relying on vision and audio alone (Table I, Vision+Audio row vs. all Original-paired rows), underscoring the importance of interoception for emotion categorization. Among the four MHNG configurations involving core affect, the symmetric Original/Original condition yielded the highest inter-agent agreement (Kappa = 0.51). The most insightful results, however, emerged from the asymmetric conditions, in which the two agents have sys- tematically different interoceptive dynamics. We focus on the âOriginal / Low arousal focusâ condition, where Agent A uses the Original core affect and Agent B uses a core affect whose arousal-axis parameters (ÎŒ a ,Ξ a ,Ï a ) are attenuated to one quarter of the Original values. The recall heatmaps in Fig.7 reveal two distinct patterns of categorical reshaping under MHNG. 1) Observation 1: Categorical structure of the agent with attenuated arousal sensitivity is sharpened un- der MHNG: Comparing the No Communication and MHNG conditions for Agent B in Fig. 7b, the diagonal recall values increase for all eight emotion categories: Neutral 25% â 32%, Calmness 32% â 42%, Hap- piness 39% â 53%, Sadness 24% â 38%, Anger 31%â 53%, Fear 27%â 43%, Disgust 28%â 31%, and Surprise 31% â 40%. The largest gains occur for Anger (+22 points) and Fear (+16 points), the two categories whose RAVDESS reference labels lie at the high-arousal end of the valence-arousal space (cf. Fig. 4) and whose interoceptive signatures are therefore most affected by Agent Bâs reduced arousal sensitivity. The overall sharpening is also reflected in the corresponding ARI for Agent B in the Original/Low arousal focus row of Table I, which improves from 0.09 ± 0.01 (No Com.) to 0.20± 0.04 (MHNG). MHNG-mediated communication thus enables Agent Bâwhose interocep- IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS11 tive readout provides only weakly differentiated arousal informationâto develop categorical assignments that more closely correspond to the RAVDESS reference labels. 2) Observation 2: Categorical boundaries of the other agent are reshaped, not uniformly improved: The change in Agent Aâs recall pattern under MHNG is not a uniform improvement or degradation but a category- specific reorganisation. Comparing Agent A in Fig. 7a (symmetric Original/Original under MHNG) and Fig. 7b (asymmetric Original/Low arousal focus under MHNG), the diagonal recall values move in opposite directions for different categories: six categories show reduced diagonal recall (Neutral 72%â 53%, Calmness 76%â 64%, Sadness 62% â 52%, Anger 74% â 68%, Fear 69% â 64%, Disgust 41% â 33%), while two show small increases (Happiness 46% â 48%, Surprise 43% â 48%). The corresponding ARI for Agent A in Table I shifts modestly, from 0.41± 0.08 in Original/Original to 0.34± 0.07 in Original/Low arousal focus. Together, Observations 1 and 2 show that asymmetric in- teroception leads not to a one-directional transfer of structure from one agent to the other, but to a category-specific bidi- rectional reshaping of both agentsâ categorical systems. Even under such reshaping, the inter-agent Kappa under MHNG remains high (0.39 ± 0.04 for Original/Low arousal focus, comparable to the other asymmetric MHNG conditions in Table I), indicating that the two agents still converge on a shared symbolic system. The interpretation of these reshaping patternsâspecifically, the constructionist reading that intero- ceptive heterogeneity is a constitutive feature of emotional life rather than a deficit to be correctedâis developed in Section VI-C. VI. DISCUSSION The experimental results in Section V can be situated in three broader theoretical contexts: the dissociation between symbolic and perceptual layers in co-construction, the role of interoceptive heterogeneity in emotion formation, and the implications for cognitive developmental robotics and symbol emergence in robotics. A. MHNG operates at the symbolic, not the perceptual layer Across all conditions, MHNG-mediated communication left the structural similarity (TopSim) of the two agentsâ la- tent spaces z A d and z B d essentially unchanged (Section V-2), while substantially improving inter-agent agreement at the symbolic level w d (Kappa rose from 0.01 to 0.51 in the Original/Original condition; Table I). This dissociation is by design: for each stimulus the speaker transmits only a single integer index w d â 1,...,K (K = 9 in our experiments, i.e. log 2 9 â 3.2 bits per stimulus), not the full multimodal observation tensor (vision: 3,815-dim, audio: 20,700-dim, interoception: 690-dim, totalling â 25,205 dimensions per sample). Because the channel transmits only this minimal categorical information, MHNG can only modify the agentsâ categorical assignments and the GMM cluster parameters, not the continuous, modality-grounded latent geometry produced by the MVAE encoders. Consequently, co-construction estab- lishes a shared symbolic layer w on top of individually embod- ied, subjective continuous representations z. This is precisely the architectural property that enables mutual understanding without requiring agents to share identical internal states or perceptual structuresâa computational analogue of the long- standing observation in linguistics that words are public while meanings are private [40]. B. Co-construction does not require interoceptive isomor- phism Across the three asymmetric conditions (Original/Happy- inverse, Original/Low-valence-focus, Original/Low-arousal- focus), MHNG produced inter-agent Kappa values of 0.39â 0.49, comparable to the symmetric Original/Original baseline (Kappa = 0.51) and dramatically higher than the No Com- munication (â 0) and All Acceptance (†0.22) baselines. This computationally instantiates Gendron and Barrettâs the- oretical claim that emotion perception relies on conceptual synchrony rather than physiological mirroring [8]: even when two agents process the same stimulus through systematically different interoceptive dynamics, communication enables them to converge on a shared categorical structure. Crucially, this convergence is not enforced uniformity. The All Acceptance conditionâin which one agent uncondition- ally adopts the otherâs signsâyields worse alignment with the reference labels (ARI †0.17 in all conditions) and worse cluster quality (DBS â„ 10) than No Communication. The selectivity of the MetropolisâHastings rejection step is what allows each agent to retain category boundaries that respect its own interoceptive evidence while still aligning categorically with its partner. In the constructed-emotion framework, this maps onto the distinction between shared categorical knowl- edge and private embodied experience: two interlocutors can agree on the applicability of the word âangerâ without having identical visceral reactions to it. C. Interoceptive heterogeneity as a feature, not a defect Real human interoception is heterogeneous: individuals dif- fer in interoceptive accuracy, attention, and valence/arousal sensitivity, with such differences linked to traits ranging from anxiety to alexithymia and contemplative expertise [37], [38], [39]. The constructed-emotion framework treats this hetero- geneity as constitutive of, rather than noise around, emotional life [7]. Our results align with this constitutive view in three ways. First, an agent learning alone (No Com.) only loosely recovers the experimenter-defined emotion structure (ARIâ 0.21â0.30; Table I), suggesting that no single body-grounded model is sufficient. Second, when two agents with non-identical interoceptive profiles communicate via MHNG, their joint categorical system reaches a higher Kappa and ARI than either does aloneâsuggesting that interoceptive diversity is itself a resource for richer category formation, not an obstacle to overcome. Third, the asymmetric reshaping of recall patterns IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS12 under asymmetric conditions (e.g., the Original/Low-arousal- focus heatmap in Fig. 7b) is reminiscent of how, in everyday social interaction, interlocutors shift their emotional vocabu- lary depending on whom they are speaking withâa flexibility that is hard to explain on essentialist (e.g., basic-emotion) accounts but natural under constructionist views. It should be emphasised that the asymmetric conditions in our experiments are not intended as models of pathological interoception; rather, they probe the broader space of intero- ceptive variation that is normal in any human population. The corresponding categorical reshaping observed under MHNG should therefore be read as a model of how diverse bodies arrive at shared concepts, not as a model of âcorrectingâ a deficient agent. D. Implications for cognitive developmental robotics and sym- bol emergence in robotics From a cognitive developmental robotics perspective [41], [42], our results suggest a concrete mechanism by which an artificial agent can acquire human-aligned emotion concepts without requiring its body to faithfully replicate human phys- iology. Prior CDR work on affective development modelled emotion formation within a single agent [17], [18]; the present work extends this to the social loop, showing that two agents with different embodiments can still converge on a shared emotional vocabulary through naming-game-like interaction. Such a mechanism is a candidate building block for caregiverâ infant emotional learning models, where caregiver and infant clearly do not share identical interoceptive states. From a symbol emergence in robotics perspective [19], [21], the present work extends the Inter-GMM+MVAE frameworkâoriginally validated on physical objects [22]âto the more abstract, body-grounded domain of emotion. The suc- cess of this extension is non-trivial: emotion categories lack the stable visual/physical regularities that ground object names, and yet MHNG still recovers shared categorical structure. This suggests that the symbol emergence framework is not restricted to perceptually grounded categories but can extend to internally grounded ones, opening a path toward modelling the emergence of social, evaluative, and abstract concepts in artificial agents. E. Limitations and future work Our model represents a deliberate simplification of real emotional co-construction, and four limitations are worth flagging. a) Stimulus ecology: RAVDESS contains posed perfor- mances by professional actors, which are known to differ from spontaneous expressions in temporal dynamics and feature distribution [43]; ecological validity is thus limited. b) Synthetic interoception: Core affect is simulated via an OrnsteinâUhlenbeck process rather than measured from physiological signals such as heart-rate variability or galvanic skin response. Integrating real interoceptive measurements is a natural next step. c) Channel bandwidth: The communication channel transmits a single discrete sign per stimulus, whereas human emotional communication is continuous and multimodalâ facial expressions, prosody, and gesture all carry affective signal. Extending the present framework to multimodal sign exchange (e.g., pairs of agents exchanging facial-expression- like vectors as well as discrete category labels) is an important direction for future work, as it would bring the model closer to the rich semiotic exchange characteristic of human emotional interaction. d) Population scale: The system uses two agents only; cultural-level emotion norms emerge in populations of many interacting agents, an extension that is straightforward in prin- ciple within the CPC framework. Future work will scale the system to larger groups to investigate how shared emotional vocabularies stabilise into culture-level norms. VII. CONCLUSION This study presented a computational instantiation of the co- construction of emotion, using the Inter-GMM+MVAE frame- work grounded in Collective Predictive Coding to simulate how two embodied agents form and align emotion categories from multimodal sensory experience. Three findings stand out. (i) Selective communication via MHNG significantly improves inter-agent agreement (Kappa) and clarity (DBS) of the emerged emotion categories, while non-selective communication degrades performance. (i) Com- munication operates primarily at the symbolic layer (w d ), leaving the modality-grounded latent geometry (z d ) largely intact âthe key property that enables agents with different embodiments to share categories without sharing internal states. (i) Asymmetric interoceptive profiles do not prevent co-construction; instead, they yield distinct, category-specific reshaping patterns that are consistent with the constructed- emotion view of interoceptive heterogeneity as constitutive of emotional life. To our knowledge, this work provides the first computa- tional validation of the co-constructionist view of emotion perception, and extends the applicability of the CPC frame- work from physical objects to the abstract, socially-grounded domain of human emotion. Future work will (a) replace simulated core affect with empirically measured physiological signals, (b) extend the communication channel from a single discrete sign to multimodal signals such as facial expressions and prosody, and (c) scale beyond two agents to populations of interacting agents, enabling investigation of how culture-level emotional norms might emerge from local CPC dynamics. APPENDIX A VAE ARCHITECTURE REFERENCES [1] R. S. Lazarus, Emotion and Adaptation. Oxford University Press, 1991. [2] K. R. Scherer, âWhat are emotions? and how can they be measured?â Social Science Information, vol. 44, no. 4, p. 695â729, 2005. [3] A. R. Damasio, The Feeling of What Happens: Body and Emotion in the Making of Consciousness. Harcourt Brace, 1999. [4] J. A. Russell, âA circumplex model of affect.â Journal of personality and social psychology, vol. 39, no. 6, p. 1161, 1980. IEEE TRANSACTIONS ON COGNITIVE AND DEVELOPMENTAL SYSTEMS13 TABLE I: The parameters of each emotion used to generate core affect Emotion ÎŒ V ÎŒ A Ï V Ï A Ξ V Ξ A Neutral0.000.000.0900.0901.51.5 Calm0.80-0.500.1350.1802.11.8 Happy0.900.500.0900.2252.72.4 Sad-0.70-0.500.1800.1352.42.1 Angry-0.600.600.2250.2701.82.7 Fearful-0.800.700.2700.3151.53.0 Disgust-0.900.200.2250.2252.12.4 Surprised0.000.800.1800.3601.21.8 [5] B. Mesquita and N. H. Frijda, âCultural variations in emotions: A review,â Psychological Bulletin, vol. 112, no. 2, p. 179â204, 1992. [6] S. Kitayama and H. R. Markus, Emotion and Culture: Empirical Studies of Mutual Influence. American Psychological Association, 1994. [7] L. F. Barrett, How emotions are made: The secret life of the brain. Pan Macmillan, 2017. [8] M. Gendron and L. F. Barrett, âEmotion perception as conceptual synchrony,â Emotion Review, vol. 10, no. 2, p. 101â110, 2018. [9] A. R. Damasio, Descartesâ Error: Emotion, Reason, and the Human Brain. G. P. Putnamâs Sons, 1994. [10] P. Ekman, âAn argument for basic emotions,â Cognition and Emotion, vol. 6, no. 3-4, p. 169â200, 1992. [11] K. A. Lindquist, T. D. Wager, H. Kober, E. Bliss-Moreau, and L. F. Barrett, âThe brain basis of emotion: a meta-analytic review,â Behavioral and brain sciences, vol. 35, no. 3, p. 121â143, 2012. [12] L. F. Barrett, âAre emotions natural kinds?â Perspectives on psycholog- ical science, vol. 1, no. 1, p. 28â58, 2006. [13] J. A. Russell, âIs there universal recognition of emotion from facial ex- pression? a review of the cross-cultural studies.â Psychological bulletin, vol. 115, no. 1, p. 102, 1994. [14] M. Gendron, D. Roberson, J. M. van der Vyver, and L. F. Barrett, âPerceptions of emotion from facial expressions are not culturally universal: evidence from a remote culture.â Emotion, vol. 14, no. 2, p. 251, 2014. [15] K. Friston, T. FitzGerald, F. Rigoli, P. Schwartenbeck, and G. Pezzulo, âActive inference: a process theory,â Neural Computation, vol. 29, no. 1, p. 1â49, 2017. [16] A. K. Seth and K. J. Friston, âActive interoceptive inference and the emotional brain,â Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 371, no. 1708, p. 20160007, 2016. [17] T. Horii, Y. Nagai, and M. Asada, âModeling development of multimodal emotion perception guided by tactile dominance and perceptual improve- ment,â IEEE Transactions on Cognitive and Developmental Systems, vol. 10, no. 3, p. 762â775, 2018. [18] C. Hieida, T. Horii, and T. Nagai, âDeep emotion: A computational model of emotion using deep neural networks,â 2018. [Online]. Available: https://arxiv.org/abs/1808.08447 [19] T. Taniguchi et al., âSymbol emergence in robotics: a survey,â Advanced Robotics, vol. 30, no. 11-12, p. 706â728, 2016. [20] Y. Hagiwara, H. Kobayashi, A. Taniguchi, and T. Taniguchi, âSymbol emergence as an interpersonal multimodal categorization,â Frontiers in Robotics and AI, vol. 6, p. 134, 2019. [21] T. Taniguchi, âCollective predictive coding hypothesis: symbol emer- gence as decentralized bayesian inference,â Frontiers in Robotics and AI, vol. 11, p. 1353870, 2024. [22] N. L. Hoang, T. Taniguchi, Y. Hagiwara, and A. Taniguchi, âEmer- gent communication of multimodal deep generative models based on metropolis-hastings naming game,â Frontiers in Robotics and AI, vol. 10, 2024. [23] A. K. Seth, âInteroceptive inference, emotion, and the embodied self,â Trends in cognitive sciences, vol. 17, no. 11, p. 565â573, 2013. [24] T. Taniguchi, Y. Yoshida, Y. Matsui, N. Le Hoang, A. Taniguchi, and Y. Hagiwara, âEmergent communication through metropolis-hastings naming game with deep generative models,â Advanced Robotics, vol. 37, no. 19, p. 1266â1282, 2023. [25] Y. Hagiwara, K. Furukawa, A. Taniguchi, and T. Taniguchi, âMultiagent multimodal categorization for symbol emergence: emergent commu- nication via interpersonal cross-modal inference,â Advanced Robotics, vol. 36, no. 5-6, p. 239â260, 2022. [26] K. Sakurai, H. Uenoyama, A. Taniguchi, and T. Taniguchi, âMh- mug: Collaborative music generation game between ai agents towards emergent musical creativity,â IEEE Access, 2026. [27] M. Wu and N. Goodman, âMultimodal generative models for scalable weakly-supervised learning,â Advances in neural information processing systems, vol. 31, 2018. [28] Y. Shi, B. Paige, P. Torr, et al., âVariational mixture-of-experts autoen- coders for multi-modal deep generative models,â Advances in neural information processing systems, vol. 32, 2019. [29] T. M. Sutter, I. Daunhawer, and J. E. Vogt, âGeneralized multimodal elbo,â arXiv preprint arXiv:2105.02470, 2021. [30] L. Hubert and P. Arabie, âComparing partitions,â Journal of classifica- tion, vol. 2, p. 193â218, 1985. [31] J. Cohen, âA coefficient of agreement for nominal scales,â Educational and psychological measurement, vol. 20, no. 1, p. 37â46, 1960. [32] L. Van der Maaten and G. Hinton, âVisualizing data using t-sne.â Journal of machine learning research, vol. 9, no. 11, 2008. [33] N. Kriegeskorte, M. Mur, and P. A. Bandettini, âRepresentational similarity analysis-connecting the branches of systems neuroscience,â Frontiers in systems neuroscience, vol. 2, p. 249, 2008. [34] D. L. Davies and D. W. Bouldin, âA cluster separation measure,â IEEE transactions on pattern analysis and machine intelligence, no. 2, p. 224â227, 2009. [35] S. R. Livingstone and F. A. Russo, âThe ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,â PloS one, vol. 13, no. 5, p. e0196391, 2018. [36] B. Tadas, Z. Amir, L. Y. Chong, and M. Louis-Philippe, âOpenface 2.0: Facial behavior analysis toolkit,â in 13th IEEE International Conference on Automatic Face & Gesture Recognition, 2018. [37] Y. Terasawa, H. Fukushima, and S. Umeda, âHow does interoceptive awareness interact with the subjective experience of emotion? an fmri study,â Human Brain Mapping, vol. 34, no. 3, p. 598â612, 2013. [38] R. Brewer, R. Cook, and G. Bird, âAlexithymia: a general deficit of interoception,â Royal Society Open Science, vol. 3, no. 10, p. 150664, 2016. [39] J. Murphy, R. Brewer, C. Catmur, and G. Bird, âInteroception and psychopathology: A developmental neuroscience perspective,â Develop- mental Cognitive Neuroscience, vol. 23, p. 45â56, 2017. [40] W. V. O. Quine, Word and Object. MIT Press, 1960. [41] M. Asada, âTowards artificial empathy,â International Journal of Social Robotics, vol. 7, no. 1, p. 19â33, 2015. [42] â, âModeling early vocal development through infantâcaregiver in- teraction,â IEEE Transactions on Cognitive and Developmental Systems, vol. 8, no. 2, p. 128â138, 2016. [43] M. G. Calvo and L. Nummenmaa, âPerceptual and affective mechanisms in facial expression recognition: An integrative review,â Cognition and Emotion, vol. 30, no. 6, p. 1081â1106, 2016.