Paper deep dive
Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition
Jing Jie Tan, Ban-Hoe Kwan, Danny Wee-Kiat Ng, Yan-Chai Hum, Shih-Yu Lo, Po-An Chen, Noriyuki Kawarazaki, Kosuke Takano, Anissa Mokraoui
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/10/2026, 5:30:09 AM
Summary
The paper introduces JAM, a theory-agnostic framework for personality recognition that replaces predefined psychological taxonomies with unified latent pseudo-facets. It leverages an Attention-Pooled Graph Prototypical Network, Cross-Theory Harmonization (combining human-guided linkage and machine-induced consensus), and an LLM-as-a-Judge mechanism to enhance robustness, data quality, and cross-framework generalization across heterogeneous personality datasets.
Entities (14)
Relation Signals (13)
JAM → employs → LLM-as-a-Judge
confidence 95% · To further improve robustness and data quality, we incorporate an LLM-as-a-Judge mechanism operating in two configurations
JAM → implements → Personality Recognition
confidence 95% · In this work, we introduce JAM... a theory-agnostic framework that shifts learning from adapting to predefined personality theories toward discovering unified latent pseudo-facets
JAM → utilizes → Attention-Pooled Graph Prototypical Network
confidence 92% · JAM achieves this through an Attention-Pooled Graph Prototypical Network that learns structured representations via clustering in embedding space
JAM → incorporates → Cross-Theory Harmonization
confidence 90% · together with a Cross-Theory Harmonization (CTH) approach that integrates (i) Human-Guided Linkage and (ii) Machine-Induced Consensus to unify heterogeneous datasets
Big-5 → isa → Personality Model
confidence 90% · Most approaches are built around specific psychological theories, such as the Big-5 or MBTI, limiting their ability to generalize across datasets and cultural contexts.
MBTI → isa → Personality Model
confidence 90% · Most approaches are built around specific psychological theories, such as the Big-5 or MBTI, limiting their ability to generalize across datasets and cultural contexts.
Cross-Theory Harmonization → integrates → Human-Guided Linkage
confidence 88% · integrates (i) Human-Guided Linkage and (ii) Machine-Induced Consensus to unify heterogeneous datasets without relying on predefined labels.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structure. This limits generalization, as personality itself is better understood as theory-invariant, while existing annotations reflect only partial and sometimes inconsistent views of the same latent traits. In this work, we introduce JAM ((J)udge for (A)daptive (M)etric-Alignment), a theory-agnostic framework that shifts learning from adapting to predefined personality theories toward discovering unified latent pseudo-facets that capture shared psychological structure. Rather than constraining the model to any personality taxonomy during training or inference, the framework learns generalizable psychological representations and can infer an individual's latent psychological profile directly from the textual samples, without requiring theory-specific labels. JAM achieves this through an Attention-Pooled Graph Prototypical Network that learns structured representations via clustering in embedding space, together with a Cross-Theory Harmonization (CTH) approach that integrates (i) Human-Guided Linkage and (ii) Machine-Induced Consensus to unify heterogeneous datasets without relying on predefined labels. To further improve robustness and data quality, we incorporate an LLM-as-a-Judge mechanism operating in two configurations, (i) LLM-before-the-loop and (ii) LLM-in-the-loop which identifies ambiguous samples to guide adaptive metric learning. Experiments show that JAM improves cross-framework generalization and performance, establishing a strong step toward theory-agnostic personality inference and supporting low-resource personality theories. The related code repository, model weights, and artifacts are available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2607.08374v1
- Canonical: https://arxiv.org/abs/2607.08374v1
Trouble viewing inline? Open PDF directly →
Full Text
99,904 characters extracted from source content.
Expand or collapse full text
1 Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition Jing Jie Tan ∗ , Ban-Hoe Kwan ∗ , Danny Wee-Kiat Ng ∗ , Yan-Chai Hum ∗ , Shih-Yu Lo † , Po-An Chen ‡ , Noriyuki Kawarazaki § , Kosuke Takano § , Anissa Mokraoui ¶ ∗ Department of Mechatronics and Biomedical Engineering, Lee Kong Chian Faculty of Engineering and Science, Universiti Tunku Abdul Rahman, Malaysia † Institute of Communication Studies, National Yang Ming Chiao Tung University, Taiwan ‡ Institute of Information Management, National Yang Ming Chiao Tung University, Taiwan § Faculty of Information Technology, Kanagawa Institute of Technology, Japan ¶ Laboratoire de Traitement et Transport de l’Information, Université Sorbonne Paris Nord, France Email: tanjingjie@1utar.my, kwanbh, ngwk, humyc@utar.edu.my,shihyulo, poanchen@nycu.edu.tw,kawara@rm, takano@ic.kanagawa-it.ac.jp, anissa.mokraoui@univ-paris13.fr Abstract—Personality recognition has traditionally been con- strained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structure. This lim- its generalization, as personality itself is better understood as theory-invariant, emerging from stable psychological patterns that should manifest consistently across different frameworks, while existing annotations reflect only partial and sometimes inconsistent views of the same latent traits. In this work, we introduce JAM ((J)udge for (A)daptive (M)etric-Alignment), a theory-agnostic framework that shifts learning from adapting to predefined personality theories toward discovering unified latent “pseudo-facets” that capture shared psychological structure. Rather than constraining the model to any personality taxonomy during training or inference, the framework learns generalizable psychological representations and can infer an individual’s latent psychological profile directly from the textual samples, with- out requiring theory-specific labels. JAM achieves this through an Attention-Pooled Graph Prototypical Network that learns structured representations via clustering in embedding space, together with a Cross-Theory Harmonization (CTH) approach that integrates (i) Human-Guided Linkage and (i) Machine- Induced Consensus to unify heterogeneous datasets without relying on predefined labels. To further improve robustness and data quality, we incorporate an LLM-as-a-Judge mechanism operating in two configurations, (i) LLM-before-the-loop and (i) LLM-in-the-loop which identifies ambiguous, mislabeled, and boundary samples to guide adaptive metric learning. Exper- iments on Essays and Kaggle personality datasets show that JAM improves cross-framework generalization and performance, establishing a strong step toward theory-agnostic personality inference and supporting low-resource personality theories. The related code repository, model weights, and artifacts are available at https://research.jingjietan.com/JAM. Index Terms—Personality Classification, Large Language Models (LLMs), N-Shot Prompting, Prototypical Networks, Fine- Tuning, Natural Language Understanding I. INTRODUCTION Personality recognition has become increasingly important, especially in recommendation systems [1]. By understanding user personalities, these systems can provide personalized suggestions, enhancing user satisfaction and trust. Under- standing user personality is crucial for delivering a superior user experience, making this an important area of study [2]. By tailoring interactions and recommendations to individual personality traits, AI systems and robots can achieve higher levels of personalization, leading to increased user satisfaction and trust [3]. Furthermore, incorporating personality traits into recommendation algorithms can help address issues such as the cold start and data sparsity problems, resulting in more ac- curate and user-centric recommendations [4], [5]. Additionally, this approach can contribute to explainable AI by providing clearer explanations for the recommendations made. Despite these advancements, current personality-recognition models remain constrained by several fundamental limitations. Most approaches are built around specific psychological theo- ries, such as the Big-5 or MBTI, limiting their ability to gener- alize across datasets and cultural contexts. This dependence is further exacerbated by the scarcity of large-scale annotated data. The widely used myPersonality dataset [6], which originally contained data from millions of users, was discon- tinued due to privacy concerns, leaving only relatively small public datasets. Moreover, existing methods treat personality recognition as a static classification task, without leveraging psychological insights, limiting their generalizability. These limitations motivate a theory-agnostic framework that eliminates dependence on predefined personality tax- onomies during both training and inference. Instead of learn- ing theory-specific representations, the model discovers latent pseudo-facets that capture underlying psychological structure across heterogeneous data sources. At inference time, the model directly infers an individual’s latent psychological profile from behavioral or textual samples without requiring theory-specific labels or adherence to a particular personality framework. This design enables the integration of heteroge- neous datasets while improving generalization across existing arXiv:2607.08374v1 [cs.CL] 9 Jul 2026 2 personality theories. A. Contributions In this study, we make the following key contributions: 1) We incorporate a psychology-based methodology for in- tegrating datasets constructed under different personality theories and generalize the model through human-guided linkage, supporting low-resource personality theory. 2) We enhance the prototypical network by incorporating machine-induced consensus for pseudo-facet construc- tion, improving cross-theory harmonization and mitigat- ing class imbalance, leading to improved performance. 3) We investigate the LLM-as-a-Judge mechanism, assessing its comparative effectiveness when applied during the pre- learning stage (LLM-before-the-loop) versus within the learning process (LLM-in-the-loop) for filtering ambigu- ous or noisy personality-related text samples. I. LITERATURE REVIEW A. Personality Theory Recent studies have demonstrated that humans express their personality through language use, providing a rich source of data for personality recognition [7], [8], [9], [10]. The rise of social media platforms has provided a rich source of data to understand how particular individuals often leave behind a personality footprint through their online activities [7]. This creates a significant opportunity to extract personality informa- tion from such indirect background data, which is particularly beneficial for initializing personality-aware functionalities in new robots or applications. There are two primary types of assessments to measure, quantify, and classify personality traits [11], [12]: • Self-report tests: These tests require test-takers to under- stand the provided statements and evaluate how well they describe themselves. The results can be quantitatively standardized, ensuring high reliability and validity [13]. • Projective tests: These tests involve asking test-takers to provide their interpretations of scenes, scenarios, or objects [14]. This type of test considers various aspects, such as tone, message, or body language, making it more capable of addressing potential issues like misinterpreta- tion of questions or dishonesty from the test-taker [15]. Various personality theories have been proposed to model personality using distinct sets of dimensions [16], [17]. For instance, the Myers–Briggs Type Indicator (MBTI), sometimes described in relation to 4 dichotomies, categorizes personality across paired preferences [18], [19]. The Big-5 framework (OCEAN) defines five broad traits: Openness to Experience, Conscientiousness, Extraversion, Agreeableness, and Neuroti- cism [20], [21]. The HEXACO model extends this structure by adding Honesty–Humility to Emotionality, Extraversion, Agreeableness, Conscientiousness, and Openness [22], [23]. Table I summarizes the correspondence among these three major frameworks, offering a general comparison of their dimensions. Basically, facets are commonly regarded as the lower-level constructs or descriptions that collectively define a broader personality dimension or trait. They provide a more granular representation of personality by capturing specific be- havioral and psychological tendencies within each dimension. The terminology can be summarized as in Fig. 1. Theory/ Model The model is built based on a personality theory which analyses a person through an assessment which is a test, questionnaire, or survey. Each of the personality types in the model is contributed by all dimensions of the model. The number of personality types is based on the number of classes in the model, commonly formulated as: Dimension/ Factor/ Class The number of dimensions is determined by the respective personality theory/model used. Traits Each dimension has either: Facets/ Description The facet is the description of respective traits Single trait evaluate using a level, score or scale Paired trait evaluate by categorising into 2 polar opposites Fig. 1. Overview of the relationship, terminology, and theoretical structure of personality assessment models, illustrating how theories define dimensions, traits, facets, and personality types. These mappings should be interpreted cautiously: they re- flect partial conceptual overlap rather than equivalence. For example, Big-5 Openness includes facets such as curiosity and preference for novelty, whereas MBTI’s Sensing–Intuition dimension distinguishes concrete, detail-oriented perception from abstract, future-oriented thinking. Similarly, related con- structs are distributed differently across frameworks; emotional reactivity in Big-5 Neuroticism only loosely corresponds to the MBTI Thinking–Feeling dimension, which primarily reflects decision-making style rather than affective stability. This highlights a limitation: personality models are human-constructed abstractions rather than directly ob- servable ground truth categories. Consequently, their use in machine learning can introduce systematic noise and incon- sistency, especially at the facet level where definitions differ across studies. In practice, available datasets seldom provide sufficiently fine-grained and consistently annotated signals to support reliable bottom-up learning of stable psychological components, not only due to annotation cost but also because personality is inherently composite and context-dependent. To our knowledge, no prior work has directly addressed this mis- match between psychologically defined trait hierarchies and their operationalization in data-driven settings, or proposed effective strategies to mitigate the resulting label ambiguity. B. Related Algorithms Personality recognition has evolved alongside advances in natural language processing and representation learning [24]. Early research primarily focused on developing task-specific models for predicting personality labels defined by a particular framework, such as the Big-5 or MBTI. As language mod- els became increasingly capable of capturing semantic and contextual information, researchers began adopting pretrained Transformer architectures as general-purpose representations for personality-related tasks. More recently, large language models have enabled prompting-based inference, reducing reliance on framework-specific training procedures and ex- panding the possibility of transferring personality knowledge across different theoretical formulations. 3 TABLE I MAPPING THE DIMENSIONS OF THREE PROMINENT PERSONALITY FRAMEWORKS: BIG-5, MBTI, AND HEXACO. SymbolBig-5MBTIHEXACOGeneral Description OOpenness to Experience (O)Sensors (S)-Intuitive (N)Openness (O)Creativity, curiosity, novelty CConscientiousness (C)Perceivers (P)-Judgers (J)Conscientiousness (C)Organise, responsibility, discipline EExtraversion (E)Introverts (I)-Extroverts (E)Extraversion (X)Sociability, energy, assertiveness AAgreeableness (A)Thinkers (T)-Feelers (F)Agreeableness (A)Empathy, kindness, cooperativeness NNeuroticism (N)-Emotionality (E)Sensitivity, anxiety, hostility H--Honesty-Humility (H)Sincerity, fairness, modesty Note: The correlations presented are intended as reference points, aligning these theories to the widely recognized Big-5 framework. This alignment supports evaluating model performance across diverse personality theories, fostering comparative analysis and deeper insights into their unique characteristics. To contextualize these developments, this section first re- views the language-model foundations underlying modern personality recognition, including (i) sentence transformer language models (encoder-only models) and (i) autoregressive transformer large language models (decoder-only models). It then surveys (i) prior personality recognition methods built upon these representations, followed by (iv) recent generative approaches that increasingly support flexible personality in- ference. Together, these developments provide the foundation for investigating personality recognition beyond a single pre- defined personality theory. 1) Sentence Transformer Language Model (Encoder-only Model): These models are trained on large corpora of text data, enabling them to capture complex linguistic patterns and relationships [25], [26], [27], [28], [29]. The produced embeddings represent the semantic meaning of the input text, facilitating downstream tasks such as text classification, clus- tering, and similarity analysis. MiniLM [25] compresses large Transformer models through deep self-attention distillation, where a small student model mimics the self-attention of the teacher’s last layer. It retains over 99% accuracy on SQuAD 2.0 and GLUE benchmarks while reducing Transformer pa- rameters and computation by 50%. Moreover, MPNet [26] integrates permuted language modeling (PLM) and auxiliary position information, improving dependency modeling and reducing position discrepancy. It surpasses BERT, XLNet, and RoBERTa across GLUE and SQuAD. Furthermore, Sentence- T5 [27] explores sentence embeddings from T5 models, intro- ducing an encoder-only approach that outperforms Sentence- BERT and SimCSE in STS tasks. Scaling T5 to billions of parameters further enhances performance, setting new state-of- the-art results for sentence embeddings. Given the efficiency and strong semantic representation capabilities of these mod- els, we select a relatively lightweight yet effective encoder- only model tailored for the personality domain, ensuring optimal performance while balancing computational efficiency. 2) Autoregressive Transformer Large Language Model (Decoder-only Model): Generative Pretrained Transformer 3 (GPT-3) was the first to demonstrate effectiveness in Nat- ural Language Understanding (NLU) through its few-shot learning capability, facilitated by techniques such as Chain- of-Thought (CoT) prompting. This family of decoder-only models enhances reasoning, coherence, and adaptability in tasks such as emotion detection, sentiment analysis, and per- sonality recognition [30], [31], [32]. The open-source LLaMA model introduced a family of models with varying parameter sizes, achieving competitive performance while maintaining efficiency in training and inference [33]. Following this, Mis- tral AI released the Mistral 7B model, which, despite its relatively smaller size, demonstrated remarkable capabilities in code generation, mathematics, and reasoning tasks, benefiting from an extended context window of 128k tokens [34]. Qwen models have been proposed as strong open-weight alternatives, exhibiting robust performance across reasoning, multilingual understanding tasks [35]. OpenAI’s GPT-4 further pushed the boundaries by introducing multimodal capabilities, allowing the model to process both text and images, thereby enhancing its applicability across diverse domains [36]. Later, reasoning models further extended their capabilities in scientific and mathematical contributions. OpenAI’s o1 model introduced a novel approach by generating extended chains of thought before arriving at a final answer, thereby improving the model’s reasoning depth and accuracy [37]. DeepSeek’s R1 model emerged as a notable open-source contribution, achieving performance comparable to OpenAI’s o1 model across mathematics, coding, and reasoning tasks [38] with a relatively low computational complexity. The R1 model employs a unique training methodology that emphasizes rein- forcement learning to enhance its reasoning capabilities. These developments underscore a growing trend toward integrating advanced reasoning processes within large language models, aiming to improve their problem-solving abilities and decision- making processes. 3) Prior Works for Personality Recognition: Prior research has consistently highlighted the importance of large, high- quality datasets for training machine learning models that generalize effectively [39], [40]. However, in privacy-sensitive domains, the development and utilization of such datasets are often constrained by regulatory and ethical considerations. In contrast, other research areas are better positioned to achieve scale by aggregating or integrating existing public datasets. A frequently cited example is the myPersonality Project [6], once a prominent open dataset for computational personality research, which was discontinued in 2018 due to increasing challenges related to regulatory compliance, controlled data access, and governance obligations. In personality recognition, we are limited by personality theory; hence, few researchers are tackling algorithmic improvements, including architecture, feature exploration, graph neural networks, loss reweighted contributions, etc. 4 Building on the importance of data and representation, prior work has explored a range of modeling approaches for personality classification from text. Early advances include BiLSTM-based models [41], which leverage bidirectional con- text to outperform traditional machine learning methods such as SVM, RF, and DT, highlighting the role of sequential modeling and hyperparameter tuning. To further enhance representation learning, [42] proposes Personality2vec, which integrates semantic, linguistic, and structural information from social networks through biased random walks and skip-gram modeling, demonstrating robustness particularly in data-scarce settings. After that, [40] further explored the dataset splitting and proposed a stratification feature, as well as focal loss for reweighting the contribution in the neural network, and analyzed the evaluation metrices. Subsequent studies have focused on enriching deep contex- tual representations with complementary linguistic and struc- tural information. For example, a hybrid Transformer–BLSTM framework [43] integrates psycholinguistic features with at- tention mechanisms to improve both predictive performance and interpretability. Graph-based methods have also emerged as an effective direction. KGrAt-Net [44] leverages knowledge graph attention over DBpedia entities to model richer semantic relationships, while TranSentGAT [45] combines BERT with sentiment knowledge and graph attention to enhance contex- tual representations and improve personality prediction. Nonetheless, these advances in modeling and feature in- tegration do not resolve the core bottleneck of limited and fragmented data. Performance in personality recognition remains constrained by dataset scale, diversity, and ecological validity under strict privacy and governance requirements. From this perspective, expanding datasets in a cost-efficient and flexible manner is essential to overcoming current theo- retical and empirical limits. However, collecting such data is often difficult, particularly for low-resource settings, due to time, cost, and regulatory constraints, while cross-institutional approaches such as federated learning remain limited in practice (to the best of our knowledge), partly because of heterogeneous theoretical frameworks and incompatible data assumptions. These challenges motivate a shift toward more theory-agnostic approaches. 4) Generative Approaches to Personality Recognition: Recent work on generative AI for personality inference has explored both data-centric and architecture-centric improve- ments. Early studies focus on data augmentation and hetero- geneous graph-based models for dialogue understanding. For instance, Wu et al. [46] introduce personality trait interpolation to generate synthetic training data and propose HC-GNN to capture both inter- and intra-speaker dependencies, improv- ing conversational personality recognition. Similarly, Semi- PerGCN [47] adopts a semi-supervised framework with multi- view graph augmentation and heterogeneous graph construc- tion, addressing limited labeled data and enhancing robustness in low-resource settings. Collectively, these approaches reflect a transition from sequential modeling toward graph-based and hybrid architectures, with increasing emphasis on representa- tion richness and data efficiency. More recent studies have investigated large language mod- els as direct inference engines for personality recognition. ChatGPT has been shown to exhibit strong zero-shot ca- pability and partial interpretability in personality prediction tasks [48]. Related work integrates emotional knowledge with structured prompting strategies for trait inference [32], while PICEPR [49] not only proposes embedding-based knowledge elicitation, but it also demonstrates that modular prompting can improve classification performance. Compared with traditional approaches that require task- specific feature engineering or model training, these methods leverage in-context learning, enabling models to infer person- ality labels directly from instructions, label descriptions, or a small number of exemplars. Consequently, prompting-based approaches move toward a more theory-agnostic paradigm, as the same pretrained model can be adapted to different personality frameworks by modifying the prompt and label definitions rather than redesigning or retraining the prediction architecture. While the provided labels may still originate from a specific theory, the underlying inference mechanism is not inherently bound to a specific personality theory. Through in-context learning, the same model can perform personality inference in few-shot settings using a small number of labeled examples or in zero-shot settings using only trait descriptions and task instructions, without requiring model retraining. However, these approaches also introduce important lim- itations. First, decoder-only large language models rely on inference-time reasoning, which is computationally expen- sive and often impractical for large-scale deployment. Second, they do not produce explicit intermediate representations of psychological states, limiting their usefulness for downstream modeling of structured cognitive or behavioral processes. More critically, since these models are trained on internet-scale corpora, there is a non-trivial risk of implicit exposure to similar evaluation data, weakening the assumption of strict zero-shot generalization and introducing potential dataset leakage concerns. These issues suggest that encoder-based models remain necessary for stable and controllable represen- tation learning in personality recognition systems. I. METHODOLOGIES Fig. 2 illustrates the proposed JAM architecture, which targets tailoring generalization across datasets annotated un- der different psychological theories, thus achieving a theory- agnostic model. A. Datasets We utilized Tan’s train-validation-test split algorithm for the standard Essays and Kaggle datasets [40]. This stratified algorithm ensures an even distribution of personality traits across all dimensions within each split, promoting fairness in comparison and validating the effectiveness of our model. Table I summarizes the datasets used in this study. Note that, due to the nature of the JAM’s prototypical learning, the training sets are resampled into support and query sets for each episode in training; while evaluating, the existing validation set will serve as the support set, and the existing test set will be the query set. 5 Embedding Encoder Layer Encoder Layer Encoder Layer Encoder Layer Embedding Embedding Embedding ... Attention-Pooled Graph Neural Network Language Embeddings Backbone Large Language Model Embedding Model Architecture Potential Problematic Dataset Fully Connected Layer Support & Query set Contrastive Loss B. LLM-in-the-loop A. LLM-before-the-loop Attention Pooling Meta-Training Processed & Format Ready Integrated Dataset Large Language Model Support set I can't wait until friday because I am... O=1C=0E=1A=0N=1 Possible/ Impossible/ Ambiguous? Decision Possible/ Impossible/ Ambiguous? I can't wait until friday because I am... O=1C=0E=1A=0 A.1 A.2 A.3 B.1 B.2 B.3 remove/ remain sent to Dataset FC Cross Entropy Loss Personality Kaggle Dataset Meta-Testing Essays Dataset Testing Few-shot Learning Training Kaggle Dataset Essays Dataset Decision N=0 FC Cross Entropy Loss sent to Human-Guided Linkage O ↔ S/N CP/J EI/E AT/F N - Truth Label Query set Adapter Layer Machine-Induced Consensus 3 33 2 1 1 1 c c|d d 5 2|5 25 52 64 a b e Big-4 (MBTI)Big-5 (OCEAN)Shared Yes Continue No Support set Proposed Cross-Theory Harmonization Collect Fig. 2. Overview of the proposed JAM architecture for theory-agnostic personality recognition. The framework integrates a language embedding backbone with an attention-pooled graph prototypical network to learn representations that capture latent pseudo-facets through clustering in the embedding space. A Cross- Theory Harmonization module is introduced to bridge heterogeneous psychological annotations and consists of Human-Guided Linkage and Machine-Induced Consensus for constructing consistent pseudo-facet structures. Under this harmonization process, Essays and Kaggle datasets are organized into support–query pairs for prototypical learning. Two training pipelines are explored: (A) LLM-before-the-loop, where an LLM filters and refines data prior to training, and (B) LLM-in-the-loop, where the LLM dynamically assesses sample quality during training to identify ambiguous, mislabeled, or boundary cases. The resulting embeddings are attention-pooled and passed through a projection layer. The model is optimized using a meta-learning objective with contrastive loss, followed by a classification head trained with cross-entropy, enabling the formation of machine-induced pseudo-facets. As in Eq. 1, we treated the dataset,D standard , as a multi-label (or multi-task) classification problem instead of a multi-class classification problem, where: n is the total number of samples in the dataset, x represents the sample text, and y represents the labels (the number of labels depends on the dimensions of the personality theory, d =d 1 ,d 2 ,...,d |d| ) associated with x. This is because personality theory treats each dimension as independent from the others [52], [53]. However, certain research argues that there are interdependencies or correlations between personality dimensions, suggesting that traits might not be entirely independent but could interact in complex ways [16], [53]. Nevertheless, we adopt a multilabel classification approach to ensure that the model outputs a probability dis- tribution over the dimensions. This approach avoids framing the task as a binary classification problem, instead allowing the intermediate layers of the neural network to automatically capture potential correlations between the labels via gradient backpropagation. In addition to the labels y, we include a vector of sample-specific confidence z, where each z i adjusts the influence of the corresponding label y i during training. D standard = (x,y,z) n x = sample text y = (y 0 ,y 1 ,...,y |d|−1 ), z = (z 0 ,z 1 ,...,z |d|−1 ), y i ∈0, 1, z i ∈R ∀i (1) B. Proposed Algorithms As illustrated in Fig. 2, the JAM framework comprises three major components designed to achieve the overall ob- jectives. We design an Attention-Pooled Graph Prototypi- cal Network to learn representations that capture underlying pseudo-facets through clustering in the embedding space. JAM incorporates a Cross-Theory Harmonization module, which includes Human-Guided Linkage and Machine-Induced Consensus, enabling the model to derive pseudo-facets from 6 TABLE I DATASETS AND THEIR DESCRIPTIONS WITH THE NUMBER OF SAMPLES IN TRAIN, VALIDATION, AND TEST SPLITS DatasetTheorySplitting DistributionDescription Essays (Public) [50] Big-51578, 395, 494 (Train, Validation, Test) This dataset was collected in a controlled environment where volunteers were instructed to write down whatever came to mind over a 20-minute period. It includes both self-reported ratings (via questionnaire) and projective ratings (by 18 observers). Kaggle (Public) [51] MBTI5552, 1388, 1735 (Train, Validation, Test) This dataset comprises data crawled, labeled, and filtered from PersonalityCafe, an online community forum where users share self-reported personality test results. learned representations rather than relying solely on prede- fined human annotations, as discussed in Section I-A. We further investigate an LLM-as-a-Judge mechanism operating under two configurations: LLM-before-the-loop and LLM-in- the-loop. These configurations differ in the stage and manner in which the LLM influences training. The mechanism is incorporated to assess training sample quality and identify hard examples, including ambiguous, mislabeled, or boundary- adjacent instances, thereby guiding more focused learning. 1) Attention-Pooled Graph Prototypical Network:We adopt the Longformer model as the language embedding backbone due to its ability to efficiently process long textual sequences, which is particularly important for personality recognition tasks involving extensive user-generated content. Longformer employs a combination of local and global atten- tion mechanisms, enabling scalable encoding of long docu- ments while maintaining computational efficiency [54]. Given an input text x, we first tokenize it and feed it into the Longformer encoder to obtain contextualized hidden represen- tations. The encoder produces layer-wise hidden states across all transformer layers. From these, we derive a set of L pooled representations, where each node corresponds to a layer-wise pooled representation of the input inR v obtained via pooling over token embeddings within that layer. The resulting node feature matrix is defined in Eq. 2. H = Longformer(x), H∈R L×v (2) Later, the output H is used as input to the graph neural network, which refines the representations by modeling inter- actions among embedding vectors derived from the language encoder. We construct a normalized adjacency matrix ̃ A from an adjacency matrix A that defines the relationships between nodes. In this work, we adopt a fully connected weighted graph, where each node is connected to all other nodes, to enable unrestricted and symmetric information exchange across all representation nodes. The number of nodes in the graph neural network corresponds to the number of layers (L) in the selected encoder, and the adjacency matrix remains shared across all layers. The node aggregation process is illustrated in Eq. 3. ̃ A = D −1/2 AD −1/2 , D i = X j A ij , A ij = 1, ∀i̸= j, A i = 0, ̃ A∈R L×L . (3) This design treats all layer-wise representations as mutually interacting components without imposing predefined hierar- chical or locality constraints, thereby allowing the model to learn how information should be integrated across different abstraction levels in a data-driven manner. While this choice provides a simple and uniform mechanism for cross-layer fusion, it does not explicitly encode heterogeneous or sparse inter-layer dependencies. Investigating adaptive or learned graph structures that could more finely capture layer- specific relationships is therefore left as a promising direction for future work. Through matrix multiplication with the normalized adja- cency matrix ̃ A, information is aggregated from neighboring nodes to update node embeddings. This propagation is per- formed iteratively across GNN layers, where H (k) denotes the node representations at the k-th layer. Eq. 4 illustrates the node update process. Here, σ(·) denotes the LeakyReLU activation function with negative slope coefficient α = 0.01, and W (k) is the trainable weight matrix at the k-th layer. H (k) = σ ̃ AH (k−1) W (k) , σ(x) = max(αx,x), W (k) ∈R v×v , H (k) ∈R L×v . (4) Lastly, we apply attention-based pooling to obtain the final graph representation, denoted as h graph , as shown in Eq. 5. The attention mechanism computes a scalar importance score for each node based on its interaction with a trainable query vector q. These scores are normalized via a softmax function to obtain attention weights α i , which determine the contribution of each node to the final representation. This operation is fully differentiable, allowing gradients to propagate not only to the query vector q, but also to each dimension of the node embeddings H (k) i . This enables the model to jointly learn which nodes are important and how the inferred embeddings should be adjusted to optimize personality trait prediction. h graph = L X i=1 α i H (k) i ,α i = exp H (k) i q P L j=1 exp H (k) j q ,q ∈R v (5) 2) Cross-Theory Harmonization (CTH): In order to achieve theory-agnostic modeling, we aim to ensure that the model captures the core personality features embedded within the text. Given the presence of personality theories across different datasets, this setting provides an opportunity for the model to learn not only surface-level patterns but also to structure the representation space in a more disentangled manner. Fig. 3 illustrates the overall idea behind this motivation. Under a standard cross-entropy (CE) formulation, different personality theories are treated as independent labels. As a result, the model tends to prioritize dominant and easily sepa- rable signals during optimization, while underutilizing or com- pletely ignoring subtler but potentially shared representational 7 Big-4 (MBTI)Big-5 (OCEAN)Shared a 5 6 5 1|3 3|2 2|1 4 b|4 e c d (a) Regular CE Big-4 (MBTI)Big-5 (OCEAN)Shared a5 6 5 1|3 3|2 2|1 b|4 4 e c d (b) Weighted CE Big-4 (MBTI)Big-5 (OCEAN)Shared 1 a 2 5 3 3 2 1 6 c|d 5 1|3 3|2 2|1 b|4 4 4 b e c d (c) Prototypical Finetuning (PF) 1 a 2 5 3 3 3 2 1 2|5 5 64b e c|d Big-4 (MBTI)Big-5 (OCEAN)Shared (d) PF + HGL Big-4 (MBTI)Big-5 (OCEAN)Shared 1 a 2 5 c 32 1 c|d 1 3 d 5 64b e (e) PF + MIC 3 33 2 111 c c|d d 5 2|5 25 52 64 a b e Big-4 (MBTI)Big-5 (OCEAN)Shared (f) PF + HGL + MIC (Full CTH) Fig. 3.The conceptual schematic visualization comparing the effects of the baseline (prior) algorithm, the proposed CTH algorithm, and its ablation variants. Each particle represents a pseudo-facet that belongs to the respective cluster of the personality theory: red and blue correspond to the respective personality theories used during inference. The numbers inside each particle represent shared features that can potentially be learned and aligned. Identical numbers (e.g., 1, 2, 3, . . . ) indicate the same feature across theories, while the numerical values themselves are only used as indices and do not carry intrinsic meaning. Alphabetic labels denote features unique to each personality theory; letters (a, b, c, . . . ) are also meaningless sequential identifiers. Purple particles indicate shared features that have been successfully learned and aligned, while grey particles represent incorrectly aligned features. The values are presented by ·|· ⃝, where the left side corresponds to the Big-5 representation and the right side corresponds to the MBTI representation. facets across theories. Although weighted CE (eg., focal loss) can rebalance gradient contributions and improve sensitivity to harder samples, these approaches primarily amplify already- learned discriminative cues rather than fundamentally restruc- turing the representation space. Consequently, the underlying bias toward dominant features remains largely unchanged. To address this limitation, we adopt prototypical fine- tuning (PF). The class prototypes are computed from the support set D (S) , where each prototype s i is defined as the (weighted) centroid of embeddings belonging to a given class. As shown in Eq. 6, each prototype is obtained as a weighted mean of support embeddings, where sample weights are denoted by z (set to z = 1 in our current setting). The encoder f φ (x) maps each input x into an embedding space parameterized by φ, and classification is performed by comparing query embeddings against class prototypes in this space. In particular, the model minimizes the distance between a query embedding and the prototype of its ground-truth class while maximizing its distance to all other prototypes. A softmax over negative Euclidean distances is then used to define the probability of assigning a query q to class d i . The training objective is formulated as a weighted negative log- likelihood over the query set D (Q) , where each query is also associated with a sample-specific weight. L PF =−w (Q) · log exp −∥f φ (q)− s i ∥ 2 P j exp (−∥f φ (q)− s j ∥ 2 ) , s i = P (x,w)∈D (S) i w· f φ (x) P (x,w)∈D (S) i w (6) Through sampling combinations of instances and proto- types, PF naturally alleviates data imbalance and reduces representational bias. This induces a structured clustering behavior in which distinct facets are progressively aligned with their corresponding theory-specific prototypes, while latent shared structures begin to emerge. However, at this stage, certain facets remain under-trained or ambiguously repre- sented, which may still lead to misclassification. Overall, PF encourages the representation space to organize around theory- aware prototypes, as illustrated in Fig. 3(c). We extend this with Human-Guided Linkage (HGL), where external human-defined supervision is introduced to ex- plicitly align shared facets across personality theories, encour- aging the emergence of a more unified shared representation space. As in Fig. 3(d), while this strategy improves cross- theory alignment, it may also introduce noise due to imperfect or overly rigid mappings that do not always reflect true seman- tic correspondence. Nevertheless, these guided linkages are particularly valuable for handling edge cases: even when the provided alignments are partially inaccurate, they still gently steer the representations toward shared regions, facilitating the re-discovery and re-association of related facets. In parallel, we explore Machine-Induced Consensus (MIC). We employ a cross-entropy objective (Eq. 7) to guide cross-dataset adaptation via a lightweight adaptation layer that projects embeddings into a shared space. This layer is not treated as part of the core model architecture; instead, it serves as a training-time mechanism to facilitate convergence toward representations that remain discriminative across different per- sonality theories. L (D∈Essays,Kaggle) MIC (y, ˆy) =− " y log ˆy + (1− y) log(1− ˆy) # (7) As illustrated in Fig. 3(e), consensus is learned through joint optimization over paired tasks, enabling the model to infer shared structure from agreement signals rather than explicit manual supervision. This process helps refine and further “clean” the shared space established by prototypical fine-tuning (PF), while also revealing latent pseudo-facets that are consistently aligned across theories. However, MIC exhibits a conservative alignment tendency, prioritizing high- confidence correspondences and potentially down-weighting weaker but still informative relationships. Consequently, it is most effective when applied after Human-Guided Linkage (HGL), where it acts as a stabilizing mechanism that regular- izes and consolidates the previously introduced human-guided alignments. Lastly, we integrate PF, HGL, and MIC, termed Cross- Theory Harmonization (CTH). This hybrid design leverages the complementary strengths of each component: PF provides 8 stable prototypical anchors for organizing the representation space, HGL introduces external guidance for soft alignment of potentially shared facets across theories, and MIC fur- ther refines these relations through data-driven consensus. Together, these mechanisms progressively reshape the repre- sentation space from one dominated by isolated theory-specific signals into a coherent shared manifold, where both distinct and overlapping pseudo-facets are more faithfully encoded, and previously under-represented or unlearned features are systematically recovered and integrated. 3) LLM-as-a-Judge (LAJ): While CTH aims to align repre- sentations across personality theories, residual noise may still hinder effective learning and lead to incorrect connections, as shown in Fig. 3(c), despite MIC efforts, which cannot fully address noise arising from mislabeled data. Moreover, while LLM-based augmentation has been shown to improve performance [49], it introduces potential risks of data leakage. We therefore argue that LLMs are better positioned as auxiliary evaluators rather than primary predictors. In particular, encoder-based architectures remain the backbone for representation learning and classification, while LLMs are employed as reasoning-based judges for data quality assess- ment rather than direct personality inference. They evaluate label consistency, contradictions, and noisy or implausible samples, improving dataset integrity and reducing leakage into core predictions. This preserves the theoretical grounding of encoder-based personality modeling while leveraging LLMs’ world knowledge and reasoning capabilities. In this work, we adopt LAJ to evaluate data cor- rectness, specifically using (i) OpenAI’s ChatGPT model (gpt, gpt-4o-2024-08-06), (i) Alibaba Qwen model (qwen, Qwen3.6-35B-A3B), (i) Meta Llama model (llama, Llama-3.1-8B-Instruct), and (iv) DeepSeek model (ds, DeepSeek-V4-Flash). We employ Chain-of-Thought (CoT) prompting to guide the LLMs in analyzing and deter- mining whether a labeled sample is Possible, Impossible, or Ambiguous with respect to its associated personality label, as in Fig. 4. 1 messages = [ 2 3"role": "system", 4"content": ( 5"Evaluate whether the given written content is consistent with the specified personality traits using the Big Five (OCEAN) framework. " 6"Your task is to assess if someone with the provided personality labels could plausibly have written the content. " 7"Use psychological reasoning and textual analysis to determine whether the content reflects each trait. " 8"Input: " 9" 'user_personality': ['High Openness', 'Low Conscientiousness', ...], " 10" 'user_content': '...' # The written content to analyze " 11"Output format (JSON): " 12" " 13" 'openness': 'analysis': ′, 'judgment': ′, " 14" 'conscientiousness': 'analysis': ′, 'judgment': ′, " 15" 'extraversion': 'analysis': ′, 'judgment': ′, " 16" 'agreeableness': 'analysis': ′, 'judgment': ′, " 17" 'neuroticism': 'analysis': ′, 'judgment': ′ " 18" " 19"Use: 'possible', 'impossible', or 'ambiguous' for each judgment." 20) 21, 22 23"role": "user", 24"content": processed_text 25 26 ] Fig. 4. The CoT System Prompt generates output in a structured JSON schema. It performs an analysis of text samples and provides recommendations on whether to ‘maintain’, ‘mute,’ or ‘lower’ contributions for training. We further conduct QLoRA fine-tuning experiments on open-source LLMs. Training is performed exclusively on the training split using the reasoning-data synthesis method proposed by [49], where ground-truth labels are leveraged to generate reasoning traces that supervise the model’s decision- making process. However, the original dataset consists only of prototypical binary personality labels and lacks ambiguous cases that require reasoning over mixed or conflicting per- sonality signals. To address this limitation, we augment the training data by constructing ambiguous samples through the combination of two distinct personality profiles, randomly se- lecting sentences from each profile to form a single input. This augmentation exposes the model to more realistic borderline cases and enables us to investigate whether LLM-based judges exhibit systematic failures or introduce new prediction biases when personality evidence is ambiguous. The augmentation is applied exclusively to the training data, while the test set remains untouched, ensuring that no data leakage occurs. As shown in Fig. 2, we study two pipelines. In the LLM- before-the-loop (LBL) approach (Algorithm 1), the LLM is employed prior to training to assess and refine training samples based on label plausibility. In contrast, the LLM- in-the-loop (LIL) strategy (Algorithm 2) integrates the LLM during training, where it dynamically evaluates samples in real time based on their loss values. Algorithm 1 LLM-before-the-loop (LBL) Require: D standard = (x,y,z), where z i = 1 by default, sys- tem_prompt, JSON_output_schema Ensure: Updated D standard with modified z 1: for all (x,y,z) in D standard do 2:for i = 0 to |d|− 1 do 3:Build x and y i content and format as the "user" role, and Merge with system_prompt 4:repeat 5:Send to LLM using JSON_output_schema 6:Attempt to parse LLM output 7:until valid JSON is returned 8:Extract LLM judgment and update the corresponding z. 9:end for 10: end for 11: return D standard Algorithm 2 LLM-in-the-loop (LIL) Require: D standard =(x,y,z), where z i = 1 by default, model f φ , system prompt, JSON_output_schema, threshold τ Ensure: Updated model parameters and possibly revised labels and weights 1: for each training episode do 2:Sample support and query sets (D (S) ,D (Q) ) from D standard , 3:Compute prototypes s j from D (S) 4:for all (x,y,w (Q) ) in D (Q) do 5:Compute query embedding f φ (q) for L JAM (LIL) 6:if L JAM (LIL) > τ then 7:Build x and y i content and format as the "user" role, and Merge with system_prompt 8:repeat 9:Send message to LLM with JSON_output_schema 10:Receive and attempt to parse response 11:until valid JSON is returned 12:Extract LLM judgment and update the corresponding z. 13:end if 14:end for 15:Aggregate and minimize total JAM (LIL) loss over query set 16:Update model parameters φ using gradient descent 17: end for Regardless of the pipeline, the outcomes, depending on the LLM’s judgment, update z in D standard . This updated z 9 thereby influences each example’s contribution during proto- type computation and training. The resulting examples are later sampled to formD prototypical in each episode and are organized into Support (D (S) ) and Query (D (Q) ) sets for training. We implemented an N -way K-shot classification task. In our setting, D (S) = K∗N and D (Q) = q∗N , where N =|d| and q depends on the GPU RAM capacity. These were designed to facilitate representational learning, enabling the model to generalize to new personality theories. Eq. 8 illustrates the structure of the dataset for prototypical model training, while Eq. 9 shows the weighting in this experiment (modified later for sensitivity testing). We also used the weighted mean of the N embeddings as the support set for few-shot learning. D prototype = D (S) = N [ i=1 n x (S) ij ,y (S) ij ,z (S) ij j = 1,...,K o , D (Q) = N [ i=1 n x (Q) ij ,y (Q) ij ,z (Q) ij j = 1,...,q o (8) z = 1.0 if ζ ∈ Possible, the LLM judg- ment is Possible (valid); maintain con- tributions. 0.2 if ζ ∈ Ambiguous, the LLM judg- ment is Ambiguous; lower the contribu- tion. 0if ζ ∈Impossible, the LLM judg- ment is Impossible (invalid); mute the contribution. (9) Finally, we incorporate the LLM-as-a-Judge mechanism to weight the contributions of each loss term. We then summarize the overall formulation in Eq. 10, where the joint loss com- bines multiple components from different datasets and tasks. L JAM = φ L (Essays) PF +L (Kaggle) PF +ψL (Essays ⇔ Kaggle) HGL +ρL MIC (10) C. Experiment Design We train the model individually on the Essays-only dataset and the Kaggle-only dataset to establish the baseline reference for the ablation study. Next, we train on the combined dataset to evaluate whether the JAM can generalize to new person- ality theories. Subsequently, we study the JAM algorithm to determine whether it can improve dataset quality and model performance. The experiments use a batch size of 32, a learning rate of 1 × 10 −5 , a maximum of 30,000 training episodes with early stopping, a random seed of 42, and 4 NVIDIA A100 GPUs. Table I presents the acronyms used and the corresponding experiment configurations. D. Evaluation To evaluate the performance, we adopted the following met- rics: Regular Accuracy (RA) (Eq. 11) to determine the overall TABLE I OVERVIEW OF EXPERIMENTAL CONFIGURATIONS AND NOTATIONS NotationDescription CE † Regular cross-entropy serves as a standard classification baseline. PO † Off-the-shelf model as the few-shot prototypical baseline. (No training) PF † Fine-tuning the model using regular meta-learning for the few-shot prototypical baseline.(φ = 1; ψ = 0; ρ = 0; Eq. 6| z=1 ) HGL ‡ PF with HGL. (φ = 1; ψ = 1; ρ = 0; Eq. 6| z=1 ) MIC ‡ PF with MIC. (φ = 1; ψ = 0; ρ = 1; Eq. 6| z=1 ) CTH ‡ PF with (HGL + MIC). (φ = 1; ψ e+1 ≤ ψ e ,∀e ≥ 1; ρ = 1; Eq. 6| z=1 ). JAM (LBL) <model> Proposed JAM approach with the LBL pipeline on CTH. JAM (LIL) <model> Proposed JAM approach with the LIL pipeline on CTH. [Essays]Training using only the Essays Dataset. [Kaggle]Training using only the Kaggle Dataset. [Both]Training using both Essays Dataset and Kaggle Dataset. Notes: † Baseline performance without any CTH modules. ‡ Ablation setting experiments for CTH modules of the JAM approach. † ‡ Experiments conducted without any weighting (no LAJ involved). <model> Indicates which LLM was used as judge in the JAM experiments. [·] Indicates the training dataset, attached after the notation. accuracy of the model, Balanced Accuracy (BA) (Eq. 12) to assess performance on imbalanced datasets, and the F1 Score (Eq. 13) to measure the model’s bias. Here, TP , FP , TN , and FN represent true positives, false positives, true negatives, and false negatives, respectively. RA = TP + TN TP + TN + FP + FN .(11) BA = 1 2 TP TP + FN + TN TN + FP .(12) F1 = 2· TP 2· TP + FP + FN .(13) IV. RESULTS A. Baseline Acquisition First, we establish a baseline for comparison by training the model separately on the Essays dataset, the Kaggle dataset, or both. Fig. 5 illustrates the balanced accuracy of each approach under different dataset settings: Cross-Entropy (CE), Off-the- shelf Prototypical (PO), and Fine-tuned Prototypical (PF). In the CE experiment, there is no doubt that when the model is trained and evaluated on the same personality theory dataset, it achieves relatively good performance. However, when the model is trained on one dataset and evaluated on another, there is a significant drop in performance, indicating that the model struggles to generalize across different personality theories. We also observe that a combination of datasets does not improve performance; rather, it degrades it. This is likely because the model is confused by the conflicting signals from the two different personality theories. On the other hand, in the PO experiment, we observe that the Essays dataset shows very poor performance, while the Kaggle dataset is still able to capture relationships. This can 10 TABLE IV PERFORMANCE COMPARISON ON THE ESSAYS DATASET, INCLUDING PRIOR WORK, THE BASELINE, ABLATION, AND LAJ MECHANISMS. Experiment O - OpennessC - ConscientiousnessE - ExtraversionA - AgreeablenessN - Neuroticism BAF1RABAF1RABAF1RABAF1RABAF1RA Psycholinguistic MLP [55]0.64600.59200.60000.58800.6050 BERT MLP [55]--0.6040-0.5730-0.5690-0.5700-0.5980 CoT with Emotion [32]-0.60930.6102-0.68640.6800-0.63020.6201-0.65010.6498-0.56000.5600 Baseline Representative [Essays]0.60300.62780.60320.53460.52870.53440.60360.61170.60320.57600.61900.57490.55360.62580.5573 HGL [Both]0.51440.54820.51620.50030.49070.50000.54240.57940.54450.55750.56800.55670.50610.53960.5061 MIC [Both]0.63880.65640.63970.54260.54070.54250.58430.60190.58500.59280.59760.59110.60320.61570.6032 CTH [Both]0.67340.68740.67410.61180.60000.61130.62780.63200.62750.63890.63410.63560.66400.66930.6640 JAM (LBL) gpt [Both]0.71630.72550.71660.64220.63050.64170.64230.64240.64170.65110.64630.64780.68420.69170.6842 JAM (LIL) gpt [Both]0.67520.69110.67610.60360.59340.60320.61620.61380.61540.61920.61070.61540.67810.67620.6781 TABLE V PERFORMANCE COMPARISON ON THE KAGGLE DATASET, INCLUDING PRIOR WORK, THE BASELINE, ABLATION, AND LAJ MECHANISMS. Experiment O - OpennessC - ConscientiousnessE - ExtraversionA - Agreeableness BAF1RABAF1RABAF1RABAF1RA BERT MLP [55]--0.6840--0.6440--0.7830--0.7440 TrigNet [56]-0.6717--0.6769--0.6954--0.7906- TAE [57]-0.8117--0.7020--0.7090--0.6621- DGCN [58]-0.6719--0.6816--0.6952--0.8053- Baseline Representative [Kaggle]0.81630.93970.89740.78370.73920.79140.81820.72180.87200.84120.85520.8427 HGL [Both]0.59010.67170.54990.64090.58690.63980.62320.43330.62310.75830.76260.7556 MIC [Both]0.76950.90390.84090.69020.63900.68930.68250.50100.70490.80580.81770.8058 CTH [Both]0.81440.92600.87610.76340.71770.76600.77390.64800.83340.84080.85590.8427 JAM (LBL) gpt [Both]0.81330.92870.88010.77210.72800.77350.78720.66590.84030.84980.86480.8519 JAM (LIL) gpt [Both]0.79410.91940.86510.75380.70720.75560.73560.58180.78790.81710.83720.8202 be attributed to the nature of how the datasets are collected. The Essays dataset comes from a constrained environment, whereas the Kaggle dataset is sourced from social media, which is more informal and diverse. This provides confidence that few-shot learning can be effective in personality recogni- tion tasks. One interesting observation is that it naturally solves the class imbalance problem, as the algorithm focuses on learning the class prototypes rather than being biased towards the majority class. When the model is fine-tuned using PF, we see a signif- icant improvement in performance within the same dataset (training and evaluation). This is especially evident for the Kaggle dataset, giving confidence that the model can learn the underlying personality patterns in the data. An interesting observation is that a model trained on the Kaggle dataset can generalize to the Essays dataset, but not vice versa, further demonstrating the quality of the Kaggle dataset. Additionally, the results show that the model trained on the Kaggle dataset achieves similar results to the CE approach trained on the Essays dataset, further justifying that a prototypical network has potential in capturing personality features from text. B. Performance 1) Cross-Theory Harmonization (CTH) Performance: Ta- ble IV and Table V tabulate the performance of each CTH module under different dataset settings with its ablation study. To provide a reference baseline, we include the row corre- sponding to the highest performance (BA) from the afore- mentioned experiments (as shown in Fig. 5) as representative baseline. We also include prior work; however, it is not fully comparable since they focus on training on a single dataset, whereas we train 1 model for 2 personalities theories. Never- theless, our performance remains competitive and significant in most cases, highlighting the advantage of our approach for generalization. By observing the results on the Essays dataset, we find that the model improves on average by more than 9% in balanced accuracy across all traits compared to the best-performing baseline. We further analyze the contribution of each com- ponent. Across both datasets, when only HGL is included, performance actually degrades regardless of the dataset. This supports the aforementioned hypothesis that human knowledge is limited; in the Essays dataset, due to its constrained data collection environment, such limitations may hinder the model and expose it to suboptimal supervision, obscuring the ultimate learning objective through self-consensus. Considering MIC alone, it only outperforms the baseline on the Essays dataset but not on Kaggle, further strengthening the observation that the combination of Essays data introduces additional noise. Nevertheless, the combination of both components (CTH) is able to restore performance to a level comparable with the baseline on Kaggle and further improve results on the Essays dataset, suggesting that integrating HGL and MIC allows the model to better merge complementary knowledge sources, achieving theory-agnostic modeling. 2) LLM-as-a-Judge (LAJ) Performance: Next, we study the impact of the LAJ mechanism. Generally, the JAM (LBL) gpt [Both] approach consistently outperforms the JAM (LIL) gpt [Both] approach across all personality traits in both datasets. This suggests that pre-evaluating and refining the dataset before training is more effective than dynamically assessing sam- ples during training. The JAM (LBL) gpt [Both] model converges 11 OCEAN 0.40 0.42 0.44 0.46 0.48 0.50 0.52 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 Balanced Accuracy (BA) 0.548 0.499 0.470 0.508 0.569 0.510 0.546 CE [Essays] CE [Kaggle] CE [Both] PO PF [Essays] PF [Kaggle] PF [Both] (a) Essays Dataset O (S/N)C (J/P)E (E/I)A (S/N) 0.50 0.52 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 0.70 0.72 0.74 0.76 0.78 0.80 0.82 0.84 0.86 0.88 0.90 0.92 0.94 0.96 0.98 Balanced Accuracy (BA) 0.709 0.727 0.645 0.568 0.573 0.815 0.656 CE [Essays] CE [Kaggle] CE [Both] PO PF [Essays] PF [Kaggle] PF [Both] (b) Kaggle Dataset Fig. 5. Visualization of balanced accuracy for regular classification using cross-entropy (CE), prototypical few-shot learning using off-the-shelf ready model (PO), and fine-tuned prototypical few-shot learning (PF) on the Essays and Kaggle datasets. Each line in the legend corresponds to an individual experiment, with the dataset used for training indicated in the respective brackets, while the dotted lines with numbers represent the average value across dimensions of the respective experiment. within approximately 3000 episodes, whereas JAM (LIL) [Both] requires around 7000 episodes to converge. This difference arises because the JAM (LIL) gpt [Both] approach only gradually obtains a cleaner dataset over multiple iterative episodes, and not all samples are evaluated by the LLM due to the thresh- olding mechanism applied during training. As a result, an ini- tially noisy dataset may mislead the model toward suboptimal local minima, which helps explain why the JAM (LIL) gpt [Both] approach is less effective than JAM (LBL) gpt [Both]. This trend is also reflected in the Essays dataset, where the performance of JAM (LBL) gpt [Both] is comparable to CTH [Both]. Fig. 6 illustrates the distribution of Possible and Impossible judgments across personality traits for both datasets. The Essays dataset contains a higher proportion of Possible judg- ments (85.7%) compared to the Kaggle dataset (70.8%). Given that only a limited proportion of the dataset is estimated to be noisy (approximately > 8%), the observed 2% improvement is within a reasonable range. This indicates that the filtering process effectively reduces the influence of noisy samples, while also suggesting that the achievable performance gain is naturally bounded by the proportion of removable noise in the dataset. O_1 O_1O_1 O_1 O_1 C_1 C_1C_1 C_1 C_1 E_1 E_1E_1 E_1 E_1 A_1 A_1A_1 A_1 A_1 N_1 N_1N_1 N_1 N_1 O_0 O_0O_0 O_0 O_0 C_0 C_0C_0 C_0 C_0 E_0 E_0E_0 E_0 E_0 A_0 A_0A_0 A_0 A_0 N_0 N_0N_0 N_0 N_0 possible (85.7%) possible (85.7%)possible (85.7%) possible (85.7%) possible (85.7%) impossible (5.5%) impossible (5.5%)impossible (5.5%) impossible (5.5%) impossible (5.5%) ambiguous (8.8%) ambiguous (8.8%)ambiguous (8.8%) ambiguous (8.8%) ambiguous (8.8%) (a) Essays Dataset O_1 O_1O_1 O_1 O_1 C_1 C_1C_1 C_1 C_1 E_1 E_1E_1 E_1 E_1 A_1 A_1A_1 A_1 A_1 O_0 O_0O_0 O_0 O_0 C_0 C_0C_0 C_0 C_0 E_0 E_0E_0 E_0 E_0 A_0 A_0A_0 A_0 A_0 possible (70.8%) possible (70.8%)possible (70.8%) possible (70.8%) possible (70.8%) impossible (15.7%) impossible (15.7%)impossible (15.7%) impossible (15.7%) impossible (15.7%) ambiguous (13.4%) ambiguous (13.4%)ambiguous (13.4%) ambiguous (13.4%) ambiguous (13.4%) (b) Kaggle Dataset Fig. 6.Distribution of LLM gpt judgments on personality trait implications across datasets. The figure illustrates the proportion of LLM-judged out- comes—Possible, Impossible, and Ambiguous—based on the binary labels for each of the personality traits. Labels in the form d_y represent trait d with label y, where y = 1 denotes a positive class (presence of the trait) and y = 0 denotes a negative class. The divergence in distributions highlights the influence of annotation conditions and data origin on interpretability judgments made by language models. Fig. 7 shows the performance of different LLMs under varying hyperparameter settings. Overall, gpt performs the best. On the Kaggle dataset, LLMs generally do not improve over the baseline, which is consistent with the CTH stage (without LAJ). In contrast, the Essays dataset shows a different trend: with the involvement of LAJ, performance generally improves over CTH alone. This supports the generalisability claims for the low-resource theory (Essays dataset) compared with the Kaggle dataset, showing that in most cases, model selection does not degrade performance. We observed that setting z ζ∈Impossible = 0 (muting the Impossible samples) often leads to better performance in the Essays dataset, supporting the generalisation claim. However, this effect is less consistent in the Kaggle dataset and depends more on the choice of LLM. From the results of the Kaggle dataset, z ζ∈Impossible = 0 generally has relatively low performance compared to other hyperparameters, suggesting that the model is not only incapable of filtering data but also introduces more noise by removing some important contribu- tions. In addition, we conduct an experiment in which 80% of the minor (relatively underrepresented) labels/classes are removed (in gpt) using the setting z ζ∈Ambiguous = 0.2 and z ζ∈Impossible = 0, to simulate a worse model scenario. The results show a significant drop across all dimensions, despite this being the best observed hyperparameter setting for gpt. This again supports the claim that a worse LLM can negatively affect training, and shows that z ζ∈Impossible is sensitive. Across different models, z ζ∈Ambiguous does not sig- nificantly affect performance (evaluated under settings z ζ∈Ambiguous = 0.2 and z ζ∈Impossible = 0.8, as well as z ζ∈Impossible = 0). This suggests that it primarily acts as a tunable non-sensitive hyperparameter rather than a structural factor. On the other hand, in terms of their fine-tuned model, the results show that although there is improvement in certain dimensions for some LLMs on the Kaggle dataset, the gains are not statistically significant. In some cases, performance even drops on the Essays dataset, likely due to imbalance, as Kaggle is more dominant and relatively easier to learn. These findings suggest that well-trained LLMs are already 12 0.46 0.48 0.50 0.52 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 0.70 0.72 0.74 0.76 0.78 0.80 Balanced Accuracy (BA) OCEAN CTH (No LAJ) z = 1 z ζ∈ Ambigous Imposible = 0 z ζ∈ Ambigous Imposible = 0.5 z ζ∈Ambigous = 0.8 z ζ∈Imposible = 0.5 z ζ∈Ambigous = 0.8 z ζ∈Imposible = 0 z ζ∈Ambigous = 0.2 z ζ∈Imposible = 0 Modified model z ζ∈Ambigous = 0.2 z ζ∈Imposible = 0 0.6985 ± 0.0145 gpt 0.6897 ± 0.0163 qwen 0.6840 ± 0.0052 llama 0.6866 ± 0.0129 ds 0.6351 ± 0.0082 gpt 0.6172 ± 0.0068 qwen 0.6262 ± 0.0128 llama 0.6095 ± 0.0054 ds 0.6371 ± 0.0062 gpt 0.6238 ± 0.0095 qwen 0.6241 ± 0.0057 llama 0.6169 ± 0.0068 ds 0.6443 ± 0.0055 gpt 0.6292 ± 0.0068 qwen 0.6278 ± 0.0078 llama 0.6291 ± 0.0088 ds 0.6765 ± 0.0058 gpt 0.6700 ± 0.0067 qwen 0.6664 ± 0.0044 llama 0.6660 ± 0.0067 ds 0.67340.61180.62780.63890.6640 (a) Essays Dataset 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Balanced Accuracy (BA) OCEA CTH (No LAJ) z = 1 z ζ∈ Ambigous Imposible = 0 z ζ∈ Ambigous Imposible = 0.5 z ζ∈Ambigous = 0.8 z ζ∈Imposible = 0.5 z ζ∈Ambigous = 0.8 z ζ∈Imposible = 0 z ζ∈Ambigous = 0.2 z ζ∈Imposible = 0 Modified model z ζ∈Ambigous = 0.2 z ζ∈Imposible = 0 0.8131 ± 0.0073 gpt 0.7925 ± 0.0145 qwen 0.7927 ± 0.0113 llama 0.7893 ± 0.0089 ds 0.7653 ± 0.0105 gpt 0.7364 ± 0.0134 qwen 0.7402 ± 0.0201 llama 0.7443 ± 0.0125 ds 0.7659 ± 0.0212 gpt 0.7448 ± 0.0190 qwen 0.7487 ± 0.0246 llama 0.7362 ± 0.0279 ds 0.8371 ± 0.0025 gpt 0.8089 ± 0.0167 qwen 0.8110 ± 0.0170 llama 0.8126 ± 0.0190 ds 0.81440.76340.77390.8408 (b) Kaggle Dataset Fig. 7. Visualization of different large language models (gpt, qwen, llama, ds) and their performance under respective hyperparameter settings on ζ to z values in JAM (LBL) gpt [Both], including a reduced contribution of the ambiguous and impossible flags by down-weighting them at certain levels. We further incorporate fine-tuning for open-source models using QLoRA, while for gpt we conduct an experiment in which 80% of the minor (relatively underrepresented) labels/classes are removed to address class imbalance, under the assumption that the LLM significantly underperforms on the task, in order to study the potential effects of such imbalance induction. sufficiently capable, and additional fine-tuning is unnecessary and induces higher costs for the judging task. C. Statistical Findings We conducted the McNemar test to support our findings, as illustrated in Fig. 8. The results indicate its potential utility in constrained environments, while the CTH method addresses the challenges of merging datasets, suggesting a pathway toward a theory-agnostic model. Furthermore, this demonstrates that our proposed approaches are able to reduce noise that appeared in the native collection of the dataset, where it statistically further improves through LAJ (an average of 2% improvement in the Essays dataset). To visualize the effectiveness of the proposed algorithm, we visualized the personality embeddings using t-SNE in Fig. 9. Compared to the off-the-shelf PO approach, the embeddings produced by JAM (LBL) gpt [Both] are noticeably more structured and exhibit clearer separation between clusters correspond- ing to distinct personality trait combinations. This improved separation indicates that the proposed method captures and preserves subtle personality cues from textual data more G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 PF: TP PF: TPPF: TP PF: TP PF: TP PF: FN PF: FNPF: FN PF: FN PF: FN PF: TN PF: TNPF: TN PF: TN PF: TN PF: FP PF: FPPF: FP PF: FP PF: FP MIC: TP MIC: TPMIC: TP MIC: TP MIC: TP MIC: FN MIC: FNMIC: FN MIC: FN MIC: FN MIC: TN MIC: TNMIC: TN MIC: TN MIC: TN MIC: FP MIC: FPMIC: FP MIC: FP MIC: FP p-value= 1.969e-06 (a) Essays Dataset (PF→MIC)[Both] G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 MIC: TP MIC: TPMIC: TP MIC: TP MIC: TP MIC: FN MIC: FNMIC: FN MIC: FN MIC: FN MIC: TN MIC: TNMIC: TN MIC: TN MIC: TN MIC: FP MIC: FPMIC: FP MIC: FP MIC: FP CTH: TP CTH: TPCTH: TP CTH: TP CTH: TP CTH: FN CTH: FNCTH: FN CTH: FN CTH: FN CTH: TN CTH: TNCTH: TN CTH: TN CTH: TN CTH: FP CTH: FPCTH: FP CTH: FP CTH: FP p-value= 3.364e-22 (b) Essays Dataset (MIC→CTH)[Both] G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 CTH: TP CTH: TPCTH: TP CTH: TP CTH: TP CTH: FN CTH: FNCTH: FN CTH: FN CTH: FN CTH: TN CTH: TNCTH: TN CTH: TN CTH: TN CTH: FP CTH: FPCTH: FP CTH: FP CTH: FP JAM: TP JAM: TPJAM: TP JAM: TP JAM: TP JAM: FN JAM: FNJAM: FN JAM: FN JAM: FN JAM: TN JAM: TNJAM: TN JAM: TN JAM: TN JAM: FP JAM: FPJAM: FP JAM: FP JAM: FP p-value= 0.008246 (c) Essays Dataset (CTH→ JAM (LBL) gpt )[Both] G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 PF: TP PF: TPPF: TP PF: TP PF: TP PF: FN PF: FNPF: FN PF: FN PF: FN PF: TN PF: TNPF: TN PF: TN PF: TN PF: FP PF: FPPF: FP PF: FP PF: FP MIC: TP MIC: TPMIC: TP MIC: TP MIC: TP MIC: FN MIC: FNMIC: FN MIC: FN MIC: FN MIC: TN MIC: TNMIC: TN MIC: TN MIC: TN MIC: FP MIC: FPMIC: FP MIC: FP MIC: FP p-value= 1.074e-128 (d) Kaggle Dataset (PF→MIC)[Both] G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 MIC: TP MIC: TPMIC: TP MIC: TP MIC: TP MIC: FN MIC: FNMIC: FN MIC: FN MIC: FN MIC: TN MIC: TNMIC: TN MIC: TN MIC: TN MIC: FP MIC: FPMIC: FP MIC: FP MIC: FP CTH: TP CTH: TPCTH: TP CTH: TP CTH: TP CTH: FN CTH: FNCTH: FN CTH: FN CTH: FN CTH: TN CTH: TNCTH: TN CTH: TN CTH: TN CTH: FP CTH: FPCTH: FP CTH: FP CTH: FP p-value= 8.748e-268 (e) Kaggle Dataset (MIC→CTH)[Both] G: 0 G: 0G: 0 G: 0 G: 0 G: 1 G: 1G: 1 G: 1 G: 1 CTH: TP CTH: TPCTH: TP CTH: TP CTH: TP CTH: FN CTH: FNCTH: FN CTH: FN CTH: FN CTH: TN CTH: TNCTH: TN CTH: TN CTH: TN CTH: FP CTH: FPCTH: FP CTH: FP CTH: FP JAM: TP JAM: TPJAM: TP JAM: TP JAM: TP JAM: FN JAM: FNJAM: FN JAM: FN JAM: FN JAM: TN JAM: TNJAM: TN JAM: TN JAM: TN JAM: FP JAM: FPJAM: FP JAM: FP JAM: FP p-value= 0.1011 (f) Kaggle Dataset (CTH→ JAM (LBL) gpt )[Both] Fig. 8. The Sankey diagram illustrates the transitions between 4 approaches: from the PF [Both] to the MIC-involved method (first MIC [Both], and then to CTH [Both] method, and to the proposed full approach JAM (LBL) gpt [Both]). Note that the HGL-only is excluded from this test since it functions as early generalization guidance. It encapsulates noise and performs worse when functioning alone. Each method’s outcomes are compared against the original ground truth (G). Statistical significance is determined using the McNemar p- value test. Due to the fact that personality encompasses multiple dimensions, we flatten and concatenate these dimensions to facilitate clearer visualization. effectively. In particular, the cluster centers in the JAM (LBL) gpt [Both] visualizations are more distinct and less overlapping, suggesting that the model can better differentiate between similar personality profiles. This pattern is consistent across both datasets, with the effect being particularly pronounced in the Kaggle dataset, where clusters are clearer due to a larger number of samples and relatively higher classification accu- racy, further highlighting the robustness and generalizability of JAM (LBL) gpt [Both]. D. Computational Analysis Fig. 10 compares the proposed JAM with previous ap- proaches in terms of Floating Point Operations (FLOPs) and the corresponding computational cost. This comparison applies to both LLM-before-the-loop (LBL) and LLM-in-the-loop (LIL) settings. Technically, the former achieves lower com- putational cost because the filtering mechanism depends only on the loss values; therefore, not all samples in the dataset require full inference. Overall, the proposed method achieves the lowest infer- ence time compared to the PICEPR (Embeddings) method [49]. Although a slight additional overhead is introduced due to prototype retrieval, this overhead is negligible because it only involves retrieving and averaging 4–5 embedding vectors. Despite this, the method still achieves approximately 8× lower training FLOPs. Furthermore, the proposed approach demonstrates the feasibility of incorporating LLMs to improve performance, while simultaneously reducing the risk of data leakage, which is difficult to achieve with purely decoder- only model-driven approaches such as PICEPR (Contents) [49]. Overall, JAM maintains very low inference latency, which is particularly important for test-time deployment and real-world 13 40200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 00000 (0) 00001 (1) 00010 (2) 00011 (3) 00100 (4) 00101 (5) 00110 (6) 00111 (7) 01000 (8) 01001 (9) 01010 (10) 01011 (11) 01100 (12) 01101 (13) 01110 (14) 01111 (15) 10000 (16) 10001 (17) 10010 (18) 10011 (19) 10100 (20) 10101 (21) 10110 (22) 10111 (23) 11000 (24) 11001 (25) 11010 (26) 11011 (27) 11100 (28) 11101 (29) 11110 (30) 11111 (31) (a) Essays Dataset PO [Essays] 40200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 00000 (0) 00001 (1) 00010 (2) 00011 (3) 00100 (4) 00101 (5) 00110 (6) 00111 (7) 01000 (8) 01001 (9) 01010 (10) 01011 (11) 01100 (12) 01101 (13) 01110 (14) 01111 (15) 10000 (16) 10001 (17) 10010 (18) 10011 (19) 10100 (20) 10101 (21) 10110 (22) 10111 (23) 11000 (24) 11001 (25) 11010 (26) 11011 (27) 11100 (28) 11101 (29) 11110 (30) 11111 (31) (b) Essays Dataset JAM (LBL) gpt [Both] 40200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0000 (0) 0001 (1) 0010 (2) 0011 (3) 0100 (4) 0101 (5) 0110 (6) 0111 (7) 1000 (8) 1001 (9) 1010 (10) 1011 (11) 1100 (12) 1101 (13) 1110 (14) 1111 (15) (c) Kaggle Dataset PO [Kaggle] 40200204060 t-SNE Component 1 40 20 0 20 40 t-SNE Component 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0000 (0) 0001 (1) 0010 (2) 0011 (3) 0100 (4) 0101 (5) 0110 (6) 0111 (7) 1000 (8) 1001 (9) 1010 (10) 1011 (11) 1100 (12) 1101 (13) 1110 (14) 1111 (15) (d) Kaggle Dataset JAM (LBL) gpt [Both] Fig. 9. The t-SNE visualization of personality embeddings on both Essays and Kaggle datasets. Subplots 9(a) and 9(c) show embeddings using the off- the-shelf model (PO), while subplots 9(b) and 9(d) show embeddings using the proposed JAM (LBL) gpt [Both] approach. Each color represents a distinct personality trait combination, with annotated labels indicating cluster centers. Compared to the PO approach, JAM (LBL) gpt [Both] produces embeddings that are more structured and better separated, reflecting improved modeling of personality cues obtained from text. inference scenarios, while also reducing inference FLOPs, thereby lowering energy consumption and carbon footprint and improving the environmental sustainability especially on large- scale deployment. E. Limitations and Future Works Although JAM demonstrates consistent improvements across heterogeneous personality datasets, it is currently val- idated on only two personality theories (MBTI and Big-5), and further evaluation on additional theories and culturally diverse datasets is needed to validate the robustness of the proposed Cross-Theory Harmonization (CTH). In addition, as personality recognition relies on sensitive personal data, future work should prioritize privacy, informed consent, and fairness. One promising direction is to integrate JAM with a federated learning framework, enabling collaboration across institutions or countries while keeping data local and aggregating only model updates, thereby improving generalizability without compromising data privacy. V. CONCLUSION In this work, we introduced Large-Language-Models-as-a- Judge in Theory-Agnostic Adaptive Metric-Alignment for Pro- totypical Networks in Personality Recognition (JAM), a novel prototypical framework designed to enhance personality pre- diction and generalizability by leveraging Cross-Theory Har- monization (Human-Guided Linkage and Machine-Induced Consensus) and LLM-as-a-Judge mechanisms, enabling the learning of latent pseudo-facets that capture shared behavioral 012T24T36T48T PICEPR (Contents) PICEPR (Embeddings) EERPD TAE JAM Classify's COT (I)/ (O) Summary's COT (I)/ (O) Psycho's COT (I)/ (O) Inference Summary's COT (I)/ (O) Mimic's COT (I)/ (O) Encoder Model Training (Task-Focus) Classify's COT (I)/ (O) Sentences Categorisation (I)/ (O) Retrieval Reference Library (Single)/ (Whole Training Dataset) Reasoning's COT (I)/ (O) InferenceRetrieval Reference Library (Single)/ (Whole Training Dataset) Inference Retrieval Reference Library (Single)/ (Whole Training Dataset) 24P48P FLOPs Encoder Model Training (Task-Focus) Encoder Model Training (Task-Focus) 74P101P 23.78T FLOPs $0.0019 96.02P FLOPs $7.5018 33.69P FLOPs $2.6322 16.72P FLOPs $1.3063 11.91P FLOPs $0.9301 00.00090.00190.00280.0037 Cost ($) 1.883.755.817.88 Inference timePre-inferenceTraining time Fig. 10.Visualization of Cost and FLOPs Comparison. In this analysis, we assume that input and output tokens incur identical costs, adopting the estimated FLOPs per token and the corresponding pricing model used in [49] to enable direct comparison. We assume a constant inference cost of 3.2 GFLOPs per token and a pricing rate of $0.25 per 10M tokens. The visualization is based on three components: (i) single inference (grey, with pink representing search time, e.g., retrieval-augmented generation lookup or similarity measurement, quantified according to the respective approach used to achieve the reported performance), (i) training time (purple, covering all experiments), and (i) pre-inference (yellow and orange). Pre-inference (e.g., retrieval-augmented generation indexing or prototype preparation) is difficult to standardize because greater involvement can improve performance for certain methods. To address this, we divide the pre-inference stage into yellow and orange segments to represent an estimated ratio of involved classes or prototypes, where this is quantified according to the respective approach used to achieve the reported performance. Note that this visualization excludes the training cost of the pretrained decoder-only model, as it is common across all approaches and assumed to be comparable. structure across heterogeneous personality theories under an embedding space. Through experiments on the Essays and Kaggle datasets, we demonstrated that JAM consistently outperforms existing baselines and prior work across multiple personality traits, with no class imbalance issues. Specifically, JAM achieves an average BA improvement of 12% on the Essays dataset and 14% on the Kaggle dataset compared to the regular prototyp- ical few-shot learning approach. When combining datasets, it becomes possible to leverage the dataset for generalization in low-resource situations, where the performance of the Kaggle dataset is not affected. Moreover, incorporating the LLM-as-a- Judge mechanism further improves performance on the Essays dataset by an additional 2.4%, by reweighting the contribution of each data row during model training. These results indicate that JAM is particularly effective in constrained environments while maintaining robust performance across diverse datasets. VI. ACKNOWLEDGMENTS This research was funded by the Universiti Tunku Abdul Rahman Research Fund (IPSR/RMC/UTARRF/2021-C1/K03). The authors also appreciate the support of Grid5000, France, for providing the computational resources used in this study. The authors thank the Advanced Artificial Intelligence Re- search Center at the Kanagawa Institute of Technology, Japan, for supporting OpenAI’s model inferences. The first author, Jing Jie Tan, also appreciates National Yang Ming Chiao Tung University for providing the Research Scholarship to establish a professional connection for guidance related to 14 psychology. Additionally, he appreciate the Embassy of France to Malaysia for the Doctoral Research Mobility Grant and the Japan Student Services Organization (JASSO) Scholarships, which facilitated the research collaboration between Universiti Tunku Abdul Rahman, Malaysia, Université Sorbonne Paris Nord, France, and Kanagawa Institute of Technology, Japan. REFERENCES [1] S. Dhelim, N. Aung, M. A. Bouras, H. Ning, and E. Cambria, “A survey on personality-aware recommendation systems,” Artificial Intelligence Review, vol. 55, no. 3, p. 2409–2454, Sep. 2021. [Online]. Available: http://dx.doi.org/10.1007/s10462-021-10063-7 [2] T. Ait Baha, M. El Hajji, Y. Es-Saady, and H. Fadili, “The power of personalization: A systematic review of personality-adaptive chatbots,” SN Computer Science, vol. 4, no. 5, Aug. 2023. [Online]. Available: http://dx.doi.org/10.1007/s42979-023-02092-6 [3] Z. Li, S. Yang, and S. Wang, “Exploring personality-driven personalization in xai: Enhancing user trust in gameplay,” 2024. [Online]. Available: https://arxiv.org/abs/2408.04778 [4] S. Omidvar and T. Tran, “Tackling cold-start with deep personalized transfer of user preferences for cross-domain recommendation,” International Journal of Data Science and Analytics, Nov. 2023. [Online]. Available: http://dx.doi.org/10.1007/s41060-023-00467-9 [5] S. Dhelim, L. L. Chen, N. Aung, W. Zhang, and H. Ning, “Big-five, mpti, eysenck or hexaco: The ideal personality model for personality-aware recommendation systems,” 2021. [Online]. Available: https://arxiv.org/abs/2106.03060 [6] D. Stillwell and M. Kosinski, “mypersonality project website,” 2015. [7] L. V. Phan and J. F. Rauthmann, “Personality computing: New frontiers in personality assessment,” Social and Personality Psychology Compass, vol. 15, no. 7, Jun. 2021. [Online]. Available: http: //dx.doi.org/10.1111/spc3.12624 [8] M. Dalvi-Esfahani, A. Niknafs, Z. Alaedini, H. Barati Ahmadabadi, D. J. Kuss, and T. Ramayah, “Social media addiction and empathy: Moderating impact of personality traits among high school students,” Telematics and Informatics, vol. 57, p. 101516, Mar. 2021. [Online]. Available: http://dx.doi.org/10.1016/j.tele.2020.101516 [9] G. Lampropoulos, T. Anastasiadis, K. Siakas, and E. Siakas, “The impact of personality traits on social media use and engagement: An overview.” International Journal on Social and Education Sciences, vol. 4, no. 1, p. 34–51, 2022. [10] E. Ahmed and S. Ahmed, “Social media addiction, personality traits, and disorders: an overview of recent literature,” Current Opinion in Psychiatry, vol. 38, no. 1, p. 72–77, Sep. 2024. [Online]. Available: http://dx.doi.org/10.1097/YCO.0000000000000969 [11] R. M. Spielman, K. Dumper, W. Jenkins, A. Lacombe, M. Lovett, and M. Perlmutter, “Personality assessment,” Introduction to Psychology (A critical approach), 2021. [12] R. W. Robins, J. L. Tracy, and J. W. Sherman, “What kinds of methods do personality psychologists use?: A survey of journal editors and editorial board members,” 2023. [13] M. H. Waugh, C. M. McClain, E. C. Mariotti, A. L. Mulay, E. N. DeVore, K. A. Lenger, A. N. Russell, A. R. Florimbio, K. C. Lewis, J. M. Ridenour, and L. G. Beevers, “Comparative content analysis of self-report scales for level of personality functioning,” Journal of Personality Assessment, vol. 103, no. 2, p. 161–173, Jan. 2020. [Online]. Available: http://dx.doi.org/10.1080/00223891.2019.1705464 [14] I. MUKHTASAR and M. MAVLUDA, “The study of projective methods in psychology,” JournalNX, vol. 7, no. 02, p. 66–69, 2021. [15] R. Tett and D. Simonet, “Applicant faking on personality tests: Good or bad and why should we care?” Personnel Assessment and Decisions, vol. 7, no. 1, May 2021. [Online]. Available: http://dx.doi.org/10.25035/pad.2021.01.002 [16] B. W. Roberts and H. J. Yoon, “Personality psychology,” Annual Review of Psychology, vol. 73, no. 1, p. 489–516, Jan. 2022. [Online]. Available: http://dx.doi.org/10.1146/annurev-psych-020821-114927 [17] D. Cervone and L. A. Pervin, Personality: Theory and research. John Wiley & Sons, 2022. [18] I. B. Myers, “The myers-briggs type indicator: Manual (1962).” 1962. [19] A. Furnham, “Myers-briggs type indicator (mbti),” Encyclopedia of Personality and Individual Differences, p. 1–4, 2017. [20] L. R. Goldberg, “The structure of phenotypic personality traits.” Amer- ican Psychologist, vol. 48, p. 26–34, 1993. [21] X. Luo, Y. Ge, and W. Qu, “The association between the big five personality traits and driving behaviors: A systematic review and meta-analysis,” Accident Analysis and Prevention, vol. 183, p. 106968, 2023. [Online]. Available: https://w.sciencedirect.com/ science/article/pii/S0001457523000155 [22] M. C. Ashton, K. Lee, M. Perugini, P. Szarota, R. E. de Vries, L. D. Blas, K. Boies, and B. D. Raad, “A six-factor structure of personality- descriptive adjectives: Solutions from psycholexical studies in seven languages.” Journal of Personality and Social Psychology, vol. 86, p. 356–366, 2004. [23] P. William and A. Badholia, “Analysis of personality traits from text based answers using hexaco model,” in 2021 International Conference on Innovative Computing, Intelligent Communication and Smart Elec- trical Systems (ICSES), 2021, p. 1–10. [24] J. J. Tan, B.-H. Kwan, D. W.-K. Ng, and Y. C. Hum, “Psychology- informed natural language understanding: Integrating personality and emotion-aware features for comprehensive sentiment analysis and depression detection,” Pertanika Journal of Science and Technology, vol. 33, no. S4, Jun. 2025. [Online]. Available: http://dx.doi.org/10. 47836/pjst.33.S4.04 [25] W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33.Curran Associates, Inc., 2020, p. 5776–5788. [Online]. Available: https://proceedings.neurips.c/paper_ files/paper/2020/file/3f5e243547dee91fbd053c1c4a845a-Paper.pdf [26] K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu, “Mpnet: Masked and permuted pre-training for language understanding,” in Advances in Neural Information Processing Systems, H. Larochelle, M.Ranzato,R.Hadsell,M.Balcan,andH.Lin,Eds., vol.33.CurranAssociates,Inc.,2020,p.16 857–16 867. [Online].Available:https://proceedings.neurips.c/paper_files/paper/ 2020/file/c3a690be93a602e2dc0ccab5b7b67e-Paper.pdf [27] J. Ni, G. H. Ábrego, N. Constant, J. Ma, K. B. Hall, D. Cer, and Y. Yang, “Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models,” 2021. [Online]. Available: https://arxiv.org/abs/2108.08877 [28] D. Bear and P. Cook, “Fine-tuning sentence-RoBERTa to construct word embeddings for low-resource languages from bilingual dictionaries,” in Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP), M. Mager, A. Ebrahimi, A. Oncevay, E. Rice, S. Rijhwani, A. Palmer, and K. Kann, Eds.Toronto, Canada: Association for Computational Linguistics, Jul. 2023, p. 47–57. [Online]. Available: https://aclanthology.org/2023. americasnlp-1.7/ [29] J. Han and L. Yang, “Sentence embedding generation framework based on kullback–leibler divergence optimization and roberta knowledge distillation,” Mathematics, vol. 12, no. 24, p. 3990, Dec. 2024. [Online]. Available: http://dx.doi.org/10.3390/math12243990 [30] S. G. Tesfagergish, J. Kapo ˇ ci ̄ ut ̇ e-Dzikien ̇ e, and R. Damaševi ˇ cius, “Zero-shot emotion detection for semi-supervised sentiment analysis using sentence transformers and ensemble learning,” Applied Sciences, vol. 12, no. 17, 2022. [Online]. Available: https://w.mdpi.com/ 2076-3417/12/17/8662 [31] M. Shafikuzzaman, M. R. Islam, A. C. Rolli, S. Akhter, and N. Seliya, “An empirical evaluation of the zero-shot, few-shot, and traditional fine-tuning based pretrained language models for sentiment analysis in software engineering,” IEEE Access, vol. 12, p. 109 714–109 734, 2024. [32] Z. Li, D. Zhu, Q. Ma, W. Xiong, and S. Li, “Eerpd: Leveraging emotion and emotion regulation for improving personality detection,” 2024. [Online]. Available: https://arxiv.org/abs/2406.16079 [33] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2302.13971 [34] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” 2023. [Online]. Available: https://arxiv.org/abs/2310.06825 [35] J. Bai et al., “Qwen technical report,” 2023. [Online]. Available: https://arxiv.org/abs/2309.16609 [36] OpenAI et al., “Gpt-4 technical report,” 2023. [Online]. Available: https://arxiv.org/abs/2303.08774 [37] T. Zhong et al., “Evaluation of openai o1: Opportunities and challenges of agi,” 2024. [Online]. Available: https://arxiv.org/abs/2409.18486 15 [38] DeepSeek-AI et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025. [Online]. Available: https://arxiv.org/abs/2501.12948 [39] J. J. Tan, B.-H. Kwan, D. W.-K. Ng, Y.-C. Hum, N. Kawarazaki, K. Takano, and A. Mokraoui, “Cross-lingual attention distillation with personality-informed generative augmentation for multilingual personality recognition,” IEEE Transactions on Cognitive and DevelopmentalSystems,p.1–16,2026.[Online].Available: http://dx.doi.org/10.1109/TCDS.2026.3682672 [40] J. J. Tan, B.-H. Kwan, D. W.-K. Ng, and Y.-C. Hum, “Adaptive focal loss with personality stratification for stably mitigating hard class imbalance in multi-dimensional personality recognition,” Scientific Reports, vol. 15, no. 1, Nov. 2025. [Online]. Available: http: //dx.doi.org/10.1038/s41598-025-22853-y [41] A. Khattak, N. Jellani, M. Z. Asghar, and U. Asghar, “Personality classification from text using bidirectional long short-term memory model,” Multimedia Tools and Applications, vol. 83, no. 10, p. 28849–28873, Sep. 2023. [Online]. Available: http://dx.doi.org/10.1007/ s11042-023-16661-7 [42] Z. Guan, B. Wu, B. Wang, and H. Liu, “Personality2vec: Network representation learning for personality,” in 2020 IEEE Fifth International Conference on Data Science in Cyberspace (DSC), 2020, p. 30–37. [43] E. Kerz, Y. Qiao, S. Zanwar, and D. Wiechmann, “Pushing on personality detection from verbal behavior: A transformer meets text contours of psycholinguistic features,” in Proceedings of the 12th Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis, J. Barnes, O. De Clercq, V. Barriere, S. Tafreshi, S. Alqahtani, J. Sedoc, R. Klinger, and A. Balahur, Eds.Dublin, Ireland: Association for Computational Linguistics, May 2022, p. 182–194. [Online]. Available: https://aclanthology.org/2022.wassa-1.17/ [44] M. Ramezani, M.-R. Feizi-Derakhshi, and M.-A. Balafar, “Text-based automatic personality prediction using kgrat-net: a knowledge graph attention network classifier,” Scientific Reports, vol. 12, no. 1, Dec. 2022. [Online]. Available: http://dx.doi.org/10.1038/s41598-022-25955-z [45] S. S. Bajestani, M. M. Khalilzadeh, M. Azarnoosh, and H. R. Kobravi, “Transentgat: A sentiment-based lexical psycholinguistic graph attention network for personality prediction,” IEEE Access, vol. 12, p. 59 630– 59 642, 2024. [46] M. Wu, Z. Xu, and L. Zheng, “Heterogeneous graph contrastive learning with adaptive data augmentation for semi-supervised short text classification,” Expert Systems, vol. 42, no. 2, Oct. 2024. [Online]. Available: http://dx.doi.org/10.1111/exsy.13744 [47] H. Zhu, X. Zhang, J. Lu, Y. Wu, Z. Bai, C. Min, L. Yang, B. Xu, D. Zhang, and H. Lin, “Enhancing textual personality detection toward social media: Integrating long-term and short-term perspectives,” 2024. [Online]. Available: https://arxiv.org/abs/2404.15067 [48] Y. Ji, W. Wu, H. Zheng, Y. Hu, X. Chen, and L. He, “Is chatgpt a good personality recognizer? a preliminary study,” 2023. [Online]. Available: https://arxiv.org/abs/2307.03952 [49] J. J. Tan, B.-H. Kwan, D. W.-K. Ng, Y.-C. Hum, A. Mokraoui, and S.-Y. Lo, “Prompting-in-a-series: Psychology-informed contents and embeddings for personality recognition with decoder-only models,” IEEE Transactions on Computational Social Systems, p. 1–15, 2025. [Online]. Available: http://dx.doi.org/10.1109/TCSS.2025.3593323 [50] J. W. Pennebaker and L. A. King, “Linguistic styles: Language use as an individual difference.” Journal of Personality and Social Psychology, vol. 77, p. 1296–1312, 1999. [51] M. J., “(mbti) myers-briggs personality type dataset,” 2017. [Online]. Available: https://w.kaggle.com/datasets/datasnaek/mbti-type [52] I. Zettler, I. Thielmann, B. E. Hilbig, and M. Moshagen, “The nomological net of the hexaco model of personality: A large-scale meta-analytic investigation,” Perspectives on Psychological Science, vol. 15, no. 3, p. 723–760, Apr. 2020. [Online]. Available: http://dx.doi.org/10.1177/1745691619895036 [53] I.Thielmann,M.Moshagen,B.Hilbig,andI.Zettler, “Onthecomparabilityofbasicpersonalitymodels:Meta- analytic correspondence, scope, and orthogonality of the big five and hexaco dimensions,” European Journal of Personality, vol. 36, no. 6, p. 870–900, Jun. 2021. [Online]. Available: http://dx.doi.org/10.1177/08902070211026793 [54] I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long- document transformer,” 2020. [Online]. Available: https://arxiv.org/abs/ 2004.05150 [55] Y. Mehta, S. Fatehi, A. Kazameini, C. Stachl, E. Cambria, and S. Eetemadi, “Bottom-up and top-down: Predicting personality with psy- cholinguistic and language model features,” in 2020 IEEE International Conference on Data Mining (ICDM), 2020, p. 1184–1189. [56] T. Yang, F. Yang, H. Ouyang, and X. Quan, “Psycholinguistic tripartite graph network for personality detection,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli, Eds.Online: Association for Computational Linguistics, Aug. 2021, p. 4229–4239. [Online]. Available: https: //aclanthology.org/2021.acl-long.326 [57] L. Hu, H. He, D. Wang, Z. Zhao, Y. Shao, and L. Nie, “Llm vs small model? large language model based text augmentation enhanced personality detection model,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, p. 18 234–18 242, Mar. 2024. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/ view/29782 [58] T. Yang, J. Deng, X. Quan, and Q. Wang, “Orders are unwanted: Dynamic deep graph convolutional network for personality detection,” 2022. [Online]. Available: https://arxiv.org/abs/2212.01515 AUTHOR BIOGRAPHIES Jing Jie Tan received his Bachelor of Computer Science (Hons) and PhD in Engineering, specializing in machine learning, from Universiti Tunku Abdul Rahman (UTAR). He is currently a Research Fellow at the National University of Singapore (NUS). His research interests include affective computing, health intelligence, natural language processing, computer vision, deep learning, large language models, and vi- sion transformers. He is passionate about translating research into practical applications, with the goal of leveraging technology to better serve society. Ban-Hoe Kwan received the Bachelor of Engineer- ing (Electrical), Master of Engineering Science, and Ph.D. degrees in Engineering from the University of Malaya (UM). He is currently an Associate Professor at Universiti Tunku Abdul Rahman (UTAR). His research interests include image processing, artificial intelligence, medical signal processing, the Internet of Things, and robotics. Danny Ng Wee Kiat (Senior Member, IEEE) re- ceived the Ph.D. degree in Engineering and is a reg- istered Professional Engineer with the Board of En- gineers Malaysia. He is currently an Assistant Pro- fessor at Universiti Tunku Abdul Rahman (UTAR). His research focuses on robotics and artificial intel- ligence, particularly AI-driven robotics, generative AI, and autonomous AI agents for industrial and enterprise applications. In addition to his academic role, he serves as the Chief Executive Officer and Chief Technology Officer of Netizen Robotics and as the Technical Director of Netizen Experience, where he leads initiatives in advanced robotics and digital transformation. 16 Yan Chai Hum (Senior Member, IEEE) received the Ph.D. degree in Engineering with specialization in artificial intelligence from Universiti Teknologi Malaysia. He is currently an Associate Professor with the Department of Mechatronics and Biomed- ical Engineering, Lee Kong Chian Faculty of Engi- neering and Science, Universiti Tunku Abdul Rah- man (UTAR). His research interests include artifi- cial intelligence in healthcare, biomedical imaging, computer vision, Internet of Things systems, and intelligent sensing technologies. His recent work focuses on AI-driven diagnostic systems, multimodal medical data analysis, and autonomous sensing platforms. He is also the founder of Promptiq Enterprise, an AI consultancy specializing in generative AI and intelligent systems. Shih-Yu Lo is an Associate Professor at the Institute of Communication Studies, National Yang Ming Chiao Tung University, Taiwan. His research inte- grates cognitive psychology and human–computer interaction, focusing on how emerging technologies such as AI, virtual reality, and social robots shape memory, decision-making, empathy, and social cog- nition. Po-An Chen is a Professor and Director at the Institute of Information Management, National Yang Ming Chiao Tung University, Taiwan. He is gener- ally interested in economics and computation, arti- ficial intelligence, and operations research, specifi- cally including algorithmic game theory, theoretical online/reinforcement learning, social networks, and multiagent and distributed systems. Noriyuki Kawarazaki received the Ph.D. degree in Engineering from Kyushu University. He is currently a Professor and Department Chair of Information Systems at Kanagawa Institute of Technology, Japan. His research interests include human-robot interac- tion, image processing, and artificial intelligence in robotics. Kosuke Takano received a B.A. degree in Environ- ment and Information Studies from Keio University, Japan, and his M.A. and Ph.D. degrees in Media and Governance from Keio University. He is cur- rently a Professor in the Department of Information and Computer Sciences at Kanagawa Institute of Technology, Japan. His research interests include emotional AI and multimodal affective computing, data management systems for real-world monitoring, and AI-driven educational systems. Anissa Mokraoui received the state engineering degree in electrical engineering from national school of telecommunications in 1989 from Algeria, the M.S degree in information technology in 1990 and the Ph.D. degree in 1994 both from University Paris 11, Orsay France. From 1992 to 1994, she worked at the National Institute of Telecommunications (INT, at Evry France) where her research activities were on digital signal processing, fast filtering algorithms and implementation problems on DSP. In 1997, she was appointed as assistant professor and in December 2011 as associate professor at Galilé Institute of University Paris 13, France. Since 2013, she is full professor at Galilée Institute of Université Sorbonne Paris Nord (USPN), France. From 2016 to 2024, she was the director of the L2TI laboratory of USPN. Her current research interests include source coding (image, video, multi-view, stereoscopic); joint source-channel-protocol decoding, robust mobile transmission, MIMO-OFDM channel estimation (massive), computer vision, few-shot object detection, cross-domain. She is co-author of more than 150 contributions to journals and conference proceedings. She served on program committees for conferences. She acts as a reviewer for many IEEE and EURASIP conferences and journals.