Paper deep dive
Revisiting Northrop Frye's Four Myths Theory with Large Language Models
Edirlei Soares de Lima, Marco A. Casanova, Antonio L. Furtado
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 3:04:44 AM
Summary
This paper proposes a character function framework based on Jungian archetypes to analyze Northrop Frye's four fundamental narrative genres (comedy, romance, tragedy, satire). The authors map four universal character functions (protagonist, mentor, antagonist, companion) to sixteen genre-specific roles. They validate this framework using six Large Language Models (LLMs) on a dataset of 40 narrative works, achieving a mean balanced accuracy of 82.5% and demonstrating that LLMs can effectively recognize and reject character-role correspondences, supporting the use of computational narratology for narrative analysis.
Entities (22)
Relation Signals (11)
Macbeth → exemplifiesgenre → Tragedy
confidence 95% · Autumn maturity tragedy Macbeth
Le Bourgeois Gentilhomme → exemplifiesgenre → Comedy
confidence 95% · Spring adolescence comedy Le Bourgeois Gentilhomme
Ramayana → exemplifiesgenre → Romance
confidence 95% · Summer youth romance Ramayana
1984 → exemplifiesgenre → Satire
confidence 95% · Winter senescence satire 1984
Northrop Frye → proposedtheory → Four Myths Theory
confidence 95% · Northrop Frye's theory of four fundamental narrative genres... has profoundly influenced literary criticism
Large Language Models → validated → Character Function Framework
confidence 92% · To validate this framework, we conducted a multi-model study using six state-of-the-art Large Language Models (LLMs)
Jungian Archetype Theory → influenced → Character Function Framework
confidence 90% · Drawing on Jungian archetype theory, we derive four universal character functions
Protagonist → specializedin → Satire
confidence 85% · Satire nonconformist... protagonist
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Northrop Frye's theory of four fundamental narrative genres (comedy, romance, tragedy, satire) has profoundly influenced literary criticism, yet computational approaches to his framework have focused primarily on narrative patterns rather than character functions. In this paper, we present a new character function framework that complements pattern-based analysis by examining how archetypal roles manifest differently across Frye's genres. Drawing on Jungian archetype theory, we derive four universal character functions (protagonist, mentor, antagonist, companion) by mapping them to Jung's psychic structure components. These functions are then specialized into sixteen genre-specific roles based on prototypical works. To validate this framework, we conducted a multi-model study using six state-of-the-art Large Language Models (LLMs) to evaluate character-role correspondences across 40 narrative works. The validation employed both positive samples (160 valid correspondences) and negative samples (30 invalid correspondences) to evaluate whether models both recognize valid correspondences and reject invalid ones. LLMs achieved substantial performance (mean balanced accuracy of 82.5%) with strong inter-model agreement (Fleiss' $\kappa$ = 0.600), demonstrating that the proposed correspondences capture systematic structural patterns. Performance varied by genre (ranging from 72.7% to 89.9%) and role (52.5% to 99.2%), with qualitative analysis revealing that variations reflect genuine narrative properties, including functional distribution in romance and deliberate archetypal subversion in satire. This character-based approach demonstrates the potential of LLM-supported methods for computational narratology and provides a foundation for future development of narrative generation methods and interactive storytelling applications.
Tags
Links
- Source: https://arxiv.org/abs/2602.15678v1
- Canonical: https://arxiv.org/abs/2602.15678v1
Trouble viewing inline? Open PDF directly →
Full Text
67,997 characters extracted from source content.
Expand or collapse full text
REVISITING NORTHROP FRYE’S FOUR MYTHS THEORY WITH LARGE LANGUAGE MODELS Edirlei Soares de Lima Academy for AI, Games and Media Breda University of Applied Sciences Breda, The Netherlands soaresdelima.e@buas.nl Marco A. Casanova Department of Informatics PUC-Rio Rio de Janeiro, Brazil casanova@inf.puc-rio.br Antonio L. Furtado Department of Informatics PUC-Rio Rio de Janeiro, Brazil furtado@inf.puc-rio.br ABSTRACT Northrop Frye’s theory of four fundamental narrative genres (comedy, romance, tragedy, satire) has profoundly influenced literary criticism, yet computational approaches to his framework have focused primarily on narrative patterns rather than character functions. In this paper, we present a new character function framework that complements pattern-based analysis by examining how archetypal roles manifest differently across Frye’s genres. Drawing on Jungian archetype theory, we derive four universal character functions (protagonist, mentor, antagonist, companion) by mapping them to Jung’s psychic structure components. These functions are then specialized into sixteen genre-specific roles based on prototypical works. To validate this framework, we conducted a multi-model study using six state-of-the-art Large Language Models (LLMs) to evaluate character-role correspondences across 40 narrative works. The validation employed both positive samples (160 valid correspondences) and negative samples (30 invalid correspondences) to evaluate whether models both recognize valid correspondences and reject invalid ones. LLMs achieved substantial performance (mean balanced accuracy of 82.5%) with strong inter-model agreement (Fleiss’κ= 0.600), demonstrating that the proposed correspondences capture systematic structural patterns. Performance varied by genre (ranging from 72.7% to 89.9%) and role (52.5% to 99.2%), with qualitative analysis revealing that variations reflect genuine narrative properties, including functional distribution in romance and deliberate archetypal subversion in satire. This character-based approach demonstrates the potential of LLM-supported methods for computational narratology and provides a foundation for future development of narrative generation methods and interactive storytelling applications. Keywords Jungian Archetypes·Computational Narratology·Character Functions·Literary Genres·Large Language Models 1 Introduction As part of our long-term project in computational narratology and interactive story composition, we return to Northrop Frye’s seminal work [10], focusing on his book’s third essay entitled “Archetypal Criticism: Theory of Myths”. Our objective is to model the character functions underlying the four fundamental genres that Frye associated with the seasons’ cyclic succession. We propose to complement our previous approach to Frye’s theory, formulated in terms of narrative patterns [6], by developing a character-based approach grounded in Carl Gustav Jung’s analytical psychology [12]. Similarly to the seasonal succession – starting with merry spring, followed by the plenitude of summer, then by the falling decline of autumn, and ending in winter’s desolation – Frye’s fundamental genres (comedy, romance, tragedy, satire) tell stories whose protagonists exhibit different degrees of power to influence action, combined with growing or decreasing happiness. The power of this seasonal metaphor extends beyond literary analysis: cyclic history researchers such as Oswald Spengler and Arnold Toynbee applied similar frameworks to civilizational rise and decline [23], demonstrating how recurring structural patterns can be identified across different domains through archetypal analysis. Just as these historians sought to identify the characteristic features distinguishing each historical phase, our character arXiv:2602.15678v1 [cs.CL] 17 Feb 2026 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models function approach seeks to identify the distinctive ways archetypal roles manifest across Frye’s fundamental genres. This requires moving beyond recognizing that genres differ to specifying how character functions operate differently within each genre’s structural logic. To achieve this specification, we identified four character functions through correspondence with Jungian psychic structure: protagonist, mentor, antagonist, and companion. We then specialized these functions by devising genre- specific role names that express how each function operates within a particular genre’s conventions. For each genre, we selected one prototypical work, identified characters embodying each function, and formulated role names expressing the specialized actions inherent to each function in that narrative context. To evaluate whether these character-role correspondences capture systematic patterns beyond the four prototypical works, we conducted a multi-model validation study using six state-of-the-art Large Language Models (LLMs) as analytical instruments. The validation employed both positive samples (valid character-role correspondences across 40 works) and negative samples (systematically introduced invalid correspondences) to assess whether the framework exhibits both descriptive adequacy (recognizing valid implementations across diverse narratives) and discriminative power (rejecting functionally incorrect assignments). Analysis of model reasoning patterns in error cases provides additional insight into the interpretive challenges inherent in character function identification and the boundaries of the functional role framework. The paper is organized as follows. Section 2 reviews Frye’s genre characterization and presents our character functions proposal grounded in Jungian archetypes. Section 3 describes the validation methodology and presents results examining overall performance, inter-model agreement, and performance patterns across genres and roles, followed by discussion of the findings. Section 4 provides concluding remarks. 2 The Character Functions Approach In the third essay of his book Anatomy of Criticism [10], entitled “Archetypal Criticism: Theory of Myths”, Northrop Frye associates each of four fundamental literary genres with a season of the year. The seasonal succession, in turn, corresponds to the progression of human life from youth to old age. Frye’s schema, together with exemplary works for each genre [16, 26, 22, 19], is presented in Table 1. Table 1: Frye’s four fundamental genres with their seasonal and life-stage correspondences, and examples of prototypical works. SeasonAgeGenreInstance SpringadolescencecomedyLe Bourgeois Gentilhomme SummeryouthromanceRamayana AutumnmaturitytragedyMacbeth Wintersenescencesatire1984 Comedies typically concern the efforts of an inexperienced young man to conquer a beloved damsel. A pompous older figure, known as an alazon in Greek comedy, attempts to block his pursuit, while another figure, remarkable for creative practical talent and known as an eiron, comes to his aid (Figaro, the resourceful servant, serving as a well-known example). These stories follow the spring phase of renewal and integration, leading to a happy ending in which the young protagonist overcomes old-fashioned privileges and social prejudices. Romances, in the original epic sense, involve a quest whose objective is to counteract a villainy or obtain something marvelous. A hero is called to undertake the quest, but may require the guidance or magical instruments that a wise sage provides. The hero’s inspiration and final reward is often the love of a highly virtuous woman. Tragedies are the sequel of a nefarious act (hamartia) that may be either mysteriously dictated by destiny or due to arrogant pride (hybris) deserving dire punishment (nemesis). Previous well-being is interrupted by the recognition of the act (anagnorisis) followed by reversal of fortune (peripeteia), as interpreted by oracular manifestation. The broken world order is painfully restored at the end. Satires denounce dystopian world order, in the extreme case being characterized by the “disappearance of the heroic” [10]. Protesters and their followers are punished for merely trying to behave normally by obedient guardians acting in the name of a supreme authority that can never be questioned. To identify what might be the decisive character functions, ideally common to all genres, we looked at what Jung claimed to be the parts of every individual’s psyche [12], namely: ego, persona, shadow, anima. The first two belong to consciousness – the individual is continuously aware of existence through the ego part, while the persona “is a compromise between the individual and society based on that which one appears to be” [12]. The last two are originally 2 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Figure 1: Diagram of Jung’s psychic structure showing the four components (ego, persona, shadow, anima) and their integration through the individuation process [12]. unconscious, the shadow constituting an alter ego wherein is kept whatever is rejected as negative, while the anima represents – for a male – a commonly negated feminine side. A crucial resource of Jung’s therapy is what he calls the individuation process, whereby the level of consciousness is raised in order to allow the patient to take fuller advantage of originally underdeveloped potentials. The first step is to break the persona limitations. Quite often one’s profession may exclusively determine what one believes to be. Next, shadow and anima, treated as archetypes, are brought to balanced consideration and their constructive traits are incorporated into consciousness. Another archetype then emerges as if by magic, the Wise Old Man, which “represents the cold and objective truth of nature ... in a variety of shapes ... well known from the world of primitives and from mythology” [12]. At this point the individuation process concludes, with the introduction of a new central personality focus, the self, which from then on supplants the ego (as graphically suggested in Figure 1). We are now in a position to introduce our proposed character functions, by simply attributing to separate characters what we have described as parts of the psyche of a single individual, as shown in Table 2. Note that in place of the persona – the discarded masking part of the psyche (recalling that “persona” meant the mask used by actors to project the sound of their voices and identify their role to the audience of ancient theater performances) – we have placed the revealing Wise Old Man archetype, where the “old” qualifier denotes superior wisdom not necessarily marked by advanced age. Table 2: Character functions derived from Jungian psychic structure. Part of psycheCharacter function Egoprotagonist Persona (→ Wise Old Man)mentor Shadowantagonist Animacompanion In summary, the protagonist is recognized and oriented by the mentor, opposed by the antagonist, and inspired by the companion. The terms "mentor" and "shadow" also figure in Campbell and Vogler [1,25], both of whom, like Frye himself [10], acknowledge their debt to Jung. The "anima" archetype, described as a soul image whose "eternal feminine" figure shows Gretchen raising Faust’s spirit to Heaven, is evoked in both Jungian studies [12] and in Frye’s book [10]. Additionally, Jung-oriented researchers Emma Jung and Marie-Louise von Franz [13] have argued that Grail romance narratives symbolically describe the entire individuation process, identifying the Grail as symbol of the Self. We finally proceeded to specialize these character functions into typical character roles, phrased to suggest how the functions might work differently in narratives pertaining to the different genres. To achieve this objective, we started from the prototype theory principle that classification is often achieved not through strict, necessary, and sufficient definitions but rather through similarity to a prototypical example [14]. The titles shown earlier in the “instances” column of Table 1, namely Le Bourgeois Gentilhomme, Ramayana, Macbeth, and 1984, were selected as prototypical based on their status as widely recognized exemplars of each genre. 3 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models The formulation of genre-specific role names required interpretive judgment about how character functions manifest in each prototype work. This interpretive dimension aligns with reader-response literary theory’s recognition [11] that literary analysis involves subjective engagement with texts rather than purely objective description. Based on analysis of the four prototypical works, we formulated role names expressing how each character function operates within genre-specific narrative contexts. The resulting role-character correspondences for the four prototypical works are listed below (in the format role:character): •Le Bourgeois Gentilhomme (Comedy) – lad in love:Cleonte; troubleshooter:Covielle; pompous blocker:Jourdain; witty damsel in love:Lucile • Ramayana (Romance) – hero:Rama; donor:Visvamitra; villain:Ravana; faithful victim:Sita • Macbeth (Tragedy) – ill-fated adventurer:Macbeth; soothsayer:three witches; order restorer:Macduff; ill-fated partner:Lady Macbeth •1984 (Satire) – nonconformist:Winston; inquisitor:O’Brien; dystopian idol:Big Brother; rebellion partner:Julia Table 3 presents these specialized role names organized by their corresponding character functions. Table 3: Specialized character roles for each of Frye’s four fundamental genres. GenreProtagonistMentorAntagonistCompanion Comedylad in lovetroubleshooterpompous blockerwitty damsel in love Romanceherodonorvillainfaithful victim Tragedyill-fated adventurersoothsayerorder restorerill-fated partner Satirenonconformistinquisitordystopian idolrebellion partner The resulting framework maps four universal character functions (protagonist, mentor, antagonist, companion) to sixteen genre-specific roles, as presented in Table 3. These specialized role names are designed to capture how archetypal functions manifest differently across Frye’s four fundamental genres. The framework’s validity depends on whether these prototype-derived correspondences generalize to other narratives within each genre. A theoretical consideration arises from Frye’s recognition of “intermediate phases” consisting of narratives that combine features of more than one genre [10]. Shakespeare’s tragedies provide a well-known example, as we are told in [9] how he “learned the difficult art of intensifying tragedy with comic relief”. Genres should therefore be understood as overlapping clusters around prototypical examples rather than categories with strict boundaries, with individual works exhibiting varying degrees of alignment with genre conventions. The following section presents a systematic validation study that evaluates whether the proposed character-role correspondences capture recognizable patterns across diverse narratives beyond the four prototypical works. 3 Experimental Validation To evaluate the validity of our character function framework beyond the four prototypical works from which it was derived, we conducted a multi-model validation study. Using six LLMs as analytical instruments, we assessed whether the proposed character-role correspondences capture systematic narrative patterns through evaluation on both positive samples (valid character-role correspondences across 40 works) and negative samples (30 invalid correspondences systematically introduced in the dataset). The experiment tests both the framework’s capacity to describe character functions across diverse narratives and its discriminative power to reject functionally incorrect assignments. 3.1 Validation Dataset 3.1.1 Sample Generation and Validation To construct a dataset for validating the proposed character functions, we employed a semi-automated approach combining LLM-assisted identification with expert validation. We utilized GPT-5.2 (gpt-5.2-2025-12-11) to identify literary works and propose character-role correspondences, which were then reviewed by the authors to verify that: (1) the work was a recognized exemplar of the specified genre; and (2) the proposed character-role assignments reflected each character’s primary structural function in the narrative. We acknowledge that this validation process involves interpretive judgment; however, the purpose of this dataset is not to establish definitive “ground truth” but rather to provide a consistent set of theoretically motivated correspondences against which to evaluate whether multiple independent LLMs recognize similar functional patterns. The subsequent multi-model evaluation thus tests whether our framework-based interpretations align with patterns recognizable to different models trained on diverse literary corpora. 4 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models For each of the four fundamental genres (comedy, romance, tragedy, and satire), we identified 10 exemplar works through an iterative process. We provided the model with the genre and its four corresponding functional roles from Table 3, instructing it to propose a work from that genre and identify which characters embodied each functional role. The model was explicitly directed to focus on structural character functions (what characters do in the story) rather than psychological complexity or thematic interpretation, aligning with our emphasis on plot-level functional analysis derived from structural narratology frameworks. The complete generation prompt is provided in Appendix A. Works were identified individually to ensure independent selection without contextual influence from previously suggested examples. For each proposed work, we verified the two criteria mentioned above, repeating the process when proposed correspondences were ambiguous or incorrect. This approach ensured dataset quality and theoretical coherence while exploring the model’s domain knowledge to identify diverse works spanning different time periods, cultural contexts, and narrative media. The resulting dataset encompasses theatrical works (e.g., Shakespeare’s Much Ado About Nothing, Molière’s Le Bourgeois Gentilhomme), classic literature (e.g., Austen’s Pride and Prejudice, Orwell’s 1984), ancient epics (e.g., Valmiki’s Ramayana, Sophocles’ Oedipus Rex), and modern film and television (e.g., The Hunger Games, the 1967 TV series The Prisoner). This range ensures that our functional role framework is evaluated across varied narrative implementations rather than being constrained to specific historical or cultural conventions. The complete dataset comprises 40 works (10 per genre), yielding 160 character-role correspondences (40 works×4 roles per work). The full list of works with their validated character-role assignments is provided in Appendix B. We refer to this dataset as the positive samples, as these correspondences represent theoretically valid functional assignments according to our framework. 3.1.2 Negative Sample Construction To evaluate whether LLMs can discriminate between valid and invalid character-role correspondences, we constructed a dataset of negative samples by introducing errors into a subset of the positive correspondences. Similar to the positive sample identification process, we employed GPT-5.2 (gpt-5.2-2025-12-11) to propose incorrect character-role mappings. From the 40 works in the positive dataset, we randomly selected 20 works (5 per genre, representing 50% of the positive samples) for negative sample construction. This sampling approach ensures balanced genre representation while maintaining a substantial evaluation set for assessing discriminative ability. For each selected work, we generated an alternative set of four character-role mappings where one or two roles were reassigned to functionally incorrect characters, depending on the error strategy employed, while the remaining roles retained their correct assignments from the positive samples. We provided the model with the correct character-role correspondences from the positive sample and instructed it to create incorrect mappings by applying one of two predefined error strategies. The complete negative sample generation prompt is provided in Appendix C. The two substitution strategies were designed to create plausible but functionally incorrect correspondences that would meaningfully challenge the models’ understanding of character roles. These strategies reflect two common interpretive errors in character function analysis: 1.Primary-secondary character swap: Replacing the primary bearer of a role with a secondary or supporting character who performs related but subordinate functions within the narrative. For example, in Hamlet, substituting Rosencrantz for Ophelia as the “ill-fated partner” creates a correspondence where a minor courtier is incorrectly elevated over the character whose tragic fate is central to the protagonist’s downfall. 2. Inter-role character swap: Exchanging characters between different roles within the same story. For example, in V for Vendetta, swapping V and Evey’s functional assignments incorrectly positions Evey as the “nonconformist” (protagonist function) while demoting V to “rebellion partner” (companion function), inverting their actual structural primacy in the narrative. Similar to the positive sample curation, the authors reviewed each generated negative sample to ensure that: (1) the proposed incorrect correspondence represented a genuine functional error according to our framework; (2) the error was plausible enough to require analytical discrimination rather than being trivially detectable; and (3) the substituted character actually existed in the work and had some narrative connection to the functional domain being tested. The final negative dataset comprises 20 works with modified correspondences (5 per genre), yielding 30 incorrect character-role assignments: 20 from inter-role swap cases (10 stories×2 swapped roles) and 10 from primary-secondary character swap cases (10 stories×1 substituted role). The full list of negative samples with their incorrect character-role assignments is provided in Appendix D. 5 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models 3.2 Evaluation Protocol To ensure reliable validation of our proposed character-role correspondences, we employed a multi-LLM evaluation approach. Since recognizing character roles in narrative works is a complex analytical task requiring understanding of narrative structure and plot-level functions, we evaluated correspondences across multiple models with different architectures and training data. This approach provides diverse analytical perspectives beyond the authors’ interpretations while avoiding reliance on the biases or limitations of any single model. Consistent recognition or rejection of character-role correspondences across architecturally diverse models strengthens confidence that our framework captures systematic structural patterns rather than model-specific artifacts. 3.2.1 Model Selection We selected six state-of-the-art LLMs for the validation experiment. The evaluated LLMs encompassed both proprietary and open models from multiple providers, including Anthropic, DeepSeek, Google, OpenAI, Meta, and Alibaba, representing diverse architectural approaches and training methodologies. The complete list of models, including their providers and specific versions, is presented in Table 4. Table 4: Overview of the LLMs used in the validation study. ModelDescriptionVersion Claude Opus 4.5Flagship model for complex tasks by Anthropic.claude-opus-4-5-20251101 DeepSeek R1 (671B)671-billion-parameter reasoning model by DeepSeek.deepseek-r1:671b Gemini 2.5 ProAdvanced multimodal model by Google.gemini-2.5-pro GPT OSS (120B)120-billion-parameter open model by OpenAI.gpt-oss:120b Llama 4 (128B-A17B)128-billion-parameter mixture model by Meta.llama4:128x17b Qwen 3 (235B)235-billion-parameter model by Alibaba.qwen3:235b Proprietary models were accessed via their respective APIs, while open-source models were executed locally using a self-hosted Ollama server. 1 All models were queried using identical parameters to ensure uniform evaluation conditions: temperature was set to 0.0 to promote deterministic outputs and maximize reproducibility, while all other parameters retained their default values. Notably, we excluded GPT-5.2 from the validation analysis despite its use in positive sample identification. While the expert curation process ensured that accepted correspondences reflected real examples rather than GPT-5.2’s specific patterns, we conservatively excluded this model from validation to eliminate any potential for circularity bias. 3.2.2 Prompt Design The evaluation task required LLMs to analyze character-role correspondences and determine whether each assignment could be justified based on the character’s functional role in the narrative. The prompt was designed to focus model attention on structural functional analysis rather than thematic or psychological character interpretation, with particular emphasis on identifying the primary bearer of each role rather than accepting any character who performs related actions. LLMs were provided with the story title, genre classification, and four character-role correspondences to evaluate. For each correspondence, models were instructed to determine whether the character served as the primary bearer of that role’s function in the narrative structure. Responses included a binary justification judgment and a brief reasoning explanation (2-3 sentences). This structured output format enabled both quantitative performance measurement through the binary judgments and qualitative analysis of model reasoning patterns through the explanations. The complete evaluation prompt and system prompt, including the specific output format specification, are provided in Appendix E. 3.2.3 Data Collection and Processing For each of the 40 works in the positive dataset and 20 works in the negative dataset, we generated character-role evaluations from all six models, resulting in a total of 360 evaluation instances (240 from positive samples and 120 from negative samples). Each positive sample evaluation consisted of four individual character-role assessments (one per 1 https://ollama.com/ 6 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models functional role), yielding 960 judgments across all models (40 works×4 roles×6 models). For negative samples, we evaluated only the altered character-role correspondences: in inter-role swap cases (10 stories), both swapped roles were evaluated (20 roles×6 models = 120 judgments), while in primary-secondary character swap cases (10 stories), only the single substituted role was assessed (10 roles×6 models = 60 judgments). This yielded 1,140 total role-specific judgments for analysis (960 + 120 + 60). We automatically parsed model responses to extract the structured justification judgment for each character-role correspondence, along with the accompanying reasoning explanation. For positive samples, a response was scored as correct (True Positive) if the model judged the correspondence as justified, indicating recognition of a valid functional assignment. For negative samples, only the altered correspondences were evaluated: a correct response (True Negative) required the model to judge the correspondence as not justified, indicating successful discrimination of the functionally incorrect assignment. The unchanged correspondences in each negative sample were excluded from analysis to prevent repeated evaluation of valid correspondences already assessed in the positive samples. 3.2.4 Evaluation Metrics To assess model performance in validating our functional role framework, we computed classification metrics that evaluate both the ability to recognize valid character-role correspondences and to reject invalid ones. We evaluated three primary metrics: • Recall: The proportion of valid correspondences correctly identified, measuring the model’s ability to recognize functional role fulfillment. Formally,Recall = T P T P+F N , whereTPrepresents true positives and FN represents false negatives. •Specificity: The proportion of invalid correspondences correctly rejected, measuring the model’s ability to discriminate functionally incorrect assignments. Formally,Specificity = T N T N+F P , whereTNrepresents true negatives and FP represents false positives. •Balanced Accuracy: The mean of recall and specificity, providing an overall performance measure that treats both recognition and discrimination capabilities equally. Formally, Balanced Accuracy = Recall+Specificity 2 . Balanced Accuracy serves as our primary performance indicator, as it captures both essential aspects of framework validation: recognizing valid correspondences and rejecting invalid ones. Both false acceptances (incorrectly accepting invalid correspondences) and false rejections (incorrectly rejecting valid correspondences) equally undermine confidence in the theoretical framework, making it important to weight recall and specificity equally regardless of sample sizes. Note that random guessing would yield a balanced accuracy of 50%, providing a baseline for interpretation. Performance was analyzed at three levels: (1) overall performance across all correspondences; (2) genre-specific performance to identify whether certain narrative structures present greater validation challenges; and (3) role-specific performance to determine whether particular functional roles prove more difficult to validate. Additionally, we conducted qualitative analysis of model reasoning patterns to understand the factors contributing to incorrect classifications. 3.3 Results and Discussion 3.3.1 Overall Performance Table 5 presents the overall validation performance across all six evaluated models. All models demonstrated strong discriminative ability, with balanced accuracy ranging from 79.9% to 85.3%, substantially exceeding the random baseline of 50% by at least 29 percentage points. Table 5: Overall performance metrics for all evaluated models. Each model evaluated 160 positive correspondences and 30 negative correspondences (190 total per model, 1,140 across all models). ModelRecall (%)Specificity (%)Balanced Accuracy (%)TPFNTNFP Claude Opus 4.580.690.085.312931273 Qwen 3 (235B)76.993.385.112337282 Gemini 2.5 Pro76.290.083.112238273 GPT OSS (120B)71.990.080.911545273 Llama 4 (128B-A17B)78.183.380.712535255 DeepSeek R1 (671B)73.186.779.911743264 Mean± SD76.1± 3.286.5± 3.681.4± 2.0 7 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Claude Opus 4.5 achieved the highest overall performance (balanced accuracy of 85.3%), demonstrating both strong recall (80.6%) and specificity (90.0%). This represents a near-optimal balance between recognizing valid correspon- dences and rejecting invalid ones. Qwen 3 (235B) followed closely with 85.1% balanced accuracy, while Gemini 2.5 Pro achieved 83.1%. The remaining models ranged from 79.9% to 80.9%. An interesting pattern emerges when examining the recall-specificity trade-off across models. Most models exhibited higher specificity than recall, correctly rejecting invalid assignments (mean specificity of 88.9%) more often than correctly identifying all valid ones (mean recall of 76.1%). This pattern is most pronounced in Qwen 3 (235B), which achieved 93.3% specificity while maintaining 76.9% recall, and GPT OSS (120B) with 90.0% specificity and 71.9% recall. In contrast, Claude Opus 4.5 and Llama 4 maintained more balanced profiles, with recall rates closer to their specificity rates. Performance remained consistent across the six models, with a standard deviation in balanced accuracy of 2.1%. This consistency across architecturally diverse models suggests that evaluation patterns may reflect systematic properties of the narrative data. To examine this hypothesis more rigorously, we next analyze the extent to which models agree on individual character-role assessments. 3.3.2 Model Agreement and Consensus Table 6 presents inter-model agreement metrics across all 190 character-role evaluations. Table 6: Summary of inter-model agreement metrics. MetricValue Mean pairwise agreement82.0%± 3.4% Pairwise agreement range76.3% - 87.4% Unanimous agreement (6/6)113/190 (59.5%) Majority agreement (4+/6)159/190 (83.7%) Split decisions (3/3)11/190 (5.8%) Fleiss’ Kappa (κ)0.600 Models demonstrated strong agreement across all metrics. Pairwise agreement averaged 82.0% (SD = 3.4%), with all 15 model pairs exceeding 75% agreement (see the complete table in Appendix F). Examining consensus across all six models simultaneously, 113 judgments (59.5%) achieved unanimous agreement, with an additional 46 judgments (24.2%) reaching majority consensus where 4-5 models agreed. Only 11 judgments (5.8%) produced split decisions with models dividing evenly. Overall, 83.7% of judgments showed clear majority agreement (4 or more models concurring). To quantify inter-rater reliability, we calculated Fleiss’ Kappa, which adjusts for chance agreement when multiple raters evaluate categorical data. The analysis yieldedκ = 0.600, indicating substantial inter-rater reliability according to standard interpretation guidelines [15]. This level of agreement substantially exceeds what would be expected by random chance (expected: 55.0%, observed: 82.0%). 3.3.3 Performance by Genre The substantial inter-model agreement establishes that models converge on similar character-role assessments. Therefore, we aggregate evaluations across all six models to examine whether performance varies systematically across Frye’s four fundamental genres. Table 7 presents the aggregated metrics for each genre. Table 7: Performance metrics by genre, aggregated across all six models. Each genre includes 240 positive judgments (10 works× 4 roles× 6 models) and 42-48 negative judgments depending on error distribution. GenreRecall (%)Specificity (%)Balanced Accuracy (%)PositiveNegativeTotal Tragedy82.197.689.924042282 Comedy79.690.585.024042282 Romance67.989.678.824048288 Satire75.079.277.124048288 Performance varied notably across genres, spanning 12.8 percentage points from 77.1% (satire) to 89.9% (tragedy). Tragedy and comedy, representing Frye’s autumn and spring phases respectively, achieved the highest performance with both strong recall and high specificity, suggesting that character-role correspondences in these genres present well-defined functional patterns. Romance maintained high specificity (89.6%) but exhibited the lowest recall (67.9%), 8 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models indicating that models were conservative in accepting correspondences – correctly rejecting invalid assignments but missing some valid ones. This pattern suggests that while incorrect romance correspondences are readily identifiable, valid correspondences may involve more subtle or varied functional implementations. Satire stands out with relatively balanced recall (75.0%) and specificity (79.2%), but both metrics substantially lower than other genres, suggesting that character functions in satirical narratives may be more ambiguous or context-dependent, challenging models to discriminate between valid and invalid correspondences. 3.3.4 Performance by Role The substantial inter-model agreement also enables examining performance patterns across the 16 specialized character roles derived from the four functional archetypes. Table 8 presents performance metrics aggregated across all six models for each role. Table 8: Performance metrics by role, aggregated across all six models. Each role includes 60 positive judgments (10 genre-specific works× 6 models). Negative judgments vary by role (6-18) based on error distribution in the negative sample construction. RoleGenreRecall (%)Specificity (%)Balanced Accuracy (%)PositiveNegative Protagonist function NonconformistSatire98.3100.099.26012 Ill-fated adventurerTragedy93.3100.096.7606 Lad in loveComedy85.0100.092.56012 HeroRomance78.3100.089.2606 Mentor function TroubleshooterComedy85.083.384.26012 DonorRomance61.7100.080.86012 SoothsayerTragedy70.083.376.7606 InquisitorSatire78.372.275.36018 Antagonist function Order restorerTragedy83.3100.091.76012 Dystopian idolSatire76.7100.088.3606 VillainRomance75.0100.087.56012 Pompous blockerComedy78.383.380.8606 Companion function Ill-fated partnerTragedy81.7100.090.86018 Witty damsel in loveComedy70.091.780.86012 Faithful victimRomance56.772.264.46018 Rebellion partnerSatire46.758.352.56012 Performance varied substantially across roles, with balanced accuracy ranging from 52.5% to 99.2%. Protagonist- function roles consistently achieved the highest performance across all genres, with all four exhibiting perfect specificity (100.0%) and strong recall (78.3% - 98.3%). In contrast, companion roles exhibited the widest performance variation (52.5% - 90.8%), spanning 38.3 percentage points. Eight roles achieved perfect specificity (100.0%), all belonging to either protagonist or antagonist functions. However, several roles exhibited notable recall-specificity gaps, with “donor” (romance) showing the largest disparity at 38.3 percentage points (61.7% recall vs. 100.0% specificity). This pattern was also observed in “villain” (romance) and “dystopian idol” (satire), where models more consistently rejected invalid correspondences than identified all valid ones. The lowest-performing roles cluster within specific genre-function combinations: “rebellion partner” (satire, 52.5%), “faithful victim” (romance, 64.4%), and “inquisitor” (satire, 75.3%). These three roles all belong to either companion or mentor functions in romance and satire genres, suggesting differences in performance across genre-function pairings. 3.3.5 Discussion The validation results support the theoretical premise that Jungian archetypes, when specialized according to Frye’s genre distinctions, provide a promising framework for analyzing character functions in narrative works. LLMs achieved substantial performance (mean balanced accuracy of 82.5%) with strong inter-model agreement (Fleiss’κ= 0.600 and mean pairwise agreement of 82.0%). The consistency of this performance across six architecturally diverse models suggests that the observed patterns reflect structural properties of the narrative data rather than model-specific biases. Moreover, the character-role correspondences derived from four prototypical works generalized effectively across forty 9 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models diverse narratives spanning different time periods, cultural contexts, and media, suggesting that the framework captures fundamental rather than idiosyncratic patterns of character function. The genre-level performance differences align with Frye’s characterization of these narrative forms. Tragedy and comedy follow relatively structured narrative arcs with clearly defined character functions: tragedy progresses from prosperity through hamartia to catastrophe and the restoration of moral order, while comedy moves from social obstruction to integration and renewal [10]. In tragedy, the order restorer serves a clearly delineated function – restoring the moral order disrupted by the protagonist’s transgression – while in comedy, the troubleshooter operates within well-established conventions to overcome the blocking figure’s obstruction. These established structural patterns create consistent expectations for how characters fulfill their functions, facilitating reliable character-role identification. Romance and satire exhibit structural properties that introduce ambiguity in character-role identification. Romance’s quest structure exhibits the variability that Propp observed in folktale morphology, where not all narrative functions manifest in every story and helper figures take diverse forms [21]. A hero may receive aid from a sage who provides counsel, from magical objects without a clear donor figure, or may overcome obstacles through inherent prowess without distinct antagonistic characters. When functions can be distributed across multiple characters, absent entirely, or fulfilled through non-character elements, determining which character serves as the primary bearer of a role becomes ambiguous. Qualitative examination of model reasoning patterns for the “donor” role (which achieved perfect specificity but only 61.7% recall) provides concrete evidence for the functional distribution characteristic of romance narratives. When evaluating Mrs. Fairfax as donor in Jane Eyre, multiple models identified alternative donors serving different quest phases: Claude Opus 4.5 noted “Miss Temple (who nurtures Jane’s education)” and “Jane’s uncle who leaves her the inheritance grants her independence”, while Qwen 3 emphasized that “the primary donor role belongs to Jane’s uncle John Eyre, who grants her financial independence”. Similarly, for Visvamitra in the Ramayana, Gemini 2.5 Pro observed he provides “divine weapons” but is “preliminary to the main plot”, with “more central donors appear[ing] later, such as Sugriva who provides the army... and Vibhishana who provides the critical intelligence”. When the donor function is distributed across multiple characters and quest stages, determining which character serves as the primary bearer of the role becomes inherently ambiguous. Satire exhibits a different form of ambiguity: Frye characterizes it as marked by the “disappearance of the heroic” [10], where protagonists lack the agency of traditional heroes and archetypal distinctions deliberately blur. Characters often fulfill functions in subverted or merged ways: a figure labeled “inquisitor” (our mentor-function role for satire) may simultaneously guide and manipulate the protagonist, as O’Brien does to Winston in 1984. This functional ambiguity creates legitimate disagreement about whether a character truly fulfills a role: is O’Brien primarily an inquisitor who reveals truth, or has the mentor function collapsed entirely into the antagonistic dystopian system? When archetypal functions are deliberately subverted or merged, even human readers might disagree on character-role assignments, making lower model performance an expected reflection of the genre’s inherent interpretive ambiguity. Examination of model reasoning for the “inquisitor” role confirms this interpretation. When evaluating Effie Trinket as inquisitor in The Hunger Games, multiple LLMs reasoned that while she enforces Capitol norms, the inquisitorial function more properly belongs to other entities: Claude Opus 4.5 identified “characters who question or probe the protagonists’ motives”, DeepSeek R1 noted “the primary inquisitor function belongs to antagonists like Seneca Crane (Gamemaker surveillance) and President Snow himself”, and Gemini 2.5 Pro observed that “the inquisitorial function... is performed by Capitol officials such as President Snow and the Gamemakers”. For Clevinger in Catch-22, Gemini 2.5 Pro noted he is “the victim of an inquisition, not the inquisitor himself”, while GPT OSS observed he is “more a victim of bureaucratic interrogation than the agent who conducts it”. LLMs disagreed not because they failed to understand character functions, but because satire’s deliberate subversion creates legitimate ambiguity about which character (if any single character) primarily embodies a given function. The role-level patterns reveal that performance differences stem from both narrative structure and framework design choices. The consistently high performance of protagonist and antagonist roles across genres likely reflects their structural necessity: protagonists drive narrative action while antagonists create conflict, making these functions plot-central and relatively invariant across different stories. Companion functions, conversely, operate more at the relational and thematic level. A companion’s primary contribution often lies in emotional support, moral influence, or thematic counterpoint rather than plot mechanics, allowing greater variation in how and whether these functions manifest as distinct character roles. The clustering of low performance in romance and satire companion roles (faithful victim with 64.4% and rebellion partner with 52.5%) may therefore reflect genuine structural differences in how these genres employ companion functions rather than inconsistent role characterizations. Performance differences also reflect our framework design choices in formulating role names. The specialized role names vary considerably in their semantic specificity and clarity. Terms such as “hero” and “villain” invoke well- established archetypal concepts with relatively stable meanings, while “rebellion partner” or “witty damsel in love” combine multiple qualifiers that narrow the concept but potentially introduce interpretive ambiguity. The “donor” 10 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models role borrows directly from Propp’s established terminology [21], carrying the weight of a defined narratological function, while “dystopian idol” represents a more novel construction. These differences in label clarity may contribute to performance variation: roles with precise, established labels may facilitate recognition, while roles requiring interpretation of multiple qualifiers or novel combinations may be more difficult to generalize in multiple narrative implementations. Examination of model reasoning for the “rebellion partner” role (52.5% balanced accuracy, the lowest overall) illustrates this challenge. When evaluating Clarisse McClellan in Fahrenheit 451, LLMs consistently reasoned that while she “catalyzes Montag’s transformation” (Claude Opus 4.5, DeepSeek R1, Gemini 2.5 Pro), she “disappears early and does not actively participate in his rebellion” (DeepSeek R1), with Faber serving as the “operational partner” (DeepSeek R1) or “actively collaborat[ing] with Montag, provid[ing] a strategic plan” (Gemini 2.5 Pro). For Evey Hammond in V for Vendetta, Gemini 2.5 Pro noted her “primary function is not that of a partner but of a successor”, while Llama 4 observed she is “more supportive and developmental” rather than an equal collaborator. Models grappled with whether partnership requires sustained collaboration, ideological equality, and shared agency, all questions without clear and consistent answers. Notably, in Brave New World, five of six models accepted Helmholtz Watson instead of the designated Bernard Marx, with reasoning emphasizing that Helmholtz “joins John in both intellectual dissent and active, physical rebellion” (Gemini 2.5 Pro) while Bernard acts from “self-interest and cowardice” (Claude Opus 4.5, Gemini 2.5 Pro), revealing that models evaluated partnership quality (authenticity of commitment and depth of collaboration) rather than simply matching character to role labels. This reflects an inherent trade-off in our approach: genre-tailored role names capture the nuances Frye identifies in how archetypes manifest differently across genres, but this specificity comes at the cost of reduced generality. 4 Concluding Remarks This work establishes a character function framework that bridges Frye’s literary genre theory with computational narrative analysis. By deriving four universal functions from Jungian archetypes and specializing them into sixteen genre-specific roles, we provide a formalized structure for representing how characters operate within different narrative genres. The validation demonstrates that these character-role correspondences capture systematic patterns across diverse works. LLMs successfully recognized valid correspondences while rejecting invalid ones, with performance variations reflecting genuine structural differences in how character functions manifest across genres. This framework offers computational narratology a practical resource for narrative analysis and generation, with the specialized roles providing explicit functional specifications that can guide character behavior in interactive storytelling systems, inform automated story analysis, or serve as constraints in narrative generation algorithms. While the validation demonstrates the framework’s viability, several considerations define its current scope and indicate possible directions for future research. Although our validation included narratives with both male and female protagonists, the theoretical foundation in Jung’s male-centered archetype theory raises questions about gender-specific manifestations of character functions. Recent research has demonstrated that LLMs exhibit gender biases in narrative generation, often using protagonist gender as a heuristic for interpreting narrative structures [24]. In this context, Jung identifies different psychic structures for women (animus rather than anima, Great Mother rather than Wise Old Man), which could theoretically affect how archetypal functions map to character roles in female-centered narratives. Whether these theoretical differences manifest as systematic patterns in actual narrative analysis remains an empirical question for future investigation, potentially drawing on feminist revisions of archetypal theory such as Murdock’s heroine’s journey [17] or Pratt’s work on archetypal patterns in women’s fiction [20]. Additionally, the specialized role names reflect interpretive choices made through analysis of the four prototypical works. While the validation demonstrates that these prototype-derived roles generalize across diverse works within each genre, alternative prototype selections might yield different role formulations that prove equally valid. The observed trade-off between genre-specific precision and cross-genre generality suggests that role name specificity, while capturing important genre distinctions, may limit broader applicability. Future work could also expand the functional scope to encompass additional character types that Frye identifies as intensifying genre-specific emotional effects. While our framework focuses on the four archetypal functions necessary to drive narrative action (protagonist, mentor, antagonist, companion), Frye notes that each genre employs supplementary characters who amplify its characteristic mood. In comedy, buffoons and churls complement the eiron (our troubleshooter) and alazon (pompous blocker) to polarize the comic atmosphere [10]. In romance, magical helpers beyond the primary donor perform extraordinary feats to render the hero’s achievement more remarkable [21]. Tragedy may feature denouncers or supernatural agents who intensify the emotions of fear and pity that Aristotle identified as essential to tragic catharsis [18], while satire often employs hate symbol or scapegoat figures who concentrate blame and deflect attention from systemic failures, intensifying the genre’s critical mood. Investigating whether these mood-intensifying characters follow systematic patterns across genres, and how they interact with the core functional 11 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models roles, represents a promising direction for enriching the character function framework while maintaining its theoretical grounding in archetypal structures. The proposed framework provides computational narratology with a theoretically grounded resource for character-based narrative generation and analysis. The sixteen specialized roles offer explicit functional specifications that can support the development of computational narrative generation methods and interactive storytelling applications. Previous work has demonstrated that character-based interactive storytelling can be achieved through multi-agent planning systems, where autonomous character agents interact within plot management frameworks to generate coherent narratives in highly interactive game environments [3,2]. The framework proposed in this work can extend these approaches by providing genre-specific functional roles that LLM-based character agents could embody. This character function approach complements our recent work on LLM-based narrative generation methods that employ semiotic reconstruction [7,4] and narrative pattern guidance [5,8], enabling systems to operate at both the pattern level (plot structure) and the character level (functional roles). The integration of character functions with pattern-based methods advances toward more sophisticated computational narrative systems that maintain both structural coherence and genre-appropriate character behavior. This work demonstrates that computational methods can engage productively with established literary theory, not merely applying humanistic concepts to computational tasks but using computational validation to illuminate the theoretical constructs themselves. The convergence of multiple independent LLMs on similar character-role assessments, combined with the interpretive insights revealed through error analysis, suggests that LLM-supported methods offer computational narratology valuable instruments for testing, refining, and extending narratological frameworks. As these methods mature, the integration of theoretically grounded frameworks with computational validation may bridge the longstanding gap between humanistic literary scholarship and computational approaches to narrative generation. References [1] J. Campbell. The Hero with a Thousand Faces. New World Library, Novato, California, 2008. [2]E. S. de Lima, B. Feijó, and A. L. Furtado. A character-based model for interactive storytelling in games. In 2022 21st Brazilian Symposium on Computer Games and Digital Entertainment (SBGames), pages 1–6, 2022. doi: 10.1109/SBGAMES56371.2022.9961071. [3]E. S. de Lima, B. Feijó, and A. L. Furtado. Managing the plot structure of character-based interactive narratives in games. Entertainment Computing, 47:100590, 2023. ISSN 1875-9521. doi: 10.1016/j.entcom.2023.100590. [4]E. S. De Lima, B. Feijó, M. A. Cassanova, and A. L. Furtado. ChatGeppetto - an AI-powered storyteller. In Proceedings of the 22nd Brazilian Symposium on Games and Digital Entertainment, page 28–37, New York, NY, USA, 2024. ACM. doi: 10.1145/3631085.3631302. [5]E. S. de Lima, M. M. E. Neggers, M. A. Casanova, B. Feijó, and A. L. Furtado. A pattern-oriented AI-powered approach to story composition. In P. Figueroa, A. Di Iorio, D. Guzman del Rio, E. W. Gonzalez Clua, and L. Cuevas Rodriguez, editors, Entertainment Computing – ICEC 2024, pages 1–16. Springer Cham, 2024. doi: 10.1007/978-3-031-74353-5_10. [6] E. S. de Lima, M. M. E. Neggers, and A. L. Furtado. Multigenre ai-powered story composition, 2024. [7]E. S. de Lima, M. M. Neggers, B. Feijó, M. A. Casanova, and A. L. Furtado. An AI-powered approach to the semiotic reconstruction of narratives. Entertainment Computing, 52:100810, 2025. doi: 10.1016/j.entcom.2024. 100810. [8]E. S. de Lima, M. M. E. Neggers, M. A. Casanova, and A. L. Furtado. From images to stories: Exploring player-driven narratives in games. In A. Marto, R. Prada, P. Gouveia, R. C. Espinosa, A. Gonçalves, E. Abrantes, and R. Ribeiro, editors, Videogame Sciences and Arts, pages 228–242, Cham, 2025. Springer Nature Switzerland. [9]W. Durant and A. Durant. The Age of Reason Begins: A History of European Civilization in the Period of Shakespeare, Bacon, Montaigne, Rembrandt, Galileo, and Descartes: 1558–1648. Simon and Schuster, New York, 1961. [10] N. Frye. Anatomy of Criticism: Four Essays. Princeton University Press, Princeton, New Jersey, 2020. [11]W. Iser. The Act of Reading: A Theory of Aesthetic Response. Johns Hopkins University Press, Baltimore, Maryland, 1978. [12] J. Jacobi. Psychology of C. G. Jung. Routledge, London, 2013. [13] E. Jung and M.-L. von Franz. The Grail Legend. Princeton University Press, Princeton, New Jersey, 1998. [14] G. Lakoff. Women, Fire, and Dangerous Things: What Categories Reveal about the Mind. University of Chicago Press, Chicago, Illinois, 1990. 12 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models [15]J. R. Landis and G. G. Koch. The measurement of observer agreement for categorical data. Biometrics, 33(1): 159–174, March 1977. doi: 10.2307/2529310. [16] Molière. Le Bourgeois Gentilhomme. Folio, Paris, 2013. [17] M. Murdock. The Heroine’s Journey: Woman’s Quest for Wholeness. Shambhala, Boston, Massachusetts, 1990. [18]P. Murray and T. S. Dorsch, editors. Classical Literary Criticism. Penguin Classics. Penguin Books, London, 2001. [19] G. Orwell. Nineteen Eighty-Four. Wordsworth Editions, Ware, Hertfordshire, 2021. [20] A. Pratt. Archetypal Patterns in Women’s Fiction. Indiana University Press, Bloomington, Indiana, 1981. [21] V. Propp. Morphology of the Folktale. University of Texas Press, Austin, Texas, 1968. [22] W. Shakespeare. Macbeth. SeaWolf Press, 2022. [23] P. A. Sorokin. Social Philosophies of an Age of Crisis. Beacon Press, Boston, Massachusetts, 1950. [24]I. C. van Blerck, E. S. de Lima, M. M. Neggers, and T. Calders. Unveiling gender bias in LLM-generated hero and heroine narratives. Entertainment Computing, 55:100972, 2025. ISSN 1875-9521. doi: 10.1016/j.entcom. 2025.100972. [25]C. Vogler. The Writer’s Journey: Mythic Structure for Writers. Michael Wiese Productions, Studio City, California, 2007. [26]V ̄ alm ̄ ıki. The R ̄ am ̄ ayan . a of V ̄ alm ̄ ıki: The Complete English Translation. Princeton Library of Asian Translations. Princeton University Press, Princeton, New Jersey, 2022. 13 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix A Work Identification Prompt A.1 System Prompt You are a narratology expert identifying works that exemplify specific genres through character functions. Focus on structural roles in the story, not character depth or complexity. A.2 Task Prompt Identify a work (novel, play, film, or TV series) in the G genre where characters clearly fulfill these four character functions: 1. R 1 2. R 2 3. R 3 4. R 4 Focus on what characters do in the story structure (their character function), not their psychological complexity. Example format: Title: Le Bourgeois Gentilhomme - lad in love: Cleonte - troubleshooter: Covielle - pompous blocker: Jourdain - witty damsel in love: Lucile Provide your response in this exact format: OUTPUT_START TITLE: [Title] - R 1 : [Character name] - R 2 : [Character name] - R 3 : [Character name] - R 4 : [Character name] OUTPUT_END 14 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix B Positive Sample Dataset Table 9 presents the complete set of 40 works with their validated character-role assignments across the four fundamental genres. These correspondences were identified through a semi-automated process combining LLM-assisted generation with expert validation, as described in Section 3.1.1. Table 9: Complete character-role assignments for the positive sample dataset. Genre TitleRole 1Role 2Role 3Role 4 Comedy lad in lovetroubleshooterpompous blocker witty damsel in love Le Bourgeois GentilhommeCleonteCovielleJourdainLucile Much Ado About NothingClaudioDon PedroLeonatoBeatrice The Importance of Being Earnest Jack WorthingAlgernon MoncrieffLady BracknellGwendolen Fairfax She Stoops to ConquerYoung MarlowTony LumpkinMr. HardcastleKate Hardcastle The Rivals Captain Jack Absolute Fag Sir Anthony Absolute Lydia Languish A Midsummer Night’s DreamLysanderPuckEgeusHermia The Barber of SevilleCount AlmavivaFigaroDoctor BartoloRosina The School for ScandalCharles SurfaceSir Oliver SurfaceJoseph SurfaceMaria The Taming of the ShrewLucentioTranioBaptista MinolaBianca The Marriage of FigaroCount AlmavivaFigaroDoctor BartoloSusanna Romance herodonorvillainfaithful victim RamayanaRamaVisvamitraRavanaSita Pride and PrejudiceElizabeth BennetMr. DarcyGeorge Wickham Lydia Bennet Jane EyreJane EyreMrs. FairfaxBertha MasonRochester Beauty and the Beast (1991)BelleEnchantress/Mrs. Potts GastonBeast Cinderella (1950)Prince CharmingFairy GodmotherLady TremaineCinderella The Princess BrideWestleyMiracle Max Prince Humperdinck Buttercup Pretty WomanEdward LewisBarney ThompsonPhilip StuckeyVivian Ward Notting HillWilliam ThackerSpikeAnna ScottWilliam Thacker Sabrina (1954)Linus LarrabeeBaron St. FontanelDavid LarrabeeSabrina Fairchild Wuthering HeightsEdgar LintonNelly DeanHeathcliffIsabella Linton Tragedy ill-fated adventurer soothsayerorder restorerill-fated partner MacbethMacbethThree WitchesMacduffLady Macbeth Julius CaesarBrutusSoothsayerOctaviusCassius Oedipus RexOedipusTiresiasCreonJocasta Romeo and JulietRomeoFriar LaurencePrince EscalusJuliet Antony and CleopatraMark AntonyThe SoothsayerOctavius CaesarCleopatra HamletHamlet The Ghost of Hamlet’s Father FortinbrasOphelia OthelloOthelloEmiliaLodovicoDesdemona King LearKing LearThe FoolEdgarCordelia A Streetcar Named DesireBlanche DuBoisMitchStanley Kowalski Stella Kowalski Doctor FaustusFaustusOld ManGood AngelMephistopheles Satire nonconformistinquisitordystopian idolrebellion partner 1984WinstonO’BrienBig BrotherJulia Brave New WorldJohn the SavageMustapha MondHenry FosterBernard Marx Brazil (1985)Sam LowryJack LintMr. Helpmann Archibald “Harry” Tuttle The Handmaid’s TaleOffred/JuneAunt LydiaSerena JoyMoira V for VendettaVInspector FinchAdam SutlerEvey Hammond Fahrenheit 451Guy MontagCaptain BeattyMildred MontagClarisse McClellan The Prisoner (1967)Number SixNumber TwoNumber OneNadia The Hunger GamesKatniss EverdeenEffie TrinketPresident SnowPeeta Mellark Catch-22YossarianClevingerColonel Cathcart Orr A Clockwork OrangeAlex DeLargeDr. Brodsky Minister of the Interior Pete 15 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix C Negative Sample Generation Prompt C.1 System Prompt You are a narratology expert creating systematically incorrect character-function mappings for research validation. Focus on creating plausible-but-wrong associations that test analytical discrimination. C.2 Task Prompt You are helping create test cases for validating character function analysis in narratology research. Given this CORRECT character-role mapping for T ( G genre): - R 1 : CH 1 - R 2 : CH 2 - R 3 : CH 3 - R 4 : CH 4 Generate ONE INCORRECT mapping by applying this error type: E ERROR TYPE DEFINITIONS: - role_swap: Swap two characters between their roles (maintaining same characters, wrong functions) - minor_character: Replace one major character with a minor/insignificant character from the work IMPORTANT: - Use actual character names from the work - Make it plausible enough that it requires analysis to detect the error - Only change what is needed for the specified error type Provide your response in this exact format: NEGATIVE_CASE_START ERROR_TYPE: error_type - [role name]: [character name] - [role name]: [character name] - [role name]: [character name] - [role name]: [character name] EXPLANATION: [brief 1-sentence explanation of why this is incorrect] NEGATIVE_CASE_END 16 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix D Negative Sample Dataset Table 10 presents the complete set of 20 negative samples across the four fundamental genres with their incorrect character-role assignments highlighted in red and bold. These invalid correspondences were generated through a semi-automated error introduction process with expert validation to ensure plausibility, as described in Section 3.1.2. Table 10: Complete character-role assignments for the negative sample dataset. Incorrect character-role assignments are highlighted in red and bold. Genre TitleRole 1Role 2Role 3Role 4 Comedy lad in lovetroubleshooterpompous blocker witty damsel in love The Importance of Being Earnest Jack WorthingAlgernon MoncrieffLady BracknellMiss Prism Le Bourgeois GentilhommeCovielleCleonteJourdainLucile The School for ScandalCharles SurfaceSir Oliver Surface Sir Benjamin Backbite Maria The Taming of the ShrewTranioLucentioBaptista MinolaBianca The RivalsCaptain Jack Absolute Fag Sir Anthony Absolute Lucy Romance herodonorvillainfaithful victim Pride and PrejudiceElizabeth BennetGeorge WickhamMr. DarcyLydia Bennet Cinderella (1950)Fairy GodmotherPrince CharmingLady TremaineCinderella Beauty and the Beast (1991)Belle Enchantress/Mrs. Potts GastonMaurice The Princess BrideWestleyMiracle MaxButtercup Prince Humperdinck RamayanaRamaVisvamitraRavanaUrmila Tragedy ill-fated adventurersoothsayerorder restorerill-fated partner Antony and CleopatraMark AntonyThe SoothsayerCleopatraOctavius Caesar MacbethMacbeththree witchesMacduffFleance HamletHamlet The Ghost of Hamlet’s Father FortinbrasRosencrantz Romeo and JulietBalthasarFriar LaurencePrince EscalusJuliet A Streetcar Named DesireBlanche DuBoisStanley KowalskiMitchStella Kowalski Satire nonconformistinquisitordystopian idolrebellion partner The Handmaid’s TaleOffred/JuneAunt LydiaMrs. PutnamMoira Fahrenheit 451Guy MontagClarisse McClellanMildred Montag Captain Beatty The Prisoner (1967)Number TwoNumber SixNumber OneNadia A Clockwork OrangeDr. BrodskyAlex DeLarge Minister of the Interior Pete Brave New WorldJohn the SavageMustapha MondHenry Foster Helmholtz Watson 17 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix E Validation Prompt E.1 System Prompt You are a narratology expert evaluating character-role correspondences based on character function analysis. Base your analysis on structural narratology frameworks and focus on plot-level functions rather than thematic or psychological interpretations. When multiple characters perform similar functions, identify which one is the primary bearer of that role (the character who most centrally drives that function in the plot structure). E.2 Task Prompt Analyze the following character-role correspondences for their character function justification: Title: T Genre: G Character-Role Mappings: - R 1 : CH 1 - R 2 : CH 2 - R 3 : CH 3 - R 4 : CH 4 For each character-role correspondence listed above, evaluate whether the character is the primary bearer of that role function in the narrative. A correspondence is justified only if this character is the main or central character performing that specific function, not a secondary or supporting character who also performs related actions. Provide your analysis in the following structured format: ANALYSIS_START — ROLE: R 1 CHARACTER: CH 1 JUSTIFIED: [YES or NO] REASONING: [Your explanation in 2-3 sentences] — ROLE: R 2 CHARACTER: CH 2 JUSTIFIED: [YES or NO] REASONING: [Your explanation in 2-3 sentences] — ROLE: R 3 CHARACTER: CH 3 JUSTIFIED: [YES or NO] REASONING: [Your explanation in 2-3 sentences] — 18 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models ROLE: R 4 CHARACTER: CH 4 JUSTIFIED: [YES or NO] REASONING: [Your explanation in 2-3 sentences] — ANALYSIS_END 19 Revisiting Northrop Frye’s Four Myths Theory with Large Language Models Appendix F Pairwise Model Agreement Table 11 presents the complete pairwise agreement rates between all model pairs across 190 character-role evaluations. Table 11: Detailed pairwise agreement rates between all model pairs. Model 1Model 2AgreementsAgreement (%) Claude Opus 4.5Llama 4 (128B-A17B)16687.4 Claude Opus 4.5Gemini 2.5 Pro16586.8 Claude Opus 4.5Qwen 3 (235B)16586.8 Llama 4 (128B-A17B)Qwen 3 (235B)16184.7 Gemini 2.5 ProQwen 3 (235B)15883.2 Gemini 2.5 ProLlama 4 (128B-A17B)15782.6 GPT OSS (120B)Qwen 3 (235B)15782.6 DeepSeek R1 (671B)Gemini 2.5 Pro15682.1 DeepSeek R1 (671B)Llama 4 (128B-A17B)15581.6 DeepSeek R1 (671B)Qwen 3 (235B)15481.1 DeepSeek R1 (671B)GPT OSS (120B)15179.5 Claude Opus 4.5GPT OSS (120B)15078.9 Claude Opus 4.5DeepSeek R1 (671B)14978.4 GPT OSS (120B)Llama 4 (128B-A17B)14877.9 Gemini 2.5 ProGPT OSS (120B)14576.3 20