Paper deep dive
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
SungHo Kim, Juhyeong Park, Eda Atalay, SangKeun Lee
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/15/2026, 1:37:47 AM
Summary
SCRIPT is a model-agnostic, plug-and-play module designed to inject subcharacter compositional knowledge (Jamo) into Korean pre-trained language models (PLMs). By utilizing a dual-channel strategy that fuses subcharacter-level structural representations with original subword embeddings, SCRIPT enhances performance across various Korean NLU and NLG tasks without requiring architectural changes or additional pre-training.
Entities (5)
Relation Signals (3)
Jamo → composes → Hangul
confidence 100% · each character is systematically composed of subcharacter units known as Jamo.
SCRIPT → enhances → Korean PLMs
confidence 98% · SCRIPT allows to enhance subword embeddings with structural granularity
SCRIPT → integrates → Jamo
confidence 95% · SCRIPT, a module that injects subcharacter-level structural knowledge
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure of Korean but also encode frequent and linguistically meaningful morphophonological processes. However, most current Korean language models (LMs) are based on subword tokenization schemes, which are not explicitly designed to capture the internal compositional structure of characters. To address this limitation, we propose SCRIPT, a model-agnostic module that injects subcharacter compositional knowledge into Korean PLMs. SCRIPT allows to enhance subword embeddings with structural granularity, without requiring architectural changes or additional pre-training. As a result, SCRIPT enhances all baselines across various Korean natural language understanding (NLU) and generation (NLG) tasks. Moreover, beyond performance gains, detailed linguistic analyses show that SCRIPT reshapes the embedding space in a way that better captures grammatical regularities and semantically cohesive variations. Our code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.12377v1
- Canonical: https://arxiv.org/abs/2604.12377v1
Trouble viewing inline? Open PDF directly →
Full Text
102,659 characters extracted from source content.
Expand or collapse full text
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models SungHo Kim 1 , Juhyeong Park 1 , Eda Atalay 1 , SangKeun Lee 1,2 1 Department of Artificial Intelligence, Korea University, Seoul, South Korea 2 Department of Computer Science and Engineering, Korea University, Seoul, South Korea sungho3268, johnida, edaatalay, yalphy@korea.ac.kr Abstract Korean is a morphologically rich language with a featural writing system in which each char- acter is systematically composed of subchar- acter units known as Jamo. These subcharac- ters not only determine the visual structure of Korean but also encode frequent and linguis- tically meaningful morphophonological pro- cesses. However, most current Korean lan- guage models (LMs) are based on subword tokenization schemes, which are not explicitly designed to capture the internal compositional structure of characters. To address this limi- tation, we proposeSCRIPT, a model-agnostic module that injects subcharacter compositional knowledge into Korean PLMs.SCRIPTallows to enhance subword embeddings with struc- tural granularity, without requiring architec- tural changes or additional pre-training. As a result,SCRIPTenhances all baselines across various Korean natural language understand- ing (NLU) and generation (NLG) tasks. More- over, beyond performance gains, detailed lin- guistic analyses show thatSCRIPTreshapes the embedding space in a way that better cap- tures grammatical regularities and semantically cohesive variations. Our code is available at https://github.com/SungHo3268/SCRIPT. 1 Introduction In human writing systems, the grapheme, the small- est unit of written language that encodes linguis- tic information, plays a crucial role in shaping how meaning is represented and processed (Coul- mas, 2003; Sampson, 2015; Daniels and Bright, 1996). In many alphabetic systems, such as En- glish, graphemes typically correspond to atomic letters (e.g., a, b, c), and words are formed through linear combination (e.g., “cat” consists of c, a, and t). However, not all alphabetic systems operate in such a linear manner, nor do they necessarily treat characters as minimal units of written composition. Korean, in particular, employs a unique featural writing system, Hangul, in which each character is a structured composition of smaller subcharac- ter units known as Jamo. As illustrated in Fig- ure 1(a), each character consists of three Jamo units: Choseong (initial consonant), Jungseong (vowel), and Jongseong (final consonant), following fixed spatial arrangements and a strict compositional order. These principles were explicitly defined in Hunminjeongeum 1 (National Hangeul Museum, 2018, 2021), the historical document detailing the invention principles of Hangul, including the de- sign and combination rules for the subcharacters. Crucially, this compositional structure is not merely orthographic. As a morphologically rich and agglutinative language, Korean exhibits exten- sive morphophonological alternations across mor- pheme boundaries (Lee and Chung, 2003; Matteson et al., 2018; Jun, 2018). Predicate inflection, for ex- ample, often triggers systematic subcharacter-level alternations, such as the addition of the final conso- nant ‘ᄊ’ to mark past tense or ‘-ᄆ’ for nominaliza- tion, as shown in Figure 1(b). Additionally, phono- logical assimilation between adjacent syllables fre- quently alters subcharacters to facilitate natural pro- nunciation (Sohn, 2001; Shin et al., 2012). These phenomena highlight that subcharacter-level fea- tures in Korean are tightly linked to grammatical, semantic, and morphophonological functions (Lee and Ramsey, 2001). Despite this linguistic reality, most contempo- rary Korean PLMs, including advanced off-the- shelf LLMs (Yoo et al., 2024; LG AI Research et al., 2024), rely almost exclusively on subword- based tokenization. While subword modeling ef- fectively captures lexical semantics from large cor- pora, it struggles to reflect Hangul’s compositional structure, limiting sensitivity to fine-grained mor- phosyntactic variations (Albright and Kang, 2009; Kim et al., 2025). In contrast, a few subcharacter- 1 Hunminjeongeum explains the letter design and well- formed combinations; modern encoding schemes and key- board input sequences arise from contemporary standards. arXiv:2604.12377v1 [cs.CL] 14 Apr 2026 ᄎ ᅮ ᄇ ᄃ ᅡ ChoseongJungseong Jongseong (empty) Character (a) §Emphatic morphological extension 춥다 be cold à추 cold +ᄇ (suffix) à우 (epenthetic vowel) +어 is + 죽겠다 will die à추워죽겠다 is really cold 춥다 be cold à추 cold +ᄇ (suffix) à우 (epenthetic vowel) +었 (past tense marker) + 다 be à추웠다 was cold §Tense inflection Epenthesis AlternationElision Contraction §Part-of-speechshift 춥다 be cold à추 cold +ᄇ (suffix) à우 (epenthetic vowel) +다 be + -ᄆ (noun-forming suffix) à추움 coldness §Honorificandpolitenessinflection 춥다 be cold à추 cold +ᄇ (suffix) à우 (epenthetic vowel) +세요 (honorific ending) + 다 be à추우세요 is cold (b) Figure 1: (a) Examples of the components of Hangul. This figure illustrates two characters, such as ‘춥 cold ’ and ‘다 ending suffix ’, with each subcharacter highlighted in blue. (b) Examples of linguistic phenomena arising from the inflection of predicate ‘춥다 be cold ’ at the subcharacter-level, with the transformed subcharacters highlighted in red. based LMs (Moon and Okazaki, 2020; Cognetta et al., 2023; Kim et al., 2024) show strong robust- ness to such variations but often underperform on downstream tasks due to weaker semantic repre- sentations and increased computational cost. To leverage complementary strengths, we pro- poseSCRIPT, a lightweight, plug-and-play mod- ule that injects subcharacter-level structural knowl- edge directly into existing subword-based PLMs. SCRIPTattaches to the embedding layer of a PLM and operates through a dual-channel strat- egy. It compresses subcharacter sequences into structure-aware subword representations grounded in Hangul’s compositional principles, and then fuses them with the PLM’s original subword em- beddings. This design enables the model to cap- ture fine-grained subcharacter compositionality while preserving the rich semantic information learned from large-scale corpora, without modi- fying the PLM architecture or requiring additional pre-training. The main contributions of this work are summarized as follows: •We empirically show that most Korean mor- phological variations occur at the subcharacter- level, motivating subcharacter-aware modeling. •We introduceSCRIPT, a model-agnostic module that injects structure-aware subcharacter com- positional representations into existing PLMs via embedding-level integration. •We show thatSCRIPTimproves performance across a wide range of Korean NLU and NLG benchmarks, while effectively capturing key morphosyntactic phenomena. 2 Motivation In this section, we present our empirical observa- tions on the pervasiveness of diverse morphophono- logical changes at the subcharacter-level in real Korean usage, highlighting the importance of mod- eling the subcharacter structure of Hangul. 2.1 Setup Specifically, we conduct a large-scale corpus-based analysis quantifying their frequency. We used the Korean Part-of-Speech Tagged Corpus 2 , which con- tains 3M words annotated with morpheme-level and POS information. This corpus provides a solid basis for estimating the frequency of subcharacter- level alternations in real usage. To systematically capture these alternations, we adopted a simple annotation scheme, following Matteson et al. (2018), which marks how each character relates to its base form (lemma). Each character is assigned to one of three categories: • KEEP : unchanged with respect to the base form. • MOD: modified from the base form. • NOOP: omitted in the base form. For example, in ‘했다 did ’, whose base form is ‘하 다 do ’, the segment ‘했 did ’ is labeled asMODbecause it reflects a tense change arising from the combina- tion of the verb stem ‘하’ and the past tense marker ‘-었-’, which undergoes phonological contraction (‘하+었→ 했’). In contrast, ‘다’ is labeled as KEEP since it remains unchanged. 2 Part-of-Speech Tagged Corpus (v1.1) from ModuCorpus, provided by the National Institute of Korean Language (2020) To p-10 MOD Samples Usage past tense, past connective form, ... past adnominal form, participial modifier,... connective ending, conditional clause,... adnominal form, nominal clause, conditional form,... compound modifier, experiential participle, ... future adnominal form, speculative expression, ... compound verb connective form, ... passive past adnominal form, result expression, ... passive past tense, perfective aspect, ... declarative connective form, cause conjunction, ... 하à했 하à한 하à해 이à인 하à한 하à할 하à해 되à된 되à됐 이à라 MOD Subcharacter-level MOD Character-level MOD Distribution of MOD -tagged characters 7.25% 92.75% Character Tag MOD-(하,았) MOD-(하,ᄂ) MOD-(하,ᅡ) MOD-(이, ᄂ) MOD-(하,ᄂ) MOD-(하,ᄅ) MOD-(하,ᅡ) MOD-(되, ᄂ) MOD-(되,었) MOD-(이, 라) MOD-Level Subchar (V+F) Subchar (F) Subchar (V) Subchar (F) Subchar (F) Subchar (V) Subchar (V) Subchar (F) Subchar (V+F) Character Original Char Modified Char Figure 2: Morphological modifications in a large-scale Korean POS-tagged corpus: the left panel distinguishes subcharacter- and character-levelMODcases, and the right lists the ten most frequentMODtypes with raw target characters and corresponding factors. In this paper, we focus our analysis on charac- ters tagged asMOD, as they directly encode morpho- logical alternations. EachMODcharacter is further classified by its level of granularity 3 : •Subcharacter-levelMOD: when only part of a character is altered. (e.g., the change from ‘하’ to ‘한’, adding the final consonant ‘ᄂ’). • Character-levelMOD: when the entire character is changed into a different character. (e.g., ‘이’ is replaced with ‘라’). 2.2 Observation As shown in Figure 2, the overwhelming ma- jority of characters tagged asMOD(92.75%) in- volved subcharacter-level modifications, while only a small fraction (7.25%) represented character- level changes. These findings underscore the impor- tance of modeling subcharacter-level alternations for a deeper and more comprehensive understand- ing of Korean, especially in adapting to diverse usage and morphological variation. This aligns with prior work (Kim et al., 2024; Lee et al., 2025), which observed that jamo-based language model- ing demonstrates robustness in handling character- level conjugation changes and exhibits strong per- formance on noisy, real-world data such as offen- sive content. Motivated by this, we aim to explicitly encode subcharacter compositional knowledge in a principled manner, thereby incorporating this lin- guistic information into language models that have previously overlooked it. 3 Methodology Based on our observation (§2), we proposeSCRIPT (SubcharacterCompositionalRepresentation 3 Detailed tagging procedures are provided in Appendix B Injection Module for Korean Pre-Trained Lan- guage Model), a module that enhances PLM’s embeddings with subcharacter compositional knowledge. In this section, we instance Jamo as the subcharacter unit inSCRIPTand apply it to subword-based PLM, aligning with standard practices in modern Korean PLMs.We also provide extensions for alternative subcharacter units, such as BTS units (Appendix D). 3.1 Overall Framework SCRIPT is attached to PLMs at the embedding layer, as illustrated in Figure 3(a). Given a Korean text, the model employs two parallel tokenization paths: (1) a subword tokenizer that produces the PLM’s original subword sequence, and (2) a sub- character tokenizer that generates fine-grained in- put forSCRIPT. The subword sequence is projected through the PLM’s original embedding layer, while SCRIPTconstructs an alternative subword-level rep- resentation by compressing the subcharacter se- quence (§3.2). These two subword representations are then integrated into a unified subword represen- tation (§3.3). This dual-channel strategy allows us to leverage the strengths of both approaches. The full algorithm for synthesizing subword representa- tions with SCRIPT is provided in Table 6. 3.2 SCRIPT Figure 3(b) illustrates the detailed architecture of SCRIPT, which compresses subcharacter represen- tations into subword representations in two stages. 3.2.1 Stage 1: Subcharacter-to-Character The first stage of deriving subword representations from subcharacter representations is to compress subcharacter representations into character repre- Transformer Stacks Pre-trained Subword Embedding Layer Text Input Pre-trained LM (e.g., BERT, GPT, ...) SCRIPT SubwordsequenceSubcharactersequence Representation Fusing Layer (a) Reshape Convolution (2x1) (I+V) + F Sequential Composition Encoding I+V Pooling Sequential Composition 대한민국 Projection SCRIPT query key value Pre-trained Embedding Layer 대한민국 Transformer Stacks Ⓐ Ⓑ Ⓒ 대 ▃ 하 ᄂ ᄆ ᅵᄂᄀ ᅮ ᄀ 대미구하ᄂᄀ ▃ 대미구하ᄂᄀ ▃ 대미구하 ▃ ᄂᄀ 대한민국 대한민국 대한민국 대한민국 Cross- Attention (b) Figure 3: (a) Overall illustration of the PLM enhanced withSCRIPT. (b) Detailed architectural example of the SCRIPT, starting from the word ‘대한민국 South Korea ’, which is tokenized into 12 Jamo units: I : [ᄃ,ᄒ,ᄆ,ᄀ], V: [ᅢ,ᅡ,ᅵ,ᅮ], F: [,ᄂ,ᄂ,ᄀ]. Each sub-process (A-C) represents a successive fusion step: (A) fusion of Choseong and Jungseong (§3.2.1), (B) addition of Jongseong (§3.2.1), and (C) character-to-subword (§3.2.2). sentations. Inspired by Kim et al. (2024), to effec- tively model the compositional structure of Hangul, we explicitly incorporate three fundamental compo- sitional principles into our methodology (National Hangeul Museum, 2021; Yeon and Brown, 2013; Unicode Consortium, 2019): 1. Composition: A character is composed of up to three Jamo: Choseong and Jungseong are essential components, whereas Jongseong is not mandatory. 4 2. Spatial arrangement: Within a syllable block, Choseong is placed either above or to the left of Jungseong, while Jongseong, if present, is always positioned beneath them. 3.Sequential order: Jamo consistently follow a prescribed order: Choseong→Jungseong→ Jongseong. SCRIPTadopts a hierarchical compression archi- tecture grounded in the design principles of Hangul, to better capture linguistic features of Korean. This step underpins the ‘structure-aware’ subword em- beddings inSCRIPT, as it explicitly encodes the subcharacter compositional knowledge. Subcharacter Representation. Given an input texts, each character is first decomposed into se- quential three subcharacters, I (short for initial con- sonant, Choseong), V (short for vowel, Jungseong), and F (short for final consonant, Jongseong), fol- lowing Principles 1. If a character lacks a final consonant, a special empty token () is inserted in 4 A detailed explanation is provided in Appendix A its place. The resulting subcharacter embeddings are denoted ase ∈R 푁×퐷 , where푁is the num- ber of subcharacter tokens and퐷is the embedding dimension. Furthermore, to clarify the ordered ar- rangement of I, V, and F, we denote the sequential subcharacter representation푒 푖 ∈ eas푒 I,푘 ,푒 V,푘 , and 푒 F,푘 for each integer 푘 in [1, N/3]: 푒 푖 = 푒 I,푘 if 푖= 3푘− 2 푒 V,푘 if 푖= 3푘− 1 푒 F,푘 if 푖= 3푘 (1) Fusion of Choseong and Jungseong. To accu- rately reflect the sequential order (Principle 3), we follow the fixed composition sequence (I→V→ F). The entire subcharacter sequence is first en- coded using a GRU-based sequential composition layer. Following this order, the I and V components are merged via element-wise summation to obtain a combined representation, h I+V ∈R 푁 3 ×퐷 . Addition of Jongseong.After forming the inter- mediate representation of I and V, we incorporate the V to complete the character representation. Re- flecting the visual structure of Hangul, where F is always positioned below I and V (Principle 2), we model this arrangement by vertically concatenating the F representation,h F =ℎ F,푘 , withh I+V . This composition is formally expressed as follows: h R = h I+V h F ∈R 2× 푁 3 ×퐷 (2) To merge these vertically aligned subcharacters into characters, we apply a convolutional layer cap- turing the relative positional information. We then finalize the character representations by applying average pooling overh R , yielding dense charac- ter representations,h C ∈R 푁 3 ×퐷 , grounded in the compositional principles of Hangul. 3.2.2 Stage 2: Character-to-Subword The second main stage ofSCRIPTcompresses char- acter representations into subword representations, aligning their granularity with that of the original subwords used in PLMs. Our goal was to aggregate character representations within each subword to form a unified subword representation. However, directly averaging or summing these character rep- resentations often led to unstable training. To mit- igate this issue, and in line with Principle 3, we apply the sequential composition layer once more to capture the compositional order of characters within each subword. Then, we apply a simple pooling operation, specifically, selecting the final character representation at each subword boundary, to obtain the subword representation: h S = POOLING(GRU(h C )) ∈R 푁 ′ ×퐷 (3) where 푁 ′ denotes the subword sequence length. 3.3 Fusion of Two Subword Representations Despite these structured, dense subcharacter-level linguistic features, the resulting subword represen- tations, compressed from subcharacters alone, lack semantic expressiveness, as they are not pre-trained on large-scale Korean corpora. To address this lim- itation, we fuse them with semantically richer sub- word embeddings obtained from the existing PLM. Specifically, we introduce a fusion mechanism that integrates two complementary representations: the synthesized subword representation fromSCRIPT, denoted ash S , and the original pre-trained sub- word embedding,e S ∈R 푁 ′ ×퐷 , projected into the same embedding space. A cross-attention layer is employed to combine these sources, yielding the final structure-aware subword representation e F ∈R 푁 ′ ×퐷 , which is then used as input to the subsequent Transformer layers: e F = CROSSATTN(Q= e S , KV= h S )(4) ThroughSCRIPT, we construct a fused subword representation that integrates the compositional knowledge of Hangul with the semantic richness of pre-trained subword embeddings. This dual- channel encoding enhances the language model’s ability to capture Korean character structure while preserving subword-level semantic content. These fused representations are then fed into the PLM’s Transformer stacks, allowing downstream tasks to benefit from this linguistically enriched input. 4 Experiments In this section, we evaluateSCRIPTon a range of Korean NLU and NLG tasks across strong PLMs (§4.2). We further conduct ablation studies to ana- lyze the contribution of Hangul-specific structural knowledge and key design choices (§4.3). Additionally, we show that this efficiency comes with minimal computational overhead, which re- mains comparable to standard subword-based mod- els (see Appendix J). 4.1 Experimental Settings Baselines. We appliedSCRIPTto four Korean subword-based PLMs (KoGPT2 base , KoGPT3- 1.2B, EXAONE-3.5-2.4B-Instruct, BERT base ), and additionally compare with a state-of-the-art Jamo- based encoder model, KOMBO base (Kim et al., 2024). Detailed specifications and implementa- tional details are provided in Appendix C, E. Tasks.We evaluateSCRIPT-enhanced models on nine Korean NLU tasks, including four standard benchmarks (KorNLI, KorSTS, NSMC, PAWS-X) and five KoBEST tasks designed to assess diverse linguistic and cognitive capabilities. To evaluate generative performance, we additionally consider KoCommonGen for commonsense reasoning, XL- Sum for summarization, and Korean GEC for gram- matical error correction. Detailed dataset statistics and explanations are provided in Appendix E.3. 4.2 Experimental Results Korean Standard NLU Tasks.As shown in Ta- ble 1,SCRIPTimproves performance across all baselines, yielding average gains of up to 1.6%p. Compared to the Jamo-based baseline KOMBO base , our BERT base +SCRIPTmodel achieves superior per- formance despite using the same underlying ar- chitecture and a comparable model size. Unlike KOMBO base , which directly processes raw subchar- acters and relies on costly full pre-training,SCRIPT is applied as a plug-in module during only fine- tuning, enabling efficient incorporation of Hangul structure. Notably,SCRIPTis applicable to both en- coder and decoder architectures, overall improving performance across model types. ModelKorNLIKorSTSNSMCPAWS-X KoBEST BoolQCOPAWiCHellaSwagSentiNeg KOMBO base 75.9777.2888.3473.4061.4061.0068.9163.8079.07 BERT base 75.8576.7288.9672.3860.7560.9073.1463.2083.12 BERT base + SCRIPT76.4977.6888.9673.6862.3261.3074.3064.4083.38 KoGPT2 base 72.2473.8288.9076.3367.2268.9067.0769.1088.50 KoGPT2 base + SCRIPT72.4774.2788.8076.6168.2870.9068.1872.4089.47 KoGPT3-1.2B80.1176.1490.5177.4077.3282.8072.7878.9096.31 KoGPT3-1.2B + SCRIPT80.3979.6090.5379.9577.6382.8074.6579.3096.48 EXAONE-2.4B83.9985.0890.0485.2492.5993.3082.1485.6094.21 EXAONE-2.4B + SCRIPT85.7785.2790.8985.9093.3093.3082.4686.0094.96 Table 1: Performance on nine Korean NLU tasks. The evaluation metrics for each task are as follows: KorSTS is evaluated using Spearman correlation×100, while other tasks are evaluated based on accuracy (%). The best results in each family of models are highlighted in boldface. Model KoCommonGen BLEU 3BLEU 4ROUGE-2ROUGE-LMETEORmBERTScoreKoBERTScore KoGPT2 base 18.2910.3344.2454.5040.0583.3791.21 KoGPT2 base + SCRIPT25.0115.5747.4260.0042.5384.6791.46 KoGPT3-1.2B26.1917.2058.8562.5352.1185.4191.17 KoGPT3-1.2B + SCRIPT28.8919.5859.2864.8052.3786.2691.78 EXAONE-2.4B40.1128.4162.2564.8454.8487.6593.12 EXAONE-2.4B + SCRIPT41.4831.8071.0372.1661.2788.1293.95 Table 2: Performance on KoCommonGen generative task. We use eight automatic evaluation metrics, including n-gram based measures like BLEU, ROUGE, and METEOR, and two BERT-based scores for semantic similarity. The best results in each family of models are highlighted in boldface. Korean Advanced NLU Tasks.SCRIPTalso outperforms baselines on knowledge-intensive tasks in KoBEST, including reading compre- hension (KB-WiC) and commonsense reasoning (KB-HellaSwag). It proves more effective than KOMBO base , which lags on complex tasks due to its reliance on subcharacter-only inputs. By in- tegrating subword and subcharacter information, SCRIPToffers robust and architecture-agnostic en- hancements across task complexities. Korean Generation Tasks. Our method also delivers consistent gains on Korean generative tasks. As shown in Table 2, on KoCommonGen, a task that involves transforming and combining given morphemes to generate plausible sentences, SCRIPTimproves across all seven generative met- rics, with gains (an average of 1.4-3.5%p) depend- ing on model size. The improvements are particu- larly pronounced in n-gram metrics such as BLEU, METEOR, and ROUGE, indicating enhanced mod- eling of local compositional patterns in morpho- logically rich Korean. This trend extends to other generation tasks as shown in Appendix F. Notably, on the Korean GEC task Kor-Learner,SCRIPTex- ceeds the best-performing baseline by over 3.2%p on average. According to Yoon et al. (2023), Kor- Learner includes a high concentration of errors involving particles, endings, and conjugations com- pared to Kor-Native. One possible explanation for the relatively larger gains on generative tasks, compared to NLU tasks, is that Hangul’s sequential compositional structure (Choseong→Jungseong→Jongseong) aligns nat- urally with token-by-token decoding, allowing sub- character information to more directly influence generation decisions. We consider this a promis- ing direction for further analysis, as a deeper un- derstanding of this phenomenon requires further investigation. 4.3 Ablation Study for SCRIPT Architecture Table 3 presents an ablation study examining the core design choices of SCRIPT. Alternative Tokenization Methods forSCRIPT. Across different granularities, Jamo-basedSCRIPT achieves the best overall performance. In contrast, SCRIPT Fusion w/ PLM KoBEST Avg. Initial Token UnitCompressionBoolQCOPAWiCHellaSwagSentiNeg JamoPrinciplesCrossAttention68.2870.9068.1872.4089.4773.85 StrokePrinciplesCrossAttention68.3565.2067.8171.0089.4672.36 CjiPrinciplesCrossAttention68.8265.2067.9871.4088.0172.28 BTSPrinciplesCrossAttention67.9964.7066.6771.5088.9771.97 CharacterPrinciplesCrossAttention66.0059.3063.3769.4088.4069.29 SubwordPrinciplesCrossAttention59.1954.5067.6269.7088.5767.92 WordPrinciplesCrossAttention66.4861.1062.8272.0088.7170.22 JamoLinearCrossAttention67.0457.7064.3569.3088.1169.30 JamoAttentionCrossAttention68.1464.7067.2871.6088.3272.01 JamoPrinciplesSummation68.2765.5067.9870.6089.0772.28 JamoPrinciplesConcatenation66.5561.1066.6169.9089.0670.64 Table 3: Ablation results for various architecture ofSCRIPTapplied to KoGPT2 base . The first row presents the best-performing variant. Cells corresponding to ablated components are highlighted inlight blue. The global-best results are highlighted in boldface. using excessively fine-grained units such as BTS leads to a slight performance drop, suggesting limi- tations in compressing overly fine-grained represen- tations into higher-level coarse units. Furthermore, when we extended the comparison to larger units beyond the subcharacter level, including character, subword, and word units, performance degraded substantially. This finding suggests that our pro- posed method is specifically designed to preserve the compositional structure of subcharacters and is therefore less suited to larger linguistic units. No- tably, using subword units, also employed in the base PLM, resulted in the largest performance drop. This result indicates that the observed gains are not simply attributable to increased parameter count or additional fusion capacity, but are instead driven by Jamo-level structural information. Compression Method of Subcharacters in SCRIPT. Replacing the proposed composition- principled compression with generic pooling meth- ods (Attention (Dai et al., 2020) or Linear (Nawrot et al., 2022)) leads to a 1.8–4.6%p performance drop, underscoring the importance of preserving the hierarchical compositional structure of Hangul during subcharacter aggregation. Integration Method of Subword Representa- tions. We further compare integration strategies between subcharacter and subword representations. CrossAttention (Vaswani et al., 2017) yields the strongest results, outperforming Summation and Concatenation, suggesting that dynamic alignment with PLM representations is key to integrating two heterogeneous subword representations effectively. Overall, these results demonstrate thatSCRIPT’s gains arise from explicitly encoding Hangul’s com- positional structure and aligning it with pre-trained representations, rather than from any single archi- tectural choice. 4.4 Effect of Tokenization Granularity on SCRIPT Beyond conducting ablations on individual com- ponents of theSCRIPTarchitecture (§4.3), we also observed that the integration and effectiveness of compositional knowledge vary depending on the PLM’s tokenization scheme. To investigate this, we compared four tokenization strategies: word, mor- pheme, subword, and character. As the final com- positional token units changed accordingly, we also adjustedSCRIPT’s compressed output token unit to match. Thus, instead of the original “Character-to- Subword” setting described in Section 3.2.2, we ex- perimented with “Character-to-Word,” “Character- to-Morpheme,” and “Character-to-Character.” 5 As shown in Table 4,SCRIPTproved effective when applied with larger units, such as Word and Morpheme, compared to Subword. This suggests that when the base PLM has already captured suf- ficient semantic meaning (Aguilar et al., 2021; Kaushal and Mahowald, 2022), integrating syntac- tic compositional knowledge leads to a synergistic improvement. In contrast, with the smaller Char- 5 To minimize OOV occurrences, we constructed vocabu- laries based on prior work (Park et al., 2020; Kim et al., 2024), setting vocabulary sizes to 64k for Word, 32k for Morpheme and Subword, and 2k for Character. Model PLM Tokenization KoBEST Avg. BoolQCOPAWiCHellaSwagSentiNeg BERT base Word 60.0457.6062.7055.0052.3957.55 BERT base + SCRIPT Jamo 60.4757.6064.7657.2052.3958.48(▲ 0.94) BERT base Morpheme 63.7558.5071.7561.8078.5966.88 BERT base + SCRIPT Jamo 65.0360.0072.0662.2080.3567.93(▲ 1.05) BERT base Subword 67.2267.1068.9069.1088.5072.16 BERT base + SCRIPT Jamo 68.2868.2070.9072.4089.4773.85(▲ 1.69) BERT base Character 62.8961.0071.3548.6078.8464.54 BERT base + SCRIPT Jamo 61.0459.3071.3549.4078.3463.89(▼ 0.65) Table 4: Comparison of the effectiveness of compositional knowledge integration into PLMs across different tokenization methods. The global-best results are highlighted in boldface and local-best results for each section are highlighted in underline, respectively. acter unit, where the base PLM primarily learns syntactic rather than semantic knowledge (Aguilar et al., 2021; Mielke et al., 2021), applyingSCRIPT introduced noise and hindered performance. 5 In-Depth Analysis Beyond the quantitative results, this section offers a detailed linguistic analysis of howSCRIPToperates in Korean. We examine how SCRIPT enriches sub- word representations and enables the base model to more effectively capture key linguistic phenomena. 5.1 Impact on Morphological Variations 잤다 slept 자다 sleep 자다 sleep 눕다 lie 누웠다 lay 누웠다 lay 잤다 slept 눕다 lie Figure 4: PCA visualization of subword embeddings for word pairs exhibiting subcharacter-level alternations. Each pair (e.g.,자다 sleep –잤다 slept ,눕다 lie –누웠다 lay ) shares the same root meaning but differs in tense. ToassesshowwellSCRIPTcaptures subcharacter-level morphological alternations, we compare two subword representations: one from the PLM’s original subword embeddings and the other fromSCRIPT’s subcharacter-based representations. Using these, we represent morpho- logically related word pairs, such as tense-inflected forms. Figure 4 provides a qualitative geometric illustration using mean-centered embeddings projected via PCA. In the projected space,SCRIPT places morphologically related forms in closer angular proximity, indicating a more structured encoding of tense relationships, while the original PLM embeddings appear more dispersed. To verify that this pattern is not an artifact of 2D projection, we additionally compute cosine similarity in the original embedding space over 50 verb–past tense pairs, observing a consistent increase from 0.71 to 0.80 (+11%).These results suggest that subcharacter compositionality improves the model’s ability to capture fine-grained grammatical variations in Korean. 5.2 Impact on Word Embedding Cohesion As shown in Figure 5, we examine how Korean LMs organize semantically related predicate inflec- tions in embedding space using five forms of the predicate ‘춥다 be cold ’. Larger subword-based LMs show increasingly cohesive clustering, while the Jamo-based model, KOMBO (Kim et al., 2024), ex- hibits a more scattered distribution, likely due to its extreme focus on syntactic granularity. For clarity, we visualize KoGPT2 as a representative backbone, whereSCRIPTproduces the most compact group- ing even at a small scale. Consistent patterns are observed across larger backbones. Together with ablation results showing degraded cohesion when compositional encoding is removed or altered, this suggests that the observed structure primarily arises from subcharacter compositional modeling rather than normalization effects. KoGPT2 base + SCRIPT KOMBO base 추움 coldness 춥습니다 is cold 추워죽겠다 is really cold 춥다 be cold 추웠다 was cold [Word variations] [Baselines] KoGPT2 base KoGPT3-1.2B EXAONE-3.5-7.8B Principal Component 1 Principal Component 2 Figure 5: PCA visualization of word embeddings averaged over tokens for five semantically related Korean predicate inflections derived from the predicate ‘춥다 be cold ’. The dashed boundaries indicate the dispersion ranges of the smallest baseline, KoGPT2 base , and its SCRIPT-augmented counterpart. 6 Related Work 6.1 Korean Pre-trained Language Models Most off-the-shelf Korean PLMs employ subword- based tokenization (Yoo et al., 2024; LG AI Re- search et al., 2024), which has proven effective for handling Korean’s rich morphology. However, such subword-based tokenizations do not explicitly model subcharacter-level structure, where many morphophonological processes in Korean occur. To address this limitation, several studies have ex- plored Jamo-level modeling of Korean (Moon and Okazaki, 2020; Cognetta et al., 2023; Kim et al., 2024). In particular, Kim et al. (2024) explicitly en- codes the compositional structure of Hangul to en- rich character representations. However, these ap- proaches typically rely on non-standard or encoder- only architectures and require full pre-training from scratch, which limits their applicability to general- purpose PLMs. 6.2 Multi-Granular Representations Some prior work has attempted to incorporate multiple granularities in Korean language model- ing, such as combining Jamo and word embed- dings (Kwon et al., 2021) or switching between Jamo and subwords depending on context (Lee et al., 2025). However, these methods typically alternate between token levels rather than struc- turally integrating them. Moreover, they often lack architectural generality or remain limited to encoder-based tasks. Despite growing evidence that combining fine-grained morphological cues with higher-level representations improves perfor- mance (Lai et al., 2021; Zhao et al., 2023; Wang et al., 2024), Korean PLMs still underexplore this integration in a principled and efficient manner. 7 Conclusion This work presentsSCRIPT, a modular framework for injecting subcharacter compositional knowl- edge into Korean PLM. Through a structure-aware compression mechanism grounded in the compo- sitional principles of Hangul,SCRIPTcaptures morphophonological variations at the subcharacter- level and enriches coarse PLM’s token representa- tions. Our experiments demonstrate thatSCRIPT generally improves model performance across a wide range of Korean NLU and NLG tasks, enrich- ing both conventional subword- and subcharacter- based approaches. Beyond quantitative gains, our linguistic analyses show thatSCRIPTenhances the semantic and grammatical organization of the em- bedding space, enabling more cohesive clustering of inflected predicates and more faithful modeling of Korean linguistic phenomena. These findings underscore the limitations of subword tokenization in morphologically rich language and advocate for subcharacter-aware modeling as a necessary exten- sion for Korean NLP. Limitations This work focuses on improving Korean language understanding by explicitly modeling structural characteristics specific to Korean. Accordingly, SCRIPTis primarily designed and evaluated within the Korean linguistic context, and its effectiveness for other languages is not systematically validated in this study. Although the modular design may be adaptable to languages with rich internal char- acter structure or complex morphology, such ex- tensions are beyond the scope of this paper. As a minimal proof of concept, we show thatSCRIPT can be integrated into a multilingual pre-trained model and still improves Korean task performance (Appendix G). However, this experiment does not establish general cross-lingual applicability. Second, whileSCRIPTdemonstrates consistent improvements across models ranging from approx- imately 100M to 2.4B parameters, we do not eval- uate models at larger scales (e.g., 7B+). As larger models develop stronger internal representations, the relative benefit of explicitly modeling subchar- acter structure may vary. EvaluatingSCRIPTon larger-scale models remains an important direction for future work. Finally,SCRIPTintroduces additional parame- ters and sequence-length-dependent computations in the embedding layer, particularly when using cross-attention. Although this leads to improved convergence and performance, it also increases inference cost. Future work may explore more lightweight variants that better balance efficiency and effectiveness. Acknowledgments This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Ko- rea government (MSIT) (No.RS-2025-00517221 and No.RS-2024-00415812) and Institute of In- formation & communications Technology Plan- ning & Evaluation (IITP) grant funded by the Ko- rea government (MSIT) (No.RS-2024-00439328, Karma: Towards Knowledge Augmentation for Complex Reasoning (SW Starlab), No.RS-2024- 00457882, AI Research Hub Project, and No.RS- 2019-I190079, Artificial Intelligence Graduate School Program (Korea University)). References Gustavo Aguilar, Bryan McCann, Tong Niu, Nazneen Rajani, Nitish Shirish Keskar, and Thamar Solorio. 2021. Char2Subword: Extending the subword em- bedding space using robust character compositional- ity. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 1640–1651, Punta Cana, Dominican Republic. Association for Compu- tational Linguistics. Adam Albright and Yoonjung Kang. 2009. Predict- ing innovative alternations in korean verb paradigms. Current issues in unity and diversity of languages: Collection of the papers selected from the CIL 18, held at Korea University in Seoul, pages 1–20. Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large anno- tated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empiri- cal Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Compu- tational Linguistics. Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez- Gazpio, and Lucia Specia. 2017. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada. Association for Computational Linguistics. Marco Cognetta, Sangwhan Moon, Lawrence Wolf- sonkin, and Naoaki Okazaki. 2023.Parameter- efficient Korean character-level language modeling. In Proceedings of the 17th Conference of the Euro- pean Chapter of the Association for Computational Linguistics, pages 2350–2356, Dubrovnik, Croatia. Association for Computational Linguistics. Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross- lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium. Association for Computational Lin- guistics. Florian Coulmas. 2003. Writing Systems: An intro- duction to Their Linguistic Analysis. Cambridge University Press. Daniel Dahlmeier and Hwee Tou Ng. 2012. Better evaluation for grammatical error correction. In Pro- ceedings of the 2012 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 568–572, Montréal, Canada. Association for Compu- tational Linguistics. Zihang Dai, Guokun Lai, Yiming Yang, and Quoc V. Le. 2020. Funnel-transformer: filtering out sequential redundancy for efficient language processing. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, Red Hook, NY, USA. Curran Associates Inc. Peter Daniels and William Bright. 1996. The World’s Writing Systems. Oxford University Press. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language under- standing. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, and 516 others. 2024. The Llama 3 herd of models. Preprint, arXiv:2407.21783. Jiyeon Ham, Yo Joong Choe, Kyubyong Park, Ilji Choi, and Hyungjoon Soh. 2020. KorNLI and KorSTS: New benchmark datasets for Korean natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 422–430, Online. Association for Computational Lin- guistics. Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Is- lam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL- sum: Large-scale multilingual abstractive summariza- tion for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computa- tional Linguistics. Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-rank adap- tation of large language models. arXiv preprint arXiv:2106.09685. Myeongjun Jang, Dohyung Kim, Deuk Sin Kwon, and Eric Davis. 2022. KoBEST: Korean balanced eval- uation of significant tasks. In Proceedings of the 29th International Conference on Computational Lin- guistics, pages 3697–3708, Gyeongju, Republic of Korea. International Committee on Computational Linguistics. Jongho Jun. 2018. Morpho-phonological processes in korean. Ayush Kaushal and Kyle Mahowald. 2022. What do tokens know about their characters and how do they know it? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 2487–2507, Seattle, United States. Association for Computational Linguistics. Nayeon Kim, Jun-Hyung Park, Joon-Young Choi, Eojin Jeon, Youjin Kang, and SangKeun Lee. 2022. Break it down into BTS: Basic, tiniest subword units for Korean. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7007–7024, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. SungHo Kim, Nayeon Kim, Taehee Jeon, and SangKeun Lee. 2025. Polishing every facet of the GEM: Test- ing linguistic competence of LLMs and humans in Korean. In Proceedings of the 63rd Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9955–9984, Vienna, Austria. Association for Computational Linguistics. SungHo Kim, Juhyeong Park, Yeachan Kim, and SangKeun Lee. 2024. KOMBO: Korean character representations based on the combination rules of subcharacters. In Findings of the Association for Computational Linguistics: ACL 2024, pages 5102– 5119, Bangkok, Thailand. Association for Computa- tional Linguistics. Ohjoon Kwon, Dohyun Kim, Soo-Ryeon Lee, Junyoung Choi, and SangKeun Lee. 2021. Handling out-of- vocabulary problem in hangeul word embeddings. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Lin- guistics: Main Volume, pages 3213–3221, Online. Association for Computational Linguistics. Yuxuan Lai, Yijia Liu, Yansong Feng, Songfang Huang, and Dongyan Zhao. 2021. Lattice-BERT: Leverag- ing multi-granularity representations in Chinese pre- trained language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1716–1731, Online. Association for Computational Linguistics. Iksop Lee and S Robert Ramsey. 2001. The korean language. Junyoung Lee, Marco Cognetta, Sangwhan Moon, and Naoaki Okazaki. 2025. Jamo-level subword tokeniza- tion in low-resource Korean machine translation. In Proceedings of the Eighth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2025), pages 66–80, Albuquerque, New Mexico, U.S.A. Association for Computational Lin- guistics. Kyong-Nim Lee and Minhwa Chung. 2003. Modeling cross-morpheme pronunciation variations for korean large vocabulary continuous speech recognition. In INTERSPEECH, pages 261–264. LG AI Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeong- won Jo, Hyunjik Jo, Jiyeon Jung, Yountae Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, and 14 others. 2024. EXAONE 3.5: Series of large lan- guage models for real-world use cases. Preprint, arXiv:2412.04862. Andrew Matteson, Chanhee Lee, Youngbum Kim, and Heuiseok Lim. 2018. Rich character-level informa- tion for Korean morphological analysis and part-of- speech tagging. In Proceedings of the 27th Inter- national Conference on Computational Linguistics, pages 2482–2492, Santa Fe, New Mexico, USA. As- sociation for Computational Linguistics. Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, and 1 oth- ers. 2021. Between words and characters: A brief history of open-vocabulary modeling and tokeniza- tion in nlp. arXiv preprint arXiv:2112.10508. Sangwhan Moon and Naoaki Okazaki. 2020. Jamo pair encoding: Subcharacter representation-based extreme Korean vocabulary compression for effi- cient subword tokenization. In Proceedings of the Twelfth Language Resources and Evaluation Confer- ence, pages 3490–3497, Marseille, France. European Language Resources Association. Courtney Napoles, Keisuke Sakaguchi, Matt Post, and Joel Tetreault. 2015. Ground truth for grammatical error correction metrics. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Confer- ence on Natural Language Processing (Volume 2: Short Papers), pages 588–593, Beijing, China. Asso- ciation for Computational Linguistics. National Hangeul Museum. 2018.A guide to Hunminjeongeum.https://hangeul.go. kr/user/synapView.jsp?filename=BBS/ 2BD8647E-0CBE-42A9-EEF9-8C5AD8AC70CE.pdf. National Hangeul Museum. 2021.Easy read- ing of Hunminjeongeum.https://hangeul. go.kr/user/synapView.jsp?filename=BBS/ A0479188-1C12-D328-AE4D-B4DF9C279181.pdf. Piotr Nawrot, Szymon Tworkowski, Michał Tyrolski, Lukasz Kaiser, Yuhuai Wu, Christian Szegedy, and Henryk Michalewski. 2022. Hierarchical transform- ers are more efficient language models. In Find- ings of the Association for Computational Linguis- tics: NAACL 2022, pages 1559–1571, Seattle, United States. Association for Computational Linguistics. Kyubyong Park, Joohong Lee, Seongbo Jang, and Da- woon Jung. 2020. An empirical study of tokenization strategies for various Korean NLP tasks. In Proceed- ings of the 1st Conference of the Asia-Pacific Chap- ter of the Association for Computational Linguistics and the 10th International Joint Conference on Nat- ural Language Processing, pages 133–142, Suzhou, China. Association for Computational Linguistics. Lucy Park. 2016.Naver sentiment movie corpus. https://github.com/e9t/nsmc. Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, and 24 others. 2025. Qwen2.5 technical report. Preprint, arXiv:2412.15115. Melissa Roemmele, Cosmin Adrian Bejan, and An- drew S Gordon. 2011. Choice of plausible alter- natives: An evaluation of commonsense causal rea- soning. In 2011 AAAI spring symposium series. Geoffrey Sampson. 2015. Writing Systems. Equinox Publishing Limited. Jaehyung Seo, Seounghoon Lee, Chanjun Park, Yoonna Jang, Hyeonseok Moon, Sugyeong Eo, Seonmin Koo, and Heuiseok Lim. 2022. A dog is passing over the jet? a text-generation dataset for Korean common- sense reasoning and evaluation. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 2233–2249, Seattle, United States. Asso- ciation for Computational Linguistics. Jiyoung Shin, Chi-y ̆ ong Sin, Jieun Kiaer, Chae- ̆ un Ch’a, and Jaeeun Cha. 2012. The sounds of Korean. Cam- bridge University Press. Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Anastasia Kozlova, Vladislav Mikhailov, and Tatiana Shavrina. 2024. mGPT: Few-shot learners go multi- lingual. Transactions of the Association for Compu- tational Linguistics, 12:58–79. Ho-Min Sohn. 2001. The Korean language. Cambridge University Press. Unicode Consortium, editor. 2019. The Unicode Stan- dard, Version 12.0 – Core Specification. Unicode, Inc., Mountain View, CA. Core specification PDF. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Pro- cessing Systems, volume 30. Curran Associates, Inc. Yilin Wang, Xinyi Hu, and Matthew Gormley. 2024. Learning mutually informed representations for char- acters and subwords. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 3201–3213, Mexico City, Mexico. Association for Computational Linguistics. Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sen- tence understanding through inference. In Proceed- ings of the 2018 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguis- tics. Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversar- ial dataset for paraphrase identification. In Proceed- ings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter- national Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguis- tics. Jaehoon Yeon and Lucien Brown. 2013. Korean: A Comprehensive Grammar, 2 edition. Routledge. Kang Min Yoo, Jaegeun Han, Sookyo In, Heewon Jeon, Jisu Jeong, Jaewook Kang, Hyunwook Kim, Kyung-Min Kim, Munhyong Kim, Sungju Kim, Donghyun Kwak, Hanock Kwak, Se Jung Kwon, Bado Lee, Dongsoo Lee, Gichang Lee, Jooho Lee, Baeseong Park, Seongjin Shin, and 377 others. 2024. HyperCLOVA X technical report. Preprint, arXiv:2404.01954. Soyoung Yoon, Sungjoon Park, Gyuwan Kim, Junhee Cho, Kihyo Park, Gyu Tae Kim, Minjoon Seo, and Alice Oh. 2023. Towards standardizing Korean gram- matical error correction: Datasets and annotation. In Proceedings of the 61st Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 6713–6742, Toronto, Canada. Association for Computational Linguistics. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Wein- berger, and Yoav Artzi. 2019. BERTScore: Evalu- ating text generation with BERT. arXiv preprint arXiv:1904.09675. Shan Zhao, ChengYu Wang, Minghao Hu, Tianwei Yan, and Meng Wang. 2023. MCL: Multi-granularity con- trastive learning framework for chinese ner. In Pro- ceedings of the AAAI Conference on Artificial Intelli- gence, pages 14011–14019. A Composition of Hangul Characters Hangul characters are constructed by combin- ing initial consonants (Choseong), medial vow- els (Jungseong), and, optionally, final consonants (Jongseong). The following outlines the possible components for each position: Initial Consonants (Choseong) The initial position can be occupied by one of the following 19 consonants: ᄀ ᄁ ᄂ ᄃ ᄄ ᄅ ᄆ ᄇ ᄈ ᄉ ᄊ ᄋ ᄌ ᄍ ᄎ ᄏ ᄐ ᄑ ᄒ Medial Vowels (Jungseong) The medial position consists of 21 vowels, which can be categorized based on their placement rela- tive to the initial consonant: •Vertical vowels (placed to the right of the initial consonant):ᅡ,ᅣ,ᅥ,ᅧ,ᅵ,ᅢ,ᅤ, ᅦ,ᅨ •Horizontal vowels (placed below the initial consonant): ᅩ,ᅭ,ᅮ,ᅲ,ᅳ •Complex vowels (combining both vertical and horizontal elements):ᅪ,ᅫ,ᅬ,ᅯ,ᅰ, ᅱ,ᅴ These three categories correspond to the examples illustrated in Figure 6. This classification results in three distinct types of character shapes, each reflecting the spatial arrangement dictated by the vowel’s orientation. Final Consonants (Jongseong) The final position may be unoccupied or contain one of the following 27 consonant combinations. means the empty final consonant and it is also considered a valid configuration: ᄀ ᄁ ᆪ ᄂ ᆬ ᆭ ᄃ ᄅ ᆰ ᆱ ᆲ ᆳ ᆴ ᆵ ᄚ ᄆ ᄇ ᄡ ᄉ ᄊ ᄋ ᄌ ᄎ ᄏ ᄐ ᄑ ᄒ Syllable Structure A typical Hangul syllable block is formed in one of the following structures: • CV: Consonant + Vowel (e.g.,가) •CVC: Consonant + Vowel + Consonant (e.g.,갈) The positioning of vowels within the syllable block depends on their type: •Vertical vowels are placed to the right of the initial consonant. •Horizontal vowels are placed below the ini- tial consonant. •Complex vowels may occupy both right and bottom positions relative to the initial conso- nant. This systematic arrangement allows for the con- struction of 11,172 possible syllable combinations in Hangul. B Details for Inflection Frequency Evaluation Action Definition.The alignment oracle (Matte- son et al., 2018) aligns surface forms with lemma sequences by assigning one or more actions to each input character: • KEEP: retain the character unchanged. •MOD: modify the character into a different form (e.g., adding a final consonant, changing a vowel, or altering the initial consonant). •NOOP: drop the character, i.e., it does not appear in the lemma. Each action is augmented with BIO prefixes: “B- ” marks the beginning of a morpheme, while “I-” denotes continuation within a morpheme. For in- stance,B-KEEPindicates the start of a morpheme where the character is preserved, whileB-MOD-ᄂ signals a morpheme-internal modification introduc- ing the consonant “ᄂ.” Corpus Preprocessing.We first parse an oracle- aligned action file in which each line contains a single input character and its action sequence (possibly multiple actions for one character), sepa- rated by a fixed delimiter. Non-Hangul characters are filtered out using a Hangul checker (we use hgtk.checker.is_hangul 6 ). For each Korean character, we collect the raw action string and in- crement counters forKEEP,MOD, orNOOPdepending on whether the action string contains these tokens. BIO prefixes (B-/I-) are preserved to indicate mor- pheme boundaries. (Matteson et al., 2018) 6 https://pypi.org/project/hgtk/ C Type 1 I F V Type 2 I F V Type 3 I V (e.g.,)| | F 관,될,의금,물,초감,집,하 Choseong (Initial consonant) Jungseong (Vowel) Jongseong (Final consonant) I V F C Character Subcharacter(Jamo) Figure 6: Three structural types of Korean syllable blocks, classified by the spatial arrangement of Choseong, Jungseong, and Jongseong. Input CharacterOracle ActionsOutput Lemma Units 런 (reon)B-MOD-럽, I-MOD-ᄂ럽 (reob),ᄂ (n) 했 (haess)B-MOD-하, I-MOD-았하 (ha),았 (ass) 다 (da)B-KEEP다 (da) Table 5: Examples of oracle actions aligned with lemma units. Each input character may correspond to multiple oracle actions (e.g., B-MOD, I-MOD) depending on the type and structure of the morphological transformation. The output shows the resulting lemma units produced by these actions. Why Focus on MOD.Our analysis concentrated on “MOD” actions because they directly represent the sites of morphophonological change. “KEEP” characters reflect unchanged segments and “NOOP” characters denote deletions, both of which con- tribute little information about how Korean mor- phology operates. In contrast, “MOD” characters capture precisely the subcharacter or character- level transformations (e.g., tense, honorifics, ad- nominal endings) that are central to Korean gram- mar. By quantifying only “MOD,” we obtain a clearer picture of where and how morphophonolog- ical alternations occur. Granularity Classification (Subcharacter vs Character). We classify eachMODinstance into two levels: • Subcharacter-level: exactly one of Choseong (I), Jungseong (V), Jongseong (F) differs be- tween the input and the aligned output, or a cross-syllable transfer/merge occurs, consis- tent with Korean fusion rules across character boundaries. (Matteson et al., 2018) •Character-level: entire syllable box is replaced wholesale (non-comparable at the subcharacter- level). In our implementation, we filter Korean characters (hgtk.checker.is_hangul) and compute counts ofKEEP/MOD/NOOP. ForMODcases, we determine the granularity by comparing the input character’s Unicode-decomposed (I, V, F) tuple with that of its aligned outputs. When a character yields mul- tiple actions, we first reconstruct the immediate output units produced by that character’s actions and then compare at the subcharacter-level. If only one subcharacter differs (or a Jongseong-Choseong transfer is observed), we mark it as subcharacter- level; otherwise, character-level. NOOP Handling.NOOPmarks deletions (input characters with no aligned lemma). As deletions indicate absence rather than transformation, we excludeNOOPfrom granularity statistics; counts are still reported for completeness. (Alternatively, NOOPcan be treated as character-level; we opted for exclusion.) C Implementation Details and Model Considerations As shown in Figure 3(b), our proposed method, SCRIPT, can be simply plugged into the embedding layer, making it easy to apply to existing PLMs. This provides a model-agnostic advantage, allow- ing seamless integration with various architectures. As an extra implementation detail, unlike the de- coder model, such as GPT, the encoder model, like BERT, has a particular design of the input sequence. BERT adds a special token[CLS]at the beginning of the input sequence, and this token is used as the representation of the sequence at the last-layer Algorithm 1: Subword Representation with SCRIPT Input: s (raw input sentence), embedding dimension 퐷. Output: e F (fused subword-level embeddings). 1 t subchar ← TOKENIZE subchar (s)// subcharacter sequence 2 e← EMBEDDING subchar (t subchar )// subcharacter representation 3 if first token in e is [CLS] then// (for encoder-only model) e CLS ← e[1] e← e[2 : 푁] 4 h← GRU(e)// sequential composition 5 for each for each integer 푘 in [1, N/3]: h I ← h[3푘− 2]// Choseong h V ← h[3푘− 1]// Jungseong h F ← h[3푘]// Jongseong 6 h I+V ← GRU(h 퐼 +h 푉 )// fusion of Choseong and Jungseong 7 h R ← STACK(h 퐼 +h 푉 , h 퐹 )// reshape subcharacter sequence 8 h C ← AVGPOOL(CONV 2×1 (h R ))// character representation 9 h S ← POOLING(GRU(h C ))// subword representation of SCRIPT 10 t subword ← TOKENIZE subword (s)// subword sequence 11 e S ← EMBEDDING subword (t subword )// original subword representation 12 e F ← CROSSATTN(Q=e S , K=h S , V=h S )// fusion of subword representations 13 if e CLS exists then e F ← [e CLS ; e F ] return e F Table 6: The input textsis tokenized into subcharacter tokens (line 1) and subword tokens (line 10), respectively (§ 3.2). While structure-aware subword embeddings are built bottom-up from subcharacter to character to subword (lines 2-9), the original subword embeddings are looked up (lines 10-11). These two subword representation streams are fused by cross-attention to yield the final subword representation e F (lines 12-13) (§ 3.3). hidden state. To preserve the special structure and meaning of the special token during the compres- sion stage, we first separate the[CLS]token from the rest of the tokenized sequence (line 3 in Ta- ble 6). Then, we only use the remaining subcharac- ter sequence as input toSCRIPT. After processing throughSCRIPT, we add the[CLS]hidden state back to the output ofSCRIPT, the compressed sub- word representation (line 13 in Table 6). Formally, this can be represented as h S ∈R (푁 ′ +1)×D : h S = EMB subchar (t 0 )⊕ SCRIPT(t 2:푁+1 )(5) This trivial technique provides the versatility for our proposed methodology,SCRIPT, to be easily applied across all existing PLMs. In the following experimental section, we demonstrate the utility of our methodology by applyingSCRIPTto both pre-trained encoder-only and decoder-only models. Another important consideration is the integra- tion of the two distinct subword representations, which come from different levels of granularity. This process is highly sensitive to normalization, as it can easily disrupt compatibility between the pre-trained subword embeddings and those derived from subcharacter representations. Notably, our empirical analysis shows that the pre-trained em- beddings are approximately 30 times larger in size than those generated bySCRIPT. Normalizing these vectors to a common scale can distort the distribu- tion of the pre-trained embeddings, with the ran- domly initialized subcharacter representations in- troducing significant noise and resulting in the loss of crucial learned information. Our analysis shows that structure-aware subword knowledge is gradu- ally fine-tuned and transferred into the pre-trained embeddings during integration. Therefore, nor- malizing the two representations to the same scale Transformer Stacks ‘ 대한민국 South Korea ’ Fusion of Two Distinct Subword Representations 니 ▃ ᄋ l ᄂ ᄆ ᅵᄂᄀ ᅳ ᄀ 대한민국 대한민국 Subcharacter(BTS Unit)Embeddings Addition of Jongseong Pooling to Subword Pre-trained Subword Embeddings SCRIPT [ I 1 ][ V 1 ][ F ][ I 1 ][ V 1 ][ F ][ I ][ V ][ F ][ I ][ V 1 ][ F ][ S ][ S ] - - [ I 2 ][ I 2 ] · [ V 2 ] ᅵ [ V 3 ] · [ V 2 ] Fusion of Choseong and Jungseong 니 ▃ 이 ᄂ ᄆ ᅵ ᄂᄀ ᅳ ᄀ- -· ᅵ· 대구 ▃ ᄀ 미ᄂ ᄂ하 대한민국 대한민국 · [ V 2 ] · Figure 7: Implementation ofSCRIPTfor the BTS unit. This illustrates the hierarchical integration of subword representations derived from BTS unit inSCRIPT, using the example word ‘대한민국 South Korea ’. The word ‘대한민 국 South Korea ’ consists of two subwords ([S]:대한,민국), four characters (대,한,민,국), and eighteen subcharacters: initial consonants ([ I ]:ᄂ, -,ᄋ, -,ᄆ,ᄀ), vowels ([V]:ᅵ,·,ᅵ,ᅵ,·,ᅵ,ᅳ,·), and final consonants ([F]:, ᄂ,ᄂ,ᄀ). before fusion risks compromising the effectiveness of the pre-trained subword embeddings; it is empir- ically more effective to use them as they are. D Implementation Details for BTS Units Unlike Section 3.2, where we introduced the archi- tecture ofSCRIPTusing Jamo as the base unit, this section extends the discussion to describeSCRIPT in terms of other subcharacters, such as BTS units. As noted in Section 3, BTS units are alternative sub- character representations of Hangul, decomposing each character into even finer subcomponents than Jamo. According to Kim et al. (2022), consonants can be split into up to four subcomponents (e.g., the consonant ‘ᄍ’ decomposes into ᄉ, -,ᄉ, - , while vowels can be split into up to five subcom- ponents (e.g., the vowel ‘ᅫ’ decomposes into ·, ᅳ,ᅵ,·,ᅵ. Depending on the selective decom- position of consonants or vowels, there are three distinct types of units: consonant-only decompo- sition (denoted as Stroke), vowel-only decomposi- tion (denoted as Cji, short for Cheonjiin), and both consonant and vowel decomposition (denoted as BTS). As a result, the maximum number of tokens per character varies depending on the decomposi- tion type: the Stroke decomposition results in up to 9 subcharacter tokens, the Cji decomposition yields up to 7 tokens, and the BTS decomposition produces up to 13 tokens. To incorporate the information of BTS units into the original subword representation of PLMs, we follow the progressive steps outlined in Section 3.2 in a similar manner. For simplicity, we employ BTS as the initial subcharacter unit ofSCRIPTin the following explanation. We first tokenize the input textsinto subcomponents, such as BTS, then project the resulting subcharacter sequencet subchar into a subcharacter embedding space. A GRU layer is applied sequentially for contextualization: t subchar = TOKENIZE subchar (s) ∈R 푁 (6) e= EMB subchar (t) ∈R 푁×퐷 (7) h= GRU(e) ∈R 푁×퐷 (8) Next, we merge the subcomponent tokens to con- struct representations for Choseong, Jungseong, and Jongseong. For each integer푘 ∈ [1, 푁/13], the representations of Choseong (ℎ I,푘 ), Jungseong (ℎ V,푘 ), and Jongseong (ℎ F,푘 ) are defined as follows: h I,푘 = 4 ∑︁ 푗=1 h 13(푘−1)+푗 ∈R 푁 13 ×퐷 (9) h V,푘 = 9 ∑︁ 푗=5 h 13(푘−1)+푗 ∈R 푁 13 ×퐷 (10) h F,푘 = 13 ∑︁ 푗=10 h 13(푘−1)+푗 ∈R 푁 13 ×퐷 (11) Next, we combine the representation of Choseong h I and Jungseong h V : h I+V = h I + h V ∈R 푁 13 ×퐷 (12) After that, we vertically concatenate the combined representationh I+V with Jongseong representation h F : h R = h I+V h F ∈R 2× 푁 13 ×퐷 (13) To generate the dense character representation, we merge these vertically aligned representations by applying a convolution and a pooling layer: h C = AVGPOOL(CONV(h R )) ∈R 푁 13 ×퐷 (14) Finally, a GRU layer is applied to the character representationsh C , followed by a pooling layer to compress and align the granularity of the character representations with that of the original subword representations: h S = POOLING(GRU(h C )) ∈R 푁 ′ ×퐷 (15) where푁 ′ represents the number of tokens of the original subword sequence. This compressed sub- word representationh S is fused with the original subword representatione S through cross-attention, as described in Section 3.3, to produce the final subword representation. E Experimental Settings E.1 Baselines To evaluate the effectiveness of our proposed method,SCRIPT, we apply it to various PLMs listed below. Models equipped with our method are denoted as ‘model+SCRIPT’. When needed, the subcharacter type (e.g., Jamo) used inSCRIPTis indicated as a subscript, as inSCRIPT Jamo . If no subscript is provided, Jamo is used by default. •BERT base (Devlin et al., 2019): A bidirec- tional language model based on the Trans- former architecture.It consists of multi- ple Transformer encoder layers. It is pre- trained in a self-supervised manner, enabling it to learn without labeled data.We uti- lize the BERT base model, which includes 12 Transformer encoder layers. It has a total of 110 million parameters. We employ a morpheme-aware tokenizer (Park et al., 2020) with a vocabulary size of 32k.We pre- train the BERT base model for 1 million steps on Masked Language Modeling (MLM) and Next Sentence Prediction (NSP) tasks, using a corpus of 6.2GB consisting of the Korean Wikipedia and Namuwiki. 7 • KOMBO base (Kim et al., 2024): A Jamo- based Korean encoder-only PLM that lever- ages the invention principles of Hangul to rep- resent characters. While the architecture of KOMBO base is designed based on BERT base , it differs in that it uses subcharacter-level to- kens instead of subwords. It also includes a combination layer below and a restoration layer above its 12 Transformer blocks, intro- ducing additional computational cost and over- head. However, it achieves better performance on NLU tasks than BERT base . Moreover, since both BERT base and KOMBO base models are encoder-only models, they do not apply to gen- erative tasks. We pre-train KOMBO base for 1 million steps on MLM and NSP tasks using a 6.2GB corpus from Korean Wikipedia and Namuwiki, following the same configuration as BERT base . • KoGPT2 base 8 : A generative language model composed of multiple Transformer decoder blocks. Unlike the BERT model, which is trained on MLM and NSP tasks, GPT is trained on a next token prediction task, en- abling it to generate contextually relevant text. KoGPT2 is a Korean variant of the GPT model, following the GPT2 base configuration with 12 Transformer decoder blocks and 125 million parameters. It has been pre-trained on the Korean Wiki and Korpora datasets, to- tally over 40GB. KoGPT2 employs subword tokenization with a vocabulary size of 51.2k. 7 https://namu.wiki/ 8 https://github.com/SKT-AI/KoGPT2 •KoGPT3-1.2B 9 : A large-scale Transformer decoder model containing 1.2 billion parame- ters and 24 Transformer decoder blocks. This model follows the GPT-3 architecture. The model is trained on Ko-DAT, a large-scale, curated Korean dataset created by SK Tele- com with 35 billion tokens, using the next token prediction task over 72k training steps. Similar to KoGPT2 base , KoGPT3-1.2B uses subword tokenization with a 51.2k vocabulary size. •mGPT-1.3B (Shliazhko et al., 2024): The mul- tilingual extension of GPT-3 which is pre- trained across 61 languages. It consists of 24 Transformer decoder layers with a total of 1.3 billion parameters. Pre-training was con- ducted on the Wikipedia and C4 corpora, over 600GB in total, for 600k steps. The model employs a 100k size vocabulary and utilizes Byte-level Byte Pair Encoding (BBPE) as its default tokenization strategy, enhancing its multilingual capabilities. •EXAONE-3.5-2.4B-Instruct (EXAONE-2.4B for short) (LG AI Research et al., 2024): A bilingual (Korean and English) instruction- tuned language model developed by LG AI Research. It uses a decoder-only Transformer architecture and is part of the EXAONE 3.5 series. The model has 30 Transformer decoder layers and 2.41 billion parameters total. It was trained under a causal/next-token prediction objective using a bilingual corpus curated by LG AI Research, supports a maximum con- text length of 32,768 tokens, and employs a shared vocabulary of 102,400 tokens using a BBPE tokenizer. E.2 Implementation Details We utilized a series of pre-trained GPT models, in- cluding KoGPT2 base , KoGPT3-1.2B, mGPT-1.3B, and EXAONE-2.4B all sourced from the Hugging- face library 10 , and fine-tuned them using LoRA (Low-Rank Adaptation) (Hu et al., 2021). How- ever, for BERT-based models, such as BERT base and KOMBO base , training with LoRA showed insta- bility, so we opted for full fine-tuning exclusively for these models. As noted by Hu et al. (2021), there was no significant difference in performance 9 https://huggingface.co/skt/ko-gpt-trinity-1. 2B-v0.5 10 https://huggingface.co/models between LoRA and full fine-tuning. Since we uti- lize the DeepSpeed library 11 for models larger than 1B parameters, all models were trained on a single NVIDIA RTX 3090 GPU. We set the default maxi- mum sequence length to 256 for the original PLM’s subword tokenizer, and to 2048 for the subcharacter tokenizer used in KOMBO base andSCRIPT Jamo . For tasks requiring longer inputs, such as HellaSwag and XL-Sum, we used sequence lengths of 512 and 3072, respectively. We use theAdamWoptimizer and cosine learning rate scheduler. For most other experimental settings, we used the default config- urations of each pre-trained model. The detailed hyperparameter settings are summarized in Table 7. All experiments are repeated over 3 random seeds (42–44), and we report the mean. E.3 Tasks Korean NLU Tasks. To investigate the perfor- mance of our proposed method on Korean NLU tasks, we evaluated baselines on nine distinct Korean NLU datasets. Four of these, KorNLI, KorSTS, NSMC, and PAWS-X, are widely used benchmarks for Korean NLU tasks, which we refer to as “Korean Standard NLU Tasks” (Jang et al., 2022). The remaining five datasets belong to the KoBEST benchmark (abbreviated as KB), which is designed to evaluate Korean language models on more complex linguistic understanding. We refer to these as “Korean Advanced NLU Tasks” (Jang et al., 2022). •KorNLI (Ham et al., 2020): A dataset com- prising 943k train, 25.5k validation, and 5k test samples for NLI, derived from the SNLI (Bowman et al., 2015), MNLI (Williams et al., 2018), and XNLI (Conneau et al., 2018) datasets. The data is labeled across three classes: entailment, neutral, and contradic- tion. • KorSTS (Ham et al., 2020): A dataset devel- oped to assess the semantic similarity between sentence pairs, adapted from the Korean STS- B dataset (Cer et al., 2017). KorSTS consists of 5,749 train samples and 2,879 evaluation samples, each labeled with a similarity score from 0 to 5, indicating the degree of semantic similarity between the sentences. • NSMC (Park, 2016): A dataset sourced from NAVER is used for sentiment analysis of Ko- 11 https://github.com/deepspeedai/DeepSpeed TaskEpoch Batch Size Learning Rate Dropout Ratio Warmup Ratio LoRA rLoRA 훼 KorNLI564 BERT: 5e-05, 1e-04 GPT : 1e-05, 5e-05, 1e-04, 1e-03, 3e-03, 1e-02 0.030.132128 KorSTS1564 NSMC564 PAWS-X1064 BoolQ108 BERT: 1e-05, 5e-05 GPT : 1e-05, 5e-05, 1e-04, 1e-03, 3e-03, 1e-02 0.030.132128 COPA1516 WiC 1516 HellaSwag108 SentiNeg1064 KoCommonGen1564 GPT : 1e-04, 1e-03, 1e-02, 2e-02, 3e-02, 4e-02 0.030.132128 XL-Sum 1064 Kor-Learner1064 Kor-Native 1064 Table 7: Hyperparameters used in all experiments in this paper for each task. For the learning rate, we select the value that yields the best performance for each baseline. Here, “BERT” refers to encoder-only models, including BERT and KOMBO, while “GPT” encompasses all decoder-only models, such as KoGPT2, KoGPT3, mGPT, and EXAONE. rean movie reviews. It includes 150k train samples and 50k test samples, with each re- view labeled as either negative or positive. • PAWS-X (Yang et al., 2019): A dataset for paraphrase identification, which includes six different language tasks. We only use the Ko- rean subset to evaluate models. It contains 53k sentence pairs (49k for train, 2k for de- velopment, and 2k for test), each data labeled with one of two values: different meanings or paraphrases. •KoBEST (Jang et al., 2022): A benchmark suite designed to evaluate broad linguistic and cognitive capabilities of Korean language models through five diverse and challenging tasks. –KB-BoolQ (Jang et al., 2022): A dataset of 3.7k train, 700 validation, and 1.4k test instances. The task is a true/false question and answer format based on paragraphs, with sources from Korean Wikipedia. –KB-COPA (Jang et al., 2022): A dataset includes 3.1k train, 1k validation, and 1k test instances. Models predict cause or effect given a premise, designed similarly to the English COPA dataset (Roemmele et al., 2011). –KB-WiC (Jang et al., 2022): A dataset contains 3.3k train, 1.3k validation, and 1.3k test samples, requiring models to de- termine if a target word holds the same meaning across two contexts. – KB-HellaSwag (Jang et al., 2022): A dataset composed of 2k train, 500 valida- tion, and 500 test examples, where models select the most probable sentence to follow a given context. The data is sourced from YouTube and Wikipedia. –KB-SentiNeg (Jang et al., 2022): A dataset for sentiment analysis (3.6k train, 400 val- idation, 397 test samples) focusing on the polarity of negated sentences in product reviews. Korean NLG Tasks. We used three distinct benchmarks, KoCommonGen, XL-Sum, and Ko- rean Grammatical Error Correction (GEC), to eval- uate the performance of our proposed method on Korean NLG tasks. The detailed explanations of each benchmark are provided below: •KoCommonGen (Seo et al., 2022): A genera- tive commonsense reasoning dataset compris- ing 43,188 train samples and 2,040 test exam- ples. Given a set of morphemes, the model composes a sentence that reflects commonsense knowledge. •XL-Sum (Hasan et al., 2021): A summariza- tion dataset used to evaluate models’ ability to generate concise and accurate summaries from large text bodies. We focus on the Korean sub- set, which includes 4,407 train samples and 550 validation and test samples. We are only using Model XL-Sum BLEU 3BLEU 4ROUGE-2ROUGE-LMETEORmBERTScoreKoBERTScore KoGPT2 base 7.274.9812.9126.8313.4376.2288.97 KoGPT2 base + SCRIPT7.645.2013.3027.2413.6776.3189.20 KoGPT3-1.2B9.146.2115.8830.1816.9176.9589.69 KoGPT3-1.2B + SCRIPT9.396.4015.9930.3116.9777.4689.77 Table 8: Performance on XL-Sum multilingual summarization task. We evaluate only on the Korean summarization dataset. We use seven automatic evaluation metrics, including n-gram-based measures like BLEU, ROUGE, and METEOR; and two BERT-based scores (Zhang et al., 2019), mBERTScore and KoBERTScore, for semantic similarity. The global-best results are highlighted in boldface and local-best results for each model are highlighted in underline, respectively. Model Kor-LearnerKor-Native 푀 2 푝푟푒 푀 2 푟푒푐 푀 2 퐹 0.5 GLEU 푀 2 푝푟푒 푀 2 푟푒푐 푀 2 퐹 0.5 GLEU KoGPT2 base 29.35 16.11 25.1921.6072.12 55.19 67.7661.45 KoGPT2 base + SCRIPT30.02 16.8325.3423.5472.9556.1569.0062.25 KoGPT3-1.2B47.05 23.01 38.8935.2284.76 69.54 81.2075.21 KoGPT3-1.2B + SCRIPT 49.45 26.3641.6839.4585.4970.0281.8775.40 Table 9: Performance on two Korean GEC tasks. As the evaluation metrics, we use푀 2 scorer (Dahlmeier and Ng, 2012), which measures precision, recall, and F 0.5 scores based on edits and GLEU (Napoles et al., 2015) for the simple n-gram matching. The global-best results are highlighted in boldface and local-best results for each model are highlighted in underline, respectively. the Korean subset of XL-Sum dataset for our experiments. •Korean GEC (Yoon et al., 2023): A grammat- ical error correction dataset for Korean. It consists of four sub-datasets: three standalone datasets, such as Kor-Learner, Kor-Lang8, and Kor-Native, and one aggregated dataset, Kor- Union. Kor-Learner offers a more structured and reliable dataset for the Korean GEC task, as it is annotated by Korean language tutors. In contrast, Kor-Lang8 was corrected by native speakers through an open online platform. In this paper, we focus on two orthogonal GEC tasks, such as Kor-Learner and Kor-Native, to more accurately evaluate model performance relative to dataset complexity. –Kor-Learner GEC (Yoon et al., 2023): A GEC dataset for Korean learner texts, con- taining 19,898 train sentences, 4,264 valida- tion sentences, and 4,265 test sentences. Kor- Learner contains learner-written essays that have been carefully corrected and annotated by Korean tutors. It aids in identifying and correcting grammar errors specific to Korean language learners. – Kor-Native GEC (Yoon et al., 2023): A GEC dataset targeting native Korean texts to sup- port advanced linguistic understanding. It comprises 12,292 train sentences, 2,634 vali- dation sentences, and 2,634 test sentences. F Evaluation on Generative Tasks F.1 XL-Sum As shown in Table 8,SCRIPTconsistently demon- strated its effectiveness on n-gram metrics. How- ever, the performance improvement observed in the summarization task was somewhat smaller com- pared to that in the commonsense generation task. This difference arises because the summarization task typically involves input and output sentences that are approximately five times longer than those in the commonsense generation task, such as Ko- CommonGen. This suggests thatSCRIPTis partic- ularly effective at generating concise, well-formed sentences, demonstrating its strength in handling shorter and more focused outputs. F.2 Korean GEC Our proposed method,SCRIPT, also demonstrated the strongest performance on Korean grammatical Model KoBEST Avg. BoolQCOPAWiCHellaSwagSentiNeg KoGPT3-1.2B77.3282.8072.7878.9096.3181.62 KoGPT3-1.2B + SCRIPT77.6382.8074.6579.3096.4882.17 mGPT-1.3B71.1969.7068.3876.1089.4074.95 mGPT-1.3B + SCRIPT70.7270.6069.1776.7090.4275.52 Table 10: Performance of KoGPT3-1.2B and mGPT-1.3B on KoBEST benchmark. The evaluation metrics for each task are accuracy (%). The global-best results are highlighted in boldface and local-best results for each model are highlighted in underline, respectively. Model KoCommonGen Avg. BLEU 3 BLEU 4 ROUGE-2 ROUGE-L METEOR mBERTScore KoBERTScore KoGPT3-1.2B26.1917.2058.8562.5352.1185.4191.1756.21 KoGPT3-1.2B + SCRIPT 28.8919.5859.2864.8052.3786.2691.7857.57 mGPT-1.3B15.168.0737.7750.9933.3180.1789.3444.97 mGPT-1.3B + SCRIPT16.599.1139.3152.6834.7780.8289.9546.18 Table 11: Performance of KoGPT3-1.2B and mGPT-1.3B on KoCommonGen dataset. We use eight automatic eval- uation metrics: BLEU, ROUGE, and METEOR for n-gram-based measures; and mBERTScore and KoBERTScore for semantic similarity. The global-best results are highlighted in boldface and local-best results for each model are highlighted in underline, respectively. error correction tasks. As shown in Table 9, it con- sistently outperformed the base model in both the Kor-Learner and Kor-Native tasks, showing par- ticularly strong effectiveness in the Kor-Learner task with average improvements exceeding an av- erage of 3.2%p gains over the global-best perform- ing baseline. According to Yoon et al. (2023), the Kor-Learner task contains a large proportion of er- rors related to particles, endings, and conjugations compared to the Kor-Native task. As mentioned in Section 2, linguistic variations in Korean fre- quently occur at the subcharacter-level. Therefore, the substantial performance gains observed on the Kor-Learner task demonstrate the effectiveness of our core approach: integrating subcharacter compo- sitional information into subword representations. This result further validates thatSCRIPTis highly suitable for the Korean language and effectively enriches the PLM’s subword representations. G Effectiveness of Multilingual Model Our proposed method,SCRIPT, can be seamlessly integrated into the embedding layer of any model and is broadly applicable in multilingual settings. To demonstrate its effectiveness for Korean in mul- tilingual models, we evaluated it on both a Korean monolingual model and a multilingual model that supports Korean. Specifically, we used KoGPT3- 1.2B as the monolingual baseline and mGPT-1.3B, a multilingual model with comparable architecture and scale. For the NLU task, we employed the KoBEST benchmark to assess knowledge under- standing, and for the NLG task, we used KoCom- monGen to evaluate complex knowledge genera- tion, including commonsense reasoning. As shown in Table 10 and Table 11, multilin- gual models withSCRIPTlargely outperformed their base counterparts in both KoBEST (NLU) and KoCommonGen (NLG) tasks. Notably, mGPT- 1.3B achieved a performance gain of approximately 0.6%p on KoBEST benchmarks and 1.2%p on Ko- CommonGen, closely mirroring the improvements observed in the monolingual model. Gains in gen- erative tasks were nearly twice as large as those in understanding tasks, indicating thatSCRIPTis particularly effective in enhancing generative ca- pabilities for Korean. These findings highlight the promise of scaling up to significantly larger and more extensively pre-trained multilingual genera- tive models, such as the Llama (Dubey et al., 2024) and Qwen (Qwen et al., 2025) series. In particular, this substantial improvement in multilingual mod- els is especially valuable given the current scarcity of specialized pre-trained LLMs for Korean. H Qualitative Analysis for Generations To analyze the quality of machine-generated text, in Table 12, we conducted a qualitative analysis for each generative task by model. We estab- lished two baselines for comparison: the basic KoGPT2 base model and KoGPT2 base +SCRIPT. This analysis covered all Korean NLG tasks performed in Section 4.2. Since the input text for the sum- marization task, XL-Sum, is quite long, we have included examples in the H.2 for further details if needed. H.1 KoCommonGen Given the set of morphemes as input to the model, it generates the sentence as output by including the morphemes. As a result of the experiments, shown in Table 12, we observed that KoGPT2 base +SCRIPT model correctly generated appropriate particles by identifying the characteristics and position of ob- jects, such as rail, train, and road. In contrast, the baseline model, KoGPT2 base , incorrectly generated the position of the word ‘train’ as ‘beside the tracks’ instead of ‘on the tracks’. In English, the differ- ence between ’beside’ and ’on’ involves several letters, whereas in Korean, this distinction is very subtle, differing by only a single subcharacter, ‘ᅴ’ (‘옆의’) for ‘on’ and ‘ᅦ’ (‘옆에’) for ‘beside’, which makes it more challenging to distinguish. This shows thatSCRIPTeffectively captures subtle nuances at the subcharacter-level. H.2 XL-Sum The summarization performance appears similar, but the base model tends to generate slightly longer sentences. Overall, our proposed method produces more concise summaries. For example, in this task’s sample data, while KoGPT2 base focused on the ’act of collecting samples’, including the ’lu- nar landing’, our model emphasized the ’success- ful completion of the exploration’, generating sen- tences with a clearer focus on summarization itself. H.3 Kor-Learner As mentioned earlier in Section 4.2, the Kor- Learner dataset contains a higher frequency of er- rors related to particles, endings, and conjugations compared to other Korean GEC datasets. The sam- ple in Table 12 also requires corrections for gram- matical errors in endings. While the KoGPT2 base model failed to correct these properly, our proposed method successfully adjusted endings by consider- ing their agreement with predicates. As shown in Figure 1, understanding these types of grammatical errors is particularly important for Korean. Thus, our proposed method is both suitable and essential for effectively handling Korean. H.4 Kor-Native Through examples from the Kor-Native task, we confirmed that our proposed method,SCRIPT, en- hances the ability to handle whitespace and noun recognition effectively. As shown in the sample in Table 12, it asked to identify and correct the incorrect word ‘테니그 푡푒푛푖푔 ’ to the appropriate noun ‘테니스 푡푒푛푖푠 ’. While the naive KoGPT2 base model failed to detect the error in this sentence and thus cannot make the necessary correction, the model using our proposedSCRIPTmethod success- fully identified and corrected wrong word to ‘테니 스 푡푒푛푖푠 ’. Although this adjustment involved only a subtle subcharacter-level difference, changing ‘ᄀ 푔 ’ to ‘ᄉ 푠 ’, it once again demonstrated that subword models using larger token units than character-level cannot adequately handle such distinctions. I Impact of Fused Representations To examine how compositional knowledge of sub- characters affects subword representations, we compare (i) subcharacter embeddings fromSCRIPT, (i) original subword embeddings from the PLM, and (i) fused subword embeddings augmented by SCRIPT. As shown in Figure 8,SCRIPT’s subchar- acter representations yield the highest similarity among the predicate inflected word sets sharing root semantics but differing at the subcharacter- level. Notably, this advantage transfers to the fused embeddings, which accurately capture these fine- grained relational patterns. These results under- score the effectiveness of the proposed module in modeling subcharacter-level variation through com- positional and representational fusion. TaskLangExample KoCommonGen Ko Input: 있,선로,길,옆,열차 Gold Label: 길옆옆옆의의의선로에에에열차가있다. KoGPT2: 선로가길옆옆옆의의의길옆옆옆에에에열차가세워져있다. KoGPT2 + SCRIPT: 열차들이길옆옆옆의의의선로에에에있다. En Input: be, tracks, road, beside, train Gold Label: There is a train on the tracks beside the road. KoGPT2: The train is parked beside the track beside the road. KoGPT2 + SCRIPT: The trains are on the tracks beside the road. Kor-Learner Ko Input: 그리고가장중요한영향은그앞으로그여행으로이전보다훨씬더‘처음’ 을접할거거거다다다. Gold Label: 그리고가장중요한영향은앞으로여행으로이전보다훨씬 더 ‘처음’ 을접할것것것이이이라라라는는는사사사실실실이이이다다다. KoGPT2: 그리고가장중요한영향을그앞으로그여행으로이전보다훨씬더‘처 음’을접할거거거다다다. KoGPT2 + SCRIPT: 그리고가장중요한영향은그앞으로그여행으로이전보다훨 씬더 ‘처음’을접할거거거라라라는는는것것것이이이다다다. En Input:And the most important impact would that, through that journey, they will encounter the ‘first’ much more than before. Gold Label:And the most important impact is the fact that, through future travels, they will encounter the ‘first’ much more than before. KoGPT2: And the most important impact would that, through that journey, they will encounter the ‘first’ much more than before. KoGPT2 + SCRIPT:And the most important impact is the thing that, through that journey, they will encounter the ‘first’ much more than before. Kor-Native Ko Input: 주말에함께 *테테테니니니그그그를쳐여. Gold Label: 주말에함께테테테니니니스스스를쳐요. KoGPT2: 주말에함께 *테테테니니니그그그를쳐요. KoGPT2 + SCRIPT: 주말에함께테테테니니니스스스를쳐요. En Input: Let’s play *tennig together on the weekend. Gold Label: Let’s play tennis together on the weekend. KoGPT2: Let’s play *tennig together on the weekend. KoGPT2 + SCRIPT: Let’s play tennis together on the weekend. Table 12: Examples for the qualitative analysis of four generation tasks. For each task, examples are composed of the input provided to the model, the gold label, and the predictions generated by two baselines: KoGPT2 base and KoGPT2 base applied withSCRIPT. The model outputs were generated in Korean, with English translations provided alongside for clarity. The asterisk (*) indicates a ungrammatical word. Red-colored characters represent incorrect parts, while blue-colored characters indicate correct representations. TaskLanguageExample XL-Sum Korean Input: 최종목적은2㎏정도의‘토양’표본을상승선,귀환선에전달해지구까지가 져오는것이다중국국가우주국(CNSA)은달의암석과토양표본을수집해지 구로가져오기위해출발한무인달탐사선‘창어5호’가1일밤착륙에성공했 다고 2일밝혔다.창어5호는‘폭풍의바다’(Oceanus Procellarum)라는지역내 ‘몽스륌케르’(Mons Rümker)화산지대북쪽에안착했다.이곳에서며칠간달 표면의흙과암석표본등을수집한다.창어5호탐사선에는작업을돕기위한 카메라 ,레이더,드릴,삽등이탑재돼있다.최종목적은2㎏정도의표토표본 을상승선과궤도선을거쳐귀환선에전달해지구까지가져오는것이다.달의 토양표본을지구로가져온탐사선은 44년전1976년옛소련의루나24호가마 지막으로,당시200g의토양을지구로옮기는데성공했다.창어5호프로젝트 팀이환호하는모습이날달착륙모습은일주일전발사때와달리생중계되 지않았다.중국TV채널에서는성공적인착륙이확인되고나서야정규방송 을중단하고이를녹화중계했다 .공개된착륙과정에는탐사선의다리가달의 먼지쌓인표면에그림자를드리우는장면등이포함됐다. ... KoGPT2: 중국국가우주국이달착륙에성공한창어 6호달착륙선에탑재된카메 라와레이더를통해달표면표본을지구까지운반했다. KoGPT2 + SCRIPT: 중국의달탐사프로젝트가성공적으로마무리됐다. English Input:The primary goal is to bring approximately 2 kg of lunar soil samples back to Earth by transferring them from the ascent and return modules. The China National Space Administration (CNSA) announced on the 2nd that its unmanned lunar probe, Chang’e-5, successfully landed on the night of the 1st to collect lunar rock and soil samples to return to Earth. Chang’e-5 has landed in the volcanic area north of Mons Rümker within the region known as Oceanus Procellarum. Over the next few days, it will collect samples of lunar soil and rock. Equipped with cameras, radar, drills, and shovels to aid in its operations, the ultimate goal of Chang’e-5 is to gather about 2 kg of surface samples, which will be transferred from the ascent and orbital modules to the return module for their journey back to Earth. The last mission to bring lunar soil samples to Earth was the Soviet Union’s Luna 24 in 1976, which successfully transported 200 g of lunar soil back to Earth. On the day of the landing, the Chang’e-5 team celebrated. Unlike the launch, the landing was not broadcast live, and Chinese TV channels interrupted regular programming to air recorded footage after confirming a successful landing. The released landing process included images of the lander casting a shadow on the dusty lunar surface. ... Gold Label: China has landed another probe on the surface of the moon. KoGPT2:The China National Space Administration successfully transported lunar surface samples to Earth using the camera and radar aboard the Chang’e 6 lunar lander. KoGPT2 + SCRIPT: China’s lunar exploration project has been successfully completed. Table 13: Examples for the qualitative analysis of XL-Sum task. For each task, examples are composed of the input provided to the model, the gold label, and the predictions generated by two baselines: KoGPT2 base and KoGPT2 base applied withSCRIPT. The model outputs were generated in Korean, with English translations provided alongside for clarity. SubcharacterEmbeddings (from SCRIPT ) Fused SubwordEmbeddings ( SCRIPT + PLM) SubwordEmbeddings (from the original PLM) Figure 8: Visualization of the similarities between word representations, measured from conjugated word pairs. Each word set contains five words that share the same root meaning. The first word set, ‘춥다’, ‘추움’, ‘추위’, ‘추웠어’, ‘춥디춥다’, conveys the meaning ‘cold’; the second set, ‘걷다’, ‘걷기’, ‘걸어’, ‘걸었어’, ‘걸음’, represents ‘walk’; the third set, ‘돕다’, ‘도움’, ‘도와’, ‘도왔어’, ‘돕기’, signifies ‘help’; and the final set, ‘묻 다’, ‘물어보다’, ‘물었다’, ‘물어보기’, ‘물어’, conveys the meaning ‘ask’. Using these word sets, we compared three different types of embeddings: those derived from subcharacter embeddings inSCRIPT, subword embeddings from the PLM, and fusion embeddings augmented by SCRIPT. J Computational Efficiency J.1 Computational Complexity As shown in Table 14, we quantify the overhead of each approach by breaking computation into three components: the embedding layer, the Transformer stack, and any model-specific layers. In the embedding layer, standard BERT performs a simple embedding lookup and linear projection, which has constant time complexity with respect to sequence length. AddingSCRIPTintroduces mod- est overhead: it applies a GRU-based encoder to contextualize the subcharacter sequence for each token and uses cross-attention to fuse this informa- tion with the original subword embedding. Both components are single-layer operations, in contrast to the deep Transformer stack with more than 10 layers. Moreover,SCRIPTcompresses the subchar- acter sequence before cross-attention, keeping se- quence lengths short during the expensive fusion step. By comparison, KOMBO’s embedding stage is significantly heavier (Ns≪Nj). It processes the full subcharacter sequence with a GRU, three self- attention layers, and another GRU for compression, making its embedding computation far more costly thanSCRIPT. In short, both methods add overhead beyond the base model, butSCRIPTis much lighter due to its efficient compression and fusion strategy. In the Transformer stack,SCRIPTagain aligns closely with the base model. Both the base model (BERT) and BERT+SCRIPToperate at the sub- word level throughout the Transformer layers. This means their self-attention complexity scales with the subword sequence length (푂(푁 s 2 퐷)per layer, where N s is the number of subword tokens and D is the hidden dimension), just as in the original model. In contrast, KOMBO converts inputs into much longer character-level sequences (Nc≫Ns), yielding a per-layer complexity of푂(푁 c 2 퐷). This makes KOMBO’s Transformer blocks slower and more memory-intensive for the same input. Additionally, KOMBO requires a restoration layer after the Transformer stack to convert character-level outputs back to subword repre- sentations, implemented with another GRU. Nei- ther BERT nor BERT+SCRIPTrequires such steps. Thus, aside from a small embedding-stage over- head, BERT+SCRIPTpreserves the base model’s computational profile, whereas KOMBO incurs substantial extra cost in both the Transformer and output stages. J.2 Computational Cost To assess how the computational complexity dis- cussed in the previous section translates into actual computing cost, we empirically evaluate the archi- tectural differences in terms of GPU memory usage and training time. 0100200300400500 Input Sequence Length 10 20 30 40 50 GPU Memory (GB) BERT base BERT base + SCRIPT KOMBO base (a) 0100200300400500 Input Sequence Length 50 100 150 200 250 300 Training Time (sec/epoch) BERT base BERT base + SCRIPT KOMBO base (b) Figure 9: Comparison of computational costs among the base model BERT base , BERT base withSCRIPT, and previous Jamo-based PLM, KOMBO Jamo base . The results were obtained on the KB-HellaSwag benchmark, a rep- resentative Korean NLU task, using a single NVIDIA RTX 3090 GPU. (a) Peak GPU memory usage during training with varying input sequence lengths. (b) Train- ing time per epoch with varying input sequence lengths. GPU Memory. Figure 9(a) summarizes the re- source footprint of each model across varying input sequence lengths. The GPU memory consumption of BERT base +SCRIPTis nearly identical to that of BERT base alone, and it remains far lower than that of KOMBO base , especially for longer sequences. For instance, at an input length of 256 tokens, the BERT base model uses about 7.5 GB of GPU ModelEmbedding Layer Transformer Stacks Restoration Layer BERT푂(1)푂(푁 s 2 퐷)- BERT + SCRIPT 푂(푁 j 퐷 2 + 푁 s 2 퐷)푂(푁 s 2 퐷)- KOMBO푂(푁 j 퐷 2 + 푁 j 2 퐷)푂(푁 c 2 퐷)푂(푁 c 퐷 2 ) Table 14: Comparison of the computational complexities across the three components of the model’s architecture. 푁 j is the length of subcharacter sequence,푁 s is the length of subword sequence,푁 c is the length of character sequence, and 퐷 is the hidden size. memory during training, and BERT base +SCRIPTre- quires approximately 9.2 GB, a relatively small 1.7 GB increase. In contrast, KOMBO base at the same sequence length demands roughly 17.5 GB - more than double the memory of the base model. This gap widens with longer inputs: at 512 tokens, BERT base +SCRIPTuses 18.7 GB vs. 17.6 GB for BERT (only a 6% increase), whereas KOMBO base soars to about 50 GB, nearly three times the base model’s requirement. These results confirm that plug-in design ofSCRIPTadds minimal memory overhead, while KOMBO’s character-level process- ing and extra layers drastically inflate memory us- age for large inputs. Training Time per Epoch. A similar pattern is observed in training time. As shown in Figure 9(b), SCRIPTintroduces only a moderate slowdown rela- tive to the base model, whereas KOMBO base dra- matically reduces training speed as sequence length grows. For a moderate input length (128 tokens), BERT base +SCRIPTrequires roughly 58 s per train- ing step compared to 20 s for BERT base (about 2.9×slower), and KOMBO base takes around 62 s (about 3.1×slower than base). However, as the sequence length increases, KOMBO’s runtime cost grows much more rapidly. At 512 tokens, BERT base +SCRIPTprocesses a batch in roughly 169 s (less than 2×the 87 s required by BERT base ), whereas KOMBO base requires about 308 s – over 3.5×the base model’s time. This steep slowdown for KOMBO is a direct consequence of operating over a much longer sequence with additional trans- formation layers, as discussed above. In contrast, SCRIPTmaintains a moderate runtime overhead, less than 2×the base model even at maximum se- quence lengths, making it far more practical than KOMBO in real-world training scenarios. We note that these trends hold for both encoder- based and decoder-based architectures. In our ex- periments with the decoder-only KoGPT2 model, addingSCRIPTincurred similar slowdowns and only minor memory increases, underscoring the general applicability ofSCRIPTacross model types. In summary,SCRIPToffers a significantly more efficient and practical solution for incor- porating subcharacter information than another subcharacter-based approach, such as KOMBO. By sidestepping expensive architectural changes and pre-training requirements,SCRIPTmaintains almost the same training footprint as the underly- ing base model in terms of memory. The small overhead introduced bySCRIPTis significantly out- weighed by its benefits, and it stands in stark con- trast to the heavy computational cost of KOMBO. This efficiency makesSCRIPTa highly practical plug-and-play module for real-world deployment on large-scale models and datasets. Next, we ex- amine another aspect of training efficiency, the con- vergence speed of each model during training, to further assess the practical advantages of SCRIPT. J.3 Training Efficiency In Section J.2, we analyzed the structural ef- ficiency of the proposed methodSCRIPTcom- pared to another off-the-shelf subcharacter-based model, KOMBO. In this section, we further in- vestigate training efficiency, focusing on the con- vergence speed across three different models, such as BERT base as the base model, BERT base +SCRIPT, and KOMBO base , during fine-tuning. Figure 10 compares the convergence points of each model across the nine principal Korean NLU tasks introduced in Section 4.2.As a result, BERT base +SCRIPTconsistently converges faster as both BERT base and KOMBO base , achieving su- perior performance with fewer training epochs. Specifically, it converged faster on six out of nine tasks, and matched the convergence speed on the remaining three tasks. These results demonstrate that our method significantly enhances learning ef- ficiency. Furthermore, thoughSCRIPTleverages both subword and subcharacter embeddings, it not only outperforms the model utilizing solely Accuracy Score Epochs SCRIPT Jamo (SCRIPT Jamo ) Figure 10: The graphs show the fine-tuning performances of three models, BERT base , BERT base +SCRIPT, and KOMBO base , across nine NLU tasks. The x-axis represents the number of epochs for each task, and the y-axis indicates the accuracy score. Two types of dotted vertical lines are overlaid on the line graphs for each model: the red dotted line marks the epoch at which the model with the proposedSCRIPTmodule achieves its best performance, while the gray dotted lines indicate the best-performing epochs for the other two models. 32k subwords (BERT base ) and fewer than 200 sub- characters (KOMBO base ) but also converges more rapidly, highlighting its strong adaptation to Korean language understanding. Overall, our comprehensive evaluation demon- strates thatSCRIPToffers a substantially more effi- cient and scalable alternative to prior subcharacter- based approaches. By injecting subcharacter com- positional knowledge directly into existing PLM embeddings,SCRIPTenriches the model’s rep- resentational capacity while preserving the com- putational profile of the base model. This dual advantage, greater linguistic expressiveness with only marginal computational overhead, establishes SCRIPTas a practical and robust solution for real- world deployment in Korean NLP.