Paper deep dive
The Multilingual FrameNet Corpus
Beatrice FiumanĂČ, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing existing language-specific corpora across nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian and Swedish. By training models that rely on different architectures on the mFNC, we consistently outperform existing state-of-the-art Frame Semantic Parsers in both multilingual and cross-lingual settings, underscoring the importance of multilingual training data. The mFNC and our trained FSP models are openly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.23037v1
- Canonical: https://arxiv.org/abs/2608.23037v1
Trouble viewing inline? Open PDF directly â
Full Text
63,789 characters extracted from source content.
Expand or collapse full text
The Multilingual FrameNet Corpus Beatrice FiumanĂČ 1,* , Nicolas Lazzari 1,2 , Simone Paolo Ponzetto 3 , Valentina Presutti 1 1 University of Bologna, Italy, 2 University of Pisa, Italy, 3 University of Mannheim, Germany beatrice.fiumano,nicolas.lazzari3,valentina.presutti@unibo.it,ponzetto@uni-mannheim.de * Corresponding author Abstract This paper introduces the Multilingual FrameNet Corpus(mFNC),a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing exist- ing language-specific corpora across nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian and Swedish. By training models that rely on different archi- tectures on the mFNC, we consistently outper- form existing state-of-the-art Frame Semantic Parsers in both multilingual and cross-lingual settings, underscoring the importance of multi- lingual training data. The mFNC and our trained FSP models are openly available athttps://github.com/ beatrice-f/mFNC. 1 Introduction Frame Semantic Parsing(FSP)is the task of au- tomatically identifying semantic frames in text according to Fillmoreâs Frame Semantics theory ( Fillmore,1976). FSP has proven beneficial across several NLP tasks, such as information extraction ( Li et al.,2025), Knowledge Graph construction (Alam et al.,2021), opinion mining(Recupero et al.,2015), sentiment analysis(Atzeni et al., 2018)and framing detection(Minnema et al., 2022;Coschignano et al.,2023). Despite its wide range of applications, the multi- lingual dimension of FSP remains largely underex- plored. Recent state-of-the-art(SotA)approaches ( e.g., Devasier et al.(2024)) are trained and evaluated exclusively on the Berkeley FrameNet (BFN)corpus(Baker et al.,1998), with isolated attempts at cross-lingual transfer through multi- lingual Transformer-based architectures(e.g.,Xia et al. (2021)).Unlike other tasks, however, FSP presents unique challenges due to cross-lingual variations in the grammatical, cultural and concep- tual dimensions of language. For instance, in the ExperiencerStimulus êł ì íë€. ëčììŹí ììì ëŁêł êł í„ìì êží ëŹë €ìš ìŹìë, ëȘšëì ìììŒëĄìŽê”ì êł”êČ©ì ë°êž° ì ì ï°ï· Experiencer_object Experiencer_object ïźïč Stimulus LâAfrica,quantogli Experiencer piaceva[...] ïșïž Experiencer_focused_emotion Ilikemy jobreally Experiencer [...] Content ïșïž Hostile_ encounter The missionaries against the dictates of the desert struggled valiantly Side_1Side_2 [...] Figure 1:Example of annotated sentences from the mFNC.Lexical Units are underlined. Textual spans classified by a Frame Elements are shaded using the same color. :âAfrica, how muchheliked it[...]â. :âThepersonwhorushedbackfromtheirhometown uponhearingnewsoftheemergencystruggledwithsand andnoise before facing attack from a foreign country.â. Italianâpiacereâ (to like),the thing being liked acts upon the liker, while in English it is the liker who actively likes something( Baker and Lorenzi, 2020). Similarly, the Korean expressionâêł ì í ë€â(to struggle)casts the struggling entity as acted upon by an external force, while English frames the same situation with the struggling entity as an active participant. These examples illustrate how languages can conceptualize the same situation in different ways, expressing it through distinct syn- tactic and semantic patterns. As shown in Figure 1, these differences are captured by Frame Seman- tics. The lack of a unified large-scale multilingual corpus annotated using BFN is largely responsi- ble for this under-exploration, preventing FSP sys- tems from being trained or tested on languages be- yond English. Although numerous independent language-specific FrameNets have been developed since the inception of BFN, these resources remain arXiv:2608.23037v1 [cs.CL] 24 Aug 2026 scattered and heterogeneous in their format and annotation scheme, hampering integration. Ap- proaches to automatically align them have been proposed(Baker and Lorenzi,2020), but to the best of our knowledge no unified corpus exists to date. To address this gap, we present the mFNC(Mul- tilingual FrameNet Corpus),a novel multilingual dataset that collects and harmonizes ten language- specific corpora annotated using BFN. Alongside English, the mFNC extends coverage to Brazilian Portuguese, Chinese, Dutch, French, German, Ital- ian, Korean, Latvian and Swedish. Figure1shows an example of sentences and annotations from the mFNC. Using the mFNC, we demonstrate that FSP mod- els trained on multilingual data consistently out- perform models trained on English only. Beyond FSP, the mFNC contributes to the broader goal of a unified Multilingual FrameNet( Gilardi and Baker ,2018), offering an empirical basis to inves- tigate how frame-semantic structures vary across typologically diverse languages(Ellsworth et al., 2021). In summary, our contribution is two-fold: 1.We present the mFNC, a multilingual benchmark that collects and harmonizesten language-specific FrameNets; 2.We openly share three multilingual FSP sys- tems that outperform existing SotA baselines in multilingual and cross-lingual settings. The rest of the paper is structured as follows: Section 2introduces the background on Frame Se- mantics and reviews existing work on BFN appli- cations, resources that extend BFN to other lan- guages, and existing FSP approaches. Section 3.1describes the resources used to construct the mFNC and provides insights on them. Section 4presents a comparative evaluation of SotA FSP models trained on BFN and on the mFNC, in both multilingual and cross-lingual settings. Section 5 discusses our findings and outlines directions for future work. Finally, Section6summarizes our contributions and Section7highlights limitations. 2 Related Work Before reviewing multilingual FrameNet resources and parsers, we provide an example of how a nat- ural language sentence is annotated using BFN. Consider the sentenceâI really like my jobâ,shown in Figure1. TheLexical Unit(LU)âlikeâevokesthe frame EXPERIENCER_FOCUSED_EMOTION, which repre- sents the concept ofâsomeone experiencing some emotion with respect to some contentâ.In turn, the spanâIâis classified by theFrame Element(FE) EXPERIENCER,whileâmy jobâis classified by the FE CONTENT. BFN defines a large set of frames, FEs, and En- glish LUs. For a complete treatment, we refer the reader toRuppenhofer et al.(2016). 2.1 Toward a Multilingual FrameNet While BFN originally focused on English, ex- tending it to a multilingual setting has been a central objective of the project since its incep- tion. This ambition has driven the development of numerous language-specific resources, and re- sulted in the Global FrameNet project 1 , which aims both at aligning existing datasets(Baker and Lorenzi ,2020)and at creating new multilingual ones through shared annotation tasks(Gilardi and Baker,2018). As with other lexical-semantic networks, the creation of new FrameNets follows distinct lexi- cographic strategies, such as projecting the BFN frame inventory onto another language, develop- ing a new culture-specific inventory, or combining both practices by reusing and expanding the BFN repertoire( Boas,2009). As detailed in Section3.1, we focus on resources that fully or partially reuse the English set of frames, as this enables cross- lingual alignment and the construction of a cohe- sive dataset. 2.2 Frame Semantic Parsing The FSP task consists of automatically annotating a natural language sentence with semantic frames, as shown in Figure 1. The task is usually decom- posed into four sequential sub-tasks: detecting the frame-evoking LUs, classifying them with the cor- rect frame, and identifying the spans classified by the FEs of the frame. Existing FSP systems either address each step in isolation or jointly solve mul- tiple sub-tasks. FSP ApplicationsAs introduced above, FSP en- ables structuring natural language text according to semantic frames and FEs, making its predicate- argument structure explicit. 1 https://w.globalframenet.org/ This approach has proven useful across a wide range of NLP tasks. Alongside the applications listed in Section1,Taniguchi et al.(2019)use FSP for yes/no QA to extract predicate-argument configurations from legal texts, enabling seman- tic matching between questions and candidate an- swers beyond surface string overlap. More re- cently,Li et al.(2025)proposed FrameRTE, a three-stage pipeline that combines an FSPâs out- put and LLMs for zero-shot relation triplet extrac- tion. In text summarization, frame-based graphs have been exploited to improve salience estimation for both extractive( Guan et al.,2021a)and abstrac- tive(Guan et al.,2021b)summarization. Beyond core NLP tasks, FSP has also been leveraged for applications at the discourse level. Frame-based embeddings have been used to train transformer models for metaphor detection( Li et al.,2023)and generation(Stowe et al.,2021).Remijnse et al. (2024)use frames to analyze framing and perspec- tivization of eventsâparticipants across documents, while Ryazanov et al.(2025)andFiumanĂČ et al. (2026)apply FSP to trace differences in media nar- ratives. In light of these applications, the utility of FSP clearly extends across all languages, reinforc- ing the need for multilingual systems. Multilingual FSPIn general, SotA models are trained and tested on the BFN corpus (Swayamdipta et al.,2017;Das et al.,2010; Kalyanpur et al.,2020a;Devasier et al.,2024), while models operating in multilingual or cross- lingual settings have received less attention. No- table exceptions includeJohannsen et al.(2015), where the model by Das et al.(2010)is extended with multilingual word embeddings and its cross- lingual performance is evaluated on a novel cor- pus of nine languages. Although the corpus is too small to train an FSP from scratch, the results demonstrate that a multilingual backbone is bene- ficial for cross-lingual generalization. More recently, Xia et al.(2021)build upon this finding by proposing LOME, an FSP model that fine-tunes the multilingual XLM-RoBERTa encoder model on the BFN corpus. However, the model is only evaluated on the English language. To the best of our knowledge, no FSP models trained and evaluated in a multilingual setting have been proposed to date. LLM-driven FSPBeyond specialized ap- proaches to FSP, LLMs represent valid language- agnostic tools thanks to their broad multi-language coverage. Unlike previous approaches, their appli- cation so far has been limited to individual FSP sub-tasks, such as frame identification(Chundru et al.,2025)and FE classification(Devasier et al., 2025), or in modular pipelines(Liu et al.,2024). Results have shown that LLMs underperform in zero- and few-shot settings, narrowing the performance gap with specialized FSP models only through fine-tuning. 3 mFNC: Multilingual FrameNet Corpus In this section we present the mFNC which, to the best of our knowledge, is the first large-scale mul- tilingual corpus annotated using BFN. We con- struct the mFNC by collecting ten different corpora annotated using BFN and harmonizing them in a shared format. In the next sections we first describe our data selection(Section 3.1)and harmonization (Section3.2)processes, and later provide insights on the mFNC statistics(Section 3.3)and content (Section3.4). 3.1 Data Selection Table2provides an overview of the ten original resources used to construct the mFNC. The resources vary in their development strategy and the type of language data they annotate. How- ever, they converge in their full or partial adoption of BFN frames and FEs. Bottom-up approachesWith the exception of Swedish and Korean, all resources adopt a bottom- up approach, meaning they derive frame annota- tions from existing or newly collected language- specific corpora and treebanks, with varying de- grees of adherence to BFN. Although more la- bor intensive, this method preserves the conceptual structure of the source language and avoids being constrained by the set of BFNâs LUs. Top-down approaches Korean FrameNet is in- stead developed using a top-down(or extension (DannĂ©lls et al.,2021))approach, where sentences from the BFN corpus are translated into Korean, enabling the identification of language-specific LUs. While this approach is more efficient, it forces an English-centric conceptualization on the new resource. At a later stage, the Korean FN was expanded via cross-language projection of the TokensFramesFEs# Sentences Train Valid.Test Total DE34524719688(230)35597(848)10543 1629 2989 15161 EN11290725951(788)45639(3668)3312325 1211 4848 FR2067296536(51)12514(198)3733558 1043 5334 IT255851255(192)2860(726)621105225951 KO15339121094(800)38972(4226)8407 1395 2188 11990 LV29934113527(242)24836(1131)9451 1350 2726 13527 NL211281339(200)2054(479)713104218 1035 PT282367767(519)12685(1730)1319171433 1923 SV 1229398018(954)17209(5174)5506769 1694 7969 ZH1892579107(584)23202(3660)4365786 1347 6498 total1504760 114282(1070)215568(8954)47970 7192 14074 69236 Table 1: Composition of the mFNC. Unique frames and FEs are reported within the parenthesis. Lang.SourceAcc. DE(Rehbein et al.,2012)UR EN(Baker et al.,1998)OA FR(Djemaa et al.,2016)OA IT(Venturi et al.,2009)OA KO(Kim et al.,2016)OA LV(Gruzitis et al.,2018)OA NL(Vossen et al.,2020)OA PT(Belcavello et al.,2024)OA SV(DannĂ©lls et al.,2021)OA ZH(You and Liu,2005)UR Table 2: FrameNet resources selected to construct the mFNC. Datasets marked withOAare openly accessi- ble, while datasets marked withURare available upon request. Japanese FrameNet 2 (Ohara et al.,2004). Hybrid constructionThe Swedish FrameNet is developed following a hybrid approach, first reusing BFN frames and translating English LUs into Swedish, and later defining new frames and LUs by adopting a corpus-based approach. Document genreThe ten resources also dif- fer in the genre of annotated documents, span- ning newspaper articles(German, Italian, partially French),multi-genre treebanks covering medi- cal, legal, and parliamentary texts ( French ), large general-domain and specialized corpora(Chinese), multi-genre texts reporting on selected event types 2 The resource is currently not publicly available, and we were not able to access it. (Dutch),a mixed corpus of news, fiction, legal, and spoken texts(Latvian),and informal sources such as TV series transcripts ( Brazilian Portuguese ). This diversity contributes to the richness of topics and styles in the mFNC documents. 3.2 Data Harmonization We seek to construct the mFNC to contain both plain and tokenized documents, aligning with the original BFN corpus format. However, the ten FrameNets differ in how they provide the docu- ments, sometimes offering only the plain or tok- enized text. To achieve a fully harmonized resource, we re- cover missing full-text documents from their to- kenized version using the language-specific deto- kenizers available in the SacreMoses Python li- brary 3 . Similarly, we tokenize full-text documents that are not already tokenized using the same li- brary. Finally, we collect all the annotations for each document. An annotation is defined by the LU that evokes a frame, the frame, and the list of an- notated FEs. We remove language-specific frames for those resources that follow a hybrid annotation approach, to ensure full cross-language compatibil- ity with BFN. 4 3.3 mFNC composition Table1reports an overview of the composition of the mFNC after the harmonization phase. In to- 3 https://github.com/hplt-project/sacremoses 4 In the GitHub repository, we also provide a complemen- tary version of the mFNC that preserves language-specific an- notations. tal, the mFNC includes 1.5M tokens collected from approximately 70k sentences, accumulating a total of over 100k annotated frames and 200k annotated FEs. To support reproducibility, for English we re- use the splits already available for the BFN corpus, as computed bySwayamdipta et al.(2017) 5 . Fol- lowing the same work, we only retain annotations of the full-text documents, discarding annotations of the exemplar sentences provided for each frame. Indeed,Das et al.(2014)noted that training on ex- emplar sentences hurt model performance, proba- bly due to their lack of representativeness and in- complete annotations. We compute novel splits for the remaining languages by prioritizing a balanced distribution of frames(Sechidis et al.,2011). 020040060080010001200 Frames 10 1 10 0 10 1 10 2 10 3 log num. of occurrences Train Validation Test Figure 2: Number of occurrences ofallBFN frames(in log space)in each mFNC split. Frames coverageFigure2shows that the num- ber of occurrences of each frame follows a Zipfian distribution, hinting at its highly unbalanced na- ture inherited from the resources of Table2(see also Figure 6in the Appendix for language-specific breakdown). For example, the frames CAUSATION and STATE- MENT are the two most frequently annotated frames, while CAUSATION_SCENARIO and EXPLO- SION are both annotated only once in the mFNC (see Table6in the Appendix for the most common frames annotated for each language).We report that 151 of the 1221 total frames in BFN(â12%) never occur in the mFNC annotations. Annotation densityThe resources differ in the density of available annotations. In Figure3, we report the average length, number of frames and number of annotated FEs for each document. We observe, for instance, that the English and Brazil- 5 We only include the FrameNet 1.7 corpus as it encom- passes the previous 1.5 version. 0 10 20 30 40 Avg. tokens 0 2 4 6 8 10 Avg. frames DEENFRITKOLVNLPTSVZH 0 2 4 6 8 10 Avg. roles Figure 3: Average number of tokens, frames and FEs annotated in each document. DEENFRITKOLVNLPTSVZH DE EN FR IT KO LV NL PT SV ZH 0.2 0.4 0.6 0.8 1.0 Figure 4: Cosine similarity between the centroids of each language computed using frame similarity. ian Portuguese corpora contain, on average, twice as many annotated frames as the other languages, indicating a high density of annotations. On the other hand, the French dataset contains longer doc- uments than the other languages but displays a sim- ilar number of frames, indicating a lower annota- tion density. In the next section we further explore these an- notation differences. 3.4 Similarity of annotations In this section, we investigate in more detail how the ten resources converge and diverge with respect to their frame annotations, analyzing frame similar- ity. To do so, we rely on the FFICF measure, which adapts TF-IDF to derive typicality scores for each frame(Vossen et al.,2020). The frame frequency is computed with respect to all the documents in the mFNC corpus. For each language we compute the centroid of its vectors and compare it with the centroids of other languages using cosine similar- ity. Hence, two languages are similar if they share a similar set of typical frames. Results are shown in Figure4. Topic SpecificityWe observe that Dutch and French have the most dissimilar annotations when compared to the BFN corpus and to the other resources. We speculate that this is due to the domain-specificity of annotated texts. For in- stance, the Dutch FrameNet corpus only collects documents that report on specific events(e.g., disease outbreak and wildfires) ( Vossen et al., 2020),resulting in CATASTROPHE being one of the most commonly annotated frames. Similarly, the French FrameNet corpus includes specialized documents in the medical and political domains ( Djemaa et al.,2016;Candito et al.,2014). As a consequence, these resources show a biased distribution of frames, which reflects their topic- specificity. Annotation PracticeIn contrast, the set of most typical Korean frames is similar to that of the BFN corpus, which reflects the use of a projective anno- tation practice. A similar behavior can be observed in the Swedish corpus, which also partially relies on projective annotations. 4 Experiments In this section we demonstrate the effectiveness of the mFNC by training different FSP models on both the BFN corpus and on the mFNC, and com- pare their performance. 4.1 Experimental Setting We experiment with two architectures representa- tive of SotA approaches: the multi-stage approach proposed in LOME( Xia et al.,2021) , where an XLM-RoBERTa encoder model is fine-tuned to pa- rameterize a CRF layer used to extract spans from the input sentence that are then classified using two MLP layers(one for frames and one for FEs); and a generative approach inspired by Kalyanpur et al.(2020b), where FSP is framed as a generative seq2seq task. In particular, we use the sentinel ap- proach described inRaman et al.(2022)and fine- ModelF1â Swayamdipta et al.(2017) â 0.733 Lin et al.(2021) â 0.763 Devasier et al.(2024) â 0.775 mT5 small0.560 mT5 base0.583 LOME0.800 mT5 small â 0.677 mT5 base â 0.668 LOME â 0.812 Table 3:Training on the mFNC outperforms train- ing on the BFN on the English language.Tradi- tional micro-F1 score on the target identification and classification task. Results marked with â are taken from Devasier et al.(2024). Results marked withâ are trained on the mFNC. tune the small and base versions of mT5(Xue et al., 2021). We train LOME using the default hyperparame- ters defined by the authors on an RTX3090 with 24 GB of VRAM for a maximum of50epochs, stopping the model if it does not improve its per- formance on the validation split for 3 consecutive epochs. We fine-tune the mT5 models using the hy- perparameters suggested in Raman et al.(2022)for 30epochs on an RTX6000 with 48 GB of VRAM, using the same early-stopping mechanism used for LOME. 4.2 Results on the BFN corpus Before evaluating the performance of FSP systems in a multilingual setting, we evaluate whether train- ing on the mFNC maintains competitive perfor- mance compared to training only on the BFN cor- pus. In Table3we report the traditional micro- averaged F1 score of the models trained on the BFN corpus and on the mFNC in the target classi- fication task, i.e., on frame-evoking LU detection and frame attribution. A prediction is considered correct when it fully matches the gold annotation. We compare our re- sults with those from(i) Swayamdipta et al.(2017), who frame the task as a token classification prob- lem relying on pre-trained static word embeddings and an LSTM model,(i) Lin et al.(2021),who frame the problem as a graph generation problem, fine-tuning a BERT-based encoder model, and(i) Devasier et al.(2024),who formulate the task as a QA task solved by fine-tuning a RoBERTa encoder model. We find that training on the mFNC maintains competitive performance with models trained on the BFN corpus, demonstrating that additional training data on other languages does not harm performance. Additionally, we demonstrate that training LOME on the mFNC outperforms existing SotA models. This result encourages novel FSP systems to be trained on the mFNC. 4.3 Results on the mFNC In Table4we report the performance of our mod- els trained on the BFN corpus and on the mFNC and evaluated on the testing set of the mFNC by aggregating over the ten languages. Unlike Table 3, we report micro-averaged precision, recall and F1 scores computed using the suggested configura- tion of the FairEval framework(Ortmann,2022). This allows us to account for correctly identified but mislabeled LUs and FEs, and for predictions whose boundaries partially overlap with the gold ones. Note that the metrics grouped under theFrame column measure performance in identifying and labeling frame-evoking LUs. In turn, theFEcol- umn evaluates the identification and classification of FEs. In other words, the metrics grouped un- der the FE column evaluate the end-to-end perfor- mance of the FSP model, accounting for errors propagated from the frame identification process. Multilingual performanceThe results provide strong evidence of the impact of the mFNC on the parsersâperformance. Each model, regardless of the employed architecture, greatly outperforms the corresponding variant trained only on the English corpus. In particular, LOME trained on the mFNC consistently outperforms all the other FSP models across all the evaluated dimensions, providing evi- dence that FSP benefits from being treated as a se- quence labeling task rather than as a seq2seq one. Although training on the mFNC consistently im- proves results, performance gains are not equally distributed across the ten languages in the corpus. Language-wise improvementsIn Figure5we show how F1 scores vary across different lan- guages when training on the BFN corpus or on the mFNC(see Table 7in the Appendix for a more detailed overview).Similar to the results in Table 3, performance on the English language is similar across the English-only and multilingual training 0.00 0.25 0.50 0.75 1.00 Fair Frame F1 DEENFRITKOLVNLPTSVZH 0.00 0.25 0.50 0.75 1.00 Fair FE F1 mFNC BFN Figure 5:Training on the mFNC outperforms train- ing on the BFN on all languages.F1 scores on frame and FEs performance on each language when trained on BFN vs mFNC. settings, while it improves significantly on other languages, particularly in German, French, Dutch and Latvian. The Swedish corpus is the most chal- lenging one, showing less pronounced improve- ments. While we do not have a definitive explanation for this behavior, we found that the Swedish FrameNet has the largest coverage of annotated frames, anno- tating 78% of BFN frames compared to 64% for the BFN corpus. We speculate that the unbalanced nature of the mFNC(cf. Figure2)might hamper generalization to infrequent frames, leaving open the question of whether it is possible to counter- balance this phenomenon during training(e.g., by annealing frequent frames),or pre-processing of the dataset(e.g., by performing data augmentation based on the hierarchical structure defined in the BFN). This result might also indicate that even when trained on multilingual data, the parser does not generalize to unseen languages. However, in the next section we demonstrate the opposite. 4.3.1 Cross-lingual Generalization In light of previous findings, we evaluate the im- pact that training on the mFNC has in cross-lingual settings by removing the Swedish dataset from the mFNC and training LOME from scratch on it. We follow the same experimental setting of previous experiments. Table 5reports the results, showing that models trained on the mFNC have stronger cross-lingual generalization abilities compared to training on the BFN corpus. This further demon- strates the impact that the mFNC can have on fu- ture FSP models. ModelTrain FrameFE PrecisionRecallF1PrecisionRecallF1 LOME BFN 0.25±0.230.60±0.220.32±0.220.15±0.180.33±0.200.19±0.18 mFNC0.77±0.110.65±0.200.69±0.150.61±0.150.54±0.180.56±0.15 mT5 BFN 0.14±0.150.43±0.190.19±0.150.08±0.100.19±0.130.10±0.10 mFNC 0.53±0.120.57±0.130.55±0.120.40±0.130.41±0.130.40±0.13 mT5 small BFN 0.11±0.120.37±0.170.16±0.130.06±0.080.15±0.100.08±0.09 mFNC 0.59±0.140.60±0.140.59±0.140.42±0.130.42±0.140.42±0.13 Table 4:Training on the mFNC outperforms training on the BFN on all metrics. FairEval metrics computed on the mFNC averaged over the ten languages. An annotation is defined by the span that 283 activates a frame, the activated frame, and the list 284 of annotated arguments. Each argument is defined 285 by a role and the span that it classifies. MetricBFN mFNC mFNC Frame P 0.070.30 0.55 R 0.260.28 0.42 F1 0.110.290.47 FE P 0.040.20 0.36 R 0.20 0.200.33 F1 0.070.20 0.35 Table 5:Training on the mFNC achieves better cross- lingual generalization than the BFN.Performance of LOME trained on the BFN corpus, the mFNC without Swedish data(mFNC )and on the full mFNC on the same setting as Table 4. Best results are in bold. Best between the mFNC and the mFNC are un- derlined. 5 Discussion The results described in Section4demonstrate that current SotA FSP models trained on the BFN cor- pus struggle in cross-lingual settings. Their re- sults, however, greatly improve when trained on our corpus, both in multilingual(Table4)and cross-lingual(Table 5)settings. Moreover, our results indicate that models tai- lored to the FSP task(LOME)perform sig- nificantly better than more general seq2seq ap- proaches(mT5-based models). Integrating the mFNC with other resources As addressed in Section 3.3and in Figure2, the mFNC is unbalanced with respect to the number of annotated frames. The effectiveness of FSP mod- els is directly affected by this aspect, as illustrated in Section4.3.1. Although in this paper we only collect corpora annotated using the BFN, previ- ous works(e.g.,Conia et al.(2022))have shown that integrating additional semantic resources such as PropBank(Pradhan et al.,2022)and VerbNet (Palmer et al.,2017)results in better FSP perfor- mance. We refrained from adopting this approach be- cause of the different nature of these resources compared to BFN 6 . Nonetheless, it is worth inves- tigating whether the mFNC could be extended by relying on recent efforts at aligning BFN with other resources( Lopez de Lacalle et al.,2016), which can result in broader linguistic coverage, for exam- ple by integrating the Polish( Jindal et al.,2022)or Arabic(Palmer et al.,2008)PropBank-annotated datasets. Extending the mFNCThe mFNC is the first corpus of multilingual documents annotated using BFN frames, but it is still characterized by lim- ited coverage when compared to other multilingual datasets. For example, there is a lack of repre- sentation for Middle-Eastern or African languages. Possible approaches to overcome this limitation in- clude translating texts and projecting their anno- tations( Yu et al.,2022), as discussed in Section 3.1. We remark, however, that fully automating this approach might produce imprecise annotations that do not take into account the tight relationship between the lexical, semantic and cultural dimen- sions of language. Other promising approaches include relying on(L)LMs to generate annotated sentences by explicitly defining the semantics of a frame as found in the original BFN resource 6 For example, PropBank frames are more focused on the lexical level than the semantic one, unlike BFNâs frames (Bonial et al.,2014). (Cui and Swayamdipta,2024). In this context, the mFNC can serve as a repository of multilingual ex- amples that show the linguistic diversity spanned by a frame. Frame-based Linguistic AnalysesIn this paper, we focused on the impact that the mFNC has on the training of FSP models. Nonetheless, the cor- pusâs value extends beyond this application. By harmonizing ten language-specific datasets across typologically and conceptually diverse languages, the mFNC also enables a wide range of cross- lingual analyses at scale. These include investigat- ing why semantically equivalent expressions evoke different frames across languages(Yong et al., 2022), how language-specific phenomena such as compound nouns are realized differently(Ponkiya et al.,2021), and how conceptual metaphors differ across languages(Otmakhova et al.,2026). 6 Conclusion This paper presented the mFNC, a multilin- gual dataset harmonizing ten language-specific re- sources annotated using FrameNet. This contri- bution addresses the lack of multilingual training and evaluation data for the Frame Semantic Pars- ing task, demonstrating that training on multilin- gual data substantially improves the performance of multilingual and cross-lingual systems. Beyond parsing, the mFNC allows researchers to explore new directions for comparative research on con- ceptual and frame-semantic differences across lan- guages by relying on a large, unified resource. 7 Limitations In this section we discuss the main limitations of our contributions concerning two main aspects: the mFNC construction described in Section3.1, and the experiments presented in Section4. 7.1 On constructing the mFNC Flattened linguistic diversityAs shown in Fig- ure 1, Frame Semantic annotations are highly de- pendent on the grammatical and semantic patterns of each language. In Section3.1, we combine language-specific corpora annotated using BFN, implicitly assuming that the conceptual structure of BFN correctly transfers to languages other than English. Nonetheless, we are aware that adopting this approach remains an open question in NLP and linguistic research. Recent work demonstrates its limits(Ellsworth et al.,2021;Hahm et al.,2020), particularly when translations and projective annotations are used. These works, however, do not flag the assumption as incorrect. Rather, they argue that not all the frames defined in BFN apply equally to different languages, positing the existence of a language- agnostic portion of BFN(Äulo,2013). In this context, the mFNC can serve as a research tool to identify this subset using data-driven approaches (Baker et al.,2018;Baker and Lorenzi,2020). Domain biasAs highlighted in Section3.4, the mFNC comprises some resources that are domain- specific. Specifically, the Dutch and French cor- pora differ from the other FrameNets in that they annotate documents that focus on specific topics. While this thematic diversity can be beneficial, it also induces a representational bias whereby some frames may be under-represented in a resource due to the predominant topics of its documents(see for instance the top 5 most common frames used by each corpus in Table6in the Appendix).Assess- ing the impact of this limitation and mitigating its effects is an important step toward ensuring more cross-lingual balance in the mFNC. Annotation biasRelated to the previous limita- tions, combining the corpora described in Section 3.1also assumes consistency across the human an- notations of each resource. This assumption may introduce additional biases in the mFNC. For in- stance, Dumitrache et al.(2018)andHahm et al. (2020)observed cross-annotator differences in re- sources annotated using BFN, in both expert and crowd-sourced annotations. Although this does not necessarily translate to low quality annotations, it might result in an un- even distribution of frames across corpora, similar to the domain bias described previously. Overcom- ing this limitation is not trivial, due to the inherent complexity of the annotation task. On the other hand, we argue that the mFNC can serve as a re- source for analyzing annotation perspectives and differences ( Cabitza et al.,2023) , such as cross- language and cross-cultural ones. We also note that retaining the original tok- enization of documents, as described in Section 3.2, leads to combining different(possibly in- compatible)tokenization practices( Habert et al., 1998). Similarly, recovering the full-text docu- ment through detokenization or tokenizing full- text documents using an automated tool may intro- duce noise, depending on the accuracy of the tool (van der Goot,2024). Our approach is maximally conservative with respect to the design of each cor- pus, allowing researchers to study differences or apply refinements if needed. Language coverageThe languages collected in the mFNC are mostly high-resource ones(e.g., En- glish, German, French),which leaves open the question of how well the FSP systems evaluated in Section 4generalize to low-resource languages. We already identified the extension of the mFNC to other languages as one of the main follow-ups (Section5)of this work. In this context, we high- light that there exist frame annotation efforts for low-resource languages(e.g., Arabic(Gargett and Leung ,2020)and Bengali(Datta et al.,2025)), which represent promising integrations in this di- rection. 7.2 On experimenting with the mFNC Assuming a correct frameAs reported previ- ously, different sets of frames might apply to the same sentence depending on how it is interpreted by the annotator( Dumitrache et al.,2018;Hahm et al.,2020). This identifies an important limita- tion of FSP models as well. From a general per- spective, the models of Section4treat the FSP task in a discriminative fashion by assuming that an LU is classified by a single, correct frame(and simi- larly for FEs).It follows that those models inherit the(unknown)biases induced by the annotatorâs perspective. Ongoing research on how to tackle this limitation, which is shared by other common NLP tasks( Frenda et al.,2025), can benefit from the mFNC as an additional corpus for experimen- tation. Recall vs precisionThe limitation discussed above also raises the question of whether the eval- uation metrics used in Section 4over- or under- estimate the applicability of the tested FSP sys- tems. For instance, a system that reaches a high precision at the expense of a low recall might re- flect a conservative parsing behavior. This might be beneficial in some settings, but for many of the FSP applications reviewed in Section 2, a low vol- ume of predictions translates to a limited amount of available information. By relying on the FairEval framework( Ortmann,2022), we partially address this problem, so that correctly identified but mis- labeled spans are treated ashalf-correctpredic- tions. Nonetheless, evaluating whether the metrics of Section4improve performance on downstream FSP applications remains an open problem that re- quires further research. Diverse morphologiesThe models evaluated in Section4rely on Transformer-based multilingual models, which assumes that their internal repre- sentations act as a cross-lingual bridge. It is well known, however, that the tokenization phase of those models might disfavor some languages ( Petrov et al.,2023). Coupled with the limita- tions discussed previously, this poses further ques- tions on the abilities of FSP models to process low- resource languages. Although out of scope for this paper, experimenting with backbone models that rely on a different tokenization strategy(e.g., ByT5 (Xue et al.,2022))might result in better cross- lingual generalization. Acknowledgments We wish to thank the anonymous reviewers and area chairs for their valuable comments; Ines Re- hbein for providing us with access to the Salsa dataset; Arianna Graciotti for proofreading an early draft of the paper; Carmelo Caruso, Ludovica Pannitto and theLaboratorio Sperimentaleof the Department of Modern Languages, Literatures and Cultures(University of Bologna)for providing ac- cess to their computational resources. Beatrice FiumanĂČ and Valentina Presutti are supported by INFINITY: a EU Horizon Europe project under Grant Agreement No 101233051. Beatrice Fiu- manĂČ is funded by the National Recovery and Re- silience Plan(NRRP),funded by the European Union â NextGenerationEU - Mission 4 âEduca- tion and Researchâ, Component 1 âEnhancement of the offer of educational services: from nurs- eries to universitiesâ- Investment 4.1âExtension of the number of research doctorates and innova- tive doctorates for public administration and cul- tural heritageâ.(DM 118/2023). References Mehwish Alam, Aldo Gangemi, Valentina Presutti, and Diego Reforgiato Recupero. 2021.Semantic Role Labeling for Knowledge Graph Extraction from Text .Progress in Artificial Intelligence, 10(3):309â320. Mattia Atzeni, Amna Dridi, and Diego Refor- giato Recupero. 2018.Using frame-based re- sources for sentiment analysis within the finan- cial domain.Progress in Artificial Intelligence, 7(4):273â294. Collin F. Baker, Michael Ellsworth, Miriam R. L. Petruck, and Swabha Swayamdipta. 2018. Frame Semantics across Languages: Towards a Multilingual FrameNet. InProceedings of the 27th International Conference on Computa- tional Linguistics: Tutorial Abstracts, pages 9â 12, Santa Fe, New Mexico, USA. Association for Computational Linguistics. Collin F. Baker, Charles J. Fillmore, and John B. Lowe. 1998.The Berkeley FrameNet Project. In36th Annual Meeting of the Association for Computational Linguistics and 17th Interna- tional Conference on Computational Linguis- tics, COLING-ACLâ98,August 10-14, 1998, UniversitĂ© de MontrĂ©al, MontrĂ©al, Quebec, Canada. Proceedings of the Conference, pages 86â90. Morgan Kaufmann Publishers / ACL. Collin F. Baker and Arthur Lorenzi. 2020.Explor- ing Crosslinguistic Frame Alignment . InPro- ceedings of the International FrameNet Work- shop 2020: Towards a Global, Multilingual FrameNet, pages 77â84, Marseille, France. Eu- ropean Language Resources Association. Frederico Belcavello, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Maucha Ga- monal, Natalia Sigiliano, LĂvia Vicente Du- tra, Helen de Andrade Abreu, Mairon Sam- agaio, Mariane Carvalho, Franciany Campos, Gabrielly Azalim, Bruna Mazzei, Mateus Fon- seca de Oliveira, Ana Carolina Loçasso Luz, LĂvia PĂĄdua Ruiz, JĂșlia Bellei, Amanda Pestana, Josiane Costa, and 5 others. 2024. Frame2: A FrameNet-based Multimodal Dataset for Tack- ling Text-image Interactions in Video. InPro- ceedings of the 2024 Joint International Confer- ence on Computational Linguistics, Language Resources and Evaluation(LREC-COLING 2024), pages 7429â7437, Torino, Italia. ELRA and ICCL. Hans C. Boas, editor. 2009.Multilingual FrameNets in Computational Lexicography. De Gruyter Mouton, Berlin, New York. Claire Bonial, Julia Bonn, Kathryn Conger, Jena D. Hwang, and Martha Palmer. 2014.Prop- Bank: Semantics of New Predicate Types. In Proceedings of the Ninth International Confer- ence on Language Resources and Evaluation (LRECâ14), pages 3013â3019, Reykjavik, Ice- land. European Language Resources Associa- tion(ELRA). Federico Cabitza, Andrea Campagner, and Vale- rio Basile. 2023.Toward a Perspectivist Turn in Ground Truthing for Predictive Computing. InThirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Confer- ence on Innovative Applications of Artificial In- telligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7- 14, 2023, pages 6860â6868. AAAI Press. Marie Candito, Guy Perrier, Bruno Guillaume, Corentin Ribeyre, KarĂ«n Fort, DjamĂ© Seddah, and Ăric Villemonte de la Clergerie. 2014.Deep Syntax Annotation of the Sequoia French Tree- bank . InProceedings of the Ninth International Conference on Language Resources and Eval- uation, LREC 2014, Reykjavik, Iceland, May 26-31, 2014, pages 2298â2305. European Lan- guage Resources Association(ELRA). Jayanth Krishna Chundru, Rudrashis Poddar, Jie Cao, and Tianyu Jiang. 2025. Do LLMs Encode Frame Semantics? Evidence from Frame Identi- fication. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Pro- cessing, EMNLP 2025, Suzhou, China, Novem- ber 4-9, 2025, pages 29488â29500. Association for Computational Linguistics. Simone Conia, Edoardo Barba, Alessandro ScirĂš, and Roberto Navigli. 2022.Semantic Role La- beling Meets Definition Modeling: Using Nat- ural Language to Describe Predicate-Argument Structures. InFindings of the Association for Computational Linguistics: EMNLP 2022, pages 4253â4270, Abu Dhabi, United Arab Emi- rates. Association for Computational Linguis- tics. Serena Coschignano, Gosse Minnema, and Chiara Zanchi. 2023.Explaining the distribution of im- plicit means of misrepresentation: A case study on Italian immigration discourse .Journal of Pragmatics, 213:107â125. Xinyue Cui and Swabha Swayamdipta. 2024.An- notating FrameNet via Structure-Conditioned Language Generation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics(Volume 2: Short Pa- pers), pages 681â692, Bangkok, Thailand. Asso- ciation for Computational Linguistics. Dana DannĂ©lls, Lars Borin, and Karin Friberg Hep- pin, editors. 2021.The Swedish FrameNet++. John Benjamins Publishing Company. Dipanjan Das, Desai Chen, AndrĂ© F. T. Martins, Nathan Schneider, and Noah A. Smith. 2014. Frame-semantic parsing.Computational Lin- guistics, 40(1):9â56. Dipanjan Das, Nathan Schneider, Desai Chen, and Noah A. Smith. 2010.Probabilistic Frame- Semantic Parsing. InHuman Language Tech- nologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, pages 948â956, Los Angeles, California. Association for Computa- tional Linguistics. Sima Datta, Kunal Chakma, Dwijen Rudrapal, and Anupam Jamatia. 2025.Mapping the Lin- guistic Landscape: Progress Towards a Ben- gali FrameNet.Procedia Computer Science, 258:3814â3825. International Conference on Machine Learning and Data Engineering. Jacob Devasier, Yogesh Gurjar, and Chengkai Li. 2024.Robust Frame-Semantic Models with Lexical Unit Trees and Negative Samples. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6930â6941, Bangkok, Thailand. Association for Computa- tional Linguistics. Jacob Daniel Devasier, Rishabh Mediratta, and Chengkai Li. 2025.Can llms extract frame- semantic arguments?InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 30609â30622. Association for Computational Linguistics. Marianne Djemaa, Marie Candito, Philippe Muller, and Laure Vieu. 2016. Corpus Annotation within the French FrameNet: a Domain-by-domain Methodology. InProceed- ings of the Tenth International Conference on Language Resources and Evaluation LREC 2016, PortoroĆŸ, Slovenia, May 23-28, 2016. European Language Resources Association (ELRA). Anca Dumitrache, Lora Aroyo, and Chris Welty. 2018.Capturing Ambiguity in Crowdsourc- ing Frame Disambiguation. InProceedings of the Sixth AAAI Conference on Human Com- putation and Crowdsourcing, HCOMP 2018, ZĂŒrich, Switzerland, July 5-8, 2018, pages 12â 20. AAAI Press. Michael Ellsworth, Collin Baker, and Miriam R. L. Petruck. 2021.FrameNet and Typology. In Proceedings of the Third Workshop on Compu- tational Typology and Multilingual NLP, pages 61â66, Online. Association for Computational Linguistics. Charles J. Fillmore. 1976.Frame Semantics and the Nature of Language.Annals of the New York Academy of Sciences, 280(1):20â32. Beatrice FiumanĂČ, Nicolas Lazzari, Simone Paolo Ponzetto, and Valentina Presutti. 2026.Vic- tim or assailant? exploring narratives through knowledge graph queries . InProceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026), pages 40â49, Palma, Mallorca, Spain. European Language Resources Associa- tion(ELRA). Simona Frenda, Gavin Abercrombie, Valerio Basile, Alessandro Pedrani, Raffaella Panizzon, Alessandra Teresa Cignarella, Cristina Marco, and Davide Bernardi. 2025. Perspectivist Ap- proaches to Natural Language Processing: a Sur- vey.Lang. Resour. Evaluation, 59(2):1719â 1746. Andrew Gargett and Tommi Leung. 2020.Build- ing the Emirati Arabic FrameNet . InPro- ceedings of the International FrameNet Work- shop 2020: Towards a Global, Multilingual FrameNet, pages 70â76, Marseille, France. Eu- ropean Language Resources Association. Luca Gilardi and Collin Baker. 2018. Learn- ing to Align across Languages: Toward Mul- tilingual FrameNet. InProceedings of the Eleventh International Conference on Language Resources and Evaluation(LREC 2018), Paris, France. European Language Resources Associa- tion(ELRA). Normunds Gruzitis, Gunta Nespore-Berzkalne, and Baiba Saulite. 2018. Creation of Latvian FrameNet based on Universal Dependencies. InProceedings of the International FrameNet Workshop(IFNW), pages 23â27. Yong Guan, Shaoru Guo, Ru Li, Xiaoli Li, and Hongye Tan. 2021a.Frame Semantic-Enhanced Sentence Modeling for Sentence-level Extrac- tive Text Summarization. InProceedings of the 2021 Conference on Empirical Methods in Nat- ural Language Processing, pages 4045â4052, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Yong Guan, Shaoru Guo, Ru Li, Xiaoli Li, and Hu Zhang. 2021b.Frame Semantics Guided Network for Abstractive Sentence Summariza- tion .Knowledge-Based Systems, 221:106973. Benoit Habert, Gilles Adda, Martine Adda-Decker, Philippe Boula de MareĂŒil, Silvana Ferrari, Olivier Ferret, Gabriel Illouz, and P. Paraubeck. 1998. Towards tokenization evaluation. InPro- ceedings of the First International Conference on Language Resources and Evaluation, LREC 1998, May 28-30, 1998, Granada, Spain, pages 427â432. European Language Resources Asso- ciation. Younggyun Hahm, Youngbin Noh, Ji Yoon Han, Tae Hwan Oh, Hyonsu Choe, Hansaem Kim, and Key-Sun Choi. 2020. Crowdsourcing in the Development of a Multilingual FrameNet: A Case Study of Korean FrameNet. InPro- ceedingsoftheTwelfthLanguageResourcesand Evaluation Conference, pages 236â244, Mar- seille, France. European Language Resources Association. Ishan Jindal, Alexandre Rademaker, MichaĆ Ulewicz, Ha Linh, Huyen Nguyen, Khoi- Nguyen Tran, Huaiyu Zhu, and Yunyao Li. 2022. Universal Proposition Bank 2.0. In Proceedings of the Thirteenth Language Re- sources and Evaluation Conference, pages 1700â1711, Marseille, France. European Language Resources Association. Anders Johannsen, HĂ©ctor MartĂnez Alonso, and Anders SĂžgaard. 2015.Any-language frame- semantic parsing. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2062â2066, Lis- bon, Portugal. Association for Computational Linguistics. Aditya Kalyanpur, Or Biran, Tom Breloff, Jennifer Chu-Carroll, Ariel Diertani, Owen Rambow, and Mark Sammons. 2020a.Open-Domain Frame Semantic Parsing Using Transformers. CoRR, abs/2010.10998. Aditya Kalyanpur, Or Biran, Tom Breloff, Jennifer Chu-Carroll, Ariel Diertani, Owen Rambow, and Mark Sammons. 2020b.Open-Domain Frame Semantic Parsing Using Transformers. CoRR, abs/2010.10998. Jeong-uk Kim, Younggyun Hahm, and Key-Sun Choi. 2016.Korean FrameNet Expansion Based on Projection of Japanese FrameNet . InPro- ceedings of COLING 2016, the 26th Interna- tional Conference on Computational Linguis- tics: System Demonstrations, pages 175â179, Osaka, Japan. The COLING 2016 Organizing Committee. Yucheng Li, Shun Wang, Chenghua Lin, Frank Guerin, and Loic Barrault. 2023. FrameBERT: Conceptual Metaphor Detection with Frame Embedding Learning . InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 1558â1563, Dubrovnik, Croatia. Associa- tion for Computational Linguistics. Zehan Li, Fu Zhang, Wenqing Zhang, Jiawei Li, Zhou Li, Jingwei Cheng, and Tianyue Peng. 2025. Frame First, Then Extract: A Frame- Semantic Reasoning Pipeline for Zero-Shot Re- lation Triplet Extraction. InProceedings of the 2025 Conference on Empirical Methods in Nat- ural Language Processing, pages 27363â27376, Suzhou, China. Association for Computational Linguistics. ZhiChao Lin, Yueheng Sun, and Meishan Zhang. 2021.A Graph-Based Neural Model for End-to- End Frame Semantic Parsing. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3864â 3874, Online and Punta Cana, Dominican Re- public. Association for Computational Linguis- tics. Yahui Liu, Chen Gong, and Min Zhang. 2024. Leveraging LLMs for Chinese Frame Semantic Parsing. InProceedings of the 23rd Chinese Na- tional Conference on Computational Linguistics (Volume 3: Evaluations), pages 21â31, Taiyuan, China. Chinese Information Processing Society of China. Maddalen Lopez de Lacalle, Egoitz Laparra, Itziar Aldabe, and German Rigau. 2016.Predi- cate Matrix: Automatically Extending the Se- mantic Interoperability between Predicate Re- sources .Language Resources and Evaluation, 50(2):263â289. Gosse Minnema, Sara Gemelli, Chiara Zanchi, Tommaso Caselli, and Malvina Nissim. 2022. SocioFillmore: A Tool for Discovering Perspec- tives. InProceedings of the 60th Annual Meet- ing of the Association for Computational Lin- guistics: System Demonstrations, pages 240â 250, Dublin, Ireland. Association for Computa- tional Linguistics. Kyoko Hirose Ohara, Seiko Fujii, Toshio Ohori, Ryoko Suzuki, Hiroaki Saito, and Shun Ishizaki. 2004. The Japanese Framenet Project: An In- troduction. InProceedings of LREC-04 Satel- lite WorkshopâBuilding Lexical Resources from Semantically AnnotatedCorporaâ(LREC2004), pages 9â11. Katrin Ortmann. 2022.Fine-Grained Error Analy- sis and Fair Evaluation of Labeled Spans . InPro- ceedings of the Thirteenth Language Resources and Evaluation Conference, LREC 2022, Mar- seille, France, 20-25 June 2022, pages 1400â 1407. European Language Resources Associa- tion. Yulia Otmakhova, Matteo Guida, and Lea Fr- ermann. 2026.Not all ANIMALs are equal: metaphorical framing through source domains and semantic frames . InFindings of the As- sociation for Computational Linguistics: ACL 2026 , pages 38334â38355, San Diego, Califor- nia, United States. Association for Computa- tional Linguistics. Martha Palmer, Olga Babko-Malaya, Ann Bies, Mona Diab, Mohamed Maamouri, Aous Man- souri, and Wajdi Zaghouani. 2008. A pilot Ara- bic Propbank. InProceedings of the Sixth In- ternational Conference on Language Resources and Evaluation(LRECâ08), Marrakech, Mo- rocco. European Language Resources Associa- tion(ELRA). Martha Palmer, Claire Bonial, and Jena Hwang. 2017.VerbNet: Capturing English Verb Behav- ior, Meaning, and Usage. Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, and Adel Bibi. 2023.Language Model Tokenizers Introduce Unfairness Be- tween Languages . InAdvances in Neural Infor- mation Processing Systems 36: Annual Confer- ence on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. Girishkumar Ponkiya, Diptesh Kanojia, Push- pak Bhattacharyya, and Girish Palshikar. 2021. FrameNet-assisted Noun Compound Interpreta- tion . InFindings of the Association for Compu- tational Linguistics: ACL-IJCNLP 2021, pages 2901â2911, Online. Association for Computa- tional Linguistics. Sameer Pradhan, Julia Bonn, Skatje Myers, Kathryn Conger, Tim Oâgorman, James Gung, Kristin Wright-bettner, and Martha Palmer. 2022.PropBank Comes of AgeâLarger, Smarter, and more Diverse . InProceedings of the 11th Joint Conference on Lexical and Com- putational Semantics, pages 278â288, Seattle, Washington. Association for Computational Lin- guistics. Karthik Raman, Iftekhar Naim, Jiecao Chen, Kazuma Hashimoto, Kiran Yalasangi, and Kr- ishna Srinivasan. 2022. Transforming Sequence Tagging Into A Seq2Seq Task. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11856â 11874, Abu Dhabi, United Arab Emirates. Asso- ciation for Computational Linguistics. Diego Reforgiato Recupero, Valentina Presutti, Sergio Consoli, Aldo Gangemi, and Andrea Gio- vanni Nuzzolese. 2015. Sentilo: Frame-Based Sentiment Analysis.Cogn. Comput., 7(2):211â 225. Ines Rehbein, Josef Ruppenhofer, Caroline Sporleder, and Manfred Pinkal. 2012. Adding nominal spice to SALSA - frame-semantic annotation of German nouns and verbs. In11th Conference on Natural Language Processing, KONVENS 2012, Empirical Methods in Natural Language Processing, Vienna, Austria, Septem- ber 19-21, 2012, Scientific series of the ĂGAI, pages 89â97. ĂGAI, Wien, Ăsterreich. Levi Remijnse, Pia Sommerauer, Antske Fokkens, and Piek T.J.M. Vossen. 2024.Tracking Per- spectives on Event Participants: a Structural Analysis of the Framing of Real-World Events in Co-Referential Corpora. InProceedings of the First Workshop on Reference, Framing, and Perspective @ LREC-COLING 2024, pages 1â 12, Torino, Italia. ELRA and ICCL. Josef Ruppenhofer, Michael Ellsworth, Myriam Schwarzer-Petruck, Christopher R Johnson, and Jan Scheffczyk. 2016. FrameNet I: Extended theory and practice. Technical report, Interna- tional Computer Science Institute. Ilia Ryazanov, Carl Ăhman, and Jonas Björklund. 2025.How ChatGPT Changed the Mediaâs Nar- ratives on AI: A semi-automated narrative anal- ysis through frame semantics.Minds and Ma- chines, 35(2). Konstantinos Sechidis, Grigorios Tsoumakas, and Ioannis Vlahavas. 2011. On the stratification of multi-label data.Machine Learning and Knowl- edge Discovery in Databases, pages 145â158. Kevin Stowe, Tuhin Chakrabarty, Nanyun Peng, Smaranda Muresan, and Iryna Gurevych. 2021. Metaphor Generation with Conceptual Map- pings. InProceedings of the 59th Annual Meet- ing of the Association for Computational Lin- guistics and the 11th International Joint Confer- ence on Natural Language Processing(Volume 1: Long Papers), pages 6724â6736, Online. As- sociation for Computational Linguistics. Swabha Swayamdipta, Sam Thomson, Chris Dyer, and Noah A. Smith. 2017. Frame-Semantic Parsing with Softmax-Margin Segmental RNNs and a Syntactic Scaffold.arXiv preprint arXiv:1706.09528. Ryosuke Taniguchi, Reina Hoshino, and Yoshi- nobu Kano. 2019. Legal Question Answering System Using FrameNet. InNew Frontiers in Artificial Intelligence, pages 193â206, Cham. Springer International Publishing. Rob van der Goot. 2024.Where are we still split on tokenization?InFindings of the Associa- tion for Computational Linguistics: EACL 2024, pages 118â137, St. Julianâs, Malta. Association for Computational Linguistics. Giulia Venturi, Alessandro Lenci, Simonetta Mon- temagni, Eva Maria Vecchi, Maria Teresa Sagri, Daniela Tiscornia, and Tommaso Agnoloni. 2009. Towards a FrameNet resource for the le- gal domain. InProceedings of the 3rd Work- shop on Legal Ontologies and Artificial Intelli- gence Techniques: 2nd Workshop on Semantic Processing of Legal Text, pages 67â76. Piek Vossen, Filip Ilievski, Marten Postma, Antske Fokkens, Gosse Minnema, and Levi Remijnse. 2020.Large-scale Cross-lingual Language Re- sources for Referencing and Framing. InPro- ceedingsoftheTwelfthLanguageResourcesand Evaluation Conference, pages 3162â3171, Mar- seille, France. European Language Resources Association. Patrick Xia, Guanghui Qin, Siddharth Vashishtha, Yunmo Chen, Tongfei Chen, Chandler May, Craig Harman, Kyle Rawlins, Aaron Steven White, and Benjamin Van Durme. 2021. LOME: Large Ontology Multilingual Extraction . In Proceedings of the 16th Conference of the Eu- ropean Chapter of the Association for Com- putational Linguistics: System Demonstrations, pages 149â159, Online. Association for Compu- tational Linguistics. Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022.ByT5: To- wards a Token-Free Future with Pre-trained Byte-to-Byte Models .Transactions of the Asso- ciation for Computational Linguistics, 10:291â 306. Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021.mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. InProceedings of the 2021 Confer- ence of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Language Technologies, pages 483â498, Online. Association for Computational Linguistics. Zheng Xin Yong, Patrick D. Watson, Tiago Tim- poni Torrent, Oliver Czulo, and Collin Baker. 2022.Frame Shift Prediction. InProceedings of the Thirteenth Language Resources and Eval- uation Conference, pages 976â986, Marseille, France. European Language Resources Associ- ation. Liping You and Kaiying Liu. 2005.Building Chi- nese FrameNet database. In2005 International Conference on Natural Language Processing and Knowledge Engineering, pages 301â306. Xinyan Yu, Trina Chatterjee, Akari Asai, Junjie Hu, and Eunsol Choi. 2022.Beyond Count- ing Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources. InFind- ings of the Association for Computational Lin- guistics: EMNLP 2022, pages 3725â3743, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Oliver Äulo. 2013.Constructions-and-frames anal- ysis of translations: The interplay of syntax and semantics in translations between English and German . Constructions and Frames , 5(2):143â 167. A Appendix 025050075010001250 10 9 10 7 10 5 10 3 10 1 DE 025050075010001250 EN 025050075010001250 10 9 10 7 10 5 10 3 10 1 FR 025050075010001250 IT 025050075010001250 10 9 10 7 10 5 10 3 10 1 KO 025050075010001250 LV 025050075010001250 10 9 10 7 10 5 10 3 10 1 NL 025050075010001250 PT 025050075010001250 10 9 10 7 10 5 10 3 10 1 SV 025050075010001250 ZH log norm. num. of occurrences Frames Figure 6: Number of occurrences ofallBFN frames(in log space)for each resource of Table2. (a)DE # Frame 1852 CALENDRIC_UNIT 1623 TELLING 1565 PEOPLE 1368 POLITICAL_LOCALES 796 REQUEST (b)EN # Frame 842 WEAPON 672 LOCALE_BY_USE 603 STATEMENT 556 POLITICAL_LOCALES 451 LEADERSHIP (c)FR # Frame 1599 CAUSATION 786 EVIDENCE 514 COMMERCE_BUY 414 COMMERCE_SELL 318 REASON (d)IT # Frame 109 STATEMENT 47 ARRIVING 37 ATTEMPT 33 KINSHIP 33 DESIRING (e)KO # Frame 505 STATEMENT 329 EXPERIENCER_FOCUS 316 LOCALE_BY_USE 253 LEADERSHIP 248 POSSESSION (f)LV # Frame 475 TELLING 466 STATEMENT 349 POSSESSION 346 EXISTENCE 343 ARRIVING (g)NL # Frame 149 CATASTROPHE 97 SUSPICION 77 CAUSE_HARM 65 PARTICIPATION 55 COMMITTING_CRIME (h)PT # Frame 354 DEGREE 311 CARDINAL_NUMBERS 248 NEGATION 237 LOCATIVE_RELATION 157 POSSESSION (i)SV # Frame 66 EMPTYING 58 MAKE_NOISE 47 SELF_MOTION 42 EXPERIENCER_OBJ 42 PLACING (j)ZH # Frame 288 CAUSE_TO_MAKE_PROGRESS 261 CHANGE_POSITION_ON_A_SCALE 249 BEING_IN_CATEGORY 218 AMOUNTING_TO 178 CAUSATION Table 6: Top 5 most common annotated frames for each of the language-specific corpora of Table2. FrameFE PRF1PRF1 lang flagmodeldataset DELOMEBFN0.217 0.853 0.346 0.139 0.587 0.224 mFNC0.925 0.929 0.927 0.859 0.859 0.859 mT5BFN0.136 0.706 0.228 0.079 0.382 0.131 mFNC0.748 0.749 0.748 0.632 0.636 0.634 mT5 small BFN0.101 0.583 0.172 0.057 0.268 0.094 mFNC0.833 0.823 0.828 0.678 0.676 0.677 ENLOMEBFN0.7440.8520.794 0.595 0.673 0.632 mFNC0.7850.8400.812 0.666 0.705 0.685 mT5BFN0.466 0.623 0.533 0.320 0.390 0.352 mFNC0.614 0.691 0.650 0.491 0.494 0.492 mT5 small BFN0.382 0.603 0.468 0.267 0.351 0.304 mFNC0.642 0.713 0.676 0.502 0.515 0.508 FRLOMEBFN0.059 0.332 0.100 0.027 0.107 0.043 mFNC0.874 0.887 0.880 0.712 0.713 0.712 mT5BFN0.031 0.225 0.054 0.015 0.061 0.024 mFNC0.580 0.584 0.582 0.344 0.342 0.343 mT5 small BFN0.033 0.261 0.058 0.016 0.058 0.025 mFNC0.704 0.679 0.691 0.382 0.378 0.380 ITLOMEBFN0.1420.7470.238 0.084 0.348 0.136 mFNC0.7000.6900.695 0.564 0.578 0.571 mT5BFN0.051 0.341 0.088 0.022 0.090 0.035 mFNC0.438 0.533 0.481 0.296 0.318 0.307 mT5 small BFN0.041 0.272 0.071 0.018 0.069 0.029 mFNC0.530 0.534 0.532 0.330 0.323 0.327 KOLOMEBFN0.3640.6790.474 0.187 0.303 0.231 mFNC0.7910.5250.631 0.534 0.471 0.501 mT5BFN0.176 0.588 0.271 0.079 0.238 0.119 mFNC0.495 0.651 0.562 0.378 0.436 0.405 mT5 small BFN0.155 0.487 0.236 0.059 0.166 0.087 mFNC0.585 0.654 0.617 0.435 0.437 0.436 LVLOMEBFN0.1350.8060.231 0.0630.4630.111 mFNC0.8550.4350.577 0.7970.4100.542 mT5BFN0.0760.6490.136 0.033 0.310 0.059 mFNC0.6120.6120.612 0.566 0.567 0.566 mT5 small BFN0.061 0.560 0.111 0.021 0.206 0.037 mFNC0.619 0.619 0.619 0.555 0.563 0.559 NLLOMEBFN0.088 0.491 0.149 0.017 0.127 0.030 mFNC0.752 0.652 0.699 0.617 0.468 0.533 mT5BFN0.042 0.290 0.073 0.008 0.066 0.014 mFNC0.446 0.495 0.469 0.373 0.375 0.374 mT5 small BFN0.034 0.260 0.060 0.007 0.058 0.012 mFNC0.570 0.547 0.559 0.430 0.390 0.409 PTLOMEBFN0.515 0.536 0.525 0.311 0.320 0.316 mFNC0.688 0.746 0.716 0.517 0.525 0.521 mT5BFN0.326 0.336 0.331 0.174 0.189 0.182 mFNC0.574 0.618 0.595 0.397 0.394 0.396 mT5 small BFN0.262 0.267 0.264 0.131 0.148 0.139 mFNC0.561 0.626 0.592 0.388 0.393 0.390 SVLOMEBFN0.072 0.260 0.113 0.040 0.202 0.067 mFNC0.545 0.417 0.472 0.362 0.332 0.346 mT5BFN0.041 0.191 0.067 0.020 0.114 0.034 mFNC0.318 0.301 0.309 0.211 0.211 0.211 mT5 small BFN0.033 0.164 0.054 0.015 0.092 0.026 mFNC0.304 0.288 0.296 0.194 0.195 0.194 ZHLOMEBFN0.1350.4280.205 0.075 0.135 0.097 mFNC0.7470.3690.494 0.472 0.303 0.369 mT5BFN0.049 0.317 0.086 0.020 0.081 0.032 mFNC0.440 0.471 0.455 0.306 0.310 0.308 mT5 small BFN0.037 0.239 0.063 0.013 0.051 0.021 mFNC0.517 0.514 0.516 0.336 0.338 0.337 Table 7: Results obtained by each model on each language, computed using the FairEval framework.