Paper deep dive
Towards Linguistically-informed Representations for English as a Second or Foreign Language: Review, Construction and Application
Wenxi Li, Xihao Wang, Weiwei Sun
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 1:48:58 AM
Summary
The paper proposes a constructivist approach to represent English as a Second or Foreign Language (ESFL) by treating constructions as fundamental form-meaning units. It introduces the ESFL SemBank, a gold-standard syntactico-semantic resource containing 1643 annotated sentences, developed using a Synchronous Hyperedge Replacement Grammar (SHRG) framework to bridge ESFL and standard English.
Entities (4)
Relation Signals (3)
ESFL SemBank â implements â Synchronous Hyperedge Replacement Grammar
confidence 100% ¡ Building on this SHRG framework, we develop ESFL SemBank
Constructivist Theory â underpins â ESFL SemBank
confidence 95% ¡ Grounded in constructivist theories, the paper treats constructions as the fundamental units of analysis... resulting a gold-standard syntactico-semantic resource
ESFL SemBank â tests â Linguistic Niche Hypothesis
confidence 90% ¡ To demonstrate the sembank's practical utility, we conduct a pilot study testing the Linguistic Niche Hypothesis
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The widespread use of English as a Second or Foreign Language (ESFL) has sparked a paradigm shift: ESFL is not seen merely as a deviation from standard English but as a distinct linguistic system in its own right. This shift highlights the need for dedicated, knowledge-intensive representations of ESFL. In response, this paper surveys existing ESFL resources, identifies their limitations, and proposes a novel solution. Grounded in constructivist theories, the paper treats constructions as the fundamental units of analysis, allowing it to model the syntax--semantics interface of both ESFL and standard English. This design captures a wide range of ESFL phenomena by referring to syntactico-semantic mappings of English while preserving ESFL's unique characteristics, resulting a gold-standard syntactico-semantic resource comprising 1643 annotated ESFL sentences. To demonstrate the sembank's practical utility, we conduct a pilot study testing the Linguistic Niche Hypothesis, highlighting its potential as a valuable tool in Second Language Acquisition research.
Tags
Links
- Source: https://arxiv.org/abs/2604.09008v1
- Canonical: https://arxiv.org/abs/2604.09008v1
Trouble viewing inline? Open PDF directly â
Full Text
54,312 characters extracted from source content.
Expand or collapse full text
Towards Linguistically-informed Representations for English as a Second or Foreign Language: Review, Construction and Application Wenxi Li 1,2 , Xihao Wang 3 , Weiwei Sun 4* 1 School of the Chinese Nation Studies. 2 School of Liberal Arts, Minzu University of China, 27 South Zhongguancun Ave, Beijing, 100081, China. 3 Department of Chinese Language and Literature, Peking University, 5 Yiheyuan Rd, Beijing, 100871, Beijing, Chia. 4* Department of Computer Science and Technology, University of Cambridge, 15 J Thomson Ave, Cambridge, CB3 OFD, Cambridgeshire, United Kingdom. *Corresponding author(s). E-mail(s): ws390@cam.ac.uk; Contributing authors: liwenxi@pku.edu.cn; wangxihao@pku.edu.cn; Abstract The widespread use of English as a Second or Foreign Language (ESFL) has sparked a paradigm shift: ESFL is not seen merely as a deviation from standard English but as a distinct linguistic system in its own right. This shift highlights the need for dedicated, knowledge-intensive representations of ESFL. In response, this paper surveys existing ESFL resources, identifies their limitations, and proposes a novel solution. Grounded in constructivist theories, the paper treats constructions as the fundamental units of analysis, allowing it to model the syntaxâsemantics interface of both ESFL and standard English. This design captures a wide range of ESFL phenomena by referring to syntactico-semantic mappings of English while preserving ESFLâs unique characteristics, resulting a gold-standard syntactico- semantic resource comprising 1643 annotated ESFL sentences. To demonstrate the sembankâs practical utility, we conduct a pilot study testing the Linguistic Niche Hypothesis, highlighting its potential as a valuable tool in Second Language Acquisition research. Keywords: English as a Second or Foreign Language, SyntaxâSemantics Interface, Constructivist Theory, SemBank 1 arXiv:2604.09008v1 [cs.CL] 10 Apr 2026 1 Introduction In todayâs globalized world,, English as a Second or Foreign Language (ESFL), produced by non-native speakers who make up over 70% of English users worldwide, has become a primary tool for cross-lingual communication, particularly in open and informal contexts such as social media. This widespread usage necessitates a critical paradigm shift: ESFL may not be recognized as an erroneous approximation of native norms, but as a legitimate, distinct linguistic system shaped by the complex interplay of a speakerâs primary languages and English (Selinker 1972; Nemser 1991; Corder 1982). As such, ESFL requires dedicated linguistically-informed resources to accurately represent and process its unique structures just like any other natural language. However, a survey of existing representations for ESFL (see§2), ranging from POS tagging (DÄąaz-Negrillo et al. 2010) to universal syntactic frameworks (Berzak et al. 2016) and semantic annotations (Zhao et al. 2020), reveals a fundamental challenge. Because ESFL lacks a static ânativeâ anchor and is characterized by high variability and rapid evolution, developing a consistent framework is difficult. Most existing resources use standard English as a rigid reference point. This creates a persistent tension: strictly adhering to native norms risks erasing the distinctive features and cross-linguistic nuances of the multilingual speaker, while relaxing those norms often compromises the coherence of the representation. To resolve this, we adopt a multilingual-centric constructivist approach. Construc- tivist theories posit that language consists of âconstructionsâ, i.e., formâmeaning pairings that serve as the fundamental units of linguistic analysis (Goldberg 2003; Croft 2001). We therefore propose deriving sentence-level representations for both standard English and ESFL from a shared inventory of reusable syntactico-semantic constructions. This approach preserves the distinctive surface forms (syntax) generated by multilingual speakers while grounding them in the underlying semantic equivalents shared with English. By modeling the syntaxâsemantics interface as a flexible mapping, we can capture the idiosyncratic features of the ESFL grammar without sacrificing representational consistency. We implement this approach using Synchronous Hyperedge Replacement Grammar (SHRG), a formalism that pairs syntactic rules, represented as trees, with corresponding semantic rules over graphs. Together, these synchronized rules form a constructional inventory that treats ESFL sentences as systematic linguistic productions rather than as random or ill-formed errors. Building on this SHRG framework, we develop ESFL SemBank, a gold-standard syntactico-semantic resource (see§4). Its construction follows a rigorous ReviewâReviseâRebuild workflow: we manually review silver-standard semantic graphs (Zhao et al. 2020), revise rejected analyses to ensure semantic fidelity, and rebuild unparsable sentences by selecting or creating appropriate SHRG rules. ESFL SemBank provides a powerful lens for language acquisition. By explicitly encoding how ESFL speakers â whose second or foreign language is English â map linguistic form to meaning, the resource enables in-depth analyses of the cognitive mechanisms and developmental trajectories underlying ESFL production. In§5, we illustrate its utility through a case study of the Linguistic Niche Hypothesis (Lupyan and Dale 2010). Our statistical analysis shows that, while ESFL syntactic patterns largely align with those of English, ESFL often exhibits a more transparent and 2 consistent mapping between form and meaning. This finding suggests that multilingual speakers may optimize the syntaxâsemantics interface for communicative clarity, offering empirical evidence of how non-native speakers navigate typological constraints to achieve cross-lingual fluency. 2 Survey: Existing Resources for ESFL Efforts to develop knowledge-intensive, task-independent, and linguistically-informed representations for human languages have become an important research area (e.g., Oepen and Lønning 2006; Copestake 2009; Banarescu et al. 2013; Abend and Rappoport 2013). Similar efforts have also been extended to ESFL, leading to the development of several corpora specifically designed for it. Based on their strategies for referencing standard English, we group these corpora into three categories: ⢠Target-language-based: directly projecting English-based frameworks or models onto ESFL data (§2.1). ⢠Multiple-criteria-based: applying various English-derived criteria to interpret ESFL phenomena (§2.2). ⢠Universal-framework-based: using a language-neutral framework to mediate between English and ESFL (§2.3). 2.1 Target-language-based Approach One representative example using the target-language-based approach is the Konan- JIEM Learner Corpus (Nagata et al. 2011; Nagata and Sakaguchi 2016). This resource introduces a phrase-structure-based annotation scheme and a manually constructed treebank. Annotations are grounded in native English syntactic norms, consistently identifying the structure in standard English that most closely reflects the learnerâs intended meaning. When learner productions cannot be aligned with any plausible native structure, the scheme introduces specialized node labels â such as -err, in, ce, xp-ord, uk, and up â to mark errors related to omission, insertion, substitution, word order, and unknown words or phrases, respectively. Another example is the SemBank of English as a Second Language (Zhao et al. 2020), which, so far as we know, is the only effort to construct a comprehensive semantic bank for learner English. This resource adopts a target-language-based strategy by using the ACE parser, trained on English Resource Grammar (ERG; Flickinger 2000), to parse ESFL sentences. For each ESFL sentenceS, the parser derivesK-best corresponding semantic graphs. A reranking procedure then selects the candidate that best aligns with the gold-standard syntactic tree from the TLE corpus (Berzak et al. 2016). In total, 2691 ESFL sentences are successfully parsed. 2.2 Multiple-criteria-based Approach Resources under the multiple-criteria-based approach, drawing on multiple English- derived criteria, adopt layered annotation strategies to interpret ESFL data more flexibly. For instance, DÄąaz-Negrillo et al. (2010) propose a tripartite analysis for POS tagging of ESFL in the NOCE Corpus, a Spanish learner corpus of English. Based on an 3 empirical investigation of learner language collected in NOCE, the authors observe that applying single POS tagging scheme is problematic â the distributional, lexical, and morphological evidence for classifying a token in ESFL data often does not converge on a single POS tag. Therefore, they suggest annotating each word with three parallel POS tags to capture its lexical, morphological, and distributional forms, respectively. The Syntactically Annotated Learner Language of English (SALLE) project (Ragheb and Dickinson 2014; Dickinson and Ragheb 2015) also proposes a multi-layer frame- work to represent the morpho-syntactic information of learner language. Based on the CHILDES annotation framework (Sagae et al. 2004, 2007), both POS tags and dependencies are divided into two layers: one for morphological information and the other for distributional information. This corpus also includes subcategorization frames to represent dependencies that are selected for by words but not necessarily realized. 2.3 Universal-framework-based Approach The Treebank of Learner English (TLE; Berzak et al. 2016), also known as UD English- ESL, employs a universal-framework-based approach. Each ESFL sentence is manually corrected into standard English, and both versions are annotated using the Penn Treebank POS tagset (Santorini 1990) and syntactic dependencies from the Universal Dependencies framework (UD; Nivre et al. 2016). These annotation schemes are thought to be language-neutral, thus allowing comprehensive coverage of both learner ESFL sentences and their English counterparts. Besides TLE, several corpora, including the Treebank of Spoken Second Language English (SL2E Treebank; Kyle et al. 2022), as well as ESFL resources for other learner populations (Lee et al. 2017; Lin et al. 2018; Davidson et al. 2020) also adopts this approach. 2.4 Summary In our opinion, the aforementioned approaches possess distinct strengths and inherent limitations. The target-language-based approach, for instance, benefits significantly from leveraging existing English resources, offering efficiency and convenience in its application. However, a major drawback is its inability to comprehensively capture the distinctive features unique to ESFL. Conversely, the multiple-criteria-based approach offers more detailed descriptions of ESFL, which is valuable for in-depth analysis. How- ever, its primary challenge lies in scalability, as the complexity and cost of annotating large-scale ESFL data become prohibitive. Meanwhile, the universal-framework-based approach endeavors to connect ESFL with the target language (English) through their shared semantic interpretation. Yet, the finer-grained correspondences between the two languages remain undefined. We contend that these collective difficulties underscore a fundamental challenge in constructing a systematic representation framework for ESFL â unlike creoles or pidgins, ESFL cannot readily rely on native speaker intuitions to judge its well- formedness. As a result, explaining and predicting variations of ESFL in a principled way remains especially difficult. Existing approaches typically rely on standard English as a reference point, but this leads to a core dilemma: strictly following standard English 4 rules obscures the distinctive features of ESFL, whereas relaxing those constraints too much undermines systematicity and internal coherence. Therefore, striking an appropriate balance between these two poles becomes the pivotal question to address. 3 Method 3.1 Theoretical Foundation: Construcivism Faced with the challenges of representing ESFL data, we draw on insights from constructivism. Though encompassing a broad and heterogeneous body of research (see Hoffmann and Trousdale, 2013; Goldberg, 2013), constructivist theories are unified by the core assumption that words, idioms, phrases, and even full sentences are formâmeaning pairings, or constructions, that are stored in an inventory capable of exhaustively describing human language. Building on this foundation, we propose using syntactico-semantic constructions as the basic representational units for both ESFL and standard English. By decomposing sentence-level representations into finer-grained constructions and aligning ESFL and English at this micro-level, the constructivist framework offers a principled means of bridging the two languages. Figure 1 illustrates how our method addresses the so-called âerrorsâ in ESFL, i.e., its deviations from standard English. While ESFL constructions diverge in syntactic form due to token omission, insertion, or unconventional structures (see Examples 1â3), their intended semantics remains intelligible. Therefore, these ESFL constructions can be systematically linked to their standard English counterparts by identifying underlying semantic equivalences and shared distributional patterns. (1) Omission: I had to sleep in (a) tent. (2) Insertion: We contactedwith Kim. (3) Transposition: He ::::::::::: visits often Paris. We argue that the constructivist approach â which represents ESFL data through a continually expanding inventory of constructions, in contrast to lexicalist frameworks that rely on systematic rules for combining lexical units â offers two key advantages. First, it effectively accounts for the wide range of âerrorsâ typically found in ESFL. By recognizing alternative formâmeaning pairings that emerge organically from actual usage, rather than imposing rigid grammatical norms, the constructivist framework captures the internal variation of ESFL in a more nuanced and context-sensitive way. Second, while accommodating variation, this approach also maintains formal coherence across representations. ESFL constructions can be aligned with their standard English counterparts and systematically encoded into structural templates, ensuring their compatibility with subsequent syntactico-semantic composition processes. Thus, rather than requiring a trade-off between descriptive adequacy and formal systematicity, the constructivist approach achieves both as a flexible, principled, and robust framework. 5 S NP N I VP V sleep P P in NP N tent (a) ESFL sentence S NP N I VP V sleep P P in NP DET a N tent (b) English counterpart SYNTAX NP â N tent NP â DET a + N tent SEMANTICS âx.tent Ⲡ(x) (c) SYN-SEM rule Fig. 1: An illustration of our method using the example I sleep in (a) tent. The constructions highlighted by rectangles differ in whether the article is omitted, but they share similar semantics and structural encoding within the syntactic tree, allowing them to be aligned. 3.2 Modeling the SyntaxâSemantics Interface To put the constructivist approach into practice, we adopt Synchronous Hyperedge Replacement Grammar (SHRG), a formalism that models each linguistic construc- tion as a meaning-bearing unit. In SHRG, every construction consists of two tightly coupled components: a syntactic rule, represented as a tree that encodes form, and a semantic rule, represented as a graph that encodes meaning. These two components are synchronized, forming two sides of the same coin. VP V V visit Adv often NP Paris (a) Syntax Tree visitv1 ARG2 named(âPairsâ) BV properq ARG1 oftena1 (b) Semantic Graph Fig. 2: Syntactico-semantic representation of visit often Paris To illustrate this framework, consider the phrase visit often Paris. Figure 2 presents its full syntactic and semantic representations, while Table 1 decomposes the phrase into its constituent constructions. Each construction is specified by a paired set of rules: a CFG ruleAâ βcapturing its syntactic structure, and a corresponding HRG ruleAâ Gencoding its semantic composition. These paired rules jointly model the syntaxâsemantics interface. For instance, when the syntactic component combines a verb and an adverb via the ruleV â V+Adv, the corresponding semantic rule defines how the adverb modifies the underlying event structure. By explicitly aligning syntactic configurations with their corresponding semantic operations, SHRG not only captures 6 the structures characteristic of ESFL but also renders them systematic, principled, and directly comparable to those of standard English. Table 1: Component rules for visit often Paris. IndexSYNSEM 1 Vâ visitVâ visitv1 2 Advâ oftenAdvâ oftena1 3 Vâ V+AdvVâ Adv ARG1 V 4 NPâ ParisNPâ properq BV named(âParisâ) 5 VPâ V+NPVPâ VARG2 NP These SHRG rules also support the compositional derivation of syntactico-semantic representations, demonstrating the formal systematicity of this method. As exhibited in Figure 3, the syntactic derivation begins with terminal symbols, and production rules of the formAâ βare applied recursively to derive non-terminals until the start symbol is reached. Simultaneously, the semantic representation is built using parallel HRG rulesAâ G. In this process, a hyperedgeeinGis rewritten if it is labeled with a non-terminal symbolnfrom theC. The edgeeis then removed, and a copy of the previously derived hypergraphHwhose left-hand side label isn, is inserted intoG. The external nodesXofHare fused with the nodes originally connected bye, while all other hyperedges in G remain unaffected. 4 ESFL SemBank We apply the SHRG-based constructivist approach to ESFL data, resulting a syntactico- semantic ESFL SemBank. Generally speaking, its development involves three main steps: ⢠Review: silver-standard semantic graphs from Zhao et al. (2020) are manually reviewed as accepted or rejected. ⢠Revise: extract SHRG rules and those of rejected graphs are manually revised to produce their accurate semantic representations. 7 visitv1 ARG2 named(âPairsâ) BV properq ARG1 oftena1 VP oftena1 ARG1 visitv1 V visitv1 V visit oftena1 Adv often properq BV named(âParisâ) NP Paris 12 34 5 Fig. 3: Synchronous construction of syntax and semantics for visit often Paris. ⢠Rebuild: semantics of unparsable ESFL sentences in Zhao et al. (2020) are manually rebuilt by selecting appropriate SHRG rules. 4.1 Review ers, as well as their parsers. We start by manually reviewing holistic silver-standard ESFL semantic graphs from Zhao et al. (2020), determining if they could be accepted or rejected. This step establishes a reliable foundation for our sembank. To be more specific, the review is conducted by a team comprising one Ph.D. student and two undergraduate students majoring in linguistics. To ensure consistency and accuracy, they firstly undergo training in English Resource Semantics (ERS; Flickinger et al. 2014) 1 , which is the framework used to represent the previously parsed semantic graphs. During annotation, DeepBank (Flickinger et al. 2012) â a comprehensive ERS-based semantic resource known for handling a broad range of English linguistic phenomena â is used as a reference. Annotators evaluate whether each graph from Zhao et al. (2020) faithfully represents the intended semantics of the corresponding ESFL sentence and then assign one of three labels: accept, reject, or abandon based on its semantic fidelity. Following several rounds of training, comparison, and discussion, three annotators reach a high level of inter-annotator agreement (IAA). The annotation process is 1 ERS is a semantic framework developed in conjunction with English Resource Grammar (ERG; Flickinger 2000), within Head-driven Phrase Structure Grammar (HPSG; Pollard and Sag 1994). 8 then streamlined: remaining instances are either cross-validated by two annotators or annotated by a single annotator. Their annotation quality, as measured by percentage IAA, is reported in Table 2. The exceptionally high IAAs across annotators, ranging from 97% to 99%, consistently indicate strong consensus on the semantic interpretations of ESFL sentences. Table 2: Inter-annotator agree- ments (IAAs) in annotation. ESFLCEFSL Anno1-Anno2-Anno399.2999.48 Anno1-Anno298.9598.19 Anno1-Anno397.8298.77 Anno2-Anno298.2899.12 In total, 1567 ESFL and 2189 CESFL sentences are annotated. After removing records with inconsistent tags, 1543 ESFL and 2138 CESFL sentences remain. The final distribution of annotation labels is shown in Table 3: 46.92% of ESFL graphs are accepted and 53.08% rejected, while for CESFL, 61.04% are accepted and 38.96% rejected. Table 3: Numbers (#num) and percentages (#per) of valid sentences whose semantic graphs accepted (acc) or rejected (rej) by three, two, or one annotators. ESLCESL #num#per#num#per Triple-acc6445.71%12966.84% Triple-rej7654.28%6433.16% Triple-all140100.00%193100.00% Double-acc58646.03%105159.65% Double-rej68753.97%71140.35% Double-all1273100.00%1762100.00% Single-acc7456.92%12568.31% Single-rej5643.08%5831.69% Single-all130100.00%183100.00% Overall-acc72446.92%130561.04% Overall-rej81953.08%83338.96% Overall-all1543100.00%2138100.00% 9 4.2 Revise We adopt the method in Li et al. (2025) to extract SHRG rules from both accepted and rejected semantic graphs. Based on the induced rule inventory, we manually revise the compositional SHRG rules in rejected ESFL graphs to reflect their intended semantics. Number Number, as a grammatical category, is expressed by contrasts between singular and plural forms in English. However, ESFL may use various forms, particularly the zero form (see example 4), to decode number. This often leads to parsing errors with the ACE parser and thus results in inaccurate semantic interpretations. Since this phenomenon involves determiners and often lacks clear indicators of singularity or plurality, we address it with an SHRG rule that includes an abstract predicate,udefq, alongside other relevant SHRG rules (see Table 4). These rules are then syntactically recombined to produce accurate semantic representations (see Figure 4). (4) I was impressed when I heard that she liked playing puzzle alone. Table 4: Original and modified SHRG rules for Example 4. LHSN ori V ori COMP mod N mod Synthatpuzzlethatpuzzle Sem thatqdem BV genericentity puzzlev1 â udefq BV puzzlen1 VP VP P alone V puzzle NP V she liked playing N that (a) Original Syn VP VP P alone VP N puzzle V she liked playing COMP that (b) Modified Syn puzzlev1thatqdem genericentity ARG1 BV (c) Original Sem udefq puzzlen1 BV (d) Modified Sem Fig. 4: Relevant original and modified syntactic and semantic analysis for Example 4. Case Case denotes the syntactic function of a nominal constituent within a sentence. In English, case is primarily limited to pronouns and the genitive case. Example 5 is a 10 case-related divergence in ESFL, where the genitive case is not explicitly marked by âs but is instead implied through zero form and word order. To address this, we apply an SHRG rule (see Table 5) that specifies the possessive relation. Figure 5 shows the modified result. (5) I think we know a writer life. Table 5: Original and modified SHRG rules for Example 5. LHSNP ori NP mod SynN + N + NP Sem compound ARG1ARG2 N poss ARG1ARG2 NNP NP NP N life N writer DET a (a) Original Syn NP N life NP N writer DET a (b) Modified Syn compound aq udefq lifenofwriternof ARG1ARG2 BV BV (c) Original Sem poss defimplicitq aq lifenofwriternof ARG1ARG2 BV BV (d) Modified Sem Fig. 5: Relevant original and modified syntactic and semantic analysis for Example 5. Tense & Aspect Tense describes the time when an action or event occurs while aspect deals with how an action unfolds over time. These two grammatical categories are closely related as both of them encode temporal information. As shown in Example 6, ESFL may employ grammatical means distinct from English to convey tense and aspect. Figure 6 further illustrates how the zero-form verb give introduces ambiguity for the ACE parser, preventing it from correctly identifying the perfect aspect. To address this diversity, we apply SHRG rules, as detailed in Table 6, to make targeted adjustments. (6) I hope I have give to you the information you needed. Voice Voice refers to the relationship between a verb and its participants, with active and passive being the most commonly used. In English, this distinction is typically conveyed 11 Table 6: Original and modified SHRG rules for Example 6. LHSV ori VP ori V mod VP mod SynhaveV + NPhaveV + NP Sem havevcause VARG1 NP â VARG2 NP VP NP the information V VP P to you V give V have (a) Original Syn VP VP NP the information V P to you V give V have (b) Modified Syn givev1theq havevcause udefq informationnon-about pron ARG1 ARG1 ARG3 BVBV (c) Original Sem givev1thequdefq informationnon-about pron ARG2 ARG3 BVBV (d) Modified Sem Fig. 6: Relevant original and modified syntactic and semantic analysis for Example 6. through word order and inflectional markers. However, as illustrated in Example 7, ESFL exhibits notable variations in how voice is expressed, such as the zero-form verb mind. Figure 7 demonstrates how this creates challenges for the ACE parser in correctly identifying the passive voice. To address this issue, we apply SHRG rules outlined in Table 7 to make the necessary adjustments. (7)It wouldnât be mind if the restaurant had pressed in order to eat or drink something. S VP VP N mind V be V wouldnât NP N It (a) Original Syn S VP/NP VP/NP V mind V be V wouldnât NP N It (b) Modified Syn bevid udefq pronounq mindn1 pron ARG2 BV ARG1 BV (c) Original Sem pargd pronounq pron mindv1 ARG2 BV ARG1 ARG2 (d) Modified Sem Fig. 7: Relevant original and modified syntactic and semantic analysis for Example 7. Person Person is a grammatical category that distinguishes forms of different participants in an event. It could affect verb inflections: in English, the form of a verb typically 12 Table 7: Original and modified SHRG rules for Example 7. LHSV ori N ori VP ori S ori SynbemindV + NNP + VP Sem bevid udefq BV mindn1 VARG2 N VPARG1 NP LHSV mod1 V mod2 VP/NP mod S mod SynbemindV + VNP + VP/NP Semâ mindv1 parg ARG1 V ARG2 VP/NP NP ARG2 depends on the person and number of its subject. However, in ESFL, this agreement may sometimes be conveyed through a zero form, as shown in Example 8. The absence of overt marking can lead to misidentification by the ACE parser. For instance, as illustrated in Figure 8, the phrase have good meaning is incorrectly analyzed as an imperative sentence. We address this issue by applying the SHRG rules in Table 8. (8) It looks nice and have good meaning. Table 8: Original and modified SHRG rules for Example 8. LHSS ori S ori VP ori VP ori SynConj + S + SConj + VPVP + VP Sem Conj R-IND R-HND S S L-IND L-HND S Conj R-IND R-HND VP VP L-IND L-HND VP Binding Binding theory examines the constraints governing the use of anaphoric expressions, such as pronouns and reflexives, to determine how their meanings are derived from other elements in the context. According to the principles of binding theory (Chomsky 1981, 1986), a reflexive anaphora in English must be bound within its governing category or local domain. However, this principle is less strictly adhered to in ESFL, as demonstrated in Example 9, where herself is interpreted as referring to Pat, rather 13 S S S have good meaning Conj and S It looks nice (a) Original Syn S VP VP VP have good meaning Conj and VP looks nice NP It (b) Modified Syn lookv1 andc havev1 pronounq pronpronounqpron ARG1 ARG1 L-IND/HNDR-IND/HND BV BV (c) Original Sem lookv1 andc havev1 pronounq pron ARG1 ARG1 L-IND/HNDR-IND/HND BV (d) Modified Sem Fig. 8: Relevant original and modified syntactic and semantic analysis for Example 8. than the principle-sanctioned antecedent my friends (see Figure 9). We utilize the SHRG rules detailed in Table 9 to resolve this discrepancy. (9)At first, Pat denied all the things my friends told me about herself but she finally agreed with that. Table 9: Original and modified SHRG rules for Exam- ple 9. LHSPP/N ori NP ori P mod P mod SynaboutNP + NPaboutP + NP Semâ appos ARG1ARG2 NPNP aboutp PARG2 NP NP NP N herself NP S VP P/N about VP told me NP my friend N things (a) Original Syn NP S VP P NP herself P about VP told me NP my friend N things (b) Modified Syn appos tellvabout thingnof-about pron pronounq theq ARG1ARG2 BV ARG3 BV (c) Original Sem aboutp tellvabout thingnof-about pron pronounq theq ARG2 ARG1 BV ARG3 BV (d) Modified Sem Fig. 9: Relevant original and modified syntactic and semantic analysis for Example 9. Ellipsis Ellipsis refers to the omission of a previously mentioned string in subsequent structures. In ESFL, ellipsis demonstrates notable flexibility. As shown in Example 10, constituents 14 that are typically not elided in standard English may be omitted. This variability can result in inaccuracies in the ACE parserâs analysis, as illustrated in Figure 10. To mitigate this issue, we apply the SHRG rules presented in Table 10. (10) His voice is also soulful and sometimes even gave us blissful. Table 10: Original and modified SHRG rules for Example 10. LHSVP ori P ori VP ori VP mod NP mod VP mod SynV+NPADJVP+PPV+NPADJVP+NP Sem VARG2 NP ADJ subord ARG1ARG2 VPPP VARG3 NP udefq BV genericentity ARG1ADJ VPARG2 NP VP P ADJ blissful VP NP us V gave (a) Original Syn VP NP ADJ blissful VP NP us V gave (b) Modified Syn givev1 subord pronounq pron blissfula1 ARG1 ARG2 ARG2 BV (c) Original Sem givev1udefq pronounq genericentity blissfula1 pron ARG2ARG3 BV BV ARG1 (d) Modified Sem Fig. 10: Relevant original and modified syntactic and semantic analysis for Example 10. Filler-Gap Filler-Gap constructions are considered as the result of phrasal movement in generative grammar â the moved phrase serves as a filler for its canonical position, which is left unexpressed as a gap. Establishing a connection between the filler and its original in situ position is essential to ensure the utterance is interpretable. In ESFL, however, the presence and forms of fillers and gaps can differ significantly from those in English (see Example 11). As shown in Figure 11, these variations pose challenges for the ACE parser. We employ modified SHRG rules in Table 11 to address it. (11) You can find any information from it what you want. Argument Structure Argument structure refers to the syntactic configuration projected by a lexical item, particularly by a verb, within a sentence. It defines the specific requirements regarding 15 Table 11: Original and modified SHRG rules for Example 11. LHSN ori NP ori N mod NP mod SynwhatN + SwhatN + S Sem whichq BV thing NARG1 S â SARG2 N NP S S/N you want N what N P from it N information (a) Original Syn NP S S/N you want N what N P from it N information (b) Modified Syn wantv1 informationnon-about whichq pronq thing pron ARG1 ARG2ARG1 BV BV (c) Original Sem wantv1 anyq pronq informationnon-about pron ARG2 ARG1 BV BV (d) Modified Sem Fig. 11: Relevant original and modified syntactic and semantic analysis for Example 11. both the number and realization of arguments. However, in ESFL, the argument structure of the same verb can differ significantly from its English counterpart. These variations are evident not only in the number of arguments (see Example 12), but also in how they are realized (see Example 13). Figure 12 and Figure 13, along with Table 12 and Table 13, illustrate the challenges posed by these diverse argument structures for the ACE parser, as well as how we address them. (12) I hope you can enjoy. S S/N VP/N V/N enjoy V can N you NP N hope N I (a) Original Syn S VP S VP V/N enjoy V can N you V hope NP N I (b) Modified Syn compound properq udefq named(I) hopen1 pronounq enjoyv1 pron canvmodal ARG2 ARG1 ARG2ARG1 BV BV BV ARG1 (c) Original Sem hopev1 pronounq canvmodal pronpronounq ellipsisref pron ARG2ARG1 ARG1 BV BV ARG1 (d) Modified Sem Fig. 12: Relevant original and modified syntactic and semantic analysis for Example 12. (13) A telephone is used when we need to contact with someone. 16 Table 12: Original and modified SHRG rules for Example 12. LHSN ori NP ori V/N ori S ori SynhopeN + NenjoyNP + S/N Sem hopen1 compound ARG1ARG2 N enjoyv1 S/N ARG2 NP LHSV mod N mod V/N mod VP mod SynhopeIenjoyV + S Sem hopev1 pronounq BV pron ellipsisref VARG2 S Table 13: Original and modified SHRG rules for Example 13. LHSP ori N ori P ori NP ori SyntocontactwithN + P Sem usedato contactn1 withp PPARG1 N LHSCOMP mod V mod P mod VP mod SyntocontactwithV + NP Semâ contactv1 â VARG2 NP P NP P N someone P with N contact P to (a) Original Syn VP VP NP N someone P with V contact COMP to (b) Modified Syn withp someq udefq person contactn1 usedato ARG2 ARG2ARG1 BV BV (c) Original Sem contactv1 someq person ARG2 BV (d) Modified Sem Fig. 13: Relevant original and modified syntactic and semantic analysis for Example 13. 17 Broadly, we argue that the wide range of diverse form-meaning mappings in ESFL shown above, can be classified into three categories, each with a preferred handling strategy. ⢠Diversity of Grammatical Forms: ESFL demonstrates diversity in grammatical forms it employs to express grammatical meanings. Besides morphological markers, which are typically used by English, formal means like zero form, word order, and functional words, are also utilized by ESFL. This diversity is particularly evident in phenomena such as Number, Case, Tense, Aspect, Voice, and Agreement. To address this, analogous SHRG rules are adapted to account for these forms. ⢠Diversity of Syntactic Derivations: ESFL may impose unique constraints on syntactic derivations, exemplified by phenomena such as Binding, Ellipsis, and Filler- Gap. These phenomena are typically analyzed as outcomes of binding, deletion, or movement within a deep structure in mainstream generative grammar. In our frame- work, they are accommodated by establishing diverse mappings between syntactic configurations and their corresponding semantic interpretations. ⢠Diversity of Lexical Items: ESFL also exhibits significant lexical diversity, par- ticularly in verbs, which vary in both the number of semantic arguments and the ways these arguments are realized. This variation, exemplified by the phenomenon of Argument Structure, can be addressed by adapting or extending existing SHRG rules. It is also noteworthy that ESFL is grounded in English, an inflectional language with multifunctional grammatical markers, so constructions in above examples may overlap across categories. For instance, Example 8 reflects both Number and Person, while Example 12 can be analyzed in terms of either Argument Structure or Ellipsis. Despite such overlaps, we retain this set of phenomena because they are central to the acquisition of morphology, syntax, semantics, and their interfaces âwell-established topics in language acquisiton research. Therefore, covering them highlights not only the effectiveness of our SHRG-based constructivist approach but also the value of the resulting resorce for advancing research in second language acquisition. 4.3 Rebuild For unparsable sentences 2 , we manually select composition rules from the SHRG inventory to generate syntactico-semantic representations. As an initial investigation, we randomly select 100 of these unparsable sentences 3 . The annotation process is conducted in two phases. In the first phase, one annotator constructs semantic graphs by selecting appropriate SHRG rules, while the second anno- tator reviews them for acceptability. In the second phase, both annotators independently annotate the remaining 50 sentences to assess inter-annotator consistency. Annotation results from both phases are presented in Table 14. For the first 50 sentences, we report the proportion of accepted graphs to evaluate agreement between 2 According to Zhao et al. (2020), 47.50% of ESFL sentences, amounting to 2494 in total, are unparsable due to limitations of the ACE parser and inconsistent tokenization. 3 Though limited in size, the sample is methodologically sound, as this rebuilding process mirrors that of most rejected graphs â 81.2% (665 cases) involved syntactic tree reconstruction, according to statistical analysis. 18 annotators. In the second phase, we compute the average S-match score across the 50 independently annotated graphs as an element-wise measure of consistency. Table 14: Consistency scores in selecting rules. First PhaseSecond Phase Consistency Score94%90.93 4.4 Summary As Table 15 presents, following the three steps above, we develop an ESFL Sem- Bank, which contains manually annotated syntactico-semantic representations of 1643 ESFL sentences. The resource is now can be accessed through https://github.com/ MandarinMeaningBank/ESFLSemBank. Table 15: Development of ESFL Sem- Bank at each step. StepMannerNumber 1Manually Accepted724 2Manually Modified819 3Manually Composed100 Overall: 1643 5 A Study on the Linguistic Niche Hypothesis To demonstrate the practical utility of our gold-standard syntactico-semantic resource for ESFL, along with its corresponding SHRG rules, we conduct an empirical study towards the Linguistic Niche Hypothesis (LNH; Lupyan and Dale 2010). The LNH suggests that languages used in broad, communicative contexts, particularly those spoken by large populations of non-native speakers, tend to be more regular than those maintained within smaller, esoteric communities. Our resource, covering both ESFL and standard English, then provides a controlled testbed for evaluating this hypothesis. 5.1 Syntactic Complexity We begin our empirical evaluation by comparing the syntactic structures of ESFL and native English, aiming to determine whether they exhibit statistically significant differences. 19 Specifically, we analyze the distributions of CFG rules derived from SHRG-based derivations in our ESFL and English data, which involves a total of 620 unique rules. To ensure statistical robustness and avoid distortions caused by low-frequency items, we filter out CFG rules with insufficient counts that may introduce noise or violate assumptions of normality. In particular, we retain only those rules with an expected frequency greater than 4, resulting in a focused subset of 77 CFG rules for analysis 4 . Figure 14 displays the frequencies and relative ratios of the 20 most frequently used CFG rules, allowing for an initial visual comparison between ESFL and English syntactic usage. 0100200300400500600700800 N â X V â X N â N + punct AP â X ADV â X P â X NP@N â X NP@N â N + punct VP@V â X ADJ â X ADV â ADV + punct AP â AP + punct AP@ADJ â X DET â X VP â V + VP P â X S â NP + VP VP@V â V + punct VP/N@V â X ESFL English 0.60.81.01.21.41.61.82.02.2 Fig. 14: Frequency Distribution of top 20 CFG Rules and their ratios with 95% confidence intervals in ESFL and English Data 510152025 VP â V + VP S â NP + VP S/P â NP + VP/P VP â V + P S â N + VP VP â VP + VP-C VP/P â V + VP/P@VP S â NP@N + VP N â N + N VP â VP + P VP â ADV + VP VP/P â V + VP/P ROOT â S + S VP â VP@V + P VP â V + AP S/N â NP + VP/N V â V + NP VP â V + VP@V N â DET + N VP/N â V + VP/N ESFL English 0.511.522.53 Fig. 15: Frequency Distribution of top 20 non-lexical CFG Rules and their ratios with 95% confidence intervals in ESFL and English Data Furthermore, since certain CFG rules, those containing only a lexical node or a lexical node combined with a punctuation mark, primarily affect surface realizations 4 According to standard guidelines forÎą= 0.05, a minimum frequency of 4 corresponds to the safe sample size (Cohenâs d = 0.5). 20 and offer limited insight into underlying syntactic structures, we exclude them from our analysis and focus instead on the 43 retained non-lexical CFG rules. Figure 15 presents the frequencies and relative ratios of the top 20 non-lexical CFG rules across both the ESFL and English data. A Chi-Square Test is then performed to determine whether the distribution of remaining CFG rules differs significantly between ESFL and English. The results of this test, as summarized in Table 16, reveal that the syntactic profiles of the two corpora does not diverge in a statistically meaningful way. Table 16: Chi-Square Test of Inde- pendence on non-lexical CFG rule distributions between ESFL and English. Degrees of Freedom Ď 2 p-value 424.2440.999 5.2 Semantic Transparency We then turn to the concept of semantic transparency â the extent to which syntactic structures reliably map onto semantic representations â in both our ESFL and English data. This investigation addresses whether ESFL exhibits a more systematic syntaxâsemantics mapping compared to English. To be more specific, we design a targeted experiment based on semantic derivation stability â we examine the degree to which meaning is preserved when the original SHRG rules are systematically replaced with the most frequent alternatives sharing the same CFG rules (i.e., identical syntactic structures). Semantic graphs are then regener- ated from these substituted rules, and their fidelity to the gold-standard annotations is evaluated using S-match scores. Higher post-substitution S-match scores indicate greater semantic transparency, reflecting a stronger and more consistent coupling between surface-level syntactic choices and underlying semantic interpretations. Figures 16 and Table 17 summarize the experimental results for both datasets. The ESFL dataset consistently achieves higher S-match scores than the English dataset, with a mean of 0.906 compared to 0.875, and exhibits lower variability, as indicated by a smaller standard deviation (0.064 vs. 0.068). The statistical analysis reported in Table 18 further corroborates these findings, demonstrating that ESFL not only attains higher accuracy but also shows greater stability in syntaxâsemantics mappings. From a linguistic perspective, this finding aligns with the hypothesis that ESFL, as a learner variety, tends toward more transparent formâmeaning mappings, possibly as a cognitive adaptation that reduces processing complexity. Such transparency may facilitate both language comprehension and production in a second-language context, and it offers empirical support for theoretical accounts (e.g., the Linguistic Niche Hypothesis) that link communicative environments with structural regularity. 21 ESFL English 0.5 0.6 0.7 0.8 0.9 1.0 S-match score Violin Plot 0.60.70.80.91 0 50 100 150 Frequency Histogram ESFL English Fig. 16: Experimental result on the semantic derivation transparency on ESFL and English data. Table 17: Descriptive statistics of S-match scores for ESFL and English datasets. MeanMedianSDMaxMin ESFL0.9060.9130.06410.668 English0.8750.8760.06810.554 Table 18: Statistical tests com- paring ESFL and English S- match scores. t-test T-statisticsp-value 14.3438.772e-41 z-test z-scorep-value 8.4281.761e-17 6 Conclusion ESFL, used by multilingual speakers whose first language is not English, often departs from standard English through non-canonical lexical and grammatical patterns, creating challenges for systematic semantic derivation. At first glance, these deviations appear to threaten the principle of compositionality â the idea that meaning arises from the meanings of parts and their syntactic combination (Partee 1984). Under this view, ESFL semantics might seem fundamentally unstable, given its variable lexical choices and non-canonical syntax. Yet multilingual speakers and listeners routinely comprehend 22 ESFL with little difficulty, indicating that robust compositional mechanisms remain active beneath surface-level irregularities. In this paper, we reconcile this apparent paradox by integrating constructivist theories with explicit modeling. Our results show that, despite syntactic variability, ESFL exhibits systematic and reliable mappings between form and meaning that support compositional interpretation. These findings suggest that ESFL speakers do not abandon grammatical organization when using a second or foreign language; rather, they adapt and reorganize linguistic resources to maintain communicative efficiency across typologically diverse systems. To support this claim empirically, we introduce ESFL SemBank, a gold-standard syntactico-semantic resource containing 1,643 manually curated ESFL sentences with explicitly aligned syntactic and semantic representations. Given the global scale of multilingualism and the widespread use of ESFL, we argue that both our framework and ESFL SemBank have broad theoretical and practical significance. First, they provide a principled answer to a long-standing question in linguistics: whether ESFL constitutes a systematic linguistic system or merely a collection of deviations from native norms (Adjemian 1976). Our findings strongly support the former view, positioning ESFL as a structured outcome of multilingual competence rather than linguistic deficiency. Second, by offering a computationally grounded model of multilingual meaning construction, our work opens new avenues for research in applied linguistics, language acquisition, and NLP for learner language. As an initial demonstration, we evaluate the Linguistic Niche Hypothesis, illustrating its potential as a scalable testbed for future empirical investigations into multilingual language use and adaptation. References Abend O, Rappoport A. Universal Conceptual Cognitive Annotation (UCCA). In: Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Sofia, Bulgaria: Association for Computational Linguistics; 2013. p. 228â238. https://w.aclweb.org/anthology/P13-1023. Adjemian C. On the nature of interlanguage systems. Language learning. 1976;26(2):297â 320. Banarescu L, Bonial C, Cai S, Georgescu M, Griffitt K, Hermjakob U, et al. Abstract meaning representation for sembanking. In: Proceedings of the 7th linguistic annotation workshop and interoperability with discourse; 2013. p. 178â186. Berzak Y, Kenney J, Spadine C, Wang JX, Lam L, Mori KS, et al. Universal Depen- dencies for Learner English. In: Erk K, Smith NA, editors. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Berlin, Germany: Association for Computational Linguistics; 2016. p. 737â746. https://aclanthology.org/P16-1070. Chomsky N. Lectures on Government and Binding: The Pisa Lectures. Mouton de Gruyter; 1981. 23 Chomsky N. Knowledge of Language: Its Nature, Origin, and Use. Byrne D, K Ěolbel M, editors, Prager; 1986. Copestake A. Slacker Semantics: Why Superficiality, Dependency and Avoidance of Commitment can be the Right Way to Go. In: Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009) Athens, Greece: Association for Computational Linguistics; 2009. p. 1â9. http://w.aclweb.org/anthology/ E09-1001. Corder SP. Error analysis and interlanguage, vol. 198. London: Oxford university Press; 1982. Croft W. Radical construction grammar: Syntactic theory in typological perspective. Oxford University Press; 2001. Davidson S, Yamada A, Fernandez Mira P, Carando A, Sanchez Gutierrez CH, Sagae K. Developing NLP Tools with a New Corpus of Learner Spanish. In: Calzolari N, B Ěechet F, Blache P, Choukri K, Cieri C, Declerck T, et al., editors. Proceedings of the Twelfth Language Resources and Evaluation Conference Marseille, France: European Language Resources Association; 2020. p. 7238â7243. https://aclanthology. org/2020.lrec-1.894. DÄąaz-Negrillo A, Meurers D, Valera S, Wunsch H. Towards interlanguage POS anno- tation for effective learner corpora in SLA and FLT. In: Language Forum, vol. 36; 2010. p. 139â154. Dickinson M, Ragheb M. On Grammaticality in the Syntactic Annotation of Learner Language. In: Meyers A, Rehbein I, Zinsmeister H, editors. Proceedings of the 9th Lin- guistic Annotation Workshop Denver, Colorado, USA: Association for Computational Linguistics; 2015. p. 158â167. https://aclanthology.org/W15-1619. Flickinger D. On building a more efficient grammar by exploiting types. Natural Lan- guage Engineering. 2000;6(1):15â28. https://doi.org/10.1017/S1351324900002370. Flickinger D, Bender EM, Oepen S. Towards an Encyclopedia of Compositional Semantics: Documenting the Interface of the English Resource Grammar. In: Proceed- ings of the Ninth International Conference on Language Resources and Evaluation (LRECâ14) Reykjavik, Iceland: European Language Resources Association (ELRA); 2014. p. 875â881. http://w.lrec-conf.org/proceedings/lrec2014/pdf/562Paper. pdf. Flickinger D, Zhang Y, Kordoni V. DeepBank: A dynamically annotated treebank of the Wall Street Journal. In: Proceedings of the 11th International Workshop on Treebanks and Linguistic Theories; 2012. p. 85â96. Goldberg AE. Constructions: A new theoretical approach to language. Trends in cognitive sciences. 2003;7(5):219â224. 24 Goldberg AE. Argument structure constructions versus lexical rules or derivational verb templates. Mind & Language. 2013;28(4):435â465. Hoffmann T, Trousdale G. The Oxford handbook of construction grammar. Oxford University Press; 2013. Kyle K, Eguchi M, Miller A, Sither T. A Dependency Treebank of Spoken Second Language English. In: Kochmar E, Burstein J, Horbach A, Laarmann-Quante R, Madnani N, Tack A, et al., editors. Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2022) Seattle, Washington: Association for Computational Linguistics; 2022. p. 39â45. https: //aclanthology.org/2022.bea-1.7. Lee J, Li K, Leung H. L1-L2 Parallel Dependency Treebank as Learner Corpus. In: Miyao Y, Sagae K, editors. Proceedings of the 15th International Conference on Parsing Technologies Pisa, Italy: Association for Computational Linguistics; 2017. p. 44â49. https://aclanthology.org/W17-6306. Li W, Wang X, Sun W. Compositional Syntactico-SemBanking for English as a Second or Foreign Language. In: Che W, Nabende J, Shutova E, Pilehvar MT, editors. Findings of the Association for Computational Linguistics: ACL 2025 Vienna, Austria: Association for Computational Linguistics; 2025. p. 24395â24406. https: //aclanthology.org/2025.findings-acl.1252/. Lin Z, Duan Y, Zhao Y, Sun W, Wan X. Semantic Role Labeling for Learner Chinese: the Importance of Syntactic Parsing and L2-L1 Parallel Data. In: Riloff E, Chiang D, Hockenmaier J, Tsujii J, editors. Proceedings of the 2018 Conference on Empir- ical Methods in Natural Language Processing Brussels, Belgium: Association for Computational Linguistics; 2018. p. 3793â3802. https://aclanthology.org/D18-1414. Lupyan G, Dale R. Language structure is partly determined by social structure. PloS one. 2010;5(1):e8559. Nagata R, Sakaguchi K. Phrase Structure Annotation and Parsing for Learner English. In: Erk K, Smith NA, editors. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Berlin, Germany: Association for Computational Linguistics; 2016. p. 1837â1847. https://aclanthology. org/P16-1173. Nagata R, Whittaker E, Sheinman V. Creating a manually error-tagged and shallow- parsed learner corpus. In: Lin D, Matsumoto Y, Mihalcea R, editors. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies Portland, Oregon, USA: Association for Computational Linguistics; 2011. p. 1210â1219. https://aclanthology.org/P11-1121. Nemser W. Language contact and foreign language acquisition. Languages in Contact and Contrast Essays in Contact Linguistics. 1991;p. 345â364. 25 Nivre J, de Marneffe MC, Ginter F, Goldberg Y, HajiËc J, Manning CD, et al. Universal Dependencies v1: A Multilingual Treebank Collection. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LRECâ16) PortoroËz, Slovenia: European Language Resources Association (ELRA); 2016. p. 1659â1666. https://w.aclweb.org/anthology/L16-1262. Oepen S, Lønning JT. Discriminant-Based MRS Banking. In: Proceedings of the Fifth International Conference on Language Resources and Evaluation (LRECâ06) Genoa, Italy: European Language Resources Association (ELRA); 2006. http://w. lrec-conf.org/proceedings/lrec2006/pdf/364pdf.pdf. Partee B. Compositionality. In: Varieties of Formal Semantics: Proceedings of the 4th Amsterdam Colloquium, vol. 3 Foris Dordrecht; 1984. p. 281â311. Pollard C, Sag IA. Head-driven phrase structure grammar. University of Chicago Press; 1994. Ragheb M, Dickinson M. Developing a Corpus of Syntactically-Annotated Learner Language for English. In: Proceedings of the 13th International Workshop on Treebanks and Linguistic Theories T Ěubingen, Germany: University of T Ěubingen; 2014. p. 292â300. https://doi.org/10.5281/zenodo.10054513. Sagae K, Davis E, Lavie A, MacWhinney B, Wintner S. High-accuracy Annotation and Parsing of CHILDES Transcripts. In: Buttery P, Villavicencio A, Korhonen A, editors. Proceedings of the Workshop on Cognitive Aspects of Computational Language Acquisition Prague, Czech Republic: Association for Computational Linguistics; 2007. p. 25â32. https://aclanthology.org/W07-0604. Sagae K, MacWhinney B, Lavie A. Adding Syntactic Annotations to Transcripts of Parent-Child Dialogs. In: Lino MT, Xavier MF, Ferreira F, Costa R, Silva R, editors. Proceedings of the Fourth International Conference on Language Resources and Evaluation (LRECâ04) Lisbon, Portugal: European Language Resources Association (ELRA); 2004. http://w.lrec-conf.org/proceedings/lrec2004/pdf/749.pdf. Santorini B. Part-of-speech tagging guidelines for the penn treebank project (3rd revision). Philadelphia, PA: The University of Pennsylvania; 1990. Selinker L. INTERLANGUAGE. International Review of Applied Linguistics in Language Teaching. 1972;10(1-4):209â232. https://doi.org/10.1515/iral.1972.10.1-4. 209, https://doi.org/doi:10.1515/iral.1972.10.1-4.209. Zhao Y, Sun W, Cao J, Wan X. Semantic Parsing for English as a Second Language. In: Jurafsky D, Chai J, Schluter N, Tetreault J, editors. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics Online: Association for Computational Linguistics; 2020. p. 6783â6794. https://aclanthology.org/2020. acl-main.606. 26