Paper deep dive
Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact
Kiyoshige Garces, Gloria Milena Fernandez-Nieto, Linxuan Zhao, Sachini Samaraweera, Dragan Gasevic, Roberto Martinez-Maldonado, Vanessa Echeverria
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/21/2026, 6:29:07 AM
Summary
This paper investigates the optimization of Large Language Model (LLM) based dialogue coding within healthcare simulations. It specifically examines the trade-offs between coding accuracy, processing time, and environmental impact (energy consumption) by varying prompt designs and batching strategies. Using a dataset of 11,647 utterances across six dialogue constructs (task allocation, handover, sharing information, escalation, questioning, and acknowledging), the study finds that while larger batch sizes improve speed and reduce energy use, they negatively impact coding performance. The research provides practical guidance for deploying scalable, sustainable, and timely dialogue analytics in high-stakes, time-sensitive educational environments like healthcare simulation debriefing.
Entities (12)
Relation Signals (4)
Prompt Design → affects → Coding Performance
confidence 100% · investigates how prompt design and batching strategies can be optimised to balance coding accuracy, processing time, and environmental impact
Large Language Models → usedfor → Dialogue Coding
confidence 100% · LLMs offer new possibilities for automating and enhancing the dialogue layer within emerging multimodal learning analytics approaches
Batching Strategies → improves → Processing Time
confidence 90% · increasing batch size improves speed and reduces energy use, but negatively impacts coding performance.
Batching Strategies → reduces → Environmental Impact
confidence 90% · increasing batch size improves speed and reduces energy use
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Research shows that dialogue, the interactive process through which participants articulate their thinking, plays a central role in constructing shared understanding, coordinating action, and shaping learning outcomes in teams. Analysing dialogue content has been central to advancing team learning theory and informing the design of computer-supported collaborative learning environments, yet this progress has depended on labour-intensive qualitative coding. LLMs offer new possibilities for automating and enhancing the dialogue layer within emerging multimodal learning analytics approaches, with recent studies showing that they can approximate human coding through few-shot prompting. However, prior work has focused on replicating human coding accuracy for research purposes, rather than addressing a more educationally consequential question: how can we design prompts that allow an LLM to label team dialogue accurately and fast enough to be useful in real settings, such as in-person healthcare simulations, where results must be returned quickly and computational cost and sustainability also matter? This paper investigates how prompt design and batching strategies can be optimised to balance coding accuracy, processing time, and environmental impact in team-based healthcare simulation debriefing. Using a dataset of 11,647 utterances coded across 6 dialogue constructs, we compared 4 prompt designs across varying batch sizes, evaluating coding performance, processing time, and energy consumption, as well as the trade-offs between these metrics. Results indicate that increasing batch size improves speed and reduces energy use, but negatively impacts coding performance. Beyond demonstrating the feasibility of LLM-based qualitative analysis, this study offers practical guidance for scaling dialogue analytics in contexts where timeliness, privacy, and sustainability are critical.
Tags
Links
- Source: https://arxiv.org/abs/2604.23255v1
- Canonical: https://arxiv.org/abs/2604.23255v1
Trouble viewing inline? Open PDF directly →
Full Text
78,791 characters extracted from source content.
Expand or collapse full text
Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact Kiyoshige Garces RMIT University Monash University Melbourne, Australia oscar.garcesaparicio@monash.edu Gloria Milena Fernandez-Nieto Monash University Melbourne, Australia gloriamilena.fernandeznieto@monash.edu Linxuan Zhao Adelaide University Adelaide, Australia linxuan.zhao@adelaide.edu.au Sachini Samaraweera Monash University Melbourne, Australia sachini.samaraweera@monash.edu Dragan Gašević Monash University Melbourne, Australia The University of Hong Kong Hong Kong SAR, China dragan.gasevic@monash.edu Roberto Martinez-Maldonado Monash University Melbourne, Australia roberto.martinezmaldonado@monash.edu Vanessa Echeverria RMIT University Melbourne, Australia ESPOL University Guayaquil, Ecuador vanessa.echeverria@rmit.edu.au Abstract Extensive research shows that dialogue, the interactive process through which participants articulate and negotiate their thinking, plays a central role in constructing shared understanding, coordi- nating action, and shaping learning outcomes in teams. Analysing dialogue content has been central to advancing team learning the- ory and informing the design of computer-supported collaborative learning (CSCL) environments, yet this progress has depended on labour-intensive qualitative coding. Large language models (LLMs) offer new possibilities for automating dialogue analysis and for strengthening the dialogue layer within emerging multimodal learn- ing analytics approaches, with recent studies demonstrating that they can approximate human coding through zero- and few-shot prompting. However, most prior work has focused on replicating human coding accuracy for research purposes, rather than address- ing a more educationally consequential question: how can we de- sign prompts that allow an LLM to label team dialogue accurately and fast enough to be useful in real settings, such as in-person healthcare simulations, where results must be returned quickly and computational cost and sustainability also matter? Thus, this paper investigates how prompt design and batching strategies can be optimised to balance coding accuracy, processing time, and envi- ronmental impact in team-based healthcare simulation debriefing. This work is licensed under a Creative Commons Attribution 4.0 International License. L@S ’26, Seoul, Republic of Korea. © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2293-6/2026/06 https://doi.org/10.1145//3774398.3811613 Using a dataset of 11,647 utterances coded across six team learn- ing dialogue constructs (codes), we compared four prompt designs across varying batch sizes, evaluating coding performance, pro- cessing time, and energy consumption, as well as the trade-offs between these metrics. Results indicate that increasing batch size improves speed and reduces energy use, but negatively impacts coding performance. Beyond demonstrating the feasibility of LLM- based qualitative analysis, this study offers practical guidance for scaling dialogue analytics in contexts where timeliness, privacy, and sustainability are critical. CCS Concepts • Applied computing→ Collaborative learning. Keywords Qualitative Coding, Teamwork, Large Language Models, CSCL, Face-to-face, Learning Analytics ACM Reference Format: Kiyoshige Garces, Gloria Milena Fernandez-Nieto, Linxuan Zhao, Sachini Samaraweera, Dragan Gašević, Roberto Martinez-Maldonado, and Vanessa Echeverria. 2026. Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environ- mental Impact. In Proceedings of the Thirteenth ACM Conference on Learning @ Scale (L@S ’26), June 29–July 3, 2026, Seoul, Republic of Korea. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145//3774398.3811613 1 Introduction Dialogue, the interactive process through which participants ar- ticulate and negotiate their thinking, sits at the heart of effective arXiv:2604.23255v1 [cs.HC] 25 Apr 2026 L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. teamwork and small-group learning [18]. Across team science and Computer-Supported Collaborative Learning (CSCL), decades of research show that team dialogue shapes shared understanding, coordination, and learning outcomes [23,50]. Dialogue analysis has been critical for identifying key learning constructs such as explana- tion, questioning, and co-regulation [56], advancing theory and the design of CSCL environments. Yet, this progress has depended on qualitative coding of dialogue: a rigorous but time-consuming and labour-intensive method [26] that limits scale and slows feedback provision to learners or educators [69]. As collaborative learning is embedded in formal education [16,65] and remains essential for developing the collaboration skills required in the workforce [3,46], the CSCL field faces a crucial challenge in efficiently analysing team dialogue at scale and meaningfully supporting learning. Large language models (LLMs) offer new possibilities for au- tomating dialogue analysis [62]. Efforts in collaborative learning are not new; earlier work leveraged machine learning and natural language processing (NLP) techniques to classify and model dia- logue [9,10,60]. Recent research shows that LLMs can approximate human qualitative coding using zero- and few-shot prompting, sub- stantially lowering technical barriers to implementation [33]. How- ever, most prior studies have focused mainly on replicating human coding accuracy under controlled conditions, often overlooking the practical constraints of real-world deployment [33, 61, 67]. In educational settings that are not fully mediated by online platforms – such as physical classrooms [60], specialised learning spaces including laboratories [35,37], and high-fidelity simula- tion rooms [36]– where communication is central to task perfor- mance and learning, automated dialogue analysis must do more than approximate human coders. It must operate in authentic, time- sensitive environments where latency, privacy, computational effi- ciency, and resource use are critical constraints [2,51,57]. Robust and efficient dialogue analysis can also strengthen research on communication processes in emerging areas such as multimodal learning analytics, where verbal interaction is integrated with other data streams such as gesture, gaze, or physiological signals [14,45]. Dialogue-driven analytics can provide timely feedback to support reflective practice [13]. Realising these benefits at scale, however, depends on privacy-preserving deployment, and accounts for the computational and environmental impacts associated with large- scale LLM inference [5,8,20]. Advances in smaller, locally deploy- able LLMs make this shift increasingly feasible [31]. However, they introduce critical design decisions: prompt design, the crafting of instructions that guide LLMs to produce accurate and relevant outputs; batching strategies, which group multiple inputs for simul- taneous processing rather than handling them individually; and model configurations, referring to prompt design and batch settings, all of which can substantially influence coding accuracy, processing time, and energy consumption. These factors create trade-offs that must be purposely engineered rather than assumed [51]. In this work, we examine how LLM prompting designs, includ- ing batching strategies that aggregate multiple inference requests into a single run, can be optimised to balance coding performance, processing time, and environmental impact in high-stakes, time- dependent contexts, such as team-based healthcare simulation. Healthcare simulations are widely used in medical and nursing education, as well as in interdisciplinary hospital training, to nur- ture communication and clinical skills in complex clinical scenarios such as patient deterioration, emergency response, and interpro- fessional handovers [17,49]. These simulations require teams to coordinate, share critical information, anticipate risks, and make rapid decisions—processes that are enacted and observable through dialogue [52]. Decades of research in healthcare teamwork show that communication quality is directly associated with improved patient safety and clinical outcomes [58]. The ability to identify and foster effective teamwork behaviours, such as closed-loop commu- nication, information sharing, clarification, and mutual monitoring, is therefore of both educational and clinical significance [71]. A central pedagogical component of healthcare simulations is the debriefing phase, during which participants reflect on their teamwork and decision-making processes [1]. Automated dialogue coding has the potential to render patterns of communication visi- ble during or immediately after simulation, providing structured evidence to support reflection and targeted feedback [69]. However, these contexts pose distinctive constraints. Simulations are often conducted in secure university facilities or hospital environments where cloud-based services may be restricted due to connectivity limitations, data governance policies, or ethical considerations re- garding clinical scenarios [2,51,57]. Therefore, dialogue-coding tools must operate efficiently under constrained (often on-premises) deployment and produce outputs within the tight temporal bound- aries of the debriefing workflow. While prior work has explored automated dialogue coding in simulation settings using NLP models such as BERT [61,71], these approaches can be difficult to scale, adapt, and transfer across vary- ing clinical scenarios and institutional contexts. The emergence of LLMs offers greater flexibility, but raises new questions about how to balance coding performance with computational efficiency and environmental impact in authentic deployment conditions. Addressing this gap, we systematically investigated how prompt design and batching strategies affect accuracy, processing time, and energy consumption when coding 11,647 utterances across six team dialogue constructs. By explicitly examining these trade-offs, we move beyond feasibility toward actionable design guidance for deploying LLM-based dialogue analytics in high-stakes educational environments, where timely feedback, privacy, and computational and environmental impacts are critical. 2 Background and Related Work 2.1 Automated Coding of Dialogue To support qualitative coding at scale, automated analysis has be- come increasingly important [61]. Early efforts relied on traditional NLP approaches, including supervised classifiers and rule-based systems [9]. Although these methods demonstrated that dialogue could be computationally classified, they required substantial la- belled training data and were tightly coupled to specific datasets and coding schemes, limiting transferability across contexts and theoretical constructs in CSCL. Pre-trained transformer models, such as BERT, advanced this work by reducing dependence on large, task-specific annotated datasets (e.g., [10,69,71]). Fine-tuning improved dialogue coding Scalable LLM-based Coding of Team Dialogue in Healthcare SimulationL@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea. compared to earlier machine learning approaches. Yet, these mod- els still require task-specific annotated data and often struggle with the nuance, ambiguity, and interpretive flexibility inherent in qualitative coding–particularly when codes represent complex collaborative processes or context-dependent constructs [70, 71]. LLMs extend these capabilities through large-scale pre-training and inference-based modelling. Using zero-shot or few-shot prompt- ing, they can perform qualitative coding with minimal additional labelled data, making them practical for contexts where annotated corpora are scarce [19]. Recent studies illustrate this shift. Liu et al., [33] showed that GPT-4 can approximate human coding across prompting designs, though effectiveness varies by construct and context. Wang et al., [60] found that LLaMA, while less precise than BERT, still produced pedagogically meaningful insights for teachers analysing dialogic practice. In large-scale MOOC settings, Mehta et al., [39] demonstrated strong agreement between GPT- assisted coding and human annotations when linking discussion engagement to learning outcomes. Similarly, Li et al., [30] applied GPT-4 and DeepSeek to identify phases of social shared regulation in online discussions, achieving over 83% accuracy and illustrating how LLMs can operationalise theoretically rich CSCL constructs at scale. These studies prove that LLMs can generate contextually appropriate outputs for qualitative dialogue coding tasks [30, 33]. Despite these advances, evaluation of LLM outputs has focused predominantly on agreement with human coders, positioning accu- racy as the primary success indicator. Yet deployment in authentic educational settings requires broader performance considerations [36]. First, processing time is critical [6,27] as latency determines whether coded outputs can support timely reflection or interven- tion. Second, coding performance must be examined in relation to prompt and configuration choices, rather than treated as a static property of the model. Third, prompting designs—including zero- shot and few-shot approaches—shape computational load and envi- ronmental impact (CO 2 e), raising sustainability considerations for large-scale deployment [25]. Yet, despite acknowledged contextual variability in coding effectiveness [33], limited attention has been paid to how prompting configurations jointly affect processing time, coding performance, and environmental impact. These issues are particularly salient in healthcare education, where dialogue unfolds under time pressure and outputs may in- form near real-time debriefing [11]. To our knowledge, only two studies have applied LLM-based coding within authentic health- care simulation. Samaraweera et al. [51] examined prompting tech- niques for a hospital-based simulation operating under strict time constraints, while Tscholl et al. [59] used AI-generated teamwork reports to support debriefing via thematic analysis. Though both demonstrated feasibility, neither systematically examined how prompt- ing and batching configurations jointly influence processing time, coding performance, and environmental impact. 2.2 Energy and Efficiency Trade-offs in LLM Inference LLM inference refers to the serving phase in which a trained model generates outputs in response to input prompts. In educational contexts, inference enables scalable applications such as automated encoding of communication among learners and educators [28,39], feedback generation [4,43], adaptive scaffolding [44,68], and the creation of instructional or assessment materials [24,41]. In this study, coding team dialogue is itself an inference task: the model is repeatedly executed to classify utterances in near real-time. As these systems transition from controlled settings into authentic learning contexts, inference, rather than training, becomes the dominant and recurring source of computational demand. The distinction between model training and model inference matters. Although model training is energy-intensive, recent stud- ies show that the energy consumption of LLM serving has now surpassed that of training, contributing substantially to carbon footprints measured in CO2e [7,12]. Ding and Shi [12] further emphasised the limited understanding of trade-offs between per- formance and carbon in LLM serving, noting that efficiency needs, hardware configurations, and energy use are tightly interdepen- dent. For dialogue analytics deployed at scale, where inference is executed continuously across sessions and cohorts, processing time and environmental impact are not peripheral concerns but key design constraints. Moreover, constraints related to processing time and environ- mental impact during LLM inference are amplified in face-to-face authentic educational environments, where local LLM deployment (e.g., [31]) is often necessary due to privacy, governance, or connec- tivity limitations [34,51]. Parameter-efficient approaches show that LLMs can be adapted for classroom dialogue analysis with reduced fine-tuning overhead [62], yet inference itself remains computation- ally demanding. In multimodal learning analytics scenarios, where dialogue analysis may operate alongside additional data streams (e.g., indoor position or heart rate data), inefficient inference can introduce latency that undermines timely feedback [51,64], par- ticularly during simulation debriefing, where reflection depends on immediacy [1,11]. Accordingly, analysing dialogue in authentic educational environments requires more than acceptable accuracy; it demands deliberate optimisation of efficiency and environmental impact to enable feasible, ethical, and scalable deployment. 2.3 Contribution and Research Questions Situated within CSCL and team learning research, this study ad- vances automated dialogue analysis by reframing performance beyond accuracy and examining how LLM-based coding can be de- sign for authentic, high-stakes learning environments. Specifically, we investigate how prompt design and batching strategies shape the trade-offs between processing time, coding performance, and environmental impact in healthcare simulation, where timely feed- back, local deployment, and sustainability are integral to supporting reflective practice and effective teamwork. By aligning technical optimisation with pedagogical goals, this work contributes design- oriented insights for scaling dialogue analytics in face-to-face col- laborative learning contexts. Motivated by these considerations, we pose the following research questions: •RQ1: How do prompt design and batching strategies affect processing time when using LLMs for team dialogue coding in healthcare simulation? •RQ2: How do prompt design and batching strategies affect coding performance (accuracy) when using LLMs for team dialogue coding in healthcare simulation? L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. •RQ3: To what extent do prompt design and batching strate- gies influence the environmental impact of LLM inference for team dialogue coding in healthcare simulation? We pose a research question that integrates these dimensions: •RQ4: Which prompt design and batching strategies enable a favourable trade-off between coding performance (accuracy), processing time, and environmental impact for team dialogue coding in healthcare simulation? 3 Method 3.1 Learning Context and Learning Design In nursing education, healthcare simulations are used to replicate emergency scenarios in hospital wards. Nursing learners practise teamwork, coordination, and clinical communication within this face-to-face high-fidelity setting. The simulation scenario involves providing and prioritising care to four patients (i.e., manikins), recognising patient deterioration, calling for help, and managing the deterioration. As specified in the learning design, two learners, role-playing as primary nurses, initially participated in the scenario. Two additional learners, role-playing secondary nurses, entered after one primary nurse called for help. The simulation concluded once the deterioration was managed and the situation stabilised. A structured debriefing session follows each scenario to guide learners’ reflection on their teamwork and clinical decision-making. While the debriefing phase highlights the pedagogical need for timely, actionable dialogue analytics, the present study focuses exclusively on team dialogue collected during the scenario itself. The analysis of debriefing conversations falls beyond the scope of this paper. Thus, this learning setting reflects realistic constraints related to time, resources, and staff availability, commonly encoun- tered in simulation-based healthcare education. These operational conditions frame the design considerations examined in this study. 3.2 Dataset With ethics approval from Monash University Human Research Ethics Committee, all students provided consent for their data to be collected and analysed for this research. Students, aged 20 to 23, were organised into teams of four members. The dataset is formed from audio transcripts of 35 teams: 15 in 2021 and 25 in 2022. Audio from each team member was recorded using Shure PGA31 headset microphones connected via BLX/SVX 4 wireless transmit- ters to a TASCAM US-16x08 audio interface, with all channels captured on a local server. Utterance timing for each student was extracted using a Python script and a voice activity detection library 1 . Transcriptions were automatically generated with the Whisper- large model [47]. 811 utterances were manually transcribed (total duration = 5,078.01 seconds) to assess the quality of the transcrip- tion. Based on these manual transcriptions, the word error rate (WER), a widely accepted metric for transcription quality [47,54], of automatic transcriptions was 0.089. This result indicated the transcription quality was sufficient [54]. The resulting transcripts from the 35 teams were segmented into 11,647 utterances. The average number of utterances per simula- tion was 332.8 with a standard deviation of 93.9. Utterances were 1 webrtcVAD: https://github.com/wiseman/py-webrtcvad Table 1: Coding used for team dialogue analysis, with defini- tions and examples. Coding SchemeDefinition and Example task allocationA healthcare student self-assigns or assigns to another healthcare student one or more tasks. Example: "Can you start the IV line and prepare the fluids?" handoverWhen a healthcare student provides structural communication of the medical situation to the new team members. Example: "We have Jack here. He’s just been complaining of six out of ten chest pain. " sharing informationA healthcare student shares relevant clinical or situational information with one or more other healthcare professionals. Example: "His oxygen saturation has dropped to 88%." escalation A healthcare student communicates that the sit- uation exceeds their capacity and requests or suggests additional assistance, including initiat- ing or proposing a formal call for help. Example: "This is deteriorating quickly, we need to call the resuscitation team." questioningA healthcare student asks another team member to obtain information. Example: "Have the blood results come back yet?" acknowledgingA healthcare student signals receipt or recog- nition of another professional’s statement, re- quest, or presence. This is a passive action and does not necessarily indicate agreement or dis- agreement, but demonstrates awareness or un- derstanding. Example: "Okay," or "I hear you." broken within each dialogue segment into multiple sentences us- ing punctuation marks, i.e., including full stops, commas, question marks, and exclamation marks. For the dialogue analysis, we used the coding scheme described in Table 1, which was derived from team science and collaboration frameworks [40,48]. These codes were further refined with healthcare educators [69]. The detailed theoretical grounding of the coding scheme is provided in [70]. The utterances were then coded at the sentence level, with each sentence assigned to a single code, thus multiple codes can appear within a single utterance. Two researchers independently coded 20% of utterances collected in 2021 (756 utterances from 3 simulation sessions), using the definitions and examples provided in Table 1. Cohen’s Kappa was used as the metric to assess the inter-rater reli- ability [38]. A threshold of 0.61 was chosen to indicate a substantial agreement between the raters [29]. For each code, a Kappa greater than 0.7 was reached. Then, one researcher coded the remaining utterances in 2021 and 2022 datasets. Using the 11,647 coded utterances, we filtered out those that were not assigned any communication code because they were either outside the context of the simulation (e.g., prior to or after the simulation) or contained no relevant information for the simulation Scalable LLM-based Coding of Team Dialogue in Healthcare SimulationL@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea. task. The final dataset comprised 4,057 utterances, which were used to address the research questions. 3.3 Experimental Design For the experimental design, we created a set of prompt designs and batching strategies 2 . 3.3.1Prompt Design and Batching Strategies. We created four prompt designs following a few-shot classification approach (i.e., Liu et al. [33]). P.1: Prompt few-shot, consisted of a contextual descrip- tion specifying the model’s role, definitions of the six communi- cation codes plus three annotated examples per code, and task instructions framing the problem as a multi-label classification task. Each code was represented as a one-hot vector (i.e., 1 indicating the presence of a label). The model’s output consisted of seven binary values (0 or 1) and was ordered according to the predefined codes. P.2: Prompt few-shot with rules, extended the first by in- troducing explicit decision rules to support the coding of more complex dialogue codes [53,63]. In addition to code definitions and examples, this prompt included a set of decision rules. For example: If an utterance contains a first-person commitment (e.g., "I will...", "Let me..."), then code it astask allocation; if the utterance also reports patient status, additionally code it as sharing information. P.3: Prompt few-shot with rules and context, incorporated preceding utterances as conversational context, given that mod- elling dialogue history has been shown to enhance robustness in conversational tasks [21]. Specifically, this design added a confidence- based rule (If uncertainty is below 95%, leave labels as 0) and modified the task instructions to require the model to consider the three pre- ceding utterances when coding each target utterance. P.4: Prompt few-shot with rules and structured metadata, extended P.2 by explicitly including contextual metadata, such as the initiator and receiver of each utterance (i.e., primary nurse 1, secondary nurse 1), to provide additional conversational structure. In addition to prompt design, the current study investigated how batching strategies varied with the number of utterances processed simultaneously. Batching strategies are often used to analyse com- putational efficiency; for LLM inference, grouping multiple inputs into a single batch reduces repeated prompt parsing and token efficiency and can lower total inference time, thereby increasing scalability [6,27]. Batch sizes of 1, 10, 20, 30, 40, 50, 60, and 70 utter- ances were evaluated. These batch sizes were heuristically selected to systematically explore the effect of increasing the number of utterances processed per batch in increments of ten. In total, this resulted in 32 configurations (4 prompt designs× 8 batch sizes). 3.4 LLM Implementation A local deployment of an LLM was configured. All model inference was conducted using the Ollama 3 interface with the Deepseek- R1[22] model, more specifically the distilled model with 14 Bil- lion parameters,deepseek-r1:14b 4 . Ollama was used to facilitate model deployment and provide an application interface (API) for accessing the DeepSeek model. The model temperature was set 2 Prompt and batching details can be found in https://osf.io/gtx86/overview?view_only=ffdb65822459434e8f9344d8a263fb5a 3 https://w.ollama.com 4 https://ollama.com/library/deepseek-r1:14b to 0, resulting in deterministic outputs; therefore, repeated sto- chastic sampling was not required. To verify deterministic consis- tency across runs, each configuration was executed five times with batches of 20 utterances, yielding identical outputs in all cases. To further support reproducibility, a fixed random seed of 42 was used. All evaluations were conducted on the same machine, a server equipped with 128GB of 4,400MHz ECC DDR5 RAM, an NVIDIA RTX A5500 GPU with 24GB of memory, and two Intel Xeon Gold 5415+ processors (2.9GHz base frequency, 8 cores each, 22.5MB L3 cache, hyper-threading enabled; maximum turbo frequency 4.1GHz). We record the timestamps for each of the experimen- tal configurations. Running all required approximately 54 hours and 41 minutes of wall-clock time. 3.5 Analysis To address RQ1–RQ3, we analysed 4,057 utterances together with the researcher-coded ground truth to evaluate coding performance, processing time, and environmental impact. 3.5.1 RQ1 – Processing Time. We measured the total processing time required to code the complete set of 4,057 utterances under each configuration. Processing time was analysed by estimating the time required to code the utterances from a single simulation session. Specifically, the total elapsed time was divided by the num- ber of simulations in the dataset (푛=35). We analysed the total elapsed time across the 32 configurations (4 prompts×8 batch sizes) and conducted Spearman’s rank correlation analyses to identify relationships among batch size, prompt design, and per-session processing time. Additionally, we visually inspected changes in processing time across prompt designs and batch sizes to identify configurations that support feasible timely deployments. 3.5.2 RQ2 – Coding Performance. We evaluated coding perfor- mance using standard metrics, including macro-averaged F1 score, precision, recall, and accuracy [e.g.,33]. Performance was computed at the utterance level, resulting in a multi-label classification evalu- ation. F1 score was used as the primary performance metric due to its robustness to class imbalance in our dataset (see Appendix). These metrics were used to analyse the relationship between batch size, prompt design, and performance, and to determine whether variations in batch size and prompt design affected F1 scores. 3.5.3RQ3 – Environmental Impact. We measured the total energy consumed while processing the full dataset for each configuration. Energy measurements included CPU, GPU, and DRAM usage, which were aggregated to obtain total system energy consumption. Energy usage was recorded using the Zeus Python library [66] and reported in joules 5 . Rather than estimating environmental impact using carbon dioxide equivalent (CO 2 e) calculators, we report energy consumption directly as a proxy for environmental impact. This choice is motivated by the fact that carbon and water consumption calculators are derived from energy usage [15]. Furthermore, all experiments were conducted on local deployments without active water-based cooling. 5 A joule is an International System unit of energy, equal to applying a force of one newton over one metre, or roughly the energy to lift a 100 apple by 1 metre. L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. 3.5.4RQ4 – Trade-offs. We conducted a trade-off analysis based on Pareto optimality [42,55]. Pareto analysis is a widely used approach for evaluating trade-offs among conflicting objectives, where im- proving one objective may degrade others. It provides a systematic method for identifying configurations that achieve optimal trade- offs across multiple criteria. This analysis considered three objectives: (i) coding performance, measured by the macro-averaged F1 score; (i) total GPU energy consumption; and (i) total processing time. Before identifying Pareto-optimal configurations, we conducted a Spearman’s rank correlation analysis between GPU energy consumption and total processing time to examine the relationship between these dimen- sions. Pareto-optimal configurations were identified by maximising the F1 score while simultaneously minimising energy consumption and processing time. The resulting Pareto-optimal set comprises non-dominated configurations, for which no alternative configura- tion achieves better performance on one objective without degrad- ing at least one of the others. This set is commonly referred to as the Pareto front [42]. We filtered out configurations infeasible for our near real-time scenario. We further analysed per-label results within the Pareto front to identify configurations that provide a balanced trade-off across all three objectives. 4 Results 4.1 RQ1 – Processing Time The processing times (seconds) required to code a single simulation session are presented in Table 2. The results for the entire set of utterances are provided in the Appendix. Table 2 and Fig. 1a shows a clear trend, as batch size increases, total processing time decreases across all prompt designs. Notably, for P.1: Prompt few-shot, a batch size of 1 required approximately four times more processing time than a batch size of 10. For P4: Prompt few-shot with rules and structured metadata, this difference increased to approximately ten times. Similarly, P.2: Prompt few-shot with rules and P3: Prompt few- shot with rules and context required approximately fourteen times more processing time at batch size 1 compared to batch size 10. As shown in Fig. 1b, processing time began to stabilise beyond a batch size of 20 for most prompt designs, with the exception of P4. At batch size 70, P4 achieved the lowest processing time, requiring approximately half the time (0.55×) compared to the other prompts. For the remaining prompts, increasing the batch size beyond 20 results in relatively small reductions, ranging from 3 to 8 seconds. Spearman’s rank correlation between batch size and processing time was−0.895, indicating a strong negative association: larger batch sizes were consistently linked to shorter processing times. In contrast, configurations with batch size 1 (i.e., no batching) were impractical for near real-time deployment, as they required more than 13 minutes per session. Overall, these findings suggest that processing time was substan- tially affected by batching strategies. In this context, no batching is impractical because the debriefing starts immediately after the simulation ends. A batch size of 20-70 may be optimal at this point. 4.2 RQ2 – Coding Performance The results evaluating coding performance when considering prompt design and batch size (macro-averaged F1 score, precision, and recall) are reported in Table 3. These results indicate that higher performance was generally achieved at lower batch sizes. As shown in Fig. 2, there was a strong negative association between batch size and F1 score, with Spearman’s rank correlation of−0.882, indicat- ing that performance consistently declined as batch size increased, suggesting that higher batch sizes did not necessarily improve the model’s accuracy. Regarding prompt strategies, configurations with explicit deci- sion rules consistently outperformed the Prompt few-shot baseline. This finding suggests that structured decision rules enhanced the model’s ability to differentiate dialogue codes, particularly at batch sizes below 40. Given the multi-label nature of the task, we further report per- label F1 scores in Tables 4 and 5. At the code level, batch size continued to influence performance. Prompts that included decision rules generally outperformed the P.1: Prompt few-shot across most dialogue codes. In particular, P4: Prompt few-shot with rules and structured metadata achieved the strongest performance in three of Table 2: Processing Time in seconds across prompt designs and batch sizes. Highlighted values indicate the minimum value per prompt and metric, while values in bold denote the overall maximum for each metric. P.1: Prompt few-shot, P.2: Prompt few-shot with rules, P.3: Prompt few-shot with rules and context, P.4: Prompt few-shot with rules and structured metadata. Batch Prompt P.1P.2P.3P.4 1781.032,086.902,857.921,176.21 10181.23151.20198.20149.26 20125.69122.39133.96131.16 30117.20114.72128.14125.38 40117.46136.72120.66115.35 50118.66113.49121.8995.01 60113.27 108.03115.2566.17 70108.31115.35109.3161.28 Table 3: Coding performance across prompt designs. High- lighted values indicate the maximum value per prompt for each metric, while bolded values denote the overall maxi- mum for each metric. Batch F1 MacroPrecision MacroRecall Macro P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 10.610.600.610.63 0.63 0.67 0.670.680.610.59 0.610.63 100.560.600.61 0.600.640.680.680.68 0.520.590.62 0.57 200.51 0.600.61 0.59 0.610.68 0.67 0.66 0.45 0.58 0.60 0.55 300.50 0.58 0.60 0.60 0.63 0.66 0.64 0.67 0.43 0.54 0.59 0.56 400.50 0.49 0.60 0.56 0.63 0.65 0.65 0.65 0.42 0.42 0.58 0.50 500.52 0.40 0.55 0.37 0.63 0.54 0.64 0.53 0.46 0.33 0.50 0.30 600.45 0.34 0.46 0.30 0.58 0.59 0.65 0.45 0.39 0.25 0.37 0.24 700.39 0.19 0.37 0.19 0.55 0.51 0.61 0.33 0.32 0.12 0.28 0.14 Scalable LLM-based Coding of Team Dialogue in Healthcare SimulationL@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea. 010203040506070 Batch size 0 500 1000 1500 2000 2500 Processing Time (s) P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata (a) Processing Time against Batch sizes. 10203040506070 Batch size 60 80 100 120 140 160 180 200 Processing Time (s) P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata (b) Processing Time against Batch sizes for Batch ≥ 10 Figure 1: Processing Time in Seconds across prompts designs and batch sizes. 010203040506070 Batch size 0.2 0.3 0.4 0.5 0.6 F1 (macro) P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata Figure 2: Macro-averaged F1 scores and batch sizes illustrat- ing coding performance. the six codes—acknowledging, handover, and sharing information, although it showed greater instability at larger batch sizes. Beyond peak performance values, as shown in Fig. 2, we observed that prompts incorporating explicit decision rules (P.2, P.3 and P.4) demonstrate greater robustness to increasing batch sizes compared to the baseline few-shot prompt. While performance declined across all configurations as batch size increased, rule-based prompts exhib- ited slower degradation, particularly for acknowledging and escala- tion codes. Furthermore, handover remained consistently the most challenging label across all prompt designs and batch sizes, with substantially lower F1 scores compared to other codes, suggesting difficulty in modelling this dialogue code. Notably, moderate batch sizes (10-20) often yielded better performance across multiple codes, suggesting that such sizes may support contextual grouping before attention dilution affects classification accuracy [72]. In sum, these findings suggest that coding performance, as mea- sured by F1 score, is strongly influenced by batching strategies, with smaller batch sizes (푛=1-10) consistently supporting more accurate classification. Table 4: Per-label F1 scores across prompt designs and batch sizes. Batch acknowledgingescalationhandover P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 10.730.780.790.820.66 0.68 0.650.67 0.24 0.20 0.19 0.19 100.64 0.74 0.77 0.72 0.60 0.63 0.64 0.660.27 0.16 0.16 0.18 200.58 0.72 0.75 0.71 0.480.70 0.65 0.63 0.23 0.15 0.17 0.19 300.58 0.70 0.73 0.70 0.50 0.55 0.64 0.67 0.240.22 0.180.26 400.58 0.62 0.73 0.67 0.52 0.480.66 0.57 0.25 0.110.20 0.22 500.62 0.53 0.68 0.49 0.56 0.43 0.57 0.27 0.26 0.04 0.13 0.16 600.51 0.45 0.59 0.47 0.53 0.39 0.45 0.28 0.18 0.040.20 0.23 700.45 0.24 0.47 0.34 0.41 0.19 0.39 0.21 0.18 0.06 0.13 0.01 Table 5: Per-label F1 scores across prompt designs and batch sizes. Batch questioningsharing info.task allocation P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 P.1 P.2 P.3 P.4 10.85 0.80 0.810.840.55 0.53 0.560.590.63 0.58 0.670.66 100.760.820.83 0.81 0.520.570.58 0.57 0.570.690.69 0.65 200.71 0.80 0.82 0.79 0.510.570.58 0.56 0.54 0.660.69 0.64 300.71 0.80 0.80 0.78 0.510.570.58 0.56 0.48 0.62 0.66 0.61 400.70 0.73 0.79 0.76 0.48 0.50 0.55 0.54 0.48 0.49 0.65 0.58 500.73 0.62 0.77 0.54 0.50 0.44 0.53 0.41 0.42 0.31 0.59 0.37 600.65 0.52 0.61 0.37 0.47 0.38 0.43 0.30 0.38 0.25 0.45 0.14 700.55 0.30 0.53 0.22 0.41 0.20 0.37 0.27 0.35 0.11 0.35 0.09 4.3 RQ3 – Environmental Impact To address the environmental impact of prompt design and batching strategies, we analysed the effect of these configurations on energy consumption (Table 6). As expected, batch size strongly influenced energy use, closely mirroring its relationship with processing time. In general, larger batch sizes led to lower total energy consumption due to reduced computational overhead per utterance. Across prompt designs, the Prompt few-shot with rules and con- text configuration exhibited the highest energy consumption at batch size 1, consistent with its longer processing time. In contrast, the Prompt few-shot with rules showed energy consumption com- parable to the Prompt few-shot baseline across most batch sizes. L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. 010203040506070 Batch size 0 100 200 300 400 500 600 Total System Energy (J) P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata (a) Energy Consumption against Batch sizes. 10203040506070 Batch size 15 20 25 30 35 40 45 Total System Energy (J) P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata (b) Energy Consumption against Batch sizes for Batch ≥ 10 Figure 3: Energy Consumption in Joules across prompts designs and batch sizes. Table 6: GPU Energy consumption in Joules per simulation across prompt designs and batch sizes Batch Prompt P.1P.2P.3P.4 1177.46474.85652.46267.85 1041.4034.4945.3034.07 2028.7127.9730.6230.00 3026.8026.2329.3128.67 4026.8731.2827.6026.39 5027.1425.9527.8821.72 6025.90 24.7126.3715.12 7024.5926.4025.0014.01 These findings indicate that incorporating explicit decision rules did not substantially increase energy costs when combined with appropriate batching strategies. Comparing the results for Processing Time (Fig. 1) and Energy Consumption (Fig. 3), both showed highly similar patterns. Results from Spearman’s rank correlation examining the relationship be- tween energy consumption and processing time showed an almost perfect positive correlation (Spearman’s rank correlation coefficient of 0.99945), confirming that energy consumption scaled nearly pro- portionally with processing time. 4.4 RQ4 - Trade-offs Given the near-perfect correlation between processing time and energy consumption (Spearman’s rank 0.99945), we reduced the optimisation problem to two objectives: coding performance and processing time. In this context, minimising processing time effec- tively minimises energy consumption. The Pareto front is composed of 11 configurations. Among then, P4: Prompt few-shot with rules and structured metadata with batch size 1, which achieves the highest F1 score. Therefore, if coding performance alone is prioritised, this configuration would be se- lected. Nevertheless, due to its substantially higher processing time (1,176.21 seconds) and energy consumption (267.85 Joules), as re- ported in Sections 4.1 and 4.3, respectively, it is unsuitable for near real-time healthcare debriefing contexts. Excluding this configu- ration based on feasiblity grounds, we analysed the remaining 10 configurations that form the Pareto front, as shown in Fig. 4 high- lighted by black circles: P.1: Prompt few-shot with batch sizes 60 and 70; two configurations using P.2: Prompt few-shot with rules with batch sizes 20 and 30; three configurations using P3: Prompt few-shot with rules and context with batch sizes 10, 20 and 30; finally, three configurations using P4: Prompt few-shot with rules and structured metadata with batch sizes 50, 60, 70. The Pareto analysis revealed distinct deployment preferences. When prioritising computational efficiency, the most suitable con- figurations were P4: Prompt few-shot with rules and structured meta- data with batch sizes of 50, 60, and 70. When seeking a balanced trade-off between coding performance and processing time, batch sizes of 20 and 30 using rule-based prompts (P.2 and P.3) provide the most promising midpoint. Given the time-sensitive nature of healthcare simulation debrief- ings, balanced configurations may be the most appropriate. Among the Pareto-optimal set, the per-label analysis (Table 7) showed that P3: Prompt few-shot with rules and context with batch size 20 achieved the highest overall label-level performance while main- taining processing time. This configuration, therefore, represents the most favourable trade-off across performance, processing time, and environmental impact. Table 7: Per-label F1 scores for Pareto-optimal across prompt versions. Code P.2P.3 Batch 20 Batch 30 Batch 20 Batch 30 acknowledging0.720.700.750.73 escalation0.700.550.650.64 handover0.150.220.170.18 questioning0.800.800.820.80 sharing information0.570.570.580.58 task allocation0.660.620.690.66 Scalable LLM-based Coding of Team Dialogue in Healthcare SimulationL@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea. 6080100120140160180200 Total Time (s) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 F1 (Macro) Batch-60 Batch-70 Batch-20 Batch-30 Batch-10 Batch-20 Batch-50 Batch-60 Batch-70 P.1: Prompt few-shot P.2: Prompt few-shot with decision rules P.3: Prompt few-shot with decision rules and context P.4: Prompt few-shot with decision rules and structured metadata Pareto-optimal Figure 4: Pareto front. The highlighted points depict the Pareto front for the two dimensions to optimise: Processing Time and Macro-averaged F1 score. 5 Discussion 5.1 Findings by Research Question With respect to RQ1, which examined whether processing time is affected by prompt design and batching strategies when using LLMs for near real-time dialogue coding in healthcare simulation, our findings showed that both factors substantially influenced pro- cessing time, and suggested a batch size range of푛=20-70. In particular, prompts incorporating additional context and ex- plicit decision rules, when combined with batching strategies, led to marked reductions in total processing time. Increasing the number of utterances processed per batch resulted in considerable time savings, a valuable when integrating dialogue analysis into time-sensitive learning activities, where timely feedback is critical for reflective learning, such as post-simulation debriefs [1,11,34]. Our results corroborate previous work on batch prompting [6,27], confirming that increasing batch size reduces processing time. Ex- tending this line of work, our results show that similar efficiency gains can be achieved when analysing team dialogue in complex, high-stakes contexts such as healthcare simulations. Addressing RQ2, we observed that increasing batch size had a noticeable impact on coding performance. Batching improved efficiency, but performance (macro-averaged F1) generally decreased as the number of utterances processed simulta- neously increased, mirroring results from Cheng et al., [6]. This reveals an inherent trade-off between efficiency and accuracy in near real-time dialogue coding contexts. Among the dialogue codes, handover consistently performed worse than the other codes across prompt designs and batch sizes. Results for handover detection are consistent with prior BERT-based findings [70], likely reflecting the deeper conversational context required and conceptual overlap with sharing information [70]. Thus, LLMs still struggle to resolve construct-level overlaps—mirroring inconsistencies among human annotators, and echoing educator reports that handover may hold limited pedagogical value during debriefing given its overlap with adjacent codes [69]. These observations reinforce that coding per- formance is sensitive to construct interpretation [33], suggesting that refining coding schemes may improve both model performance and practical relevance—complementing evidence that prompting and batching strategies materially shape LLM coding behaviour [33,51]. Our findings also align with broader claims that LLMs struggle with long-context reasoning, particularly when relevant information is not positioned near the beginning or end of the prompt, suggesting that long, batched prompts may dilute salient signals [32]. Taken together, these findings suggest that implement- ing LLMs for team dialogue coding is plausible even in complex, constrained contexts such as healthcare simulation environments. Addressing RQ3, we found that energy consumption also de- creased as batch size increased, a trend similar to that in RQ1. By empirically demonstrating a near-perfect correlation between pro- cessing time and energy consumption, we provide evidence that the batching strategy is a direct driver of environmental im- pact in LLM-based dialogue coding. This finding establishes that processing time optimisation directly reduces energy demand and associated environmental impact [5,12,15]. Our results posi- tion batching not merely as a performance optimisation technique, but as a practical sustainability lever in LLM coding pipelines [25]. In doing so, we contribute quantitative evidence that inference energy should not be overlooked and that deployment-level opti- misation can meaningfully reduce the environmental foot print [7]. Furthermore, this study extends the current literature by explicitly quantifying how both prompt design and batching strategies influ- ence environmental impact. We provide empirical justification for using energy consumption as a proxy for environmental impact in local deployment settings, where computational resources are di- rectly tied to institutional infrastructure. To our knowledge, this is the first study to document these relationships within collaborative L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. learning contexts, where local deployment is often necessary due to privacy and infrastructural constraints [2, 57]. Collaborative learning contexts, such as healthcare simulations, are time-constrained and highly sensitive to coding performance, as debriefings often begin immediately after simulation sessions [34]. To address this tension, RQ4 examined which configurations offered a favourable balance between performance, processing time, and environmental impact. Our results indicate that balanced con- figurations, specifically those using prompts with explicit deci- sion rules, additional conversational context, and moderate batch sizes (approximately 20 utterances), provide the most suitable midpoint. These configurations maintain strong coding per- formance while substantially reducing processing time and energy consumption, thereby meeting the practical demands of pedagogi- cal instruction, such as debriefing. Our key contribution is to make these trade-offs explicit and empirically quantifiable: prompt de- sign and batching decisions must balance accuracy, efficiency, and sustainability, rather than optimising for performance alone. This study extends prior research on LLM-based automated cod- ing, such as Liu et al., [33], which primarily assessed prompt engi- neering strategies in terms of coding performance (accuracy). Their findings showed that few-shot prompting improved performance, consistent with our results, which also showed that combining few-shot prompts with explicit decision rules led to accuracy im- provements. In contrast, our work incorporates processing time and environmental impact as additional evaluation dimensions and proposes a framework to balance these trade-offs, particularly in local deployment settings critical to collaborative learning contexts where immediate analysis is required. Our findings also extend prior feasibility studies of LLM use in in-situ simulation settings [51], which primarily compared prompt- ing strategies based on inference performance. While Samaraweera et al., [51] showed that few-shot prompting improved coding perfor- mance and that selectively adding contextual information benefited certain codes accuracy, it also reported substantial increases in envi- ronmental impact. We build on this work by evaluating prompting and batching under a unified trade-off framework that integrates performance, processing time, and environmental impact. Related work on batch prompting has focused on grouping infer- ence tasks to reduce prompt length. These works were not evaluated in terms of dialogue coding or classification tasks; moreover, they were beyond the scope of healthcare simulation scenarios, nor were they systematically evaluated for multiple conflicting objectives such as accuracy, processing time, and environmental impact. 5.2 Implications for Learning at Scale This study presents methodological and practical implications for deploying scalable LLM-based dialogue coding in authentic educa- tional settings. Although contextualised in healthcare education, our results apply to other collaborative learning contexts where dialogue is central for learning [3, 46]. From a methodological standpoint, this study frames scalable LLM-based dialogue coding in authentic educational settings as a balance between processing time, performance, and environmental impact. Moving beyond prior work that has primarily reported coding accuracy [30,33,39,60], we show that automated dialogue coding should be engineered and evaluated by considering these three metrics for learning-at-scale analytics. This methodological lens guides configuration choices so they align with the learning activity, deployment constraints, and environmental costs [2,36, 51,57]. In doing so, it complements research aimed at supporting learning at scale via LLM-driven coding and learner-facing support [24,28,39,41,44] by providing a unified framework for analysing and reporting deployment trade-offs. From a practical standpoint, the findings offer guidance for de- ploying dialogue coding in face-to-face CSCL contexts [35,37,60] and healthcare simulation-based education [14,36]. Our results indi- cate that LLM-based analytics can be integrated into time-sensitive educational workflows (e.g., post-simulation debriefings), support- ing scalable use of dialogue data to inform reflective discussions and learning. The average processing time was 118.38 seconds (SD=25.12) per session for batch sizes푛>1, indicating that the coding can be generated within the timeframe of typical simulation debriefing workflows. Researchers and developers of AI-supported feedback may combine few-shot prompts with additional decision rules, and select batch sizes based on deployment priorities (e.g., coding performance, processing time, and environmental impact). If coding performance alone is priority, a batch size of 1 is prefer- able. If the focus is on processing time, larger batch sizes should be selected. For a more balanced trade-off between accuracy and pro- cessing time, batch sizes of푛=20–30 are recommended. Together, they support scalable and context-sensitive integration into collab- orative learning practice and contribute to ongoing discussions in CSCL and teamwork science [23, 50]. 5.3 Limitations and Future Work This study evaluated four prompt designs combined with batching strategies that process multiple utterances simultaneously. How- ever, alternative batching approaches were not explored. Future work could investigate distributing classification task across multi- ple specialised prompts, each identifying a specific dialogue code while processing single utterances. Such task decomposition strate- gies may yield different trade-offs between performance, efficiency, and energy consumption. Additionally, the prompts relied on ex- tended few-shot examples and contextual information, resulting in long prompt structures. While effective, these prompts may not be optimally efficient. Future research should explore prompt opti- misation techniques, such as structured formats, compressed rep- resentations, or schema-based inputs, to reduce token length and improve both computational efficiency and environmental impact. 5.4 Concluding Remarks This study researched how prompt design and batching strategies affect coding performance, processing time, and environmental impact in healthcare simulation dialogue coding. Results show a clear trade-off between efficiency and performance, with rule-based prompts and moderate batch sizes offering the most favourable balance. These findings contribute practical guidance for deploying LLMs in time-sensitive educational settings and highlight the need to optimise configurations for accuracy and sustainability. Future work should explore alternative task decomposition strategies and prompt optimisation techniques to further improve scalability. Scalable LLM-based Coding of Team Dialogue in Healthcare SimulationL@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea. References [1]Abulebda, K., Auerbach, M., Limaiem, F.: Debriefing techniques utilized in medical simulation. In: StatPearls. StatPearls Publishing, Treasure Island, FL (2025) [2]Algarni, A.M., Thayananthan, V.: Digital health: The cybersecurity for ai-based healthcare communication. IEEE Access 13, 5858–5870 (2025). https://doi.org/10.1109/ACCESS.2025.3526666 [3] An, M., Teffera, L., Mehrvarz, M., Li, B., Bogart, C., Sakr, M., M. McLaren, B.: Lever- aging intelligent tutoring systems to enhance project-based learning in work- force training at community colleges. In: Ferreira Mello, R., Rummel, N., Jivet, I., Pishtari, G., Ruipérez Valiente, J.A. (eds.) Technology Enhanced Learning for In- clusive and Equitable Quality Education. p. 57–62. Springer Nature Switzerland, Cham (2024) [4]Barno, E., Albaladejo-González, M., Reich, J.: Scaling generated feedback for novice teachers by sustaining teacher educators’ expertise: A design to train llms with teacher educator endorsement of generated feedback. In: Proceedings of the Eleventh ACM Conference on Learning @ Scale. p. 412–416. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3664677 [5]Berthelot, A., Caron, E., Jay, M., Lefèvre, L.: Understanding the environmen- tal impact of generative ai services. Commun. ACM 68(7), 46–53 (Jun 2025). https://doi.org/10.1145/3725984 [6]Cheng, Z., Kasai, J., Yu, T.: Batch Prompting: Efficient Inference with Large Language Model APIs. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track. p. 792–810. Association for Computational Linguistics, Singapore (2023). https://doi.org/10.18653/v1/2023.emnlp-industry.74 [7]Chien, A.A., Lin, L., Nguyen, H., Rao, V., Sharma, T., Wijayawardana, R.: Reducing the carbon impact of generative ai inference (today and in 2035). In: Proceedings of the 2nd Workshop on Sustainable Computer Systems. HotCarbon ’23, ACM, New York, NY, USA (2023). https://doi.org/10.1145/3604930.3605705 [8] Choi, S.P.M., Lam, S.S., Li, K.C., Wong, B.T.M.: Learning analytics at low cost: At- risk student prediction with clicker data and systematic proactive interventions. Educational Technology & Society 21(2), 273–290 (2018) [9]Crowston, K., Allen, E.E., Heckman, R.: Using natural language processing tech- nology for qualitative data analysis. International Journal of Social Research Methodology 15(6), 523–543 (2012) [10] Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirec- tional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). p. 4171–4186. Association for Computational Linguistics, Minneapolis, Minnesota (Jun 2019). https://doi.org/10.18653/v1/N19- 1423 [11] Dieckmann, P., Molin Friis, S., Lippert, A., Østergaard, D.: The art and science of debriefing in simulation: Ideal and practice. Medical Teacher 31(7), e287–e294 (Jan 2009). https://doi.org/10.1080/01421590902866218 [12]Ding, Y., Shi, T.: Sustainable llm serving: Environmental implications, chal- lenges, and opportunities : Invited paper. In: 2024 IEEE 15th International Green and Sustainable Computing Conference (IGSC). p. 37–38 (2024). https://doi.org/10.1109/IGSC64514.2024.00016 [13] Echeverria, V., Yan, L., Zhao, L., Abel, S., Alfredo, R., Dix, S., Jaggard, H., Wother- spoon, R., Osborne, A., Buckingham Shum, S., Gasevic, D., Martinez-Maldonado, R.: TeamSlides: A Multimodal Teamwork Analytics Dashboard for Teacher-guided Reflection in a Physical Learning Space. In: Proceedings of the 14th Learning Analytics and Knowledge Conference. p. 112–122. ACM, Kyoto Japan (Mar 2024). https://doi.org/10.1145/3636555.3636857 [14]Echeverria, V., Zhao, L., Alfredo, R., Milesi, M.E., Jin, Y., Abel, S., Fan, J.X., Yan, L., Dix, S., Wotherspoon, R., Li, X., Jaggard, H.A., Osborne, A., Buckingham Shum, S., Gasevic, D., Martinez-Maldonado, R.: TeamVision: An AI-powered Learning Analytics System for Supporting Reflection in Team-based Healthcare Simulation. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. p. 1–22. ACM (2025). https://doi.org/10.1145/3706598.3713395 [15]Elsworth, C., Huang, K., Patterson, D., Schneider, I., Sedivy, R., Goodman, S., Townsend, B., Ranganathan, P., Dean, J., Vahdat, A., Gomes, B., Manyika, J.: Measuring the environmental impact of delivering AI at Google Scale (Aug 2025). https://doi.org/10.48550/arXiv.2508.15734 [16]Erdoğdu, F., Kara, M., Gökoğlu, S., Telci, S.: Trends and insights in cscl research from the emergence to the present: A review through bibliometric and latent dirichlet allocation analyses. Educational Technology & Society 28(4), 166–182 (October 2025) [17]Fanning, R.M., Gaba, D.M.: The role of debriefing in simulation-based learning. Simulation in Healthcare: The Journal of the Society for Simulation in Healthcare 2(2), 115–125 (2007). https://doi.org/10.1097/SIH.0b013e3180315539 [18]Fransen, J., Weinberger, A., Kirschner, P.A.: Team effectiveness and team development in cscl. Educational Psychologist 48(1), 9–24 (2013). https://doi.org/10.1080/00461520.2012.747947 [19]Garg, R., Han, J., Cheng, Y., Fang, Z., Swiecki, Z.: Automated discourse analysis via generative artificial intelligence. In: Proceedings of the 14th Learning Analytics and Knowledge Conference. p. 814–820. LAK ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3636555.3636879 [20]Gašević, D., Dawson, S., Siemens, G.: Let’s not forget: Learning analytics are about learning. TechTrends 59(1), 64–71 (2015). https://doi.org/10.1007/s11528- 014-0822-x [21]Gekhman, Z., Oved, N., Keller, O., Szpektor, I., Reichart, R.: On the robust- ness of dialogue history representation in conversational question answer- ing: A comprehensive study and a new prompt-based method. Transactions of the Association for Computational Linguistics 11, 351–366 (Apr 2023). https://doi.org/10.1162/tacl_a_00549 [22]Guo, D., Yang, D., Zhang, H., Song, J., et al.: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645(8081), 633–638 (Sep 2025). https://doi.org/10.1038/s41586-025-09422-z [23]Hmelo-Silver, C.E., Jeong, H.: An overview of cscl methods. In: Cress, U., Rosé, C., Wise, A.F., Oshima, J. (eds.) International Handbook of Computer-Supported Col- laborative Learning, Computer-Supported Collaborative Learning Series, vol. 19. Springer, Cham (2021). https://doi.org/10.1007/978-3-030-65291-3_4 [24] Hutt, S., Hieb, G.: Scaling up mastery learning with generative ai: Explor- ing how generative ai can assist in the generation and evaluation of mas- tery quiz questions. In: Proceedings of the Eleventh ACM Conference on Learning @ Scale. p. 310–314. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3664699 [25]Inie, N., Falk, J., Selvan, R.: How co2stly is chi? the carbon footprint of generative ai in hci research and what we should do about it. In: Proc. of the CHI Conference on Human Factors in Computing Systems (CHI ’25). p. 1–29. ACM, New York, NY, USA (2025) [26]Jeong, H., Hmelo-Silver, C.E., Yu, Y.: An examination of cscl methodological practices and the influence of theoretical frameworks 2005–2009. International Journal of Computer-Supported Collaborative Learning 9(3), 305–334 (2014). https://doi.org/10.1007/s11412-014-9198-3 [27]Ji, Z., Wang, X., Luo, Z., Xie, Z., Zhang, M.: Optimized Batch Prompting for Cost-Effective LLMs. Proceedings of the VLDB Endowment 18(7), 2172–2184 (Mar 2025). https://doi.org/10.14778/3734839.3734853 [28]Jin, Y., Yu, J.: Optimizing mentor-student communication using llm-based auto- mated labeling information states. In: Proceedings of the Eleventh ACM Confer- ence on Learning @ Scale. p. 284–288. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3664691 [29]Landis, J.R., Koch, G.G.: The Measurement of Observer Agreement for Categorical Data. Biometrics 33(1), 159 (Mar 1977). https://doi.org/10.2307/2529310 [30]Li, M., Qin, W., Tang, Z., Fang, X., He, T., Cao, X.: Automating ssrl detec- tion in asynchronous ocl via llms. In: 2025 7th International Conference on Computer Science and Technologies in Education (CSTE). p. 548–551 (2025). https://doi.org/10.1109/CSTE64638.2025.11092245 [31] Liao, J., Sun, F., Liu, Y., Hu, Y.: Deepseek in education: Exploring the transfor- mative potential of ai-driven educational intelligence. Future in Educational Research n/a(n/a) (2025). https://doi.org/10.1002/fer3.70022 [32]Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: How language models use long contexts. Transac- tions of the Association for Computational Linguistics 12, 157–173 (02 2024). https://doi.org/10.1162/tacl_a_00638 [33] Liu, X., Zambrano, A.F., Baker, R.S., Barany, A., Ocumpaugh, J., Zhang, J., Pankiewicz, M., Nasiar, N., Wei, Z.: Qualitative coding with gpt-4: Where it works better. Journal of Learning Analytics 12(1), 169–185 (2025). https://doi.org/10.18608/jla.2025.8575 [34]Martinez-Maldonado, R., Echeverria, V., Fernandez-Nieto, G., Yan, L., Zhao, L., Alfredo, R., Li, X., Dix, S., Jaggard, H., Wotherspoon, R., Osborne, A., Shum, S.B., Gašević, D.: Lessons learnt from a multimodal learning analytics de- ployment in-the-wild. ACM Trans. Comput.-Hum. Interact. 31(1) (Nov 2023). https://doi.org/10.1145/3622784 [35]Martinez-Maldonado, R., Goodyear, P., Carvalho, L., Thompson, K., Hernandez- Leo, D., Dimitriadis, Y., Prieto, L.P., Wardak, D.: Supporting collaborative design activity in a multi-user digital design ecology. Computers in Human Behavior 71, 327–342 (2017). https://doi.org/10.1016/j.chb.2017.01.055 [36]Martinez-Maldonado, R., Power, T., Hayes, C., Abdiprano, A., Vo, T., Axisa, C., Buckingham Shum, S.: Analytics meet patient manikins: challenges in an authen- tic small-group healthcare simulation classroom. In: Proceedings of the Seventh International Learning Analytics & Knowledge Conference. p. 90–94. LAK ’17, ACM, New York, NY, USA (2017). https://doi.org/10.1145/3027385.3027401 [37] Martinez-Maldonado, R., Schulte, J., Echeverria, V., Gopalan, Y., Shum, S.B.: Where is the teacher? digital analytics for classroom proxemics. Journal of Computer Assisted Learning 36(5), 741–762 (2020). https://doi.org/10.1111/jcal.12444 [38]McHugh, M.L.: Interrater reliability: The kappa statistic. Biochemia Medica 22(3), 276–282 (2012). https://doi.org/10.11613/BM.2012.031 [39]Mehta, S., Srivastava, N., Liu, X., Vanacore, K., Baker, R.S.: Do mooc conver- sations matter? investigating the role of social presence and course-relevant discussion in career advancement. In: Proceedings of the Twelfth ACM Confer- ence on Learning @ Scale. p. 236–240. L@S ’25, ACM, New York, NY, USA (2025). https://doi.org/10.1145/3698205.3733930 L@S ’26, June 29–July 3, 2026, Seoul, Republic of Korea.K. Garces, G. Fernandez-Nieto, L. Zhao, S. Samaraweera, D. Gasevic, R. Martinez-Maldonado, V. Echeverria. [40]Miller, K., Riley, W., Davis, S.: Identifying key nursing and team behaviours to achieve high reliability. Journal of Nursing Management 17(2), 247–255 (Mar 2009). https://doi.org/10.1111/j.1365-2834.2009.00978.x [41] Moore, S., Schmucker, R., Mitchell, T., Stamper, J.: Automated generation and tagging of knowledge components from multiple-choice questions. In: Proceed- ings of the Eleventh ACM Conference on Learning @ Scale. p. 122–133. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3662030 [42]Ngatchou, P., Zarei, A., El-Sharkawi, A.: Pareto multi objective op- timization. In: Proceedings of the 13th International Conference on, Intelligent Systems Application to Power Systems. p. 84–91 (2005). https://doi.org/10.1109/ISAP.2005.1599245 [43]Nguyen, H., Stott, N., Allan, V.: Comparing feedback from large language mod- els and instructors: Teaching computer science at scale. In: Proceedings of the Eleventh ACM Conference on Learning @ Scale. p. 335–339. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3664660 [44]Nie, A., Chandak, Y., Suzara, M., Malik, A., Woodrow, J., Peng, M., Sahami, M., Brunskill, E., Piech, C.: The gpt surprise: Offering large language model chat in a massive coding class reduced engagement but may increase adopters’ exam performances. In: Proceedings of the Twelfth ACM Conference on Learning @ Scale. p. 376–380. L@S ’25, ACM, New York, NY, USA (2025). https://doi.org/10.1145/3698205.3733960 [45] Ouhaichi, H., Spikol, D., Vogel, B.: Rethinking mmla: Design considerations for multimodal learning analytics systems. In: Proceedings of the Tenth ACM Conference on Learning @ Scale. p. 354–359. L@S ’23, ACM, New York, NY, USA (2023). https://doi.org/10.1145/3573051.3596186 [46] Papathoma, T., Ferguson, R., Littlejohn, A., Coe, A.: Making the production of learning at scale more open and flexible. In: Proceedings of the Third (2016) ACM Conference on Learning @ Scale. p. 273–276. L@S ’16, ACM, New York, NY, USA (2016). https://doi.org/10.1145/2876034.2893432 [47] Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C., Sutskever, I.: Robust speech recognition via large-scale weak supervision. In: Proceedings of the 40th International Conference on Machine Learning. ICML’23 (2023) [48]Riley, W., Hansen, H., Gürses, A.P., Davis, S., Miller, K., Priester, R.: The nature, characteristics and patterns of perinatal critical events teams. In: Henriksen, K., Battles, J.B., Keyes, M.A., Grady, M.L. (eds.) Advances in Patient Safety: New Directions and Alternative Approaches (Vol. 3: Performance and Tools). Agency for Healthcare Research and Quality (US), Rockville, MD (Aug 2008), pMID: 21249924 [49] Rosen, M.A., Hunt, E.A., Pronovost, P.J., Federowicz, M.A., Weaver, S.J.: In Situ Simulation in Continuing Education for the Health Care Professions: A Systematic Review. Journal of Continuing Education in the Health Professions 32(4), 243–254 (2012). https://doi.org/10.1002/chp.21152 [50]Salas, E., Reyes, D.L., McDaniel, S.H.: The science of teamwork: Progress, re- flections, and the road ahead. American Psychologist 73(4), 593–600 (2018). https://doi.org/10.1037/amp0000334 [51] Samaraweera, S., Zhao, L., Echeverria, V., Alfredo, R., Chen, G., Davis, J., Leonny, S., Sevenhuysen, S., Connell, C., Gasevic, D., Martinez-Maldonado, R., Dhar- maratne, A.: From formal learning to professional practice: Automated llm-based coding and visualisation of team dialogue in in-situ healthcare simulation. In: Pro- ceedings of the 16th International Learning Analytics & Knowledge Conference (LAK ’26). ACM, New York, NY, USA (2026) [52]Sarangi, S.: Editorial: Team work and team talk as distributed and coordinated action in healthcare delivery. Communication & Medicine 13(1), 1–7 (2017). https://doi.org/10.1558/cam.32569 [53]Sivarajkumar, S., Kelley, M., Samolyk-Mazzanti, A., Visweswaran, S., Wang, Y.: An Empirical Evaluation of Prompting Strategies for Large Language Mod- els in Zero-Shot Clinical Natural Language Processing: Algorithm Develop- ment and Validation Study. JMIR Medical Informatics 12, e55318 (Apr 2024). https://doi.org/10.2196/55318 [54]Southwell, R., Pugh, S., E. Margaret Perkoff, Clevenger, C., Bush, J., Lieber, R., Ward, W., Foltz, P., D’Mello, S.: Challenges and Feasibility of Auto- matic Speech Recognition for Modeling Student Collaborative Discourse in Classrooms. In: Mitrovic, A., Bosch, N. (eds.) Proceedings of the 15th International Conference on Educational Data Mining. Zenodo (Jul 2022). https://doi.org/10.5281/ZENODO.6853109 [55]Stadler, W.: Multicriteria Optimization in Engineering and in the Sciences, vol. 37. Springer Science & Business Media (1988) [56]Stahl,G.:GroupCognition:ComputerSupportforBuilding CollaborativeKnowledge.MITPress,Cambridge,MA(2006). https://doi.org/10.7551/mitpress/3372.001.0001 [57]Stenseth, H.V., Steindal, S.A., Solberg, M.T., Ølnes, M.A., Sørensen, A.L., Strandell- Laine, C., Olaussen, C., Farsjø Aure, C., Pedersen, I., Zlamal, J., Gue Mar- tini, J., Bresolin, P., Linnerud, S.C.W., Nes, A.A.G.: Simulation-Based Learning Supported by Technology to Enhance Critical Thinking in Nursing Students: Scoping Review. Journal of Medical Internet Research 27, e58744 (Feb 2025). https://doi.org/10.2196/58744 [58]Sun, R., Marshall, D.C., Sykes, M.C., Maruthappu, M., Shalhoub, J.: The impact of improving teamwork on patient outcomes in surgery: A sys- tematic review. International Journal of Surgery 53, 171–177 (2018). https://doi.org/10.1016/j.ijsu.2018.03.044 [59]Tscholl, D.W., Ebensperger, M., RahrischRahrisch, A., Wang, H., Heckel, H., Thomasius, M., Kaserer, A., Grande, B., Seelandt, J.C., Kolbe, M.: Generative AI in simulation debriefings: An exploratory study using the Team-FIRST frame- work and qualitative feedback from simulation experts and learners. Advances in Simulation (Jan 2026). https://doi.org/10.1186/s41077-026-00407-0 [60]Wang, D., Chen, G.: Evaluating the use of bert and llama to anal- yse classroom dialogue for teachers’ learning of dialogic pedagogy. British Journal of Educational Technology 56(6), 2671–2704 (2025). https://doi.org/https://doi.org/10.1111/bjet.13604 [61]Wang, D., Tao, Y., Chen, G.: Artificial intelligence in classroom discourse: A sys- tematic review of the past decade. International Journal of Educational Research 123, 102275 (2024) [62]Wang, D., Yang, C., Chen, G.: Using lora to fine-tune large language models for analyzing collaborative argumentation in classrooms. In: Proceedings of the Twelfth ACM Conference on Learning @ Scale. p. 207–211. L@S ’25, ACM, New York, NY, USA (2025). https://doi.org/10.1145/3698205.3733924 [63]White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer- Smith, J., Schmidt, D.C.: A prompt pattern catalog to enhance prompt engineering with chatgpt. In: Proceedings of the 30th Conference on Pattern Languages of Programs. PLoP ’23, The Hillside Group, USA (2023) [64]Yan, L., Gašević, D., Echeverria, V., Zhao, L., Jin, Y., Li, X., Martinez-Maldonado, R.: In Sync or Out of Sync? Understanding Stress and Learning Performance in Collaborative Healthcare Simulations through Physiological Synchrony and Arousal. International Journal of Artificial Intelligence in Education 35(4), 2421– 2452 (Dec 2025). https://doi.org/10.1007/s40593-025-00475-9 [65] Yang, H., Alozie, N., Rachmatullah, A.: Collaboration at scale: Exploring member role changing patterns in collaborative science problem-solving tasks. In: Pro- ceedings of the Ninth ACM Conference on Learning @ Scale. p. 309–312. L@S ’22, ACM, New York, NY, USA (2022). https://doi.org/10.1145/3491140.3528319 [66] You, J., Chung, J.W., Chowdhury, M.: Zeus: Understanding and optimizing GPU energy consumption of DNN training. In: Usenix Nsdi (2023) [67] Zambrano, A.F., Liu, X., Barany, A., Baker, R.S., Kim, J., Nasiar, N.: From ncoder to chatgpt: From automated coding to refining human coding. In: Arastoopour Ir- gens, G., Knight, S. (eds.) Advances in Quantitative Ethnography, Communica- tions in Computer and Information Science, vol. 1895. Springer, Cham (2023). https://doi.org/10.1007/978-3-031-47014-1_32 [68] Zhang, A.G., Tang, X., Oney, S., Chen, Y.: Cflow: Supporting semantic flow analy- sis of students’ code in programming problems at scale. In: Proceedings of the Eleventh ACM Conference on Learning @ Scale. p. 188–199. L@S ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3657604.3662025 [69]Zhao, L., Echeverria, V., Swiecki, Z., Yan, L., Alfredo, R., Li, X., Gase- vic, D., Martinez-Maldonado, R.: Epistemic network analysis for end-users: Closing the loop in the context of multimodal analytics for collaborative team learning. In: Proceedings of the 14th Learning Analytics and Knowl- edge Conference. p. 90–100. LAK ’24, ACM, New York, NY, USA (2024). https://doi.org/10.1145/3636555.3636855 [70] Zhao, L., Gašević, D., Swiecki, Z., Li, Y., Lin, J., Sha, L., Yan, L., Alfredo, R., Li, X., Martinez-Maldonado, R.: Towards automated transcribing and coding of embodied teamwork communication through multimodal learning analyt- ics. British Journal of Educational Technology 55(4), 1673–1702 (Jul 2024). https://doi.org/10.1111/bjet.13476 [71] Zhao, L., Swiecki, Z., Gasevic, D., Yan, L., Dix, S., Jaggard, H., Wotherspoon, R., Osborne, A., Li, X., Alfredo, R., Martinez-Maldonado, R.: Mets: Multimodal learning analytics of embodied teamwork learning. In: LAK23: 13th International Learning Analytics and Knowledge Conference. p. 186–196. LAK2023, ACM, New York, NY, USA (2023). https://doi.org/10.1145/3576050.3576076 [72]Zhu, Y., Gu, J.C., Sikora, C., Ko, H., Liu, Y., Lin, C.C., Shu, L., Luo, L., Meng, L., Liu, B., Chen, J.: Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection (2024). https://doi.org/10.48550/ARXIV.2405.16178 Received 16 February 2026