Paper deep dive
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
Mohammed Nowshad Ruhani Chowdhury, Mohammed Nowaz Rabbani Chowdhury, Sakari Lukkarinen
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/27/2026, 1:16:12 AM
Summary
This study evaluates the effectiveness of fine-tuning the LLaMA 3.1-8B large language model for medical transcription in Finnish. Using a small, validated dataset of simulated clinical conversations created by students at Metropolia University of Applied Sciences, the researchers achieved strong semantic similarity (BERTScore F1 = 0.8230) despite low n-gram overlap, demonstrating the feasibility of using domain-aligned, privacy-oriented LLMs for clinical documentation in low-resource languages.
Entities (5)
Relation Signals (3)
Llama-3.1-8b → usedfor → Medical Transcription
confidence 95% · This study aims to investigate the effectiveness of a domain-aligned natural language processing (NLP); large language model for medical transcription in Finnish
Metropolia University of Applied Sciences → created → Validated Dataset
confidence 90% · simulated clinical conversations by students at Metropolia University of Applied Sciences
Hugging Face → hosts → Llama-3.1-8b
confidence 90% · The model is also accessible via platforms such as Hugging Face
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Clinical documentation is a critical factor for patient safety, diagnosis, and continuity of care. The administrative burden of EHRs is a significant factor in physician burnout. This is a critical issue for low-resource languages, including Finnish. This study aims to investigate the effectiveness of a domain-aligned natural language processing (NLP); large language model for medical transcription in Finnish by fine-tuning LLaMA 3.1-8B on a small validated corpus of simulated clinical conversations by students at Metropolia University of Applied Sciences. The fine-tuning process for medical transcription used a controlled preprocessing and optimization approach. The fine-tuning effectiveness was evaluated by sevenfold cross-validation. The evaluation metrics for fine-tuned LLaMA 3.1-8B were BLEU = 0.1214, ROUGE-L = 0.4982, and BERTScore F1 = 0.8230. The results showed a low n-gram overlap but a strong semantic similarity with reference transcripts. This study indicate that fine-tuning can be an effective approach for translation of medical discourse in spoken Finnish and support the feasibility of fine-tuning a privacy-oriented domain-specific large language model for clinical documentation in Finnish. Beside that provide directions for future work.
Tags
Links
- Source: https://arxiv.org/abs/2603.24772v1
- Canonical: https://arxiv.org/abs/2603.24772v1
Trouble viewing inline? Open PDF directly →
Full Text
51,667 characters extracted from source content.
Expand or collapse full text
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset Mohammed Nowshad Ruhani Chowdhury School of ICT and Industrial Management Metropolia University of Applied Sciences Helsinki, Finland mohammed.chowdhury@metropolia.fi 0000-0002-7429-1483 Mohammed Nowaz Rabbabi Chowdhury Department of Electrical, Computer, and Systems Engineering Rensselaer Polytechnic Institute Troy, NY, USA chowdm2@rpi.edu Sakari Lukkarinen School of ICT and Industrial Management Metropolia University of Applied Sciences Helsinki, Finland sakari.lukkarinen@metropolia.fi 0009-0007-2789-0049 Abstract—Clinical documentation is a critical factor for patient safety, diagnosis, and continuity of care. The administrative burden of EHRs is a significant factor in physician burnout. This is a critical issue for low-resource languages, including Finnish. This study aims to investigate the effectiveness of a domain-aligned natural language processing (NLP); large lan- guage model for medical transcription in Finnish by fine-tuning LLaMA 3.1-8B on a small validated corpus of simulated clinical conversations by students at Metropolia University of Applied Sciences. The fine-tuning process for medical transcription used a controlled preprocessing and optimization approach. The fine- tuning effectiveness was evaluated by sevenfold cross-validation. The evaluation metrics for fine-tuned LLaMA 3.1-8B were BLEU = 0.1214, ROUGE-L = 0.4982, and BERTScore F1 = 0.8230. The results showed a low n-gram overlap but a strong semantic similarity with reference transcripts. This study indicate that fine-tuning can be an effective approach for translation of medical discourse in spoken Finnish and support the feasibility of fine-tuning a privacy-oriented domain-specific large language model for clinical documentation in Finnish. Beside that provide directions for future work. Index Terms—LLM, Medical Transcription, NLP, Finnish Language, LLaMA. I. INTRODUCTION In a modern healthcare system, clinical documentation is absolutely essential for patient safety, correct diagnosis, sup- port of medical billing, and preservation of legal documents. However, the complexity in the health care system and the complexity in the patients’ conditions have increased the burden on the health care professionals in documenting the patients’ records, which has led to a decrease in the time spent interacting with patients, leading to burnout among physicians [1]. This burden is mainly on the doctors and nurses, who have to divide their time between interacting with the patients and documenting the patients’ records in electronic health records. M.N-R. Chowdhury is with the School of ICT and Industrial Management at Metropolia University of Applied Sciences, Helsinki, Finland (e-mail: now- shad.cse@gmail.com). This work was completed by Mohammed Nowshad Ruhani Chowdhury as part of the Metropolia University of Applied Sciences Research project: Digital scribes. No external funding was received. This has a negative impact on the quality of interaction with the patients and the quality of the patients’ records. This not only affects the quality of service provided by the health care professionals but also impacts the health care professionals’ satisfaction and well-being. It has also led to confusion in the patients’ records, which has a negative impact on the health care services provided, especially in complex patient conditions. Fig. 1. Overview of Fine-Tuned LLM model using Small Validate Dataset. Growing concerns among healthcare stakeholders have em- phasized that the increasing volume of clinical documentation arXiv:2603.24772v1 [cs.CL] 25 Mar 2026 directly affects the quality of interactions between patients and clinicians in consultations. These issues, in turn, highlight the need to adopt a structured approach to clinical documen- tation, in which patient data is systematically organized in a structured, normalized, and easily understandable format. This approach to structuring clinical data promotes greater interoperability between healthcare systems, facilitates more informed clinical decision-making, and enables secondary uses of data, such as population health management or medical research [2]. In addition, this approach ensures that vital clinical information, such as diagnoses, medications, allergies, or follow-up recommendations, is recorded in an accurate, accessible, and consistent manner to enhance the quality of patient care. The incorporation of medical scribes, whether human or digital, has surfaced as an effective strategy to alleviate admin- istrative burdens and enhance workflow efficiency [3]. Human scribes assist physicians by documenting patient encounters in real time, allowing clinicians to focus more on patient interaction rather than data entry into electronic health records. While this approach has been shown to enhance workflow and increase physician satisfaction, it also presents challenges such as high costs, limited scalability, potential inconsistencies in documentation quality, and concerns related to data privacy and standardized training. The limitations of human scribes have led to the rapid adoption of AI [4], especially in the form of automatic speech recognition and natural language processing (NLP), to facilitate the development of digital medical transcription tools. Recent developments in large language models (LLMs) have also led to the significant improvement of the ability to process and organize complex language [5]. In terms of the current study, which assesses the effectiveness of a fine- tuned LLM for Finnish medical transcription using a validated dataset, the open-source nature of LLMs provides a significant advantage, especially in terms of the ability to adapt the model to the needs of the specific domain while ensuring the privacy of the data. Fine-tuning the LLM on a limited but high-quality dataset of Finnish clinical language can facilitate the effective transformation of unstructured speech into standardized med- ical documentation. However, in non-English languages, especially those with rich morphological traits like Finnish, the construction of such systems becomes much more challenging. Finnish (Suomi) is a morphologically agglutinative language, indicating that a single word can represent several grammatical character- istics, leading to a highly sparse vocabulary and complex parsing challenges. In the medical domain, this complexity is compounded by specialized terminology, code-switching (e.g., Latin, English loanwords), and the need for high accuracy due to patient safety concerns [6]. Open-source LLMs represent a promising approach in the context of a health environment with strict privacy regulations such as GDPR in place because of the benefits of transparency and customization options along with the opportunity of local deployment in a secure manner. These models have the advantage of being adapted to specific domains using specific datasets in specific languages such as Finnish, which allows them to recognize complex linguistic features such as morphology and medical terminology. Therefore, this is in line with the objectives of this specific study, which is focused on the evaluation of a fine-tuned LLM in the context of Finnish medical transcription using a small validated dataset. This is especially relevant in the context of the development of accurate and privacy-compliant transcription systems. There is a great opportunity of minimizing documentation work with the help of this approach in order to provide better care with the assistance of such a system. Consequently, this study proposes an advance approach: fine-tune open-source LLMs to accurately the transformation of transcribed clinical speech into structured documentation, specifically for the Finnish language. Fig. 1 shown the architecture of Evaluating Fine- Tuned LLM model using Small Validate Dataset of low- resourced languages such as Finnish (Suomi). I. LITERATURE REVIEW Healthcare’s digital transformation has placed increasing pressure on the efficiency, accuracy, and structure of clinical documentation [7]. As physicians face mounting administra- tive demands, accurate record-keeping has become both a critical necessity and a major burden. Clinical documentation is essential to quality assurance, billing, diagnosis, treatment continuity, patient safety, and compliance with rules. However, because of the enormous time commitment needed, many doctors are now more burned out and dissatisfied, spending more time documenting than connecting with patients. To minimize the documentation time, the use of human medical scribes has been advocated, who assist in recording the clinical encounters in real time, thus allowing the physician to concentrate more on the patients [2] [4]. However, the use of medical scribes has also been found to have some drawbacks, such as the need to incur heavy training costs, the quality of the documentation process, and the issue of privacy, which has shown the need to explore more effective alternatives, such as the fine-tuned LLM-based Finnish medical transcription system, as researched in this study. Limitation of human scribes has led to the development of AI-based digital scribe systems, which utilize ASR and NLP techniques for clinical documentation generation [7]. The performance of ASR has improved significantly in recent times, especially for open-source tools such as GPT, Deep- Speech, and Whisper, which utilize deep learning techniques for speech recognition. This study, therefore, seeks to assess the performance of a fine-tuned LLM on a small validated dataset for the improvement of speech recognition in the con- text of Finnish medical transcription, which may be affected by domain-specific terminologies, linguistic complexity, and code-switching in noisy backgrounds. After transcription, unstructured clinical conversations must be converted into structured formats to support decision- making and system interoperability. Recent advancements in open-source LLMs have demonstrated promising results in extracting key medical information and producing structured outcomes from unstructured dialogues [8]. In this particular study, a fine-tuned LLM is assessed for medical transcription in Finnish, using a small but validated dataset, based on its capacity to transform unstructured dialogues into structured medical documents. Complexity is a fundamental attribute of clinical conversa- tions, which are often interrupted, ambiguous, and colloquial in nature, making transcription and structuring more difficult compared to regular medical records [9]. However, recent breakthroughs in the field have demonstrated the ability of fine-tuned LLMs to successfully transcribe such conversations, which are typically represented in a multi-turn dialogue for- mat, outperforming regular NLP approaches. In this study, a fine-tuned LLM is tested in the field of Finnish medical transcription, using a small validated data set. While recent studies have shown some positive results in the area of medical transcription and LLMs, most of them have focused on the English language, with little work being conducted in lesser-resourced languages such as Finnish [10]. The language of Finnish, with its complex morphology and word order flexibility, adds to the complexity of the task, especially considering the lack of resources available for the task. While tools such as FinBERT, OPUS-MT, and AaltoASR are useful, domain-specific models need to be fine-tuned. The present work aims to assess a fine-tuned model for Finnish medical transcription with a small validated data set. This study is based on the concept of transfer learning [11], where language models are fine-tuned on specific domain and language data to enhance performance. Previous studies have shown that domain-specific language models perform better than general ones in clinical tasks. This study is based on the concept, and it aims to test whether a fine-tuned LLM can produce accurate and structured Finnish medical transcription using a small validated dataset. I. METHODOLOGICAL APPROACH This section outlines the process for data creation for fine- tuning, selecting open-source LLMs, pre-train model find- ing from hugging face platform, and defining the evaluation methodology. A. Data Creation The data creation process began with the establishment of a foundational dataset comprising Finnish clinical case conversations. These baseline conversations were formatted simulated interactions between healthcare professionals such as doctors or nurses and patients. However, students from the Innovation Project [12], play those roles in creating this dataset (cf. Step-I, in Fig. 1 ). The aim was to capture authentic, context-rich dialogue reflective of real-world clinical scenarios. This baseline serves as the core reference for further development and training of conversational models or language-processing systems in the Finnish healthcare context. To build the dataset, each clinical scenario was documented in two formats: an audio recording in MP3 format and a corresponding textual transcription. The MP3 files represent the actual spoken dialogues, while the text files offer hu- man readable versions of these recordings in Finnish. These transcriptions are aligned with the audio to ensure accuracy and preserve the nuances of clinical communication, including terminology, emotional tone, and conversational flow. All records in the dataset follow the same naming pattern for traceability and for structural organisation. Files are encoded in the format I01-G01-C01.txt and I01-G01-C01.mp3, where each component possesses some metadata. The prefix I01 indicates iteration number of the batch of the dataset. The middle segment G01 delineates the specific group of students in the innovation projects who contributed to the creation of that data entry, facilitating the attribution of performance and contribution between groups of students. The final segment C01 is the clinical case number, which links the file to an individual patient scenario. This naming pattern is applied uniformly to both audio (MP3) and text (TXT) files to facilitate simple management, retrieval, and referencing within research, analysis, or model training pipelines. Students from different disciplines, like nursing, podia- try, paramedic, and gerontology, worked out to carry out the production process, and this process was monitored by their teachers. Under supervision, these students actively contributed to the creation and transcription of the clinical dialogues, guaranteeing clinical authenticity and realism. Their multidisciplinary contribution was critical in the creation of a dataset that achieves a balance between healthcare usefulness and pedagogical usefulness. The generated resource supports instructional use cases in nursing and healthcare simulation settings and offers a basis for training AI systems in clinical communication in the Finnish language. B. Hugging Face Platform The development of the so-called Transformer architectures signaled a major breakthrough in the field of NLP, allowing for a level of scalability, efficiency, and effectiveness beyond the RNN-based architectures [13]. Building upon the effectiveness of the pre-training of language models on large datasets, the study aims to fine-tune the LLMs for the task of medical transcription in Finnish, a task that requires efficiency and effectiveness. As NLP models became more advanced, the need for accessible and adaptable tools led to the rise of platforms like Hugging Face, which support open-source model sharing and fine-tuning [14]. These platforms allow researchers to easily fine-tune pre-trained models for specific tasks. In the study, these tools are used to fine-tune LLMs for medical transcription in Finnish using a small set of validated data. Hugging Face is a popular platform in the machine learning community with its open-source ecosystem providing a variety of models of the Transformer type, which can be easily fine- tuned and deployed with the help of its tools [15]. The platform is highly suitable to be used to adapt models to a particular domain or application. In the present study, the platform is used to create and test a fine-tuned LLM model for the application of medical transcription in the Finnish language with the help of a small validated set of data. In fact, Hugging Face offers a wide range of tools for devel- oping domain-specific language models through its integrated tools for datasets, pre-trained models, as well as efficient deployment of language models in a collaborative environment [16]. This study utilizes the tools provided by Hugging Face to develop and evaluate a fine-tuned LLM model for efficient transcription of medical documents in Finnish using a small validated dataset. Step-I, in Fig. 1 the approach how Hugging Face used in this research. In addition, Hugging Face contributes to both academic and practical research by supporting open-source development, making it easy to reproduce results, and providing easy access to state-of-the-art tools in NLP research [17]. The tools provided by Hugging Face, including efficient tools for fine- tuning models and deploying models, allow researchers to fine- tune pre-trained models according to specific domains or lan- guages. This research utilizes the tools provided by Hugging Face to fine-tune an LLM model for medical transcription in the Finnish language using a small validated dataset. C. Selection of LLM Model For this research, selecting the appropriate Large Language Models (LLMs) was guided by a set of criteria that align with the specific requirements of structured clinical documentation in Finnish. The goal was to identify models that not only demonstrate a high level of accuracy in medical language processing but also provide flexibility for customisation and support for the Finnish language. The core selection criteria included: a) Ability to handle medical text optimisation: Large language model parameters refer to the model’s internal vari- ables, which are the weights and biases learned during the training process. Parameters influence the effectiveness of the model in identifying linguistic patterns, sensing the context, and generating meaningful language [18]. Parameters’ size and configuration influence the complexity and effectiveness of the model. b) Support for structured clinical documentation: Struc- tured clinical documentation is the process of taking un- structured medical information such as free-text conversation or handwritten orders and placing it into a standardized, structured format [2]. It enables quicker data retrieval, clinical decision making, health system interoperability, and secondary use such as analytics and research. c) Customisation and fine-tuning capabilities: A pre- trained language model for the Finnish language is first trained on a large dataset of Finnish text to acquire language patterns, including grammar and vocabulary, and can be further fine- tuned on a specific dataset for specific tasks [19]. In the study, the method is used to fine-tune an LLM model for medical transcription in the Finnish language using a small validated dataset. d) Finnish language support, pre-training or adaptability for continued learning: An adaptable model can be modified or extended to accommodate specific research objectives or do- main requirements [20]. It includes concepts like modification of the model structure, addition of new data sets, or tuning of the training objectives for optimal task-specific performance. Although models like BioGPT and ClinicalBERT have proven efficiency in medical-related tasks, they do not have the Finnish language option. However, models like FinBERT, which have been designed for the Finnish language, have proven efficiency [21]. In this study, appropriate open-source LLMs have been identified and adapted for the purpose of balancing efficiency and flexibility in the transcription of medical texts in the Finnish language using a small validated dataset. Based on all criteria and models’ availability on Hugging face we developed a models comparison table to solve this selection stage, Table I explain it in detail. TABLE I MODELS’ COMPARISON BASED ON SELECTION CRITERIA Model NameParameters (B/M) Structured Clinical Doc- umentation Finnish (Pre- trained/ Fine-tune) Customisable Llama 3 (Meta) [22]8B,70B, 405B YesYesYes SiloGenFinnish GPT [23] 7B,13B, 33B YesYesYes Teuken-7B (OpenGPT-X) [24] 7BYes (with fine- tuning) YesYes HippocratesLLM [25] 7B, 13BYesNoYes DeepSeek [26]7B, 67BYes (with fine- tuning) YesYes TurkuNLP FinBERT [27] 110MYes (with fine- tuning) YesYes BioGPT (Microsoft) [28] 1.2BYesYesYes ClinicalBERT [29]110MYesNoYes All models support customization through fine-tuning or prompt engineering. Based on the evaluation criteria, the model chosen by the study is LLaMA 3.1-8B, as it has proven its efficiency, scal- ability, and adaptability. Being an open-source LLM model, it allows for efficient fine-tuning methods, making it more suitable for domain-specific applications, including Finnish medical transcription. The model is also accessible via plat- forms such as Hugging Face, making it more viable for use in the study. As a result, the model is fine-tuned using a small validated data set in order to produce accurate documentation in Finnish. D. Evaluation Methodology This section of the study gives a clear explanation of the method that was used to assess the performance of the fine-tuned model. Specifically, it outlines how model outputs were contrasted with reference clinical notes made by humans for the purposes of determining the quality of structured documentation that is generated from clinical conversations in the Finnish language. In the proposed study, the quality of Finnish medical tran- scriptions produced by the fine-tuned LLM will be assessed using the BLEU, ROUGE, and BERTScore methods [30]. The proposed methods will allow the assessment of the medical transcriptions’ quality by comparing them to validated medical references. The BLEU score is used to compare the lexical similarity of the Finnish medical transcription generated with the reference clinical notes through the evaluation of the precision of n-gram [31]. It is also used to determine the syntactical accuracy of the clinical terms used in fine-tuned LLM. The ROUGE metric, which includes ROUGE-N and ROUGE-L, is used to test the quality of the Finnish medical transcription texts produced by the fine-tuned LLM. This is done by comparing the texts with the reference clinical notes. The ROUGE-N metric works using n-grams, whereas the ROUGE-L metric works by finding the longest subsequence using the longest common subsequence algorithm [32]. This is important in medical contexts because it determines whether important medical details have been accurately captured de- spite the variation in the wording. To assess the semantic similarity between the produced and reference material, BERTScore is included to go beyond surface-level lexical matching. Unlike the word-based compar- ison with BLEU or ROUGE, BERTScore uses contextual word representations to compare the meaning of the output with the target output [33]. In the study, BERTScore is integrated to compare the Finnish medical transcription output generated with the LLM with the target output, thus making the model effective in capturing clinically equivalent expressions. These three indicators work together to offer a comprehen- sive framework for evaluation. ROUGE gauges content recall, BERTScore records semantic integrity, and BLEU quantita- tively assesses syntactic accuracy. This complete evaluation technique facilitates an impartial assessment of the model’s capacity to produce accurate, therapeutically pertinent, and cohesive documentation from Finnish clinical dialogues. E. K-Fold Cross-Validation Cross-validation is a popular method used to test the gen- eralizability of models in machine learning algorithms, espe- cially when working with small datasets. In the present study, the model uses the K-fold cross-validation method because of the small size of the validated Finnish medical transcription dataset [34]. In the K-fold cross-validation method, the data is divided into K-folds, and the model is trained on all the data to ensure the efficient use of the data to test the model’s performance. In this study, the data set is limited to only seven clinical dialogues in Finnish with their corresponding MP3 files and manually annotated reference documents. Due to the limited number of samples, standard single-split validation would lead to unstable and biased performance estimates. To combat this limitation, a 7-fold cross-validation (leave-one-out) method is adopted where a single conversation is employed as test case while learning is conducted on the remaining six, which is dispatch in the fig 2. This allows for a greater overall understanding of the model’s behaviour under various patient interactions and clinical subjects, which is essential in the healthcare sector where data diversity is as essential as data volume. Fig. 2. K-Fold Cross-Validation (k=7) Structure. K-fold cross-validation is advantageous when using LLMs to fine-tune domain-specific models since it reduces the vari- ance due to random sampling, guarantees the stability of model performance with different samples, and detects any inconsistencies in the dataset [35]. For a low-resource domain like Finnish medical transcription, the method is advantageous when avoiding overfitting is a priority through multiple tests of model generalizability. Together with BLEU, ROUGE, and BERTScore metrics, the model can be evaluated syntactically and semantically to ensure the scientific rigor of the model’s performance assessment. IV. EXPERIMENTAL STUDY AND ANALYSIS This section explains the experiment configuration, dataset preparation, fine-tuning process, and evaluation method of the study. It discusses the test of the fine-tuned LLM model on Finnish clinical conversations, including the results obtained as well as analysis based on automatic metrics. A. Dataset Summary As explained in Section I-A, the study utilized a custom dataset drawn from Finnish-language clinical MP3 records and their manually written textual equivalents. The dataset was specifically curated to reflect clinical conversation in Finnish, capturing a variety of medical scenarios and communication styles. Audio recordings were generated from existing text data to ensure that the resulting transcriptions maintained both clinical appropriateness and linguistic precision, supporting effective natural language processing. The dataset is comprised of 7 complete clinical dialogues. Pre-processing was applied to each sample and transferred into structured JSON format specially to cater to the training requirements of large language models. The JSON format has a clear segmentation of dialogue turns, speaker designation (e.g.,doctor or patient), and annotated fields that indicate the target output for clinical documentation. This format allowed for supervised fine-tuning, which allowed the model to learn how to convert free-form clinical discussion into well-formed, documentation that met the standards of healthcare. B. Experimental Setup In this section, the technical setup for model training and evaluation is explained, which was carried out on the CSC Puhti supercomputer to provide a secure environment for fine- tuning the large language model [36]. The setup included hardware with GPU optimization along with necessary ML libraries and frameworks to support the training of the model, particularly the transformer-based model. The setup was ef- fective for the adaptation of the LLM. a) Software: This work uses Python Ver. 3.11.5, PyTorch version≥ 2.0 with Compute Unified Device Architecture (CUDA) 11.7, and open source ML libraries on a GPU based Puhti supercomputer for fine-tuning, validation, and quantitative evaluation [37] [38]. The project leveraged basic machine learning library tools, PyTorch for model training, and CUDA for parallel processing via the supercomputer Puhti. The pre-trained LLaMA model and tokenizer are provided by the Hugging Face library, while the datasets library was used for efficient data management. Clinical voice recordings in Finnish language are initially transcribed using the Whisper LargeV3 model, followed by pre-processing in a standardized manner for fine-tuning the LLM model. b) Hardware: The fine-tuning experiments were con- ducted on a CSC Puhti supercomputer with one NVIDIA A100 GPU and 40 GB VRAM to fine-tune LLaMA 3.1-8B in 4-bit quantized form with LoRA adapters. The system had 16 CPU cores and 128 GB system memory to accommodate data preprocessing, parallel task execution, and model loading. Additionally, 200 GB disk space was available to store model weights, outputs, and data. This setup was a good compro- mise between computational efficiency and model demands to enable fine-tuning and evaluating models for Finnish medical transcription with stability and speed. C. Experimental Setup on CSC Puhti Supercomputer To begin the final experimental setup, an account was created on the CSC cloud service, then a project was created based on the title of the study. After approval of the project and receiving the project ID, as a student of Metropolia University of Applied Sciences, we were given 100,000 billing units to have full access to the Puhti supercomputer and CSC cloud services to train, tune, and test all the models without any cost for the Finnish medical transcription service. The project was executed on a GPU node that satisfied all hardware and software requirements, including a single NVIDIA A100 GPU with 40 GB of VRAM, 16 Xeon CPU cores, 128 GB of RAM, and approximately 200 GB of local storage. This provided a balance of memory and computational resources required for fine-tuning the 8-billion- parameter LLaMA 3.1 model in 4-bit quantized form with the use of LoRA adapters. The CPU cores also enabled the parallel processing of data preprocessing and I/O. Additionally, the high-speed interconnect of the Puhti cluster, along with the Lustre file system, provides a reliable level of reproducibility in the performance of the project. Before coding the project, the research team obtained the necessary access to the LLaMA 3.1-8B model through the Hugging Face library. This was done in accordance with the terms of the license, the responsible use of the model, and the ethical guidelines established by the organization. Once the necessary approval was granted, the model’s weights were downloaded into the environment, allowing for the fine-tuning of the model for the purposes of Finnish medical transcription. D. Finetuning The fine-tuning of the LLaMA 3.1-8B model marked a significant milestone in the development of a specialized LLM from a general-purpose LLM, with the broader objective of fulfilling healthcare-related educational objectives. The train- ing data comprised MP3 file formats of clinical conversations, along with their corresponding text-based versions, wherein students of the Innovation Project, act healthcare professional (e.g., doctor, nurse, therapist) [12], and patient are the roles played in the development of the dataset. In the next step, the audio recordings were transcribed using the Whisper large- v3 model. Subsequently, the audio-transcript data was pre- processed into a standard format for the development of the LLM. Five of the audio-transcript files were chosen as the ini- tial training and test files for two files each. In the subsequent step, a 7-fold cross-validation method was adopted for the de- velopment of the LLaMA 3.1-8B model. This method helped the model learn from the dataset repeatedly, thereby avoiding the chances of overfitting. This cross-validation method also helped the model attain higher reliability for its application in the development of clinical documentation and conversational AI. The fine-tuning was conducted using the meta-llama/Llama- 3.1–8B model, accessed from a local directory after obtaining the necessary license from Meta. The model was added to memory with float16 precision and device-map="auto" to utilize available GPU memory capacity on the CSC Puhti supercomputer. The LLaMA architecture-compatible tokenizer was also initialized with the original tokenizer model and utilized to convert textual training data into sequences of tokens. Records were padded or truncated up to a maximum of 512-tokens to maintain input lengths consistent during training. This 512-tokens was chosen as the input length limit to main- tain transformer structure compatibility, with tractable memory usage in training while still gathering adequate clinical context. The training involved using Hugging Face’s Trainer API in half-precision supervised fine-tuning mode. Significant hyper- parameters included a batch size of one (1), three (3) epochs of training, logging at every 10-steps (e g. 6 training samples × 3 epochs = 18 steps per fold), and saving the checkpoint at every 100-steps with a limit of 2 saved checkpoints. Gradient check pointing was not used under this setup because the training setup was made minimal for the sake of clarity and stabil- ity. The models were trained in half-precision floating point (fp16=True) to maintain the calculation below the memory constraint of the available NVIDIA V100 GPUs on Puhti. The training script included a custom tokenizer and data collator for causal language modelling. The DataCollatorFor- LanguageModeling was used with masked language modelling disabled mlm=False), aligning with the LLaMA model’s causal architecture. No evaluation was conducted during training (evaluation-strategy="no"), as the focus was on completing a robust initial finetuning pass on the full training set. Following the process of training, the fine-tuned model was saved in different forms. First, the model and tokenizer were saved in Hugging Face transformers format through save- pretrained() for convenient reloading as well as continued development. Additionally, the PyTorch state dictionary for the model was saved in .pth format for non-Hugging Face out-of- the-box applications or non-Hugging Face frameworks com- patibility. While LLaMA models themselves aren’t natively TensorFlow/Keras compatible, there was a placeholder .h5 file included to show potential future need for a conversion, though full .h5 export is not currently supported. All computations were done with SLURM job scripts on the Puhti GPU cluster, with GPU reservations requested by –gres=gpu:v100:1, 16 CPU cores, and 128 GB RAM. The virtual Python environment was also activated within the job script to assure stable dependency management and repro- ducibility. This fine-tuning process produced a domain-tuned LLaMA- 3.1–8B model with the ability to generate Finnish clinical conversation texts in the given contexts. The model can now translate spoken Finnish medical discourse between patients and healthcare professionals into neatly organized clinical conversation documents. Trained model and its training and preprocessing scripts are made available publicly at GitHub facilitating future research and deployment in clinical AI tools. V. RESULTS AND EVALUATION To evaluate the model ability after fine-tuned, research conducted automatic testing using BLEU,ROUGE-L, and BERTScore measures. These analyses were applied between generated documents and ground-true documents of Finnish clinical conversations. The evaluation process took a 7− f old cross-validation method to ensure proper performance estima- tion on different subsets of the data and prevent overfitting with respect to the finite amount of data available, Table I illustrated the results of k-fold validation. This helped in ensur- ing stringent evaluation of the model’s capability to generalize to clinical conversations generation and documentation tasks. The BLEU score was employed to evaluate the precision of n-gram overlaps between the fine-tuned LLM model output and reference Finnish clinical transcripts, which reflected the similarity between the model output and actual human conversations. The ROUGE-L score, with a focus on recall, was employed to evaluate the completeness of the generated text with respect to the reference text. Moreover, BERTScore was employed to evaluate semantic similarity, which reflected the similarity in meaning, a common phenomenon in clinical contexts. TABLE I 7-FOLD CROSS-VALIDATION RESULTS File NamesBLEUROUGE-LBERTScore I01-G01-C01.txt0.07390.31560.7642 I01-G01-C02.txt0.15790.51510.8384 I01-G01-C03.txt0.11540.57590.8621 I01-G02-C01.txt0.13300.58290.8472 I01-G02-C02.txt0.12470.45580.7905 I01-G03-C01.txt0.11510.54470.8390 I01-G03-C02.txt0.12950.49740.8197 Average0.12140.49820.8230 All data files used in k-fold validation. Across 7-fold cross-validation, the model achieved a mean BLEU score of 0.1214, ROUGE-L score of 0.4982, and BERTScore of 0.8230. These all reflect moderate syntactic similarity and high semantic matching with human-transcribed references due to the complexity and domain-specific nature of the input data. In general, 10–20 % BLEU scores are generally moderate, particularly in low resource and high- variety settings. ROUGE-L scores over 0.5 indicate adequate overlap between generated and reference text. By contrast, BERTScore scores of above 0.8 represent high semantic simi- larity, with scores of 0.85 or higher being generally regarded as representative of high-quality semantic-level correspondence. These test results confirm that fine-tune LLMs model, LLaMA- 3.1–8B’s capability to generate contextually appropriate and linguistically coherent outputs from Finnish clinical audio in- puts. The 7-fold cross-validation complemented the credibility of the performance metrics by demonstrating stable outcomes across diverse data partitions. In fig 3 illustrated the results of all dataset based on BLUE, ROUGE-L, and BERTScore. In conclusion, the evaluation supports the model’s readiness for downstream applications in Finnish-language clinical NLP systems. Fig. 3. Model results in k-fold validation. VI. DISCUSSION AND CONCLUSIONS The fine-tuning of the LLaMA-3.1-8B model for generating Finnish clinical conversation documents illustrates the utility of domain-specific fine-tuning of LLMs. Through the use of hand-curated transcripts and Whisper transcriptions, the model was able to demonstrate strong capabilities in understanding and generating clinical content in Finnish. Through evaluation metrics, it is evident that the fine-tuned model was able to successfully capture both syntactic accuracy and semantic coherence, even with the limited dataset. This illustrates the utility of specialized LLMs for tasks such as clinical docu- mentation and education in low-resource languages. The fine-tuned LLaMA-3.1–8B model was more in line with the linguistic nuances of Finnish healthcare communication. While larger models tend to outperform smaller models on general benchmarks due to large-scale pretraining, this study demonstrates that task-specific fine-tuning even on a limited dataset. The k-fold cross-validation improved the model’s strength and reduced overfitting issues, especially with domain-specific data like clinical texts. It improved the generalizability of the model towards diverse input patterns. While BERTScore (0.8230) reported high semantic retention, BLEU (0.1214) and ROUGE (0.4982) suggested there were areas to be worked on in precision and recall. Using training on the CSC Puhti supercomputer with fp16 optimization and LoRA validated the process to be sustainable using limited resources, albeit parameter fine-tuning is required for replicating the process in less-capable machines. A key limitation of this study is the small size of the dataset, which consisted of only seven Finnish clinical dia- logues, and the fact that the LLaMA-3.1–8B model which is the only model used, is not natively pre-trained in Finnish. These constraints may affect the model’s ability to learn linguistic nuances as well as clinical domain-specific patterns. Future work could involve expanding the training dataset with more authentic or synthetic Finnish clinical conversation to further enhance model generalization, and there should be train others models and comparison between the models, the baseline results (without pre-training) will also need to eval- uate. Moreover, additional architectural refinements, such as incorporating structured end-of-sequence tokens or leveraging retrieval augmented generation, may increase output accuracy and reduce hallucination. As shown in related works using models like TurkuNLP, FinBERT, or ClinicalBERT, smaller LLMs can be made more effective through such augmenta- tion, suggesting potential for lighter deployment scenarios in healthcare systems with resource limitations. In the end, while the study was based on automatic eval- uation, it shows the possibility of using human-in-the-loop feedback in the future. Healthcare professionals and Finnish language experts could be used as a validation tool, which could give better results in terms of the quality of the output of the model. The results of the study show that fine-tuned LLMs, as in the case of LLaMA-3.1-8B, could be a good basis in the development of intelligent medical scribes and conversational AI in Finnish settings. ACKNOWLEDGMENT The authors would like to thank Digital Medical Scribe project team members, Principal Lecturer Mikael Soini, Prin- cipal Lecturer Päivi Haho, Senior Lecturer Sakari Lukkarinen and supervisors of Nursing, Bachelor’s Degree, Nursing Bilin- gual Top-Up Degree, Biomedical Laboratory Science, Bache- lor’s Degree, Laboratory Science, Bachelor’s Degree, Bachelor of Health Care, Physiotherapist (top-up), and Bachelor of Health Care, Physiotherapist for their support in collecting the dataset, and direct and indirect contribution to the project. REFERENCES [1] A. J. Holmgren, N. Hendrix, N. Maisel, J. Everson, A. Bazemore, L. Rotenstein, R. L. Phillips, and J. Adler-Milstein, “Electronic health record usability, satisfaction, and burnout for family physicians,” JAMA Network Open, vol. 7, no. 8, p. e2 426 956–e2 426 956, 2024. [2] M. Sasseville, F. Yousefi, S. Ouellet, F. Naye, T. Stefan, V. Carnovale, F. Bergeron, L. Ling, B. Gheorghiu, S. Hagens, S. Gareau-Lajoie, and A. Leblanc, “The Impact of AI Scribes on Streamlining Clinical Documentation: A Systematic Review,” Healthcare, vol. 13, no. 12, p. 1447–1447, jun 16 2025. [3] T. J. Miksanek, M. R. Skandari, S. A. Ham, W.-W. Lee, V. G. Press, M. T. Brown, and N. LAITEERAPONG, “The Productivity Require- ments of Implementing a Medical Scribe Program,” Annals of Internal Medicine, vol. 174, no. 1, p. 1–7, oct 5 2020. [4] S. A. Mess, A. J. Mackey, and D. E. Yarowsky, “Artificial Intelligence Scribe and Large Language Model Technology in Healthcare Docu- mentation: Advantages, Limitations, and Recommendations,” Plastic & Reconstructive Surgery Global Open, vol. 13, no. 1, p. e6450–e6450, jan 1 2025. [5] H. Hassan, A. R. Zipursky, N. Rabbani, J. G. You, G. Tse, E. Orenstein, M. Ray, C. Parsons, S. Shin, G. Lawton, K. Jessa, L. Sung, and A. P. Yan, “Clinical Implementation of Artificial Intelligence Scribes in Health Care: A Systematic Review,” Applied Clinical Informatics, vol. 16, no. 04, p. 1121–1135, apr 30 2025. [6] M. Hämäläinen and K. Alnajjar, “The Current State of Finnish NLP,” arXiv (Cornell University), sep 23 2021. [7] A. Adedeji, S. Joshi, and B. Doohan, “The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models,” arXiv (Cornell University), feb 13 2024. [8] I. Garcia-Ferrero, R. Agerri, A. A. Salazar, E. Cabrio, I. de la Iglesia, A. Lavelli, B. Magnini, B. Molinet, J. Ramirez-Romero, G. Rigau, J. M. Villa-Gonzalez, S. Villata, and A. Zaninello, “Medical mT5: An Open- Source Multilingual Text-to-Text LLM for The Medical Domain,” arXiv (Cornell University), apr 12 2024. [9] N. A. Khatim, A. A. Irfan, and M. M. Arief, “Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia: A Proof-of-Concept Study,” arXiv (Cornell University), sep 26 2024. [10] Q. Yang, J. Chen, Y. Sun, Y. Wang, and T. Tan, “Fine-tuning medical language models for enhanced long-contextual understanding and do- main expertise,” Quantitative Imaging in Medicine and Surgery, vol. 15, no. 6, p. 5450–5462, jun 1 2025. [11] G. Vrban ˇ ci ˇ c and V. Podgorelec, “Transfer Learning With Adaptive Fine- Tuning,” IEEE Access, vol. 8, p. 196 197–196 211, jan 1 2020. [12] Metropolia University of Applied Sciences, “Innovation projects,” https: //w.metropolia.fi/en/rdi/innovation-projects, Helsinki, Finland, 2025, accessed: 2025-05-11. [13] X.-K. Wu, M. Chen, W. Li, R. Wang, L. Lu, J. Liu, K. Hwang, Y. Hao, Y.-R. Pan, Q. Meng, K. Huang, L. Hu, M. Guizani, N. Chao, G. Fortino, F. Lin, Y. Tian, D. Niyato, and F.-Y. Wang, “Llm Fine-Tuning: Concepts, Opportunities, and Challenges,” Big Data and Cognitive Computing, vol. 9, no. 4, p. 87–87, apr 2 2025. [14] U. R. Pol, “Hugging Face: Revolutionizing AI and NLP,” International Journal for Research in Applied Science and Engineering Technology, vol. 12, no. 8, p. 1121–1124, aug 27 2024. [15] A. Ait, J. L. C. Izquierdo, and J. Cabot, “HFCommunity: A Tool to Analyze the Hugging Face Hub Community,” in 2023 IEEE Interna- tional Conference on Software Analysis, Evolution and Reengineering (SANER). Taipa, Macao: IEEE, mar 2023, p. 728–732. [16] T. Susnjak, P. Hwang, N. Reyes, A. L. C. Barczak, T. McIntosh, and S. Ranathunga, “Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning,” ACM Transactions on Knowledge Discovery from Data, vol. 19, no. 3, p. 1–39, jan 31 2025. [17] Y. Shen, K. Song, T. Xu, L. Dongsheng, L. Weiming, and Y. Zhuang, “Hugginggpt: Solving AI Tasks with ChatGPT and its Friends in Hugging Face,” arXiv (Cornell University), mar 31 2023. [18] H. Zhiqiang, W. Lei, Y. Lan, W. Xu, E.-P. Lim, L. Bing, X. Xing, S. Poria, and R. K.-W. Lee, “Llm-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models,” arXiv (Cornell University), apr 5 2023. [19] S. Lankford, H. Afli, and A. Way, “adaptmllm: Fine-Tuning Multilingual Language Models on Low-Resource Languages with Integrated LLM Playgrounds,” Information, vol. 14, no. 12, p. 638–638, nov 29 2023. [20] Y. Chen, H. Zhang, X. Yang, W.-l. Zhang, and D. Qu, “Meta-Adaptable- Adapter: Efficient adaptation of self-supervised models for low-resource speech recognition,” Neurocomputing, vol. 609, p. 128 493–128 493, aug 28 2024. [21] S. Maity and M. J. Saikia, “Large Language Models in Healthcare and Medical Applications: A Review,” Bioengineering, vol. 12, no. 6, p. 631–631, jun 10 2025. [22] T. Hugo, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and Efficient Foundation Language Models,” arXiv (Cornell University), feb 28 2023. [23] AMD, “Silogen and tietoevry care are developing a finnish-speaking ai assistant for healthcare professionals,” AMD blog post, 2023, accessed: 2026-03-22. [24] A. Mehdi, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rut- mann, C. Jain, M. Lübbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, A. Hammam, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. Suarez, M. Ostendorff, A. Jude, L. Manjunath, S. Weinbach, C. Penke, O. Filatov, F. Barth, M. Paramita, L. Weber, W. Ines, R. Sifa, K. Fabian, A. Herten, R. Jäkel, G. Rehm, S. Kesselheim, J. Köhler, and N. Flores-Herr, “Teuken-7b-Base & Teuken-7b-Instruct: Towards European LLMs,” arXiv (Cornell University), oct 8 2024. [25] E. C. Acikgoz, O. B. ̇ Ince, R. Bench, A. A. Boz, ̇ I. Kesen, A. Erdem, and E. Erdem, “Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare,” arXiv (Cornell University), apr 26 2024. [26] B. Xiao, C. Deli, C. Guanting, S. Chen, D. Dai, D. Cheng-qi, D. Hong- hui, D. Kai, Du Qiushi, F. Zhe, H. Gao, G. Kaige, G. Wenjun, R. Ge, G. Kang, D. Guo, G. Jian-zhong, H. Guang-bo, H. Zhe-wen, H. Ying, W. Hu, H. PanPan, E. Li, L. Guowei, L. Jia-Shi, L. Yao, L. Y. K. , L. Wenfeng, F. Lin, A. X. Liu, L. Bo, L. Wen, L. Xiao-dong, L. Xin, L. Yiyuan, H. Lu, S. Lu, F. Luo, M. ShiRong, N. Xiaotao, P. Tian, Y. Piao, Q. Junjie, Q. Hui, T. Ren, Z. Ren, R. Chong, Z. Sha, S. Zhi-hong, S. Junxiao, S. Xuecheng, J. Sun, S. Yaofeng, T. Minghui, W. Bingxuan, W. Peiyi, S. Wang, W. Yao-hui, W. Yong-Ji, T. Wu, W. Y. , X. Xin, Z. Xie, Z. Xie, X. Yi-liang, X. Hanwei, R. X. Xu, X. Yanhong, Y. Dejian, Y. Yu-xiang, Y. Shuiping, Y. Xingkai, Z. B. , Z. Haowei, L. Zhang, Z. Li-yue, Z. Ming-Chuan, Z. Ming-hua, W. Zhang, Z. Yi- chao, Z. Chenggang, Y. Zhao, Z. Shang-yan, Z. Shunfeng, Q. Zhu, and Z. Yuheng, “Deepseek LLM: Scaling Open-Source Language Models with Longtermism,” arXiv (Cornell University), jan 8 2024. [27] A. Virtanen, J. Kanerva, R. Ilo, J. Luoma, J. Luotolahti, T. Salakoski, F. Ginter, and S. Pyysalo, “Multilingual is not enough: Bert for finnish,” arXiv preprint arXiv:1912.07076, 2019. [Online]. Available: https://arxiv.org/abs/1912.07076 [28] R. Luo, L. Sun, Y. Xia, T. Qin, Z. Sheng, H. Poon, and T.-Y. Liu, “Biogpt: generative pre-trained transformer for biomedical text gener- ation and mining,” Briefings in Bioinformatics, vol. 23, no. 6, sep 24 2022. [29] L. Matondora, M. Mutandavari, and B. Mupini, “Nlp Based Prediction of Hospital Readmission using ClinicalBERT and Clinician Notes,” International Journal of Innovative Science and Research Technology (IJISRT), p. 2549–2557, aug 10 2024. [30] Z. Tianyi, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating Text Generation with BERT,” arXiv (Cornell University), feb 28 2022. [31] T. Glushkova, C. Zerva, and A. F. T. Martins, “Bleu Meets COMET: Combining Lexical and Neural Metrics Towards Robust Machine Trans- lation Evaluation,” arXiv (Cornell University), may 31 2023. [32] M. Barbella and G. Tortora, “Rouge Metric Evaluation for Text Sum- marization Techniques,” SSRN Electronic Journal, jan 1 2022. [33] G. Tang, O. Yousuf, and Z. Jin, “Improving BERTScore for Machine Translation Evaluation Through Contrastive Learning,” IEEE Access, vol. 12, p. 77 739–77 749, jan 1 2024. [34] Department of Computer Science and Informatics, University of Energy and Natural Resources, I. K. Nti, O. N. Boateng, and J. Aning, “Perfor- mance of Machine Learning Algorithms with Different K Values in K- fold CrossValidation,” International Journal of Information Technology and Computer Science, vol. 13, no. 6, p. 61–71, dec 8 2021. [35] Z. Lin, J. Lai, X. Chen, L. Cao, and J. Wang, “Curriculum Reinforcement Learning Based on K-Fold Cross Validation,” Entropy, vol. 24, no. 12, p. 1787–1787, dec 6 2022. [36] CSC – IT Center for Science, “Puhti supercomputer,” https://docs.csc.fi/ computing/systems-puhti/, Espoo, Finland, 2025, accessed: 2025-05-18. [37] Python Software Foundation, “Python version 3.11.5,” https://w. python.org/, 2023, accessed: 2026-03-22. [38] NVIDIA Corporation, “Cuda toolkit 11.7,” https://developer.nvidia.com/ cuda-toolkit, 2022, accessed: 2026-03-22.