Paper deep dive
PLeDO: Pain Level Detection for Osteoarthritis from EMR Data
Yuhao Chen, Jiahao Cai, Nafiz Sadman, Farhana Zulkernine, John Queenan, David Barber
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/22/2026, 2:52:52 AM
Summary
This paper introduces PLeDO, an integrated tool for detecting pain levels (mild vs. moderate-to-severe) in Osteoarthritis (OA) patients using Electronic Medical Record (EMR) data. It combines SPaDe, an unsupervised synonym-based clustering method for unstructured chart notes, with structured medication data and pain scale information. The study utilizes data from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN) and demonstrates the feasibility of automated pain severity detection to improve primary care quality.
Entities (9)
Relation Signals (7)
PLeDO → detects → pain severity
confidence 98% · propose an integrated pain level detection tool for OA called PLeDO... to understand the pain severity for OA
Yuhao Chen → affiliatedwith → Queen's University
confidence 95% · Yuhao Chen... School of Computing, Queen’s University
PLeDO → includescomponent → SPaDe
confidence 95% · PLeDO... combines (a) SPaDe, the pain expression-based clustering approach
Osteoarthritis → symptom → pain
confidence 95% · Pain is the most disabling symptom of OA
PLeDO → usesdatasource → EMR
confidence 95% · understand the pain severity for OA from patients' primary care Electronic Medical Records (EMR)
SPaDe → processes → unstructured chart note
confidence 92% · based on only the pain related expressions in the unstructured chart note
CPCSSN → providesdatafor → PLeDO
confidence 85% · The patient sample included in this study is extracted from... CPCSSN
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain severity for OA from patients' primary care Electronic Medical Records (EMR), both from the structured medical data and the unstructured chart note data using information extraction, natural language processing and machine learning techniques. We propose SPaDe, a Synonym-based Pain level Detection tool to categorize patients into having mild or moderate-to-severe pain to understand diagnosis and treatment methods based on only the pain related expressions in the unstructured chart note. Expressions are subjective, objective, and influenced by cultural background and demography which poses a difficult challenge. Therefore, we improve the model by incorporating the medication information from the structured EMR data and pain scale related information from the chart note to propose an integrated pain level detection tool for OA called PLeDO. With the help of human labeled gold standard data, we demonstrate that both SPaDe and PLeDO can detect mild and moderate-to-severe pain from the EMR data to analyze and potentially improve the quality of care in primary care setting.
Tags
Links
- Source: https://arxiv.org/abs/2608.15719v1
- Canonical: https://arxiv.org/abs/2608.15719v1
Trouble viewing inline? Open PDF directly →
Full Text
80,420 characters extracted from source content.
Expand or collapse full text
PLeDO: Pain Level Detection for Osteoarthritis from EMR Data Yuhao Chen 1 Jiahao Cai 1 Nafiz Sadman 1 Farhana Zulkernine 1 John Queenan 2 and David Barber 2 Email: yuhao.chen@queensu.ca Abstract Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain severity for OA from patients’ primary care Electronic Medical Records (EMR), both from the structured medical data and the unstructured chart note data using information extraction, natural language processing and machine learning techniques. We propose SPaDe, a Synonym-based Pain level Detection tool to categorize patients into having mild or moderate-to-severe pain to understand diagnosis and treatment methods based on only the pain related expressions in the unstructured chart note. Expressions are subjective, objective, and influenced by cultural background and demography which poses a difficult challenge. Therefore, we improve the model by incorporating the medication information from the structured EMR data and pain scale related information from the chart note to propose an integrated pain level detection tool for OA called PLeDO. With the help of human labeled gold standard data, we demonstrate that both SPaDe and PLeDO can detect mild and moderate-to-severe pain from the EMR data to analyze and potentially improve the quality of care in primary care setting. keywordspain severity, osteoarthritis, primary healthcare, unstructured data, clustering †affiliation: 1School of Computing, Queen’s University, Kingston, Ontario, Canada 2Department of Family Medicine, Queen’s University, Kingston, Ontario, Canada†corresponding: Yuhao Chen, School of Computing, Queen’s University, Kingston, Ontario, Canada 1 Introduction Prevalence and incidence rates for OA are problematic to establish due to variations in diagnostic definitions [10]. Arthritis Alliance of Canada estimated that there may be over 4.4 million people living with OA in Canada and in 30 years this number can reach 10 million (1 in 4 Canadians) [1]. With a growing aging population in Canada and increasing rates of obesity and inactivity, the rate of OA is projected to increase from 13.8% to 18.6% between 2010 and 2031 [3]. Comorbidity is common in OA population and approximately 59 - 87% of people with OA have at least one other chronic illness [4]. Approximately 80% of the individuals with OA exhibit some degree of movement limitations, and 25% are not able to perform their regular daily activities of life [5]. Pain associated with OA and its severity significantly impacts the health‐related quality of life (HRQOL) and productivity of the affected population [6]. The loss in costs due to productivity or work associated with OA in Canada is substantial, which was estimated to be $12 billion in 2010 and to reach $17.5 billion Canadian dollars in 2031 [3]. Pain is the most disabling symptom of OA and a major driver of clinical decision making and heath care utilization [17]. Therefore, identifying symptomatic OA patients with moderate to severe pain in a real-world setting would be integral to understanding the burden of the disease and the treatment journey. Information about the pain medication can help pharmaceutical companies to innovate new treatment options. The Canadian Primary Care Sentinel Surveillance Network (CPCSSN) is a multi-disease EMR surveillance system [22]. It consists of over 2 million patients’ data collected from 1,500 participating primary care clinicians. CPCSSN’s data extraction algorithm only extracts the structured data items such as date of birth, gender, and disease from EMR, but not the unstructured chart note data, which often contains valuable clinical information. Traditional manual extraction/audit of unstructured and free text data from clinical notes and other narratives is time consuming, labor intensive and expensive. Information Extraction (IE) can relieve some of these problems by enabling automated extraction of the relevant information into a research database and thereby, facilitating further analysis of the data using a combination of Natural Language Processing (NLP) and Machine Learning (ML) techniques [28, 29, 30]. NLP facilitates the processing of semi-structured and unstructured text data in narrative clinical documents to identify and label medical terminology, and interpret written information [31] [52]. It also helps to extract and transform text data into numeric vectors that can be passed to ML algorithms for prediction and decision support [32, 33, 34, 58]. Currently there are no disease-modifying agents available in the market for OA. Non-pharmacologic and pharmacologic therapies used for management of OA are aimed at improving pain, disability, and quality of life [4]. The most common pharmacological treatment options for the symptomatic treatment of OA are acetaminophen, NSAIDs (topical/oral), Cox2 inhibitors, duloxetine, intra-articular (IA) corticosteroids, IA hyaluronic acid, tramadol or strong opioids [12][13]. OA embodies one of the most frequently occurring painful conditions [14]. Pathophysiology of OA pain is complex, exhibiting a combination of nociceptive and neuropathic mechanisms involved in both the local and central levels [14]. In Canada, OA pain information is unlikely to be systematically documented in the EMR. To the best of our knowledge, no studies in Canada have yet attempted to identify and characterize patients with moderate-to-severe OA pain based on EMR chart note data in the primary case setting. The use of longitudinal EMR data is very useful for surveillance of a population that is at high risk of or is diagnosed with OA [20]. CPCSSN [21][22][24] has developed and validated an algorithm that identifies patients diagnosed with OA in Canadian primary care using readily available structured data extracted from patients’ EMR [23]. However, this algorithm currently cannot classify OA pain levels into mild, moderate-to-severe based on the descriptive chart note data. Contribution: In this research, we explore and develop a variety of information extraction (IE) methods using NLP and machine learning (ML) methods to retrieve important information from the CPCSSN EMR structured and unstructured chart note data. We fabricate these methods into multiple IE pipelines to integrate and encode the data to train a classification model. The aim is to determine the severity of pain experienced by OA patients mainly based on the unstructured text data in the physicians’ notes and also the structured medication data. The contributions of this research are as follows. • To the best of our knowledge, this is the first study that utilizes integrated NLP and ML techniques to identify pain levels in patients with Osteoarthritis using physicians’ chart note data and structured medication information from EMRs. • We present an unsupervised synonym-based clustering approach, SPaDe, which extracts and clusters pain related expressions from the unstructured chart note data to categorize patients into mild and moderate-to-severe and validate it using gold standard manually labeled data. • We propose a novel integrated Pain-level Detection tool for OA, PLeDO, which combines (a) SPaDe, the pain expression-based clustering approach, (b) a pain scale-based approach, and (c) a medication-based approach to detect OA pain level using both structured medication and unstructured chart note data. PLeDO categorizes patients into mild and moderate-to-severe pain categories for studying treatment patterns in the primary care setting. • We provide a scalable, less resource-intensive methodology for pain level detection, contributing to advancements in clinical NLP. The framework can inspire further research into unsupervised methods for other healthcare text classification tasks, expanding the toolkit available to the community. • An ablation study is presented to demonstrate the improvements achieved at different stages when extending SPaDe with the additional information to build PLeDO. The rest of the paper is organized as follows. We describe the related work in Section 2. The methodology is discussed in Section 3 which provides an overview of the analytic workflow and explains the study sample. Section 4 illustrates the implementation details about SPaDe, the pain scale and the medication based approaches to pain level categorization. The experimental results are presented in Section 5 with discussions about the observations and outcomes. Finally, Section 6 concludes the paper with a list of future work directions. 2 Related Work 2.1 Background Patient reported pain experiences for knee and hip OA in the context of disease progression was categorized by OARSI/OMERACT initiative as follows [15]. • Early OA: Predictable, sharp or other pain brought on by specific triggers, eventually limiting high impact activities but affecting little the other low impact activities. • Mid OA: Predictable, more constant pain occurring in association with joint symptoms (e.g., joint locking). This pain affects daily activities such as walking. • Advanced OA: Pattern of constant dull and aching pain with intermittent unpredictable episodes of intense pain which leaves a person exhausted. This results in a significant avoidance of social and recreational activities [15]. Various patient-reported outcomes have been used to assess pain and disability in hip and knee OA [8] [18]. One of the most widely used tools is the Western Ontario and McMaster Universities OA Index (WOMAC) [16], which consists of three sub-scales: pain, stiffness, and physical function. For the evaluation of OA pain, Visual Analog Scale (VAS) or Numerical Rating Scale (NRS) assessment of pain intensity are also commonly used [7]. Patients may be asked about the experience of “pain, aching of stiffness in or around the knee” over a specific time frame. 2.2 Pain Related Classification In recent years, a growing number of studies have explored the intersection of pain assessment and machine learning to improve pain management and tailored treatment plans [49, 50, 51]. Most existing literature focuses primarily on numerical data, with limited exploration of text data using supervised learning. Lotsch et al. [45] incorporated supervised machine learning to analyze numerical data from preoperative cold pain tests to predict persistent pain after breast cancer surgery. High negative predictive values (94%) were achieved, though positive predictive values were low (10%). Results highlight the role of the endogenous pain inhibitory system in pain persistence. Similarly, Alambo et al. [48] explored the use of statistical machine learning methods, such as logistic regression and decision trees, to distinguish patterns between patients with and without pain. Their best-performing model achieved a remarkable 98% F1 score, demonstrating the potential of these methods for accurate pain classification. Beyond numerical data, several studies have investigated the use of clinical text for pain-related classification. DiMartino et al. [46] evaluated the feasibility of using NLP to detect uncontrolled symptoms (moderate or severe pain) in clinical notes from 1,644 hospital encounters for cancer patients. They used the machine learning models in Clinical Annotation Research Kit to do classification and achieved 61% accuracy and 69.5% F1 score. Their findings demonstrated the initial feasibility of NLP in identifying symptom burden but emphasized the need for further development before such tools can be reliably implemented in clinical workflows. In another study, the authors in [47] analyzed 235,789 clinical texts from an Emergency Department Information System using BlueBERT, a domain-specific BERT model pre-trained on PubMed abstracts. The model was employed for binary classification (pain or no pain), achieving an impressive 95% accuracy on the evaluation set. Recent studies also indicate that pain- and symptom-related phenotyping from clinical narratives remains feasible but highly task-specific. In a directly relevant recent study, Hughes et al. [54] investigated pain-related classification in emergency-department clinical text, showing that large-scale unstructured notes can support automated pain-focused modeling in acute-care settings. More recently, note-based pain and symptom analytics have expanded into adjacent supervised clinical-NLP tasks. For example, a hybrid machine-learning and LLM framework was reported to predict short-horizon cancer pain episodes from combined structured and unstructured EHR data [55]. Oncology note-based studies [56] have also demonstrated the feasibility of extracting symptom presence and severity directly from treatment notes using BERT-family models. Related recent work in psychiatry notes further suggests that even in contemporary clinical NLP, carefully designed rule-based systems remain competitive with or superior to larger language models when the corpus is small, expert-annotated, and clinically specialized [57]. Taken together, these studies confirm the value of numerical data, structured clinical variables, and narrative EMR text for pain- and symptom-related phenotyping. While supervised learning approaches demonstrate promise, obtaining expert annotations is often prohibitively expensive and time-consuming. Although clustering methods have been applied to pain level identification in prior research [49], these efforts have predominantly focused on numerical data, which is less complex and challenging compared to unstructured textual data. Existing studies are predominantly supervised and depend on manually labeled datasets, disease-specific annotations, or infrastructure that supports large pretrained models. In contrast, our work focuses on osteoarthritis pain severity identification from unstructured primary-care EMR notes. Gold-standard labels were unavailable at model-development time and the workflow had to operate in an air-gapped environment because of sensitive patient data. Accordingly, our contribution is distinct from recent supervised note-classification studies: we address OA pain stratification in a no-label, privacy-constrained setting using an unsupervised framework tailored to real-world primary-care EMR data. The challenge motivated us to explore and develop more scalable and cost-effective approaches for leveraging clinical text in healthcare research. 3 Methodology In this section, we discuss the key challenges that influenced the algorithms, a brief overview of our approaches, creation of the study sample, and the study environment. 3.1 Challenges Unstructured text data analytics offer many challenges. • Length of the sentences are not uniform. • Text data can contain spelling errors, domain specific terminology, ill structured data with missing “,”, “.”, “’”, and incorrect grammar and phrases. • Text data is often written in hybrid format with numeric data, dates, acronyms, emoji, and other form of literal expressions. • Text data can also contain expressions written in multiple languages in the same document. • Each chart note often contains duplicate data from previous notes because the new chart entry for a patient’s visit is created by first copying the text from the previous note and then appending new information to it. The specific type of medical chart note data that we had in this study, offered the first three challenges more frequently. The expressions, as mentioned in the 4th4^th point, were not in different languages but had wide variations for pain levels. Therefore, extraction of the words posed a serious challenge for the classification of the medical records or patients with mild, moderate, or severe pain. Each patient’s EMR is identified by an ID which helps to link personal data such as age, demography, and location, to multiple other data such as the chart notes and medications. However, a patient can have many years of notes and medications. Therefore, extracting and linearizing patients’ records especially those with many years of data can be a big challenge. When feeding this linearized data to machine learning models, the size disparity of the data creates further problems as some of the data are too long while others may have just one encounter record. Another problem with processing multiple years of data is that a patient can have different records containing expressions of varying pain levels. After treatment, pain can reduce and then increase again. How should such patients be categorized? Based on extensive discussion with the medical experts and collaborators, we decided about the following policy for classification. • If a patient has multiple notes expressing different pain levels, label the patient with the highest pain level expression. • As expressions of pain can vary for different culture, race, and ethnicity, synonyms should be considered to represent each pain level instead of specific words. • Only pain related medications should be considered for this OA pain related study to understand the treatment pattern. An overview of our methodology is presented next followed by a description of the data used in this study. 3.2 Overview of PLeDO The key objective of this study is to develop novel NLP and ML techniques to categorize OA patients based on their reported pain levels into mild and moderate-to-severe categories in the primary care EMR data. We propose an integrated Pain Level Detection (PLeDO) tool to categorize OA patients into mild and moderate-to-severe pain groups given their structured EMR data containing medications and unstructured chart note data. We use the medication information from the structured parts of the EMR data. The unstructured chart notes contain valuable information such as scale of pain level as logged by the primary care physician during the patient encounter. The notes also contain information shared by the patients about their pain which can vary widely based on patients’ demography, culture, age, and gender. We aim to extract such expressions including the pain scale information and the medication information and use the same in developing the integrated system, PLeDO, which consists of the following 3 subsystems. 1. Pain Expression-based system - Extracts patients’ expressions regarding pain levels from the chart note data by developing 3 word dictionaries representing 3 pain levels: mild, moderate, and severe. A word embedding method is also used to transform the words to vectors to facilitate word similarity calculation. The similarity of each word in patients’ notes from each dictionary is computed to classify notes and hence the patients into mild or moderate-to-severe category. 2. Pain Scale-based system - Extracts pain scale related information recorded by some physicians in the unstructured chart note data. 3. Medication-based system - Extracts and compiles prescribed medications, referral to pain clinics, surgical treatments, or use of assistive tools from all records spanning multiple years from the structured EMR data for each patient. In parallel, Pain related medications are grouped into mild, and moderate-to-severe pain levels. Based on the types of medications prescribed to them, patients are categorized to the highest level pain category. The workflow of the algorithm is given in Figure 1. Further details about each subsystem are given in Section 3. 3.3 Study Sample The patient sample included in this study is extracted from OSCAR [25] open source EMR system used by the primary case physicians participating in this study. The sample data includes primary care patient population having a diagnosis of OA based on the CPCSSN case definition algorithm. Figure 1 shows the overall workflow of this research. The complete process involved multiple steps. First, the CPCSSN data extraction algorithm customized for the various EMR systems was used to extract and store selected data items on a secured staging server, where the data is deidentified using the TiDE [35][44] by the CPCSSN data scientists. Then the deidentified data was shared with our research group on another virtual machine customized for secure data analysis. Since the chart note data often contains names and Protected Health Information (PHI), all PHI needed to be removed or deidentified, and the research ethics approval needed to be in place before the data could be shared with us. Figure 1: The Overall Workflow We created an initial study sample based on the following inclusion criteria. Then an exclusion criteria was applied to filter out specific types of patients’ records to prepare the final study sample. 3.3.1 Inclusion Criteria CPCSSN regularly collects and integrates the structured EMR data into the database but the unstructured chart notes are not included in the integrated database. Therefore, for this study we had to follow the data request procedure to extract the unstructured chart notes from the EMR systems of a regional family health team (FHT) of practitioners. All registered patients of FHT of practitioners are by default considered to have given consent to such medical studies unless they explicitly opt out via an established procedure. A total of 23,431 patients were registered in the EMR systems of the selected regional FHT whose data were extracted for this study to create the initial patient population. Out of these patients, only 2,044 patients of all ages were identified as having OA based on the CPCSSN OA case definition algorithm [23]. In a separate study, CPCSSN researchers published an algorithm to identify and label OA patients with ICD-9 disease codes based on specific criteria and the available data fields in the EMR [23]. We used that label as the gold standard label to identify and include OA patients in our study sample. We included only the OA patients’ EMR data for the period of 12 years from 2010-2022. We further narrowed down the dataset by applying the following exclusion criteria to focus on specific type of OA related pain for the study purposes. Table 1: Sample Size Description Creation of Study Sample Num. of Patients Total OA Patients (2010-2022) 23,431 After applying Inclusion Criteria 2,044 After applying Inclusion and Exclusion Criteria 797 Total Samples to Create Gold Standard Data 269 Gold Standard Samples Missing Pain Info 113 Final Gold Standard Samples 156 3.3.2 Exclusion Criteria From the OA patient population consisting of 2,044 patients, we excluded all patients having the following ICD criteria (and all ICD-9-CM subcodes). 1. Inflammatory joint diseases such as (a) Rheumatoid arthritis (ICD-9 code: 714) (b) Psoriatic arthritis (ICD-9 code: 696) (c) Ankylosing spondylitis (ICD-9 code: 720) (d) Septic arthritis (ICD code: 711) (e) Erythematous conditions (including lupus) (ICD code: 695) (f) Gout (including pseudogout) (ICD code: 274) 2. Systemic metabolic bone disease (e.g., Crystal arthropathies ICD-9: 712, Paget’s disease or Osteitis deformans without mention of bone tumor, ICD-9 code: 731; Disorders of mineral metabolism such as metastatic calcifications ICD-9 code 275) 3. Other disorders of soft tissues such as Fibromyalgia and Neuropathic pain, ICD-9 code:729) Following the application of exclusion criteria, we identified a total of 797 OA patients. We then obtained the deidentified EMR structured and unstructured data in multiple CSV files. The unstructured chart note data was processed to extract pain related expressions and information regarding the use of pain scales by the physicians to measure pain levels. A subset of this data was manually reviewed to create the gold standard data. The structured data was used for gender based analysis and mainly to obtain pain related medications prescribed to the patients. 3.3.3 Patient Demography We extracted additional structured data such as patients’ gender and age to report about the distribution of patients in our study population. 3.3.4 Gold Standard Data We randomly selected a subset of the OA patients from our study sample for manual labeling with mild, moderate, and severe pain levels. The sample size was calculated based on 26% prevalence (patients with severe pain) with an expected sensitivity of 80% and specificity of 80% at 95% confidence interval and 10% precision. We considered the manually labeled data as the gold standard data for the development and validation of our algorithms to classify OA pain levels. The manual chart validation was conducted by a person with years of experience and formal training in labeling medical data. To reduce the bias and ensure the consistency, the annotator was guided by predefined protocols developed in collaboration with medical experts. Additional, he was also supported by other medical expert collaborators who provided guidance during the annotation process. In the case of ambiguity or confusion, other experts were consulted to ensure accuracy of the labels. A total of 269 OA patients out of the 797 patients in the study sample were selected randomly for manual evaluation. Based on the results of evaluation, among these 269 patients, chart notes for 156 patients included mentions of their pain levels, while the remaining 113 patients did not have any explicit reference to pain levels in their records. Therefore, for the purpose of evaluating our algorithm, we utilized the data of the 156 patients only which had pain level information explicitly mentioned in their chart notes. 3.4 Study Environment For this study, the deidentified data was staged on a secured data analytics environment called the Restricted Data Environment (RDEN) for access, analysis, and reporting. For the 1st1^st objective of categorizing pain level, we developed an information extraction and transformation pipeline using NLP techniques to process the unstructured text data in the EMR chart notes. The overall data processing includes deidentification, information extraction and cleaning, and developing approaches to predict pain levels as mild, or moderate-to-severe based on the EMR chart note data. We combined “moderate” and “severe” into one category because the treatments heavily overlap for these two categories. However, if needed, the categories, moderate and severe, can be separated with a negative effect on the accuracy of each subcategory as the algorithms fail to determine the category correctly solely based on medication information. 4 Implementation Our integrated system (PLeDO) comprises 3 different approaches for categorizing the pain levels namely pain expression-based (SPaDe), pain scale-based, and medication-based approaches. We first explain the general data pre-processing methods that applies to one or more of our proposed approaches. The implementation of each of the approaches are described subsequently under specific subsections. 4.1 General Data Pre-processing We built the information extraction (IE) algorithms based on a manual visual inspection of the data to identify the key challenges in processing the data, positions of the desired information in the data, layout of the data, and other cues such as context, keywords, or symbols that can be used to develop the IE algorithms. The key challenges with processing EMR data were discussed under Section 3.1. Accordingly, to achieve the objectives of pain level classification, we built custom information extraction, cleaning, preprocessing, and transformation algorithms considering the structure, layout, context, and representation of the EMR structured and unstructured data. The extracted information was used for data linking and analysis, and transformed into different representations to feed into machine learning algorithms. 4.1.1 OA Paragraph Extraction: As explained under Section 3.1, each chart note often contains text from the previous chart notes and includes details about the complete historical record of the patient with multiple diseases. Consequently, extraction of pain related information returns a lot of noise from different diseases and not just OA. Therefore, we conducted a keyword search method to extract paragraphs related to OA from the unstructured chart note data. We constructed a comprehensive keyword dictionary containing keywords associated with OA. Subsequently, we searched these keywords in each chart note. If any of the keywords were identified in a note, the sentence containing the keywords, as well as two preceding and two succeeding sentences were extracted. Finally, for each patient, we combined all the OA related extracted contents from all EMRs into a combined note. 4.1.2 Data Cleaning and Preprocessing: The EMR chart note and transformed deidentified structured data in the CSV files required robust data cleaning and formatting for information extraction. Characters such as “,” were removed from medication and demographic information. As mentioned earlier, often previous notes are copied and then new notes are appended when logging the chart note data creating much duplication in the data. Eliminating duplicate data was necessary to accurately calculate frequencies from the data. Similarly, numerical data may not be germane to the analysis and can cause problems during data processing. For example, a patient’s height, weight, BMI, and other numeric information such as time, and date are not useful for the analysis. We applied Natural Language Toolkit (NLTK) [26] methods to remove the special characters, and punctuations. We developed algorithms to compare and remove duplicate records, and create a cleaner version of the data. 4.1.3 Data Feature Extraction: To extract all potentially significant data features such as verbs, adverbs, and adjectives from the EMR notes, we applied the NLTK [26] Named Entity Recognition (NER) and Parts of Speech (POS) taggers to identify and tag the significant terms for this study. Sentiments such as (positive sentiment) “improved”, “feeling better”, (negative sentiment) “worsening” are indicative of a patient’s emotional states with a focus on pain management. A regular sentiment analysis method cannot address the complex and subjective nature of expression and experiences of the patients focused on pain. Therefore, we use NER and POS to identify and tag words which are often used to describe pain such as verbs: “ache”, “hurt”, “burn”; adjectives: “sharp”, “dull”; adverbs: “intensely”, and “slightly”. After tagging, we developed algorithms to extract the words from the annotated chart note data for the different approaches. We applied regular expression based data extraction for the pain scale based approach as explained in the respective section. For expressions such as date or pain scale where the data has a specific format, regular expressions can be used to effectively extract such data. 4.1.4 Word Embedding: Word embedding [27] converts words into numerical vectors in a high-dimensional space to be processed efficiently by computer algorithms. Older approaches such as one hot encoding or TF-IDF [43] can not capture the context information. Machine learning based models [27] are able to process contextual information and create similar embeddings for synonymous words. We converted the extracted word features to vectors using GLOVE embedding [27], which gave good performance and was computationally cost-effective compared to Med-BERT [41] more complex embedding methods. Since the GloVe embeddings rely on a pretrained vocabulary, any words with typos is removed in the embedding generation process. The nature of our framework, which relies heavily on similarity comparisons between each word and each pain-level embedding. The pretrained language models such as BERT-style models are not inherently designed for direct similarity comparisons at the word level or sentence level, resulting in suboptimal performance for our specific use case [53]. While models like Sentence-BERT [53] address some limitations in similarity comparison, they are pre-trained and fine-tuned on general domain data. In contrast, the GloVe embedding version we used was trained on a substantial amount of medical-specific vocabulary, making it particularly suitable for our dataset and task requirements. This domain-specific vocabulary played a crucial role in improving the accuracy of our framework. Figure 2: The Workflow of Pain-Level Clustering 4.2 Approach I: Pain Expression-based Approach Manually labeling all chart notes for supervised learning was not feasible especially because the size of a single note varied from 1 to 9,488 words and there were 482,185 notes to label. Therefore, we developed an unsupervised learning approach to cluster the notes into mild, and moderate-to-severe categories based on words expressing different pain levels. Then we used the gold standard data to validate our approach. Patients use subjective, situational, cultural and emotional words to describe their pain level based on their background, demography, race, and gender. Therefore, we defined 3 Pain Level Word Dictionaries (PLWD), each representing a collection of words or synonyms aligned with one of the 3 pain levels, namely mild, moderate, and severe. To extract relevant expressions for OA, we first extracted OA related sentences and then applied different NLP techniques to clean and process the extracted data as explained under Section 4.1. Next we applied GLOVE embedding to convert the text to numeric vectors. Figure 2 depicts the different processing steps. The rest of this section explains the different steps in more detail such as creating the PLWDs, using them to group and extract pain related words from the chart notes, and applying unsupervised clustering to determine the centroids of the extracted word clusters. Finally, the similarity of the combined chart note of each patient, represented by the average value of the extracted word embeddings, to the centroids is used to categorize the note or the patient under either mild and moderate-to-severe pain category. 4.2.1 Pain Level Word Dictionary: We developed three Pain Level Word Dictionaries (PLWD) by compiling three distinct sets of English vocabularies with 50 synonyms in each dictionary indicating mild, moderate, or severe pain level. We used Beautiful Soup11 1 Beautiful Soup: https://tedboy.github.io/bs4_doc/, a web scraping library, to extract data from a specific website22 2 Website: https://w.thesaurus.com/ providing a digital thesaurus and a tool for identifying synonyms. After creating these dictionaries automatically, we filtered the results by ChatGPT [37][38] to eliminate irrelevant or redundant terms and thereby, ensure that the final set of synonyms accurately described the corresponding pain level. Specifically, ChatGPT was used to evaluate whether each candidate word appropriately described pain intensity corresponding to the intended severity category. Words that did not clearly reflect pain severity or that were redundant were flagged for removal. Importantly, all filtered results were manually reviewed by a medical expert before finalizing the dictionaries. The medical expert examined each retained term to ensure that it accurately described osteoarthritis-related pain severity in a clinical context. Only those terms deemed clinically appropriate were included in the final PLWDs.The final PLWD contained 32, 31, and 37 synonyms indicating mild, moderate and severe pain levels. Examples of synonyms are shown in Table. 2. Table 2: Examples of words from the 3 PLWDs indicating the three pain levels PLWD Example words from the dictionary Mild low slight mild small tingling Moderate aching tearing hurting medium sharp Severe beating grounding lancinating crushing heavy Words in the PLWDs were used as keywords to search for and extract semantically similar words from the chart notes indicating the three pain levels. 4.2.2 Computing Centroids of PLWDs: The PLWDs are instrumental for the SPaDe algorithm to categorize the chart notes to different pain levels. We applied the word embedding method to convert all the words in the PLWDs into numeric vectors and then calculated the centroids (Lcentroid,Mcentroid,HcentroidL^centroid,M^centroid,H^centroid) as the average of these vectors for each PLWD as illustrated in Algorithm 1. In the algorithm, L, M, and H correspond to ’Low’, ’Moderate’, and ’High’, denoting the set of words in text format from mild, moderate, and severe pain categories respectively. In lines 2-6, the words are transformed into word vectors using GLOVE embedding. Lines 7-9 calculates the centroids. These centroids serve as the reference points in determining the proximity of words in the chart notes to each PLWD or pain level for categorizing the chart notes. Algorithm 1 Centroid Detection for the PLWDs 1: procedure Centroid(L,M,H) 2: Initialize three empty array: Lglove,Lglove,LgloveL^glove,L^glove,L^glove 3: for Lk,Mk,HkL_k,M_k,H_k in L,M,HL,M,H do 4: Lglove.append(GLOVE(Lk))L^glove.append(GLOVE(L_k)) 5: Mglove.append(GLOVE(Mk))M^glove.append(GLOVE(M_k)) 6: Hglove.append(GLOVE(Hk))H^glove.append(GLOVE(H_k)) 7: end for 8: Lcent=Average(Lglove,axis=0)L^cent=Average(L^glove,axis=0) 9: Mcent=Average(Mglove,axis=0)M^cent=Average(M^glove,axis=0) 10: Hcent=Average(Hglove,axis=0)H^cent=Average(H^glove,axis=0) 11: return Lcent,Mcent,HcentL^cent,M^cent,H^cent 12: end procedure 4.2.3 Extracting & Grouping Pain Related Words: We developed an algorithm to search the combined chart note for synonyms or semantically similar words given the words in each PLWD, and extracted them to build mild, moderate, and severe word clusters for each patient as presented in Algorithm 2. The similarity was calculated based on the proximity of a word vector to the centroids of the 3 PLWDs. The combined chart note refers to the OA related notes extracted from multiple visits, which were then combined, preprocessed, and embedded for each patient. The algorithm calculated the Cosine Similarity of each word vector NwN_w in the processed note N to the three centroids (mild, moderate, and severe) using Eq.1 and assigned NwN_w to the cluster of the closest centroid. A larger cosine similarity value indicates a stronger semantic relationship and the closest centroid. similarity(,)=⋅‖‖=∑i=1naibi∑i=1nai2∑i=1nbi2,similarity(a,b)= a·b\|a\|\|b\|= _i=1^na_ib_i _i=1^na_i^2 _i=1^nb_i^2, (1) where w represents wthw^th word in the note, ⋅· denotes the dot product, ‖\|a\| and ‖\|b\| denote the Euclidean norms of vectors a and b, respectively, and n is the dimensionality of the vectors. Algorithm 2 Synonym-based Pain Level Grouping Algorithm 1: procedure Grouping(N,L,M,H) 2: Initialize three cluster to 0: CL,CM,CHC^L,C^M,C^H 3: Lcentroid,Mcentroid,Hcentroid=Centroid(L,M,H)L^centroid,M^centroid,H^centroid=Centroid(L,M,H) 4: for NwN_w in N do 5: SNw,Lcentroid=Similarity(Nw,Lcentroid)S^N_w,L^centroid=Similarity(N_w,L^centroid) 6: SNw,Mcentroid=Similarity(Nw,Mcentroid)S^N_w,M^centroid=Similarity(N_w,M^centroid) 7: SNw,Hcentroid=Similarity(Nw,Hcentroid)S^N_w,H^centroid=Similarity(N_w,H^centroid) 8: if SNw,LcentroidS^N_w,L^centroid is the largest & SNw,Lcentroid>TsimilarityS^N_w,L^centroid>T^similarity then 9: CLC^L += 1 10: else if SNw,McentroidS^N_w,M^centroid is the largest & SNw,Mcentroid>TsimilarityS^N_w,M^centroid>T^similarity then 11: CMC^M += 1 12: else if SNw,Hcentroid>TsimilarityS^N_w,H^centroid>T^similarity then 13: CHC^H += 1 14: end if 15: end for 16: TotalWords=CL+CM+CHTotalWords=C^L+C^M+C^H 17: CL=CLTotalWordsC^L= C^LTotalWords 18: CM=CMTotalWordsC^M= C^MTotalWords 19: CH=CHTotalWordsC^H= C^HTotalWords 20: return CLC^L, CMC^M, CHC^H 21: end procedure Lines 3-14 in Algorithm 2 show the similarity computation and assignment of words to the cluster of the closest centroid. Here, CLC^L, CMC^M, and CHC^H denote the mild, moderate, and severe cluster centroids respectively. A threshold value TsimilarityT^similarity was used to filter out word vectors that are too far (very low Cosine Similarity) from all the centroids. Based on the experiment, we found that lower thresholds (e.g., <0.6) allowed the inclusion of numerous irrelevant words, which introduced noise into the results. In contrast, higher thresholds (e.g., >0.6) excluded a significant number of words, including some that were relevant but expressed with slight variations in phrasing or language. As a result, we selected a threshold value of 0.6. Word vectors were included only if the similarity was greater than a predefined threshold value of 0.6 since they did not provide any useful information or were not relevant to pain expressions. Lines 15-18 computes the distribution of the 3 categories of word vectors for each patient’s combined and processed chart note data denoted by N. We observed that the mild pain level words occurred most frequently in most notes, followed by the moderate and then the severe pain-level words, respectively. It indicates that mild pain levels were more commonly reported than higher pain levels. 4.2.4 Unsupervised Clustering of Extracted Words: The large frequency of words indicating mild pain level created a problem in identifying the optimal thresholds of word frequency based on which notes could be categorized into mild, moderate, or severe pain levels. Therefore, we developed an unsupervised clustering approach using the K-Means clustering algorithm [36] to determine the centroids of all the word vectors extracted from the whole study sample as shown in Algorithm 3. We chose K-Means because our goal was to partition pain descriptions into stable, severity-consistent groups in a continuous embedding space and K-Means provides deterministic, interpretable centroids that reflect prototypical severity expressions. Since pain expressions in clinical notes are noisy and subjective, SPaDe cannot guarantee perfect severity inference for every expression. Using a highest-severity aggregation would make the method overly sensitive, as a single misclassified high-severity expression could dominate the patient-level label. To address this, we assign patient severity based on the most frequently inferred severity category across all chart notes, trading sensitivity for robustness and reducing the impact of occasional misclassifications. NotesNotes indicate all the notes of all patients in the study sample and NnN_n represents each patient’s note. Each note is transformed into a 3 dimensional vector representation using Algorithm 2, where each dimension represents the percentage of mild, moderate, and severe word vectors respectively in that note. Then the set of all these 3D vectors from all notes in the study sample are passed on to a K-Means clustering algorithm which divides them into 3 clusters indicating mild, moderate, and severe pain. Algorithm 3 Pain-Level Detector using K-Means Clustering 1: procedure Detector(NotesNotes,L,M,H) 2: Initialize an empty array: P 3: for NnN_n in NotesNotes do 4: CL,CM,CH=Grouping(Nn,L,M,H)C^L,C^M,C^H=Grouping(N_n,L,M,H) 5: P.append([CL,CM,CH])P.append([C^L,C^M,C^H]) 6: end for 7: prediction=KMeans(P,n_cluster=3)prediction=KMeans(P,n\_cluster=3) 8: return predictionprediction 9: end procedure The KMeans algorithm works by randomly selecting three initial centroids. Then, it iteratively assigns each data point to its nearest centroid. After each epoch, it updates the centroids by averaging the data points in each cluster. This process is repeated until the centroids converge, which means that the data points no longer change their assigned cluster. We conducted experiments with various randomly selected initial centroids, and the final outcomes were found to be similar. One example of final centroids learned by the K-Means clustering algorithm are shown in TABLE 3. Table 3: Centroid of KMeans Clustering Algorithm Mild Moderate Severe 0 0.61 0.26 0.13 1 0.51 0.28 0.21 2 0.48 0.38 0.14 The rows in the result table show 3 centroids in the extracted word vector space. Values in the columns for each row indicate the average percentage occurrence of mild, moderate, and severe pain level words in the respective cluster. Row 0 shows the highest value for mild words in the first column. Thus it represents the centroid for mild category. Similarly, row 1 has the highest value in severe column and therefore, represents the severe category. Row 2 has the highest value in the moderate column and denotes the moderate category. Thus all notes (denoted by the transformed word vectors CLC^L, CMC^M, and CHC^H) in the respective clusters of centroids 0, 1, and 2 are categorized as mild, severe, and moderate pain level. We decided a priori to focus on detecting individuals with moderate to severe pain. Our preliminary work on SPaDe showed that often similar words are used to describe moderate to severe pain and it is difficult to separate these two levels accurately. Pain is a subjective experience and different expressions are used to describe the pain in the chart note based on patients’ descriptions. Even human experts struggled to decide about moderate and severe pain levels from some of the chart notes. Therefore, we decided to develop the proof-of-concept tool from this pilot study to focus on two categories instead of three. Simplifying the analysis allowed us to streamline the process. 4.3 Approach I: Pain Scale-based Approach We preprocessed the data to extract valid pain level entries in the “score/range” format (e.g., “6/10”) while excluding unrelated patterns such as dates. All the processing were performed on the same 797 study sample. Algorithm 4 Pain Scale-based Categorization 1: procedure painscale(NnN_n) 2: Initialize one empty array: S 3: Initialize one array that contains severity related keywords: K 4: ScorelistScore^list = re.findall(``[0−9]+/[0−9]"``[0-9]+/[0-9]", NnN_n) 5: for ScorewlistScore^list_w in ScorelistScore^list do 6: Extract the number before “/” as ScoreScore 7: Extract the number after “/” as RangeRange 8: if int(Range)int(Range) == 1010 & int(Score)<=10int(Score)<=10 & ScoreScore does not prefix with “0” then 9: S.append(int(Score)int(Score)) 10: end if 11: end for 12: for KwK_w in K do 13: if KwK_w in N || max(S)>=5max(S)>=5 then 14: return ”Moderate-to-Severe Pain” 15: end if 16: end for 17: return ”Mild Pain” 18: end procedure For categorizing patients to mild and moderate-to-severe pain groups, we applied a rule-based approach. Patients whose pain scores exceeded 5 were classified as individuals experiencing a moderate-to-severe pain level, and others were placed in the mild pain group. The details of the pain scale-based algorithm are given in Algorithm 4. 4.4 Approach I: Medication-based Approach The medication records of 2,044 patients were extracted from the structured EMR data and stored in a separate CSV file in the format of the chart note data, with each record linked to the chart note by a unique patient ID. Each data row in the CSV file contains a patient ID and the list of medications prescribed to that patient for each physician encounter. Given the context of a patient having multiple records within the original dataset, we used the unique ID assigned to each OA patient. By employing this method, we combined the medications from various records pertaining to each patient into a singular record. This combined record now comprises an extensive list of prescribed medications and alternative treatments. Then all patients’ records from the 797 study sample were compiled into a multi-row dataset with each row representing a different patient ID with the corresponding medication information. The patient raw data was passed to the algorithm we developed for medication-based pain level categorization as listed in Algorithm 5. We also compiled a list of analgesics, i.e. all pain medications and treatments considered in this study were sent as the 2nd2^nd parameter to the algorithm. The analgesics were categorized based on pain levels and compiled as a dictionary of medications labeled with pain levels. This was passed to the algorithm as the 3rd parameter. The algorithm 5 delineates the procedural steps employed for extracting medication information from the structured notes and subsequently mapping them to pain levels. We first mapped pain levels to 797 patients in the study samples, and then we selected the same 156 patients in the final gold standard samples and evaluated the results. Algorithm 5 Medication-based Approach 1: procedure medication(PatientDataPatientData, AnalgesicsAnalgesics, MedDictMedDict) 2: Initialize one empty dictionary: D2D2 3: Initialize one empty dictionary: D3D3 4: for patientID,patientMedListpatientID,patientMedList in PatientDataPatientData do 5: for medicationmedication in patientMedListpatientMedList do 6: for analgesic,chemicalanalgesic,chemical in AnalgesicsAnalgesics do 7: if medicationmedication in analgesicanalgesic or medicationmedication in chemicalchemical) then 8: D2[patientID]D2[patientID].append(medicationmedication) 9: end if 10: end for 11: end for 12: for patientID,medicationpatientID,medication in D2D2 do 13: for dictPainlevel,dictMeddictPainlevel,dictMed in MedDictMedDict do 14: if medicationmedication in dictMeddictMed then 15: D3[patientID]=dictPainlevelD3[patientID]=dictPainlevel 16: end if 17: end for 18: end for 19: end for 20: return D2, D3 21: end procedure Table 4: List of Analgesics Per Class. Class Chemical Name ACETAMINOPHEN ACETAMINOPHEN TOPICAL NSAIDs DICLOFENAC ORAL NSAID ORAL NSAID, CELECOXIB, DICLOFENAL, IBUPROFEN, INDOMETHACIN, KETOPROFEN, KETOROLAC, MEFENAMIC ACID, MELOXICAM, NAPROXEN, PIROXICAM, SULINDAC, TIAPROFENIC ACID STRONG OPIOIDS CODEINE, BUPRENORPHINE, FENTANYL, OXYCODONE, HYDROCODONE, HYDROMORPHONE, MORPHINE, METHADONE, TAPENTADOL TRAMADOL TRAMADOL DULOXETINE DULOXETINE Next, we categorized the medications in the list of analgesics based on pain levels into a dictionary of medications. In accordance with the guidelines of WHO Analgesic Ladder33 3 https://w.ncbi.nlm.nih.gov/books/NBK554435/, all analgesic chemicals can be classified into three categories: mild, moderate, and severe pain. To categorize patients into mild and moderate-to-severe pain levels, we combined the moderate pain category with the severe pain category to establish the moderate-to-severe pain category as shown in Table 5. 1. Mild pain: Non-opioid analgesics with or without adjuvants. 2. Moderate Pain: Weak opioids (hydrocodone, codeine, tramadol) with or without non-opioid analgesics and with or without adjuvants. 3. Severe and persistent Pain: Potent opioids (morphine, methadone, fentanyl, oxycodone, buprenorphine, tapentadol, hydromorphone, oxymorphone) with or without non-opioid analgesics, and with or without adjuvants. Table 5: Dictionary of Medications Categorized based on Pain Levels Pain Levels Medications Mild ACETAMINOPHEN, DICLOFENAC, CELECOXIB, DICLOFENAL, DULOXETINE FLURBIPROFEN, IBUPROFEN, INDOMETHACIN, KETOPROFEN, KETOROLAC, MEFENAMIC ACID, MELOXICAM, NAPROXEN, PIROXICAM, SULINDAC, TIAPROFENIC ACID Moderate-to-Severe TRAMADOL, HYDROCODONE, CODEINE, BUPRENORPHINE, FENTANYL, OXYCODONE HYDROCODONE, HYDROMORPHONE, MORPHINE, METHADONE, TAPENTADOL Other Forms of Treatments: The lists of medications prescribed to patients were for treating numerous health problems and not only pain. To extract and consider all pain related treatments, we formulated a dictionary of keywords specifically designed to discern OA-related surgical procedures, such as ”knee surgery” and ”hip surgery,” along with OA-related mobility aids like ”walker” and ”cane”, and referral-related information, such as ”pain clinic”. Categorize Patients: If any of the medications prescribed to a patient appeared under the category of moderate-to-severe pain level, or contained any of the keywords indicating other form of treatments for moderate-to-severe pain, then the patient was categorized under moderate-to-severe pain level, and otherwise under the mild pain level as shown in lines 11-17 in Algorithm 5. 4.5 Integration Our system employs an integrated approach called PLeDO as shown in Figure 1 that combines the results of Approach I, I, and I to generalize across most chart notes, even when specific tools or scales are unavailable. While the pain scale-based systems are limited in their coverage of all possible tools and their mentions in chart notes, the integration of expression-based and medication-based methods ensures broader applicability. By leveraging this multi-method approach, we enhance the accuracy and generalizability of pain level detection across diverse chart notes. To combine the outcomes, we checked if any of the approaches indicated a moderate-to-severe pain level for a patient. If so, the patient was labeled as a moderate-to-severe pain level patient in the final assessment. Otherwise, the patient was labeled as a mild pain level patient. Table 6: Manual evaluation results and ablation study. Accuracy Precision (PPV) Recall (Sensitivity) Specificity F1 AUROC Pain Scale-based Approach 0.551 0.572 0.589 0.589 0.536 0.589 Medication-based Approach 0.551 0.560 0.575 0.575 0.532 0.575 SPaDe 0.577 0.541 0.549 0.549 0.533 0.549 PLeDO - Pain Scale-based Approach + Medication-based Approach + Pain Expression-based Approach 0.551 0.566 0.582 0.582 0.534 0.582 PLeDO + Pain Scale-based Approach - Medication-based Approach + SPaDe 0.622 0.547 0.552 0.552 0.548 0.551 PLeDO + Pain Scale-based Approach + Medication-based Approach - SPaDe 0.660 0.601 0.614 0.614 0.604 0.614 PLeDO + Pain Scale-based Approach + Medication-based Approach + SPaDe 0.660 0.605 0.621 0.621 0.608 0.621 * ”+” indicates inclusion, and ”-” indicates exclusion. 5 Validation and Results The results from experimental validation of our approaches in categorizing of the patients into mild and moderate-to-severe pain levels are discussed below. 5.1 Validation We used the gold standard data to validate our approaches. Rather than limiting the application of our clustering method to the 156 gold validation datasets, we applied it across all 797 data points. This strategy ensures that, in instances of encountering unseen data, we can easily classify it into the appropriate category by leveraging our pre-trained centroids. The results from the unsupervised clustering approach of unlabeled data was validated using the manually labeled gold standard data. The same gold standard data was used to validate the medication-based and the pain scale based approaches. However, the manual evaluation did not consider the rigorous inspection of the medication and pain scale data. We report the validation of each approach in the results table. Evaluation metrics such as accuracy, precision (PPV), recall (sensitivity), specificity, F1 score, and AU-ROC were calculated for the gold standard validation dataset to report the performance. 5.2 Results We present the pain level categorization results in TABLE 6. A simple statistical analysis was also performed to see the gender and age distribution of the patients in our sample dataset with 797 patients as shown in Figure 3. 5.3 Ablation Study To provide an ablation study, we gradually combined multiple approaches as each approach was implemented as an independent parallel system with no dependency on the other systems. Finally, we combined all three approaches to present the results of our integrated approach PLeDO, which achieved the best performance. With additional information from the other approaches, the results improved a little. Figure 3: Gender and Age based Analysis of the Study Population. 5.4 Observations PLeDO achieved the best accuracy of 0.66 with an F1-score of 0.608 and an AU-ROC of 0.621. The ablation study demonstrates the significance of each component’s inclusion on the final outcomes. Analysis of the population distribution shows the following trend (see Figure 3). • There are more female patients with OA patients compared to male patients. • Patients of age between 60-79 are more likely to have OA. 5.5 Discussion 5.5.1 Key Finding: This study allowed us to explore the quality of EMR notes and treatment patterns for OA pain in primary care setting in Canada. The aim of this study was to see if pain expressions recorded in the chart notes are consistent with the treatment patterns. The following are the key takeaways from this study. Exploration of pain related expressions from the medical chart notes proved to be extremely challenging as it varies widely based on individual nature which is influenced by demography, culture, situation, and objective assessments for reporting. Our exploration of synonyms for mild, moderate, and severe pain using online tools as described in Section 4.2.4 returned words that were in some cases very distantly related to pain expressions. By changing the number of words in each of PLWDs, we can get different results which may be explored in future studies. Too few words will evade some of the expressions found in the chart notes and too many words will introduce confusion and noise in the process. We performed many iterations testing with different sets of words in the PLWDs before finalizing the lists of words. Also the expressions can change based on treatments as the pain conditions improve or deteriorate. We considered the highest level of pain expression for each patient’s medical chart notes spanning multiple years. It is possible to categorize each note but not all notes have content indicating OA pain. Selecting good notes with pain information can introduce bias. Comorbidity in patients is another source of noise in the data. Pain can be the result of many different health problems and isolating OA related pain reports is a challenging task. Also, when other family members were mentioned or family history was narrated by the patient, it got more difficult to extract accurate information about the patient from the text data. We extracted paragraphs from the chart notes using OA related keywords but this list is not exhaustive. Any keyword related information extraction heavily relies on the set of keywords used for extracting the information and can either provide too much or too little information based on the number and quality of the keywords. For the pain scale-related approach, we found that a standard practice was not followed in the data, and the scales were not always used to report pain levels. According to the results, its standalone performance was comparable to the medication-based approach but slightly lower than SPaDe. Therefore, if a standardized pain scale were consistently used, it may lead to improved and more reliable reporting of pain levels. Regarding medications, pain is a common problem for many diseases, and the medications are also general for all types of pain management. So, for OA specific pain, medications are not the best way to categorize the pain level. We considered all medications prescribed to the patients without filtering which medication was given for OA pain management. Therefore, we see that the results are not great. Regarding medications, pain is a common problem for many diseases and the medications are also general for all types of pain management. So, for OA specific pain, medications are not the best way to categorize the pain level. We considered all medications prescribed to the patients without filtering which medication was given for OA pain management. Therefore, we see that the results are not great. When comparing individual approaches, SPaDe demonstrates the highest accuracy but exhibits lower precision and recall. This outcome stems from SPaDe’s strong performance in predicting the majority class (moderate-to-severe pain level), which constitutes a significant portion of the dataset (111 of 156 samples). However, accuracy alone does not provide insight into class-specific performance. A model can achieve high accuracy by excelling with the majority class while underperforming on the minority class (mild pain level, with only 45 samples). Lower precision suggests a higher rate of false positives, while lower recall (sensitivity) indicates a higher rate of false negatives. Interestingly, the ablation study shows that removing the pain scale-based approach results in the largest drop in accuracy, indicating its strong influence within the integrated framework. At the same time, removing SPaDe does not reduce overall accuracy in the combined setting. This highlights the strength of SPaDe, which, as an unsupervised clustering method, achieves substantial prediction consistency with both the scale-based (rule-based) approach and the medication-based approach that relies on human-defined rules and annotations. SPaDe’s generalizability is noteworthy. Unlike the pain scale-based approach, which depends on the presence of specific scales in chart notes, or the medication-based approach, which cannot cover all possible medications, SPaDe does not rely on such predefined elements. This makes it more flexible and adaptable to diverse datasets. Finally, the combination of all 3 approaches gives better results than any single approach. However, pain level detection is a difficult problem to address just using pain expressions from the unstructured data. Further exploration is needed to perhaps select a more specific population to focus on only pain management. 5.5.2 Comparison with Existing Literature: Most previous research in this domain has focused on supervised learning using image data [2][51][59], text data [47][48], or numerical data [45][49], often requiring extensive expert annotation. However, obtaining high-quality labels for medical data is both costly and time-consuming, which limits the scalability of such approaches. In recent years, large language models [9][11][19] have demonstrated exceptional performance across various domains. However, their propensity for hallucination and lack of interpretability present significant challenges for applications in sensitive areas such as the medical domain. Additionally, their substantial computational demands raise concerns about scalability and efficiency, further limiting their practicality in resource-constrained settings. This study was conducted in a secure setting without any internet access based on the established policies governing any work with real patients’ data. This limited our ability to download, fine-tune and examine some of the large language models. Furthermore, there were only few data points which were insufficient to train some of the large models. We addressed the above limitations by developing an unsupervised clustering method, which generated predictions without relying on training labels. This approach enhances robustness, particularly when dealing with noisy data such as physicians’ chart notes. Although clustering methods have been applied to pain level identification in prior research [49], these studies only focus on numerical data, which is less complex and challenging compared to unstructured textual data. Additionally, we investigate the correlation between medication usage and pain levels to determine whether this relationship can further improve pain level identification. By combining structured medication data with unstructured text from chart notes, our system offers a comprehensive and innovative solution for pain identification in OA patients. 5.5.3 Limitation: This study is conducted using real-world clinical data from a single regional primary care network. As documentation practices, prescribing patterns, and patient characteristics may vary across regions, the findings may reflect patterns specific to this dataset and introduce regional bias in the results. All EMR data used in this study were highly sensitive. For privacy and regulatory compliance, all experiments were conducted in a secure air-gapped environment with no external network access. This restriction limited the use of externally hosted large language models and certain high-compute methods, as such tools may introduce potential risks of sensitive information exposure. 6 Conclusion Pain is a critical health problem that can greatly aggravate the quality of life of OA patients. Pain can have many subjective descriptions which make automated categorization of patients’ chart notes based on pain severity a very challenging problem. Understanding the treatment pattern for different pain levels can lead to improved patient care and the discovery of new drugs. This study demonstrates the potential of using NLP and ML techniques in analyzing unstructured clinical data to extract valuable insights about pain severity among primary care patients diagnosed with OA. We developed 3 different approaches to pain level categorization using primary care unstructured chart note data and structured medication data. The main synonym based pain level detection approach, SPaDe, finds pain related expressions from chart notes using semantic similarity matching technique with the help of 3 pain level word dictionaries and then applies unsupervised clustering to group these expressions and the corresponding notes to mild and moderate-to-severe pain categories. The other two approaches, using medication and pain scale also provides similar accuracy. But our integrated approach, PLeDO, was able to achieve the best result. In the future, we plan to investigate more focused osteoarthritis patient populations with fewer comorbidities and explore advanced clinical NLP tools such as cTAKES and MetaMap [39, 40] to improve concept extraction from unstructured EMR notes. As larger annotated datasets become available, supervised and semi-supervised approaches, including domain-specific transformer models such as MedicalBERT [41], may be evaluated and compared with the proposed framework. We also plan to incorporate additional structured clinical variables, explore alternative clustering methods, and investigate privacy-preserving locally deployable foundation models. Finally, multimodal approaches that combine radiology images [42] with chart-note data may further improve osteoarthritis pain severity detection and characterization. References [1] Arthritis Alliance of Canada, ”The Impact of Arthritis in Canada: Today and Over the Next 30 years,” Fall 2011. [2] B. Guan, F. Liu, A.H. Mizaian, S. Demehri, A. Samsonov, A. Guermazi, and R. Kijowski (2022). Deep learning approach to predict pain progression in knee osteoarthritis. Skeletal Radiology, 1–11. [3] B. Sharif, R. Garner, D. Hennessy, C. Sanmartin, W. M. Flanagan and D. A. Marshall, ”Productivity costs of work loss associated with osteoarthritis in Canada from 2010 to 2031.,” Osteoarthritis and cartilage, 25(2), p. 249–258., 2017. [4] G. Hawker, ”Osteoarthritis is a serious disease,” Clinical and experimental rheumatology, 37 Suppl 120(5), p. 3–6, 2019. [5] World Health Organization, Department of Chronic Diseases and Health Promotion, ”Chronic rheumatic conditions,” 2021. [Online]. Available: https://w.who.int/chp/topics/rheumatic/en/. [6] J. E. Tarride, M. Haq, D. J. O’Reilly, J. M. Bowen, F. D. L. Xie and R. & Goeree, ”The excess burden of osteoarthritis in the province of Ontario, Canada.,” Arthritis and rheumatism, vol. 64, no. 4, p. 1153–1161, 2012. [7] T. Neogi, ”The epidemiology and impact of pain in osteoarthritis,” Osteoarthritis and cartilage, vol. 21, no. 9, p. 1145–1153, 2013. [8] J. Bedson and P. R. Croft, ”The discordance between clinical and radiographic knee osteoarthritis: a systematic search and summary of the literature,” BMC musculoskeletal disorders, vol. 9, no. 116, 2008. [9] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, and D. Bikel (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. [10] T. W. O’Neill and D. T. Felson, ”Mechanisms of Osteoarthritis (OA) Pain,” Current osteoporosis reports, vol. 16, no. 5, p. 611–616, 2018. [11] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, and J. Schulman (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. [12] R. R. Bannuru, M. C. Osani, E. E. Vaysbrot, N. K. Arden, K. Bennell and …, ”OARSI guidelines for the non-surgical management of knee, hip, and polyarticular osteoarthritis,” Osteoarthritis and cartilage, vol. 27, no. 11, p. 1578–1589, 2019. [13] S. L. Kolasinski, T. Neogi, M. C. Hochberg, C. Oatis, G. Guyatt and …, ”2019 American College of Rheumatology/Arthritis Foundation Guideline for the Management of Osteoarthritis of the Hand, Hip, and Knee,” Arthritis care & research,, vol. 72, no. 2, p. 149–162, 2020. [14] S. Perrot, ”Osteoarthritis pain,” Best Pract Res Clin Rheumatol, vol. 29, no. 1, p. 90-97, 2015. [15] G. A. Hawker, L. Stewart, M. R. French, J. Cibere, J. M. Jordan, L. March, M. Suarez-Almazor and R. Gooberman-Hill, ”Understanding the pain experience in hip and knee osteoarthritis - an OARSI/OMERACT initiative,” Osteoarthritis and cartilage, vol. 16, no. 4, p. 415–422, 2008. [16] N. Bellamy, ”Pain assessment in osteoarthritis: experience with the WOMAC osteoarthritis index,” Seminars in arthritis and rheumatism,, vol. 18, no. 4 Suppl 2, p. 14–17, 1989. [17] D. J. Hunter and S. Bierma-Zeinstra, ”Osteoarthritis,” Lancet (London, England), vol. 393, no. 10182, p. 1745–1759., 2019. [18] G. A. Hawker, R. Croxford, A. S. Bierman, P. J. Harvey, B. Ravi, I. Stanaitis and L. L. & Lipscombe, ”All-cause mortality and serious cardiovascular events in people with hip and knee osteoarthritis: a population based cohort study.,” PloS one, vol. 9, no. 3, p. e91286, 2014 [19] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F.L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, and R. Avila (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774. [20] W. Chen, K. Wei, W. Zhao and X. Zhou, ”Estimation of Key Comorbidities for Osteoarthritis Progression Based on the EMR-Claims Dataset,” IEEE Access, vol. 7, p. 72431-72442, 2019. [21] R. Birtwhistle, R. Morkem, G. Peat, T. Williamson, M. E. Green, S. Khan and K. P. Jordan, ”Prevalence and management of osteoarthritis in primary care: an epidemiologic cohort study from the Canadian Primary Care Sentinel Surveillance Network,” CMAJ open, vol. 3, no. 3, p. E270–E275, 2015. [22] Canadian Primary Care Sentinel Surveillance Network (CPCSSN), 2021. [Online]. Available: https://cpcssn.ca/. [23] T. Williamson, M. Green, R. Birtwhistle, S. Khan, S. Garies, S. Wong, N. Natarajan, D. Manca and N. Drummond, ”Validating the 8 CPCSSN case definitions for chronic disease surveillance in a primary care database of electronic health records,” Ann Fam Med, vol. 12, p. 367–372, 2014. [24] A. Kadhim-Saleh, M. Green, T. Williamson, D. Hunter and R. Birtwhistle, ”Validation of the diagnostic algorithms for 5 chronic conditions in the Canadian Primary Care Sentinel Surveillance Network (CPCSSN): a Kingston Practice-based Research Network (PBRN) report,” Journal of the American Board of Family Medicine, vol. 26, no. 2, p. 159–167., 2013. [25] McMaster University Family Medicine, ”OSCAR EMR,” 2021. [Online]. Available: https://oscar-emr.com/. [Accessed 17 March 2021]. [26] E.Loper and S.Bird. ”NLTK: The Natural Language Toolkit.” Proceedings of the ACL-02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics. 2002. [27] J.Pennington, R.Socher, and C.Manning. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics. [28] MY.Landolsi, L.Hlaoua, and LB.Romdhane. ”Information extraction from electronic medical documents: state of the art and future research directions.” Knowledge and Information Systems 65.2 (2023): 463-516. [29] H. Zafari, J. Li, F. Zulkernine, L. Kosowan, and A. Singer, “Predictive Modeling of Diabetes using EMR Data.” [Online]. Available: https://orcid.org/ [30] H. Zafari, S. Langlois, F. Zulkernine, L. Kosowan, and A. Singer, “Predicting Chronic Obstructive Pulmonary Disease from EMR data,” in 2020 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology, CIBCB 2020, Oct. 2020. [31] H. Zafari et al., “Using Deep Learning with Canadian Primary Care Data for Disease Diagnosis,” in Deep Learning for Biomedical Data Analysis, Springer International Publishing, 2021, p. 273–310. [32] H. Zafari, L. Kosowan, F. Zulkernine, and A. Signer, “Diagnosing post-traumatic stress disorder using electronic medical record data,” Health Informatics J, vol. 27, no. 4, Oct. 2021. [33] M.Abdalla, M.Abdalla, F.Rudzicz, and G.Hirst ”Using word embeddings to improve the privacy of clinical notes.” Journal of the American Medical Informatics Association 27.6 (2020): 901-907. [34] H.Zhang, AX.Lu, M.Abdalla, M.McDermott, and M.Ghassemi. ”Hurtful words: quantifying biases in clinical contextual word embeddings.” proceedings of the ACM Conference on Health, Inference, and Learning. 2020. [35] I.Pépin, and F.Zulkernine. ”A Comparative Study of De-identification Tools to Apply to Free-Text Clinical Notes.” Proceedings of the 32nd Annual International Conference on Computer Science and Software Engineering. 2022. [36] J. A. Hartigan and M. A. Wong ”Algorithm AS 136: A k-means clustering algorithm.” Journal of the royal statistical society. series c (applied statistics) 28.1, 1979: 100-108. [37] T. Brown, B. Mann, N. Ryder, and et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020, 33: 1877-1901. [38] L. Ouyang, J. Wu, X. Jiang, and et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 2022. [39] GK. Savova, J. Masanz, PV. Ogren, J.Zhang, S. Sohn, KCK. Schuler and CG. Chute. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications. J Am Med Inform Assoc. 2010;17(5):507-513.doi:10.1136/jamia.2009.001560 [40] AR. Aronson. Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program. Proc AMIA Symp. 2001;17-21. [41] L. Rasmy, Y. Xiang, Z. Xie, and D. Zhi. ”Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.” NPJ digital medicine 4.1 (2021): 86. [42] L. Shamir, SM. Ling, W. Scott, A. Bos, N. Orlov, TJ. Macura, DM. Eckley, L. Ferrucci, and LG. Goldberg. Knee x-ray image analysis method for automated detection of osteoarthritis. IEEE Trans Biomed Eng. 2009;56(2):407-415. doi:10.1109/TBME.2008.2006025 [43] S.Robertson. ”Understanding inverse document frequency: on theoretical arguments for IDF.” Journal of documentation 60.5 (2004): 503-520. [44] S.Datta, J.Posada, G.Olson, W.Li, C.Reilly, D.Balraj, J.Mesterhazy, J.Pallas, P.Desai, and NH.Shah. A new paradigm for accelerating clinical data science at Stanford Medicine, March 2020. [45] J.Lötsch, A. Ultsch, and E. Kalso. 2017. Prediction of persistent post-surgery pain by preoperative cold pain sensitivity: biomarker development with machine-learning-derived analysis. BJA: British Journal of Anaesthesia, 119(4), 821-829. [46] L.DiMartino, T.Miano, K.Wessell, B.Bohac and LC.Hanson, 2022. Identification of uncontrolled symptoms in cancer patients using natural language processing. Journal of pain and symptom management, 63(4), p.610-617. [47] JA.Hughes, Y.Wu, L.Jones, C.Douglas, N.Brown, S.Hazelwood, AL.Lyrstedt, R.Jarugula, K.Chu, and A.Nguyen. 2024. Analyzing pain patterns in the emergency department: Leveraging clinical text deep learning models for real-world insights. International Journal of Medical Informatics, 190, p.105544. [48] A. Alambo, R. Andrew, S. Gollarahalli, and et al. 2020. Measuring pain in sickle cell disease using clinical text. In 42nd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 5838–5841. IEEE. [49] J. Gálvez-Goicurĺa, J. Pagán, A.B. Gago-Veiga, J.M. Moya, and J.L. Ayala. 2021. Cluster-then-classify methodology for the identification of pain episodes in chronic diseases. IEEE Journal of Biomedical and Health Informatics, 26(5), 2339–2350. [50] J.O. Pinzon-Arenas, Y. Kong, K.H. Chon, and H.F. Posada-Quintero.2023. Design and evaluation of deep learning models for continuous acute pain detection based on phasic electrodermal activity. IEEE Journal of Biomedical and Health Informatics. [51] H. Zhu, Y. Zhao, X. Chen, F. Luo, L. Mei, S. Chen, and Y. Pan.2023. Video-based neonatal pain assessment in uncontrolled conditions. IEEE Journal of Biomedical and Health Informatics. [52] M. Moradi and M. Samwald. 2021. Explaining black-box models for biomedical text classification. IEEE Journal of Biomedical and Health Informatics, 25(8), 3112–3120. [53] N. Reimers and I. Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982–3992. Hong Kong, China: Association for Computational Linguistics. [54] J. Hughes, Y. Wu, L. Jones, C. Douglas, N. Brown, S. Hazelwood, A. Lyrstedt, R. Jarugula, K. Chu, and A. Nguyen. 2024 ”Analyzing pain patterns in the emergency department: Leveraging clinical text deep learning models for real-world insights.” International Journal of Medical Informatics 190 : 105544. [55] Y. Zhuang, Y. Guo, Y. Li, Y. Wu, P. Yu, T. Song, Z. Wang, K. Zhou, W. Wang, and L. Zhuang. 2026 AI-Driven Prediction of Cancer Pain Episodes: A Hybrid Decision Support Approach. IEEE Journal of Biomedical and Health Informatics. [56] S. Chen, M. Guevara, N. Ramirez, A. Murray, J. Warner, H. Aerts, T. Miller, G. Savova, R. Mak, and D. Bitterman.2023. Natural language processing to automatically extract the presence and severity of esophagitis in notes of patients undergoing radiotherapy. JCO Clinical Cancer Informatics, 7, e2300048. [57] B. Patra, L. Lepow, P. Kumar, and et al. 2025. Extracting social support and social isolation information from clinical psychiatry notes: comparing a rule-based natural language processing system and a large language model. Journal of the American Medical Informatics Association, 32(1), 218-226. [58] S. Lagisetty, P. Devarajulu, and A. Moka.2025. AI-Enhanced Telehealth Platforms: A Comprehensive Analysis of Automated Triage and Personalized Care Systems. In 2025 5th Intelligent Cybersecurity Conference (ICSC) (p. 370-377). IEEE. [59] V. Daniel, K. Vijayalakshmi, P. Pawar, D. Kumar, A. Bhuvanesh, and A. Christilda. 2024. Enhanced affinity propagation clustering with a modified extreme learning machine for segmentation and classification of hyperspectral imaging. E-Prime-Advances in Electrical Engineering, Electronics and Energy, 9, 100704.