Paper deep dive
A Privacy by Design Framework for Large Language Model-Based Applications for Children
Diana Addae, Diana Rogachova, Nafiseh Kahani, Masoud Barati, Michael Christensen, Chen Zhou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/20/2026, 11:54:32 PM
Summary
The paper proposes a Privacy-by-Design (PbD) framework for Large Language Model (LLM)-based applications targeting children. It maps principles from GDPR, PIPEDA, and COPPA to the LLM lifecycle stages (data collection, training, monitoring, validation) and integrates design guidelines from UNCRC and the UK's AADC. The framework aims to mitigate privacy risks like memorization and profiling through technical and organizational controls, demonstrated via a case study of an educational tutor for children under 13.
Entities (8)
Relation Signals (8)
Privacy-by-Design â appliesto â LLM
confidence 98% · We map these principles to various stages of applications that use Large Language Models (LLMs)
Privacy-by-Design â incorporatesprinciplesfrom â GDPR
confidence 95% · Our framework includes principles from several privacy regulations, such as the General Data Protection Regulation (GDPR)
Privacy-by-Design â incorporatesprinciplesfrom â COPPA
confidence 95% · and the Children's Online Privacy Protection Act (COPPA) from the United States.
Privacy-by-Design â incorporatesprinciplesfrom â PIPEDA
confidence 95% · the Personal Information Protection and Electronic Documents Act (PIPEDA) from Canada
LLM â posesrisk â Privacy Risks
confidence 95% · However, there are growing concerns about privacy risks, particularly for children... LLM-specific vulnerabilities, such as the modelsâ capacity to memorize and inadvertently reveal sensitive details
Educational Tutor â isexampleof â Privacy-by-Design
confidence 92% · we present a case study of an LLM-based educational tutor for children under 13. Through our analysis and the case study, we show that by using data protection strategies... we can support the development of AI applications for children
Privacy-by-Design â includesguidelinesfrom â UNCRC
confidence 90% · drawing from the United Nations Convention on the Rights of the Child (UNCRC)
Privacy-by-Design â includesguidelinesfrom â AADC
confidence 90% · the framework includes design guidelines for children, drawing from ... the UK's Age-Appropriate Design Code (AADC)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Children are increasingly using technologies powered by Artificial Intelligence (AI). However, there are growing concerns about privacy risks, particularly for children. Although existing privacy regulations require companies and organizations to implement protections, doing so can be challenging in practice. To address this challenge, this article proposes a framework based on Privacy-by-Design (PbD), which guides designers and developers to take on a proactive and risk-averse approach to technology design. Our framework includes principles from several privacy regulations, such as the General Data Protection Regulation (GDPR) from the European Union, the Personal Information Protection and Electronic Documents Act (PIPEDA) from Canada, and the Children's Online Privacy Protection Act (COPPA) from the United States. We map these principles to various stages of applications that use Large Language Models (LLMs), including data collection, model training, operational monitoring, and ongoing validation. For each stage, we discuss the operational controls found in the recent academic literature to help AI service providers and developers reduce privacy risks while meeting legal standards. In addition, the framework includes design guidelines for children, drawing from the United Nations Convention on the Rights of the Child (UNCRC), the UK's Age-Appropriate Design Code (AADC), and recent academic research. To demonstrate how this framework can be applied in practice, we present a case study of an LLM-based educational tutor for children under 13. Through our analysis and the case study, we show that by using data protection strategies such as technical and organizational controls and making age-appropriate design decisions throughout the LLM life cycle, we can support the development of AI applications for children that provide privacy protections and comply with legal requirements.
Tags
Links
- Source: https://arxiv.org/abs/2602.17418v1
- Canonical: https://arxiv.org/abs/2602.17418v1
Trouble viewing inline? Open PDF directly â
Full Text
146,191 characters extracted from source content.
Expand or collapse full text
A Privacy by Design Framework for Large Language Model-Based Applications for Children 1 st Diana Addae Systems and Computer Engineering Carleton University Ottawa, Canada dianaaddae@cmail.carleton.ca 2 nd Diana Rogachova School of Information Technology Carleton University Ottawa, Canada dianarogachova@cmail.carleton.ca 3 rd Nafiseh Kahani Systems and Computer Engineering Carleton University Ottawa, Canada nafisehkahani@cunet.carleton.ca 4 th Masoud Barati School of Information Technology Carleton University Ottawa, Canada masoudbarati@cunet.carleton.ca 5 th Michael Christensen Department of Law and Legal Studies Carleton University Ottawa, Canada michaelchristensen@cunet.carleton.ca 6 th Chen Zhou School of Information Technology Carleton University Ottawa, Canada chenzhou4@cmail.carleton.ca AbstractâChildren are increasingly using technologies powered by Artificial Intelligence (AI). However, there are growing concerns about privacy risks, particularly for children. Although existing privacy regulations require companies and organizations to implement protections, doing so can be challenging in practice. To address this challenge, this article proposes a framework based on Privacy-by-Design (PbD), which guides designers and developers to take on a proactive and risk-averse approach to technology design. Our framework includes principles from several privacy regulations, such as the General Data Protection Regulation (GDPR) from the European Union, the Personal Information Protection and Electronic Documents Act (PIPEDA) from Canada, and the Childrenâs Online Privacy Protection Act (COPPA) from the United States. We map these principles to various stages of applications that use Large Language Models (LLMs), including data collection, model training, operational monitoring, and ongoing validation. For each stage, we discuss the operational controls found in the recent academic literature to help AI service providers and developers reduce privacy risks while meeting legal standards. In addition, the framework includes design guidelines for children, drawing from the United Nations Convention on the Rights of the Child (UNCRC), the UKâs Age- Appropriate Design Code (AADC), and recent academic research. To demonstrate how this framework can be applied in practice, we present a case study of an LLM-based educational tutor for children under 13. Through our analysis and the case study, we show that by using data protection strategies such as technical and organizational controls and making age-appropriate design decisions throughout the LLM life cycle, we can support the development of AI applications for children that provide privacy protections and comply with legal requirements. Index TermsâLarge Language Models, Privacy-by-Design, Childrenâs Data Privacy, Data Protection Regulations, Data Governance I. INTRODUCTION The growing integration of Artificial Intelligence (AI), including Large Language Models (LLMs) into digital appli- cations for children presents both opportunities and challenges [1], [2]. On the one hand, educational tools, conversational agents, and LLM-powered storytelling bots can offer children engaging and personalized experiences [3]â[5]. On the other hand, they also pose considerable privacy risks [2]. Recent scholarship and regulatory attention have raised concerns about the large amounts of information collected for training of LLM-based applications and during user interactions, which can involve sensitive and personal data [6], [7]. Children are especially vulnerable, as they may not fully understand the risks of exposing private information about themselves in digital environments [1], [8]. Furthermore, some conversational AI systems are designed to seem friendly and trustworthy, which can encourage children to overshare personal details [9]. Additionally, parents often lack sufficient information or guidance to understand how their childrenâs data are used and processed making it difficult for them to ensure adequate privacy protections [1]. The risks are further exacerbated by LLM-specific vulnerabilities, such as the modelsâ capacity to memorize and inadvertently reveal sensitive details from training data [10], [11], and by real-world instances where companies providing AI services for children have faced regulatory action after alleged privacy law violations [12]. Therefore, it is of critical importance to design LLM-based applications with proactive privacy protections that mitigate privacy risks for children. [13]. National and transnational privacy regulations, including the General Data Protection Regulation (GDPR) [14] in the (European Union) EU, the childrenâs Online Privacy Protection Act (COPPA) [15] in the United States (US), and ongoing work to strengthen privacy law in Canada [16], agree that childrenâs data deserve special protection in digital contexts [15], [17], [18]. For example, GDPR requires stronger safeguards for arXiv:2602.17418v1 [cs.AI] 19 Feb 2026 childrenâs data, as outlined in Article 8, which focuses on parental consent, and Recital 38 [19], which acknowledges the unique privacy needs of children and their limited ability to un- derstand their rights, risks and consequences of data processing [17]. COPPA is a regulation that applies specifically to online services collecting information from and about children under 13 years of age [15]. It requires verifiable parental consent, clear explanations of data use, and limits on data retention among other requirements [15]. In Canada, PIPEDA does not have specific rules for protecting childrenâs data. However, the Office of the Privacy Commissioner of Canada (OPC) has recently launched an exploratory consultation to create a Childrenâs Privacy Code [18], affirming that children deserve special considerations and protections in digital contexts. Despite growing recognition of the need to protect childrenâs privacy when designing and deploying AI applications, to our knowledge, there is currently a lack of practical, actionable recommendations for privacy-preserving techniques specifically aimed at Children. Such guidance would help designers and developers create applications that are both legally compliant and secure for childrenâs data. We use âdevelopersâ, âdesign- ersâ, and âengineersâ to refer to companies/organizations, AI service providers, individuals, and policymakers responsible for designing, building and deploying LLM-based systems. While organizations like the United Nations Childrenâs Fund (UNICEF) have issued policy recommendations, such as the Policy Guidance on AI [1], and the Information Commissionerâs Office (ICO) in the United Kingdom has provided guidance through the Age-Appropriate Design Code (AADC) [20], they offer standards and principles, while focusing less so on practical operational guidance for designers and developers. The academic research on privacy in AI, and LLMs specifically, has also focused on general-purpose systems. In particular, there is increasing research on security and privacy threats in LLMs, including memorization risks and inference attacks [6], [10], [21]â[23]. Similarly, there are works on technical risk mitigation strategies, including machine unlearning [24], [25], differential privacy [26], [27], federated learning [28], and other approaches (we detail in Section I). Another stream of research examined the alignment of LLMs with privacy laws, particularly the GDPR in the European Union [29], [30]. Some emerging studies have integrated technical, design, and legal considerations into a framework [31]. These works provide a foundation for understanding privacy risks, challenges, and mit- igation strategies. However, when designing AI applications for children, additional considerations must be taken into account, such as the unique legal protections applicable to children (e.g., COPPA), and the unique vulnerabilities of children, tied to their cognitive, social and developmental capacities [32]. To our knowledge, there is no comprehensive framework that adapts regulatory privacy principles and technical controls to the realities of childrenâs interactions with LLM-based applications. As such, there is a current gap in ensuring that AI applications are both legally compliant and appropriate for children. To address these, we propose a framework for designing children-focused LLM applications based on the principles of Privacy-by-Design (PbD) [33]. PbD centers on the idea that service providers should incorporate privacy-preserving mechanisms from the outset and directly within the systemâs design, business practices, and infrastructures [34]. In addition, these safeguards should be maintained throughout the lifecycle, not just added later for compliance. The proposed frame- work maps the regulatory principles of GDPR, COPPA, and PIPEDA, onto the stages of the LLM lifecycle, including data minimization, purpose limitation, meaningful and verifiable consent, security by design, accountability, and user rights. The LLM lifecycle stages involve data collection, model training, operation and monitoring, and continuous validation. For each stage, we provide specific operational controls found in academic literature that translate regulatory principles into technical and organizational actions. These controls aim to mitigate privacy risks and support compliance with regulatory principles. In addition to meeting legal standards, lifecycle mapping enables us to understand how children interact with LLM applications and highlights their unique vulnerabilities and privacy needs. By incorporating privacy protections into every lifecycle stage, the goal is to guide designers and developers in designing LLM-based applications that support technical controls, legal compliance, and children-centered design considerations. The key contributions of the paper are summarized as follows. âą Analyze the privacy violations that arise in LLM-based applications for children, as well as broader risks such as regulatory non-compliance, memorization, profiling and manipulation. âąPropose a PbD-aligned framework that map the core regulatory principles of major privacy regulations into the stages of the LLM lifecycle, including data collection, model training, operation and monitoring, and continuous validation. âąConnect regulatory expectations and engineering practices by mapping regulatory expectations to practical technical and operational controls that designers and developers can implement. âąProvide an illustrative case study, the educational LLM tutor, along with design recommendations that demonstrate the frameworkâs feasibility in practice and highlight challenges in privacy-preserving LLM deployment. The remainder of this paper is organized as follows. Sec- tions I and I introduce the background and related work. Section IV discusses cases of privacy violations in LLM- based applications for children, providing justification and motivation for adopting PbD principles in applications designed for children. Section V provides the regulatory principles related to LLM-based applications for children. Section VI maps the regulatory principles to architectural features throughout the LLM lifecycle, it also discusses an illustrative case study, the educational LLM tutor for children, as a representative application of the framework. In Section VII, we provide an in- depth discussion that summarizes the findings, reiterates design recommendations, presents the limitations of the proposed approach, assesses both the regulatory suitability and the challenges of implementing the PbD principles, and outlines future work. Section VIII concludes the paper. I. BACKGROUND In this section, we provide the necessary background on PbD, privacy and data protection laws, and LLMs. A. Privacy-by-Design (PbD) PbD centers around the idea that simply meeting legal and regulatory compliance does not guarantee adequate privacy protection [33], [35]. Instead, companies and organizations must incorporate privacy as a fundamental aspect of their design processes. This view supports including privacy in organizational practices, information technology systems, and physical structures [33], [35]. PbD includes seven principles that guide organizations in implementing privacy into the design and development of information systems. (1) PbD advocates for proactive measures, which means that organizations should anticipate and reduce privacy risks before they materialize, rather than responding to violations after the fact [33]. (2) Privacy must be the âdefault settingâ, with systems automatically protecting privacy, without requiring individuals to turn on privacy protections themselves. (3) Organizations must also include privacy protections in the design and setup of their technologies and operations in a way that does not affect system functionality. (4) PbD advocates for a âpositive-sumâ approach, where privacy is seen as compatible with other organizational goals and as part of the overall functionality of a system. (5) The principle of end-to- end security requires security measures to be implemented to protect personal data throughout its lifecycle, from collection to disposal. (6) Organizations should also let stakeholders and independent reviewers evaluate their privacy practices. (7) PbD emphasizes the ârespect for user privacyâ. It states that orga- nizations should uphold individual rights and expectations by providing clear notices, establishing strong default protections, and offering options that allow users to control their personal information easily [33]. PbD emerged during a period of rapid technological devel- opment, increased system complexity, and global competition, all of which intensified risks to informational privacy and security [33]. Since 2022, the widespread development and adoption of AI technologies that collect, infer, and process data at unprecedented scale has made the PbD principles particularly relevant. The application of PbD to AI systems involves incorporating privacy standards throughout the systemâs life cycle, from data collection through training and deployment to validation [36]. In this way, AI service providers ensure they incorporate preventive measures into their system design to reduce privacy risks while supporting compliance with regulatory principles. This need for proactive integration of privacy protections reflects broader shifts in how regulatory systems have devel- oped to govern informational environments [13], [37], [38]. Some scholars have noted that traditional administrative law frameworks, built for functions of the industrial era such as rule-making and adjudication, are poorly suited to the dynamics of contemporary data infrastructures and require new governance approaches [13]. In the case of AI technologies, which are evolving at an unprecedented pace, the law must adapt rapidly to meet the challenges posed by this new paradigm in information systems. B. Overview of the Privacy and Data Protection Laws Several legal frameworks exist to govern childrenâs data. The most notable among them are GDPR [14], COPPA [15], PIPEDA [16], the California Consumer Privacy Act (CCPA) [39], the UKâs Age-Appropriate Design Code (AADC) [20], and UNICEFâs AI for Children guidance [1]. Of these, COPPA, GDPR, PIPEDA, and CCPA are binding laws; the AADC is a statutory code of practice under the UK Data Protection Act 2018 (enforceable but not a standalone statute). Although these frameworks share overlapping objectives such as protecting childrenâs personal data, ensuring meaningful consent, and providing mechanisms that ensure parental oversight, their enforcement mechanisms and technical interpretations vary across jurisdictions. Also, most of these frameworks were developed for traditional systems rather than adaptive AI models. In this paper, we focus on COPPA, GDPR, and PIPEDA because they represent three mature and complementary regulatory regimes. 1) Childrenâs Online Privacy Protection Act (COPPA): COPPA is a United States federal law that regulates how online services collect, use, and disclose personal information from children below the age of 13 [15]. It applies to operators of websites, mobile applications, AI-driven platforms, or any online service that is directed at children or knowingly collects data from them. A set of âCore Privacy Requirementsâ are articulated in the COPPA Rule (16 C.F.R. Part 312), with supporting guidance from the United States Federal Trade Commission (FTC): (i) providing clear and comprehensive privacy notices describing data collection, use, and disclosure practices; (i) obtaining verifiable parental consent (VPC) before collecting, using, or disclosing a childâs personal information [15]; (i) maintaining the confidentiality, security, and integrity of childrenâs data through reasonable procedures; (iv) limiting data collection to what is reasonably necessary for participation in a game, service, or activity; and (v) allowing parents to review, delete, or refuse further collection or use of their childâs data [15]. Under Section 312.6, COPPA empowers parents to manage their childrenâs personal information. It is a requirement for operators to provide accessible mechanisms, such as online dashboards, email verification, or secure communication chan- nels, for parents to exercise these rights [15]. However, COPPA does not explicitly include the more advanced rights found in other regulations, such as data portability or challenging automated decision-making, which limits its applicability in modern AI contexts. The law also specifies acceptable methods for obtaining verifiable parental consent, including signed consent forms, payment card verification, video calls, or government-issued ID checks. Operators are prohibited from conditioning a childâs participation in an activity on the disclosure of more informa- tion than is reasonably necessary. COPPA defines âpersonal informationâ broadly, covering not only name, address, and email, but also geolocation data, screen names, photos, audio, video, and any persistent identifier used to recognize a user over time and across websites [15]. 2) General Data Protection Regulation (GDPR): The GDPR (Regulation (EU) 2016/679) is a comprehensive legal frame- work that governs the processing of personal data across the European Union (EU) and European Economic Area (EEA) [14]. While GDPR applies to all individuals, Recital 38 explicitly recognizes that children merit specific protection, particularly in digital environments where they may have little to no knowledge of associated risks and protections. Likewise, GDPR Article 8 demands that the processing of personal data for children under the age of 16 (or 13, where allowed by Member States) be authorized by the holder of parental responsibility [14]. These provisions are particularly important for LLM-based applications, which often collect and transform user inputs into internal statistical representations of text that may be reused in opaque inference pipelines. GDPR is grounded in seven core data protection principles outlined in Article 5(1): (i) lawfulness, fairness, and trans- parency; (i) purpose limitation; (i) data minimization; (iv) accuracy; (v) storage limitation; (vi) integrity and confidential- ity; and (vii) accountability, as further detailed in Article 5(2) and Article 24. These principles impose strict obligations on AI developers to ensure that personal data is processed fairly, securely, and for clearly defined purposes. Additionally, GDPR grants individuals a broad set of rights under Articles 12â23, many of which are highly relevant to childrenâs data. In Article 12, it is stated that privacy notices should be concise, transparent, and presented in a manner easily understood by children. The rights of access, rectification, and erasure (âright to be forgottenâ) are also articulated in Articles 15â17. These rights often pose unique challenges for LLM-based applications for children, since user data may be distributed across model parameters and intermediate artifacts (e.g., embeddings, caches) rather than stored as discrete records in traditional databases [10], [40]. Again, the right to data portability is established in Article 20, while Article 21 provides the right to object to certain processing activities, including profiling. Another harmful effect of LLM-based applications that may have a detrimental influence on a childâs development or well-being arises when decisions are made solely through automated processing (including profiling) and produce le- gal or similarly significant effects. To mitigate these out- comesâparticularly critical in LLM-based educational or healthcare applicationsâadditional protective measures are provided in Article 22, subject to specific conditions and exceptions [14]. Last but not least, Article 32 further demands technical and organizational measures to ensure data security, addressing risks such as unauthorized access, inference attacks, and adversarial manipulationâthreats that are increasingly relevant in the context of generative AI and prompt-based attacks. Although GDPR contains numerous provisions, we focus on Articles 5, 8, 12â17, 20â22, 24, and 32 in this paper, as they directly address privacy and security challenges in LLM-based applications for children. These articles combined provide a rigorous framework for embedding privacy into the design, training, and deployment of AI systems, ensuring both regulatory compliance and the ethical handling of childrenâs data. 3) Personal Information Protection and Electronic Docu- ments Act: Canadaâs federal privacy law, PIPEDA, oversees the collection, use, and disclosure of personal information by private-sector organizations during commercial activity operations [16]. Unlike COPPA and GDPR, PIPEDA does not establish a fixed age threshold for consent. Instead, it relies on the principle of meaningful consent, which must be assessed based on a childâs maturity, cognitive ability, and capacity to understand the implications of data processing [16]. The OPC advises that, where a child cannot reasonably provide informed consent, organizations must obtain consent from a parent or guardian. Although the law is flexible within its applicable context and accommodates a wide range of digital services, it also introduces ambiguity when it is applied to emerging AI technologies such as LLM-based applications for children, where it is inherently challenging to determine a childâs capacity for consent. PIPEDA is embodied in ten âFair Information Principlesâ enshrined in Schedule 1 of the Act [16]. These principles include: (i) accountability, requiring organizations to designate a privacy officer and implement internal policies; (i) identifying purposes, ensuring that the reasons for data collection are communicated at or before the time of collection; (i) consent, which must be meaningful and informed; (iv) limiting collection to what is necessary for identified purposes; (v) limiting use, disclosure, and retention; (vi) accuracy; (vii) safeguards to protect data; (viii) openness about data practices; (ix) individual access; and (x) challenging compliance through accessible complaint mechanisms. These principles parallel GDPRâs data protection framework but are typically less prescriptive and more flexible in interpretation. In terms of rights, PIPEDA grants individuals the ability to access their personal data, request corrections for inaccuracies, and receive clear explanations of how their data are being collected, used, and disclosed (Schedule 1, Principles 9 and 10) [16]. However, it does not explicitly provide a right to erasure or data portability, which are central features of GDPR. For LLM-based applications for children, this gap poses challenges, particularly when personal data is embedded within model parameters and cannot easily be identified or removed. In summary, COPPA, GDPR, and PIPEDA form distinct yet overlapping regulatory frameworks designed to safeguard TABLE I KEY REGULATORY PARAMETERS IN COPPA, GDPR, AND PIPEDA AspectCOPPAGDPRPIPEDA Age ThresholdUnder 13Under 16 (may be lowered to 13 by Member States) No fixed age; capacity-based Consent RequirementVerifiable parental consent Parental consent required for children Meaningful consent required; maturity assessed Automated ProfilingNot explicitly addressed Right not to be subject to automated decision-making Not explicitly addressed Relevance to LLMsChallenges with verifying consent for dynamic interactions Difficulties with transparency, personalization, and data inference Ambiguity in consent maturity and inferred data handling TABLE I SUMMARY OF PRIVACY RIGHTS ACROSS COPPA, GDPR, AND PIPEDA Privacy RightCOPPAGDPRPIPEDA Right to Access~â Right to RectificationĂâ Right to ErasureĂâ~ Right to Restrict ProcessingĂâ~ Right to ObjectĂâĂ Right to Data PortabilityĂâĂ Right to Withdraw Consent~â Right to Challenge Automated DecisionsĂâĂ Right to Lodge Complaints or Seek Redress~â Note:ââ Explicitly required;~â Implicit or partial coverage; Ăâ Not formally required personal data, each with different approaches to consent, age thresholds, and data subject rights. Table I provides a comparative snapshot of these key parameters, establishing a baseline for understanding how these regulations apply to LLM-based applications for children. To complement this analysis, Table I provides a comparative overview of the privacy rights granted under COPPA, GDPR, and PIPEDA, highlighting areas of alignment as well as gaps that must be addressed when applying these regulations to LLM-based applications for children. This sets the stage for the next subsection, which discusses how these regulatory principles can inform privacy-preserving design strategies in LLM-based applications for children. C. Large Language Models (LLMs) This section provides a high-level overview of LLMs. The goal is to provide a broad understanding of the lifecycle stages and the data processing activities involved in the development and operation of LLMs. In addition, this section serves as a technical introduction to the following sections, which will analyze specific privacy principles, how they apply to each lifecycle stage, and the technical controls implemented to support privacy preservation. Advances in natural language processing (NLP) have led to the development of LLMs based on the Transformer architecture [41]. The Transformer architecture enabled lan- guage models to capture long-range dependencies in text [42]. This design has made LLMs effective in many NLP tasks and enabled the generation of coherent content [43]. Additionally, tuning these models on specific instructions and introducing Reinforcement Learning from Human Feedback (RLHF) have enabled LLMs to achieve strong performance across a wide range of tasks, including machine translation, text summarization, and question answering [41]. Public interest in these models surged following the release of ChatGPT in late 2022, which accelerated both their deployment and research into their capabilities and limitations [41]. More recently, LLMs have been integrated with models that process audio and visual data, resulting in multimodal systems that can understand and generate content across different modalities, such as sound, video, and images [44]. Today, LLMs are the foundation for many advanced language-processing tech- nologies, including AI-based virtual agents and multimodal applications [44]. As deep learning-based NLP systems, LLMs follow a multi-stage life-cycle [22]. Although implementation details vary depending on the deployment context and supporting infrastructure [7], for the purposes of this review, we describe the LLM life-cycle through four main stages: data collection, model training, operation and monitoring, and continuous validation. In the following, we provide a high-level overview of these stages. 1) Data Collection: Data collection is a foundational phase in the LLM life-cycle. During this phase, raw data is gathered from various sources, including books and digital content [45]. Training corpora typically consist of publicly available datasets, such as BooksCorpus, Common Crawl, and the Colossal Clean Crawled Corpus (C4), as well as additional sources such as academic texts and code repositories [45], [46]. Generally, LLMs utilize a combination of these sources [43]. Once the data are collected, they undergo pre-processing for training. This pre- processing phase includes quality filtering, data de-duplication, and privacy reduction to eliminate noisy, duplicated, and potentially harmful information. Another necessary task during pre-processing is tokenization, which involves converting raw textual data into sequences of individual tokens [43]. Finally, data scheduling is defined as the process of determining how and when subsets of data will be presented to the model during training [45]. 2) Model Training: The training phase typically comprises several sub-stages: (1) foundation model pre-training, where the model learns general language representations using self- supervised objectives such as next-token prediction; (2) domain adaptation through fine-tuning on task-specific datasets; and (3) alignment, employing techniques like RLHF or preference optimization to align the modelâs behavior with human values such as honesty, helpfulness, and harmlessness [43], [45]. 3) Operation and Monitoring: Once training is completed, LLMs are deployed in production environments to support various applications [47]. This deployment can occur on cloud services, edge devices, or local servers. After deployment, LLMs interact directly with users to provide real-time outputs [47]. Therefore, the operation and monitoring phase includes inference, in which the model generates responses based on input prompts and may collect and store user interactions for safety review, debugging, analytics, and future improvements [7]. 4) Continuous Validation: After deployment, the continuous validation phase focuses on ongoing monitoring, oversight, and governance of LLM-based applications. Continuous monitoring helps ensure that model outputs remain free of bias or inaccuracies [47]. In addition, it ensures that model performance remains optimal. This involves tracking metrics such as latency, throughput, and accuracy [47]. Updates might also include retraining the model on new data, which might come from pre- vious user interactions [7]. Continuous validation may further include red-teaming and adversarial testing, periodic audits of privacy and security controls, and updating documentation (e.g., model cards and dataset documentation) as the system evolves [48]â[50]. I. RELATED WORK Building on the preceding discussion of the PbD recom- mendations, relevant privacy regulations, and the technical foundations of LLMs, this section presents the related work. Existing research has attempted to integrate technical, design, and regulatory perspectives into a unified framework informed by the principles of PbD [31]. However, the framework proposed by Al Breiki and Mahmoud [31] is primarily oriented toward general-purpose applications and does not explicitly address child-specific considerations. To address this gap, we adopt their PbD-guided framework as a foundational structure and extend it to incorporate privacy, regulatory, and design considerations tailored to children. This extension is particularly important given growing evidence that children increasingly interact with conversational agents and AI-enabled search tools [51]. The remainder of this section reviews three main bodies of related work: (1) technical approaches to privacy protection in LLMs; (2) PbD-guided frameworks for LLM-based applications that integrate regulatory and design considerations; and (3) special vulnerabilities and considerations for children. By synthesizing previous work, we highlight both the progress and gaps, particularly the absence of a unified, comprehensive, and operational framework that integrates regulatory princi- ples, technical controls, and design considerations specific to children. A. Technical Privacy Threats and Protections in LLMs According to Ofcomâs 2024 Online Nation report, 54% of children aged 8â15 in the UK used a generative AI tool in the past year [52]. These include ChatGPT 1 , Microsoft Copilot 2 , Snapchat MyAI 3 , and Google Gemini [52], which are not designed or marketed primarily for children under 13 years of age. Similar patterns have been observed in North America [53]. In response, some companies have begun to introduce child-centred protections. For example, OpenAI has added parental controls that allow parents of teens to link their account with their teenâs account and manage safety-related settings [54]. Google implemented similar safeguards for its Gemini application. The case of Google is particularly interesting, as the service now allows children under 13 years of age to access the application with parental controls available through Family Link, which provides safeguards such as providing and removing application access and content filtering [55]. However, these measures assume parents are aware whether and how their child is using such applications, which is often not the case [53], [56]. In addition, there are different parenting approaches, with some parents avoiding the use of restrictive measures to support their childrenâs autonomy [57]. As an increasing number of people, including children, have started using LLM-based applications, it is essential to understand the privacy risks associated with these applications. Existing research on privacy in LLMs has found several types of vulnerabilities. Recent surveys classify these vulnerabilities into two main categories based on how attackers can access private information: privacy leakage and privacy attacks [22]. Privacy leakage refers to the unintentional release of sensitive information caused by the internal behavior or design of a 1 https://chatgpt.com 2 https://copilot.microsoft.com 3 https://w.snapchat.com/@myai model system. In contrast, privacy attacks involve intentional actions that take advantage of model or system weaknesses to extract sensitive data [22]. A common form of privacy leakage occurs when users inadvertently share sensitive information while interacting with LLMs. If such data are logged and later reused for fine-tuning or retraining, it may reappear in future model outputs, posing privacy risks [58]. Recent work found that LLMs can reveal private information in contexts where humans would not [59]. Separately, LLMs have also been known to remember parts of their training data, a phenomenon called memorization [10]. Under certain conditions, they can reveal this data in their responses [10], [11]. This is concerning because memorized examples may contain private details like names, addresses, social security numbers, or snippets of conversations collected from users. Memorization creates an opportunity for adversaries to exploit. In a training-data extraction attack, an attacker recovers memorized training examples verbatim from a model through targeted querying techniques. Similarly, a membership- inference attack determines whether a specific record was part of the training data, while model inversion reconstructs approximate representations of training examples; inversion is similar to data extraction but does not require verbatim recovery [10]. To minimize these vulnerabilities, researchers have proposed a range of technical strategies, including but not limited to differential privacy, machine unlearning, and federated learning [6], [25]â[27]. Differential privacy is applied to limit the influence of individual data points on model parameters [60], while machine unlearning techniques seek to enable post-hoc deletion of sensitive data from trained models [24], [25]. Federated learning allows multiple entities to learn simultaneously without sharing data and sending only model updates to the central server [6], [28]. In addition to these techniques, researchers have also explored measures such as input and output filtering [61], and safeguards against adversarial behavior [62]. Transparency artifacts such as dataset documentation [49] and model cards [50] have been introduced to make data practices and model limitations more visible to stakeholders. Lastly, organizations can identify and assess risks in LLMs through adversarial testing methods, and offer ways to address these vulnerabilities, thus strengthening accountability [48]. B. Privacy-by-Design in AI Systems Researchers have extensively explored privacy-preserving techniques in AI; however, comparatively less attention has been given to integrating these technical mechanisms with legal compliance principles. Breiki and Mahmoud [31] have developed a framework to integrate PbD into generative AI (GenAI) applications. They demonstrate the proposed framework through an AI chatbot in which privacy protections operate in three modes: strict, standard, and personalized. These modes provide different levels of privacy assurance, depending on user preferences, enabling users to have control over their data. The framework also incorporates differential privacy mechanisms to prevent the memorization of personal data and reduce the leakage of sensitive information. Additional mechanisms, such as privacy risk detection tools that warn users when entering PII, and encrypted and access-controlled logs to support regulatory compliance, add another protective layer to the framework. Given the novelty of the topic, some of the existing literature is still developing. The emerging work by Sens [63] aims to examine the application of the PbD framework in user-based machine learning (ML). The work proposes several methods for embedding PbD into ML-based software, including requirements specification with static verification during implementation and dynamic verification at runtime, as well as data minimization strategies for training ML models. The requirements specification involves creating a domain- specific language for annotating data and system components with regard to their privacy impact, to help developers apply specific measures to more privacy-impacting entities. Static and dynamic privacy verification, as well as data-efficient ML training, focus on measuring the privacy impact of specific data features and their influence on system performance. The goal is to effectively anonymize private data without compromising the utility of the model. Finally, the work proposes a modular architecture for ML-enabled systems that combines these privacy-preserving methods [63]. The study by Oh et al. [64] introduced the PbD Extract, Transform, Load, and Present (PETLP) compliance framework for embedding legal safeguards directly into data ingestion pipelines. The work specifically deals with social media information, where personal data can be easily ingested by extract, transform, load (ETL) pipelines and poses significant privacy risks. The main proposition of this work is to implement Data Privacy Impact Assessments (DPIAs) prior to data collection and to update them throughout the research lifecycle. In this way, the DPIAs serve as living documents that guide researchersâ decisions regarding data; they can help identify risks and assign mitigations early, document security measures, and document ongoing obligations. In summary, current research on privacy in LLMs focuses mainly on threats, technical security, privacy protections, and regulatory compliance. To date, one study proposes an integrated framework [31]. However, there is a lack of frameworks that combine these perspectives specifically for children, leaving gaps in operationalized approaches that simultaneously safeguard data, comply with regulations, and address child-specific design needs. This work builds on this integrated PbD-based framework [31] by adopting its core structure and operationalizing it within a child-specific context. C. Specific Considerations for Children In the previous section, we examined threats associated with general-purpose LLM-based applications identified in the literature. Although these systems are primarily designed for adult users, they are increasingly accessed by children. This section focuses on risks specific to children that have been identified in prior research. Research suggests that, due to cog- nitive, developmental, and emotional differences, children may interact with and interpret LLM-based applications differently from adults, thereby increasing their susceptibility to various harms, including privacy-related risks [9], [32]. Developmental studies summarized by Nomisha [32] indicate that young children (ages 2â7) often struggle to distinguish between living beings and machine-based social agents [65]. Older children (ages 6â11) also tend to attribute emotional states to non-living entities, for example, describing a voice assistant as âhappyâ based on its responses [66]. This tendency to ascribe human-like qualities to machines is commonly referred to as anthropomorphism, which can have significant pri- vacy implications [32]. During interactions with such systems, children may unintentionally disclose personal information, including their names, ages, and locations. When applications are designed to appear friendly or human-like, children may be further encouraged to share additional personal details [32]. Anthropomorphism, defined as the attribution of human traits, emotions, or intentions to non-living entities [67], can foster a sense of trust and understanding, thereby increasing the likelihood of information disclosure. Prior research suggests that children may, in some contexts, trust social robots more than humans [68]. In addition to anthropomorphic design, researchers have identified conversational techniques known as nudging in some chatbot systems [9], [69]. Nudging often takes the form of follow-up prompts, such as âtell me more,â intended to sustain engagement. However, in interactions with children, such prompts may inadvertently encourage the disclosure of information that children might not otherwise share [9]. Recent work has proposed policy and design recommenda- tions aimed at mitigating the effects of anthropomorphism and nudging [32], [70], particularly where these techniques may exploit childrenâs developmental vulnerabilities and undermine privacy. These considerations are increasingly important as applications designed for children, including interactive tutors, storytelling agents, and gaming applications more frequently incorporate LLMs capable of engaging in open-ended conver- sations. D. Summary The review of existing literature suggests that current research on LLM vulnerabilities, legal compliance and policy recommendations for child-specific design exists in parallel rather than in combination. This highlights the need for a comprehensive framework that integrates technical protections, legal compliance, and child-oriented design recommendations. In practice, this would involve implementing current policy recommendations and research regarding children and AI, such as designing age-appropriate interfaces and establishing parental consent workflows. Additionally, incorporating tech- nical security measures like differential privacy and machine unlearning. By combining technology-specific safeguards with child-centred approaches, we can achieve more holistic pro- tections for childrenâs privacy. Our contribution to the current body of work is to merge these established practices into a framework that ensures service providers of child-oriented AI technologies (1) comply with privacy regulations, (2) prevent privacy violations by implementing robust technical protections from the outset, (3) and uphold ethical principles that respect childrenâs best interests and rights. IV. PRIVACY VIOLATIONS IN LLM-BASED APPLICATIONS FOR CHILDREN To motivate and emphasize the need to integrate privacy protections into the design of child-focused LLM-based tech- nologies, we offer an overview of recent privacy violations and incidents resulting from noncompliance or disputed prac- tices with applicable privacy regulations across various AI applications used by or developed for children. Although not all systems discussed in this section involve LLMs, they demonstrate the harms that can result from inadequate privacy protections, particularly for vulnerable populations such as children. Further, several cases involve data flows that closely mirror LLM pipelines, including large-scale data collection for training, logging and retention of user interactions, and third- party processing. We also provide a brief explanation of what compliance might look like according to the PbD approach. (Under GDPR, consent is one lawful basis; where consent is relied upon for information society services offered directly to a child, Article 8 adds parental authorization requirements [14].) The violations we review include inadequate consent mech- anisms [12], collection of childrenâs photographs for model training [71], [72], design flaws exposing user profiles [73], and large-scale surveillance in educational environments [74]. Across these examples, safeguards are often reactive and do not address the unique vulnerabilities of children, such as their greater likelihood of disclosing personal information, the sensitivity of childrenâs data, and their susceptibility to exploitation. Furthermore, many do not take into account the special legal protections applicable to children, such as the requirement for parental consent and limits on data collection and sharing [15]. A. Data Collection without Adequate Consent Mechanisms Despite the centrality of consent in the major privacy regulations, some providers do not implement adequate mech- anisms to obtain, verify, and enforce parental consent. For example, Buddy AI [75], a popular AI-based application for children under 12 years old, was found by the Childrenâs Advertising Review Unit (CARU) to have violated COPPA and the CARU Privacy Guidelines by collecting childrenâs personal information (as defined under COPPA) without providing direct notice, posting visible privacy policies, or obtaining verifiable parental consent [12]. Although the company even- tually updated its practices, the changes were reactive and made only after regulatory scrutiny. This reactive posture contradicts the principles of PbD, which require companies to implement privacy protection mechanisms, such as consent, as default design components [33]. In LLM-based applications designed for children, compliance would look like providing detailed parental notices explaining data collection and use, a mechanism to provide and revoke consent at any time (i.e., dynamic consent, which can be granted, reviewed, and revoked over time), accessible privacy policies, and clear opt- out procedures. B. Unauthorized Training Data Collection Training LLMs often involves scraping large amounts of online content, which creates a risk of collecting data about children without their or their parentsâ knowledge or consent. For example, Meta scraped public images and posts to train its generative AI models, including photos of children taken from adult-owned accounts [76], [77]. Similarly, the LAION- 5B dataset [78] contained images of children from Brazil and Australia, many of which were labelled with names, ages, and locations [72], [79]. Researchers at the Stanford Internet Observatory also identified child sexual abuse material (CSAM) within the LAION-5B dataset [80]. Following this discovery, LAION reported that it was working to remove the harmful content [78]. These cases reveal significant shortcomings in data collection practices for AI systems. In particular, companies that scrape images from the internet may also collect content from social media platforms for training purposes, even when consent for such use has not been provided. A PbD approach would require upstream safeguards, such as obtaining explicit consent to use data for model training (e.g., via opt-in licensing or purpose-limited agreements where feasible) or implementing mechanisms to detect and exclude childrenâs images from training datasets. However, general privacy-policy terms alone may be insufficient to constitute explicit consent, and should not be treated as a substitute for transparent, purpose-limited sourcing. Existing mitigation techniques include child face detection tools, although it is still an active area of research [81], [82]. For example, the Stanford study researchers used perceptual hash-based detection, cryptographic hash-based detection, and k-nearest neighbour analysis to detect childrenâs images [80]. These automated methods also have limitations (e.g., false positives/negatives and potential bias) and therefore require governance and evaluation when used in practice. Importantly, these measures should be applied before data are incorporated into AI training. C. Accidental Exposure and Public Disclosure of Sensitive Information Poor security designs can result in accidental disclosure of sensitive information. In December 2024, Character.AI briefly exposed user account information, which âdepending on the user, could include a username, name, bio, Characters, homepage, personas, voices or chatsâ for approximately 10 minutes [73]. The company minimized the impact, noting that the breach affected fewer than 0.01% of the accounts [73]. However, given the popularity of generative AI-based applications among children aged 8â15 in the UK [52], and reports of children using this platform for therapy or emotional support [83], there is a high risk that such breaches can have serious privacy implications for children. Character.AIâs broader data-collection practices exacerbate these risks. The platform collects a significant amount of personal data, including conversations, uploaded media files, and voice recordings, and shares some information with third- party advertising and analytics vendors [84]. Additionally, their privacy policy indicates the company engages in profiling of users based on its collection of user interests and âpreferences based on account settings or feedback on the Servicesâ [84]. Embedding preventive safeguards, such as restricted collection, minimal retention, and strong encryption, would respect privacy of users, and protect sensitive data in the event of a breach. D. Surveillance in Educational Settings and Data Breaches Large-scale AI surveillance in educational settings can significantly undermine childrenâs privacy. For example, Gaggle Safety Management, a system used to monitor about six million students for indicators of cyberbullying, self-harm, violence, and other behavioral or emotional risks, had recently faced a privacy breach [85], [86]. In 2025, Vancouver Public Schools (Washington State, U.S.) inadvertently exposed nearly 3,500 sensitive student records, including poems, essays, and AI chat transcripts collected by Gaggle, through unprotected links [86], [87]. This case is an example of the dangers of large-scale surveillance and weak security controls. The absence of basic safeguards, such as password protection, encryption, and access controls, led to a significant privacy breach involving highly sensitive personal information of schoolchildren. Some of this data involved chats about being victims of violence, threats of suicide, and questions about sexuality, the disclosure of which could potentially impact the well-being of these children and have long-term consequences [86], [87]. Recent research also highlights educatorsâ concerns about the intrusive nature of monitoring tools and their implications for student trust and autonomy [88]. The cases presented in this section illustrate gaps in regula- tory compliance among many AI service providers, particularly those that market their technologies to children, those whose products are used by children, or those involving children. A common theme across these cases is that companies often implement privacy protections only after violations occur, treating these safeguards as fixes rather than as integral components of the design process. As various LLM-powered applications, such as educational tutors, companions, and storytelling bots, become more common in childrenâs daily lives, it is critical to incorporate the principles of PbD into the development, deployment, and monitoring of these systems to ensure adequate privacy protections for sensitive childrenâs data. V. REGULATORY PRINCIPLES RELEVANT TO LLM-BASED APPLICATIONS FOR CHILDREN Building on the comparative analyses of regulatory principles and privacy rights discussed in SectionII-B, along with the privacy violations outlined in Section IV, we identify seven overarching principles: Data Minimization, Purpose Limitation, Transparency and Explainability, User Rights, Accountability, Security by Design, and Meaningful Consent Mechanisms. Table I summarizes how these principles are articulated in COPPA, GDPR, and PIPEDA. Collectively, these principles form the foundation of our framework, which provides guidance on proactively embedding privacy principles throughout the LLM lifecycle âfrom data collection and training to deployment and real-time interactions. The following paragraphs elaborate on each principle in the context of LLM-based applications for children and offer a brief overview of controls presented in recent academic literature that can be implemented in practice to help organizations align their practices with the legal requirements for privacy protection. These strategies are explored in more detail and in relation to each of the LLM lifecycle stages in Section VI. Data Minimization is explicitly required under GDPR Article 5(1)(c) and PIPEDA Principle 4. Under GDPR, personal data must be âadequate, relevant, and limited to what is necessaryâ for the specified purpose for which it is processed [14]. PIPEDA similarly requires organizations to collect only the information needed for clearly defined purposes, be honest about the reasons for collecting information, and do so by lawful means [16]. Although COPPA does not state a standalone data-minimization clause, the principle is implicit and partially mentioned throughout several sections of the rule. To support compliance with the principle of data minimiza- tion in practice, service providers must enforce the collection of only data items relevant to the purpose, while excluding unrelated ones. For instance, when users are interacting with an application for educational purposes, the system might enforce collection of data from user prompts that are only relevant to that context. Live and dynamic interactions between children and LLM-based applications present unique risks, as these open-ended conversations can lead to unintended disclosures of sensitive or identifiable information. Service providers should therefore pay particular attention to ensure that the system does not collect sensitive information from children unintentionally. Technical solutions exist to support this; for example, the method proposed in [89] filters user prompts in real time to reformulate and redact potentially sensitive and unrelated information. This approach is discussed in more detail in Sec- tion VI. Service providers should apply such data minimization controls at every point where new information may be collected, including user registration, ongoing interactions, and feature upgrades that involve processing new types of data. Purpose Limitation, as outlined in GDPR Article 5(1)(b) and PIPEDA Principle 5 (âLimiting use, disclosure, and retentionâ), restricts the use of personal data to the purposes stated at the time of collection and prohibits subsequent re-use for unrelated purposes. In LLM-based applications for children, this means that data collected for one purpose, such as providing educational or interactive services, cannot later be reused for profiling, marketing, or analytics. Service providers should maintain records of the original collection purposes and enforce them whenever the system performs operations on the data, supporting compliance at all stages of the LLM lifecycle. One technical approach proposed in [90] involves tagging personal data with subjectâpurpose pairs and allowing the runtime system to verify that all data access operations comply with the consented policies. Such mechanisms can help operationalize purpose limitation along with other principles in practice. Transparency, required by Articles 12-13 of GDPR, COPPA Section 312.4, and PIPEDA Principle 8, obliges service providers to communicate clearly about their data practices, particularly to parents and guardians. We use the terms âTransparency and Explainabilityâ to describe the obligation to provide accessible explanations of how personal data are collected, stored, used, and shared, as specified across these regulatory frameworks. Under Section 312.4, service providers must ensure that parents and guardians receive clear notices about the collection and use of their childrenâs personal data before verifiable parental consent is obtained [15]. Such notices must be presented in an easily understandable format, provide complete and accurate information, and avoid unrelated, misleading, or confusing content [15]. Service providers can operationalize this by including dashboards for parents and guardians into the design of their applications; persistent privacy policy toggle that can be accessed at any time by parents and guardians to review data collection and use practices; enabling parents to ask questions and seek clarification on data practices in real-time through easily accessible AI service agent or similar service. In LLM-based applications, the transparency might also include explanations on how the system processes data and arrives at its responses. Such explanations can be provided again, in a concise and clear manner that allows parents and guardians to understand, and the underlying mechanisms. Additionally, recent research highlighted that parents and guardians struggle to understand the full range of risks that generative AI technologies present to children [83]. As a consequence, these researchers discussed that service providers include âexpert-identified risk taxonomyâ which supplies parents with detailed information on the risks posed by a technology [83]. This type of disclosure can be added to parental notices to help them better understand the risks involved with regard to the tool processing private information and make well-informed decisions. Although COPPA focuses primarily on providing information to parents and guardians, GDPR Article 12 additionally requires that information be provided to a child in a clear and accessible manner [14]. This means meeting children where they are at, providing information in a way that children can understand, and allowing children to meaningfully participate in decision- making. Recent literature offers multiple ways to make lengthy legal explanations more intelligible and participatory for children, including the implementation of visual or animated explanations [91]. In addition, research has increasingly looked at the mental models of AI in children that can support the development of explainable AI for this demographic [92]. Meaningful Consent Mechanisms are central to COPPA Rule Section 312.5, GDPR Articles 6 and 8, and PIPEDA Principle 3. In this paper, we use the term âmeaningful consent mechanismsâ as an umbrella concept encompassing verifiable parental consent under COPPA, parental authorization under GDPR TABLE I SUMMARY OF PRIVACY PRINCIPLES ACROSS COPPA, GDPR, AND PIPEDA Privacy PrincipleCOPPAGDPRPIPEDA Data Minimization~â Purpose Limitation~â Transparency and Explainabilityâ~ User Rights~â Accountability~â Security by Designâ Meaningful Consent Mechanismsâ Note:ââ Explicitly required;~â Implicit or partial coverage;Ă â Not formally required Article 8, and meaningful consent under PIPEDA. In LLM- based applications for children, parental consent interfaces and notices must be presented any time before new information from a child is collected. For example, when first registering for the app, when assessing its main functionality, and when a child is requesting other services or upgrades that involve processing new types of information (e.g., voice). Service providers must also, using reasonable methods, confirm that the person who provides the consent is the childâs parent or legal guardian (i.e., through identity card verification, knowledge-based questions) [15]. Additionally, after initial authorization, providers must offer intuitive interfaces that allow guardians to monitor, modify, or revoke permissions in real time, providing dynamic consent functionality. Dynamic consent has been explored in literature [90], which could be adapted for systems incorporating LLMs. User Rights, such as access, correction and erasure (GDPR Articles 15â17, PIPEDA Principle 9 and COPPA Rule Section 312.6), are particularly challenging to uphold in trained LLMs due to the data stored as distributed representations instead of raw records [93]. Although there are methods such as machine learning that help remove sensitive information learned from trained models, issues remain [93], [94]. They are related to the guaranty of forgetting, as well as the challenge of accessing and deleting explicit data that have been incorporated into the model parameters; as well as implicit sensitive data that do not directly identify individuals, but can reveal the identity when used in combination with other information [94]. To avoid these challenges, developers can implement tools that locate and manage stored interaction records and any user-linked data and remove private or sensitive data before it is used in AI fine- tuning or enhancement pipelines. Modern foundational model pre-processing pipelines already use training data filtering for PII detection and removal [43]. Service providers can also provide clear notices, pathways and interfaces that allow parents or guardians to exercise these rights, such as requesting deletion of their childâs data, including conversation records, voice recordings, images, and other media, in a clear and actionable manner. Security by Design, mandated by GDPR Article 32, PIPEDA Principle 7, and COPPA Rule Section 312.8, emphasizes protecting personal data against security threats. In LLM- based applications, multiple vulnerabilities have been identified that can occur at different stages of LLM development and deployment. They include adversarial attacks such as data poisoning, backdoors, gradient leakage and membership infer- ence attacks [21], [23]. Jailbreaking and prompt injection [21], [23]. Correspondingly, a range of defense mechanisms have been explored to improve model security. Recent studies and surveys highlight mitigation strategies including pruning [95], differential privacy [60], identification of malicious prompts, and the use of guardrails to prevent the generation of harmful or unsafe content [23]. The established techniques such as differential privacy can introduce noise during training or fine-tuning, significantly reducing the risk of sensitive data extraction from trained models [60]. Service providers should ensure that their models, along with the associated software and hardware, are effectively secured to protect against security attacks and unintended breaches. Accountability, as emphasized in Article 5(2) of GDPR and PIPEDA Principle 1, requires organizations to demonstrate compliance with the privacy obligations outlined above. Under PIPEDA, this principle specifically requires organizations to assume the responsibility for compliance, typically designating an individual to supervise and ensure adherence to regulatory requirements [16]. In the context of LLM-based applications, the implementation of accountability can involve the establish- ment of responsible AI (RAI) and risk management frameworks [96], which support the responsible development, deployment, and oversight of AI solutions. In addition, organizations can conduct periodic Data Protection Impact Assessments (DPIAs) to identify, evaluate, and mitigate emerging risks [97], [98]. Technical measures can also support accountability through the maintenance of detailed audit logs, documenting data sources, consent events, and model updates. The next section maps these principles to the lifecycle of an LLM-based system. In addition, it provides a detailed explana- tion of technical design patterns and organizational controls mentioned in this section and additional ones, illustrating how they can be operationalized in practice across the lifecycle stages of the LLM-based system. Data Collection Preprocessing (Text to Vectors) Model Training Cleaning, Structuring Tokenization Embedding Operation & Monitoring Evaluation (Initial Validation) Deployment (Serving/Integration) Transformer Block Stack: Shared model architecture reused across training phases Parameter icon (cylinder): Represents learned model weights during training Evaluation: Post-training performance checks for validation, fairness, and compliance Dashed arrows: Indicate feedback or retraining flows (e.g., user feedback or system monitoring signals) Pre- training RLHF Transformer Block Stack Fine- tuning (Training Phases) Shared Transformer Architecture Bias & Safety Performance Metrics Compliance Check Prompt Interface User Query Response Model Hosting Usage Monitoring Structured Sources Web sources Controlled Collection Databases Re-evaluation Continuous Validation Risk Alert Dynamic Adaptation System Observation Quality Assurance Stakeholder Alignment System Validation Retirement System Disposal Parameter Audit logging Fig. 1. LLM System Lifecycle Architecture (adapted from ISO/IEC 5338:2023) VI. MAPPING REGULATORY PRINCIPLES TO LLM ARCHITECTURE In this section, we present our proposed framework for mapping the regulatory principles summarized in Table I to four stages of the lifecycle of the LLM-based system: data collection, model training, operation and monitoring, and continuous validation. For each principle, we discuss corresponding design, technical, and organizational controls. Before examining how these regulatory principles apply at each stage, we reiterate the challenges of extending principles originally developed for traditional, deterministic information systems to the probabilistic and dynamic behaviour of LLM- based applications. A. Challenges in Applying Privacy Regulations to LLM-based Applications Ensuring compliance with privacy regulations is particu- larly challenging for LLM-based applications. Previous re- search has emphasized that unlike traditional systems such as databases, which store information in discrete records, LLMs are probabilistic and encode information patterns within high-dimensional parameters [29]. In trained LLMs, the data, including personal information, become embedded in the model parameters, making it difficult to isolate or remove specific information, which makes it particularly challenging to uphold several GDPR rights, including the Right of Access, Right to Rectification and the Right to Erasure or the âRight to Be Forgottenâ [29]. Given that previous research has identified risks associated with data memorization [11], which might later lead to disclosure of such data during inference, it would be critical to allow individuals to access, correct, or request deletion of their data from model parameters. However, as discussed in [29], this is currently difficult to achieve in practice, and requires complex technical mechanisms such as machine unlearning [25] among other approaches, which currently have some limitations [94]. Furthermore, ensuring meaningful consent across the differ- ent stages of the lifecycle of LLM-based applications remains a challenge. Children, in their interactions with LLM-based applications, can disclose personal details that services were neither expected nor intended to collect [99], [100]. Although it is possible to request renewed or dynamic consent when new categories of data are processed, in practice, this can be difficult to operationalize, particularly when consent must be managed separately for different data types, processing reasons, and downstream system uses [101]. Consent obtained at a registration may, therefore, not adequately cover the disclosures children make during subsequent interactions [101], [102]. To help mitigate this risk, researchers have proposed mechanisms that detect and reformulate sensitive or out-of-context informa- tion in user inputs before they reach the core system, effectively functioning as input filters or warnings to reduce unnecessary disclosure of private information and protect user privacy during conversational interactions with LLM-based applications [89]. Nonetheless, there is no guarantee that the system will not capture personal details in conversation histories and later incorporate them into model updates. These challenges emphasize persistent gaps when applying existing regulatory principles to LLM-based applications for children. In the proposed framework, we aim to translate regulatory principles into technical controls. At the same time, we acknowledge that, in some instances, adherence to legal requirements and data protection may remain partial. B. Lifecycle Mapping To ensure that privacy protections are proactively embedded into the design of LLM-based applications for children, in accordance with PbD approach, we mapped the regulatory principles and their respective controls to the lifecycle of LLM-based applications. Table IV illustrates our complete framework: the left column presents the four stages of the LLM-based system lifecycle, the centre column maps specific privacy principles identified in the three regulations, and the right column lists the proposed risk mitigation measures and compliance controls. The next sections discuss each of these in detail. 1) Data Collection: Data Collection is the foundational phase in the lifecycle of LLM-based applications. During this phase, systems gather raw data, including digital content, books, and code repositories. Data collection takes place before the Model Training phase and may also occur later during the Operation and Monitoring phase. The data collected during the Operation and Monitoring phase can then be used for domain adaptation or model refinement. The entire system lifecycle is presented in Figure 1, which is adapted from ISO/IEC 5338:2023 [103]. Because systems can ingest childrenâs personal and behavioural data in these corpora, the Data Collection phase must comply with the following regulatory principles: Data Minimization, Purpose Limitation, and Meaningful Consent Mechanisms. Data Minimization requires applications to limit data collec- tion to only what is necessary for their intended use [14], [16]. GDPR Article 5(1)(c), PIPEDA Principle 4 along with implicit coverage in COPPA, all prohibit gathering more information than is reasonably required to enable participation in the online activity or service. To apply this principle during the Data Collection phase, designers and developers can use strategies such as input-filtering through techniques like those deployed in [89] whereby user inputs are filtered and re-formulated to remove unrelated and potentially sensitive information. This limits the number of data fields collected. For example, systems should not collect identifiable information, such as full names, birth dates, or contact information, unless it is collected on lawful basis [14]. As discussed in earlier sections, a key vulnerability of LLMs is their ability to memorize strings of training data, which may be reproduced verbatim during inference time [11]. As such, the main goal of data collection phase with regards to privacy is to ensure that private and sensitive data do not make it to the training corpora. Furthermore, systems should not automatically save background data (such as cookies, device IDs, or metadata) or interaction data for model enhancement by default, and should only use these data when their use is clearly communicated and supported by valid consent. Purpose Limitation states that the system should collect data only for specific, legitimate purposes, and should not re-purpose them without obtaining additional consent. This principle is codified in GDPR Article 5(1)(b) and reflected in PIPEDA Principle 5. COPPA Rule Sections 312.4 and 312.5 require operators to clarify purposes and obtain new verifiable consent if those purposes change. For example, when a system collects data to customize tutoring responses, it must not use that information for behavioural advertising or analytics unless there is a valid legal basis or renewed consent is obtained. To support this principle, privacy engineering frameworks suggest various technical controls, including purpose-specific data tagging [90], scoped access controls [104], and provenance- enabled access control to preserve the initial collection purpose throughout processing [111]. These methods reduce the risk of unintended data reuse and strengthen compliance with regulatory requirements. Meaningful Consent Mechanisms require that, before data collection or processing occurs, the system obtains verifiable parental or guardian consent, as required by Section 312.5 of the COPPA Rule, Articles 6 and 8 of the GDPR, and Principle 3 of PIPEDA. While the legal responsibility rests with parents or guardians to provide consent, children should still receive age-appropriate explanations of how the system processes data to ensure transparency and respect for their developing autonomy. Consent interfaces should, therefore, be child- friendly; they should use simplified language, visualizations or video explanations to convey how the system will use their data [91]. The system must also include verification methods to ensure that consenting adults are of legal age (e.g., knowledge- based questions, verification through a bank account, or a digital signature) [15]. Furthermore, consent should remain dynamic [90], [116]. Dynamic consent allows parents or guardians to review, update, or revoke consent. The formal framework for consent management presented in [90] demonstrates how dynamic consent can be achieved in practice. Their framework enables users (e.g. parents or guardians) to access and view current privacy settings in real-time through a user-friendly interface, like a privacy dashboard, and update these settings when their preferences change. 2) Model Training: The Model Training phase turns raw data into learned representations. This phase presents additional privacy risks if earlier safeguards fail and personal or sensitive information, or data that can indirectly identify someone, enters model parameters. To reduce these risks, developers TABLE IV MAPPING PRIVACY PRINCIPLES AND CONTROLS ACROSS THE LLM LIFECYCLE LLM Lifecycle StageRelevant Regulatory PrinciplesTechnical and Organizational Controls Data CollectionData Minimization; Purpose Limitation; Meaningful Consent Mechanisms Input filtering to remove unnecessary or sensitive data [89]; limited collection of identifiers and metadata; purpose-subject data tagging [90]; scoped access controls [104]; verifiable parental consent; age-appropriate consent interfaces and parental dashboards. Model TrainingData Minimization; Security by Design; Accountability Task-specific fine-tuning with minimal datasets [105]; PII removal or anonymization [106], [107]; encrypted data storage and access controls [108], [109]; validation against data poisoning [21]; differential privacy [110]; gradient-based pruning to mitigate backdoors [95]; data provenance records [111]; DPIAs [97]; Datasheets [49] and Model Cards. Operation and MonitoringMeaningful Consent; Transparency and Explainability; Purpose Limitation; Security by Design Real-time input filtering [89]; parental dashboards and consent controls [90]; age-appropriate explanations [91], [92]; purpose-restricted data use [90]; ephemeral session memory [112]; default exclusion of interaction data from training; detection of prompt injection, jailbreaks, unsafe outputs [21]. Continuous ValidationAccountability; User Rights; Meaningful Consent Mechanisms; Security by Design Periodic audits and DPIAs [113]; logging of system and consent changes; clear interfaces to access, modify or delete data [114]; revalidation of parental consent; adversarial testing [115]; integrity checks and model audits [113]. Note: Left Column â LLM Lifecycle Stage; Middle Column â Regulatory Principles; Right Column â Technical and Organizational Controls. should follow the principles of Data Minimization, Security by Design, and Accountability to ensure that sensitive or personal information does not enter the training process. Data Minimization requires using only the smallest and most relevant subset of data required [105]. Service providers can align their practices with legal requirements by fine-tuning a foundational model on datasets directly applicable to the task. For example, a math tutoring model should rely on math- related educational content. During further refinement with user interactions, developers must include only the interactions for which explicit consent was given by parents to allow use in model refinement. These data should have PII and other sensitive content removed. Techniques for achieving this include input-filtering, anonymization, and PII reduction methods [89], [106], [107]. Security by Design requires safeguards at every stage of model training. Article 32 of GDPR, Principle 7 of PIPEDA, and Section 312.8 of the COPPA Rule collectively mandate the implementation of appropriate technical and organizational measures to protect personal data. During model training, datasets should be encrypted in both transit and rest, access to training data must be restricted through access control mechanisms, and data sources should be thoroughly analyzed and validated to mitigate data poisoning attacks, in which adversaries deliberately manipulate training data to compromise model decision-making [21]. Security teams must also protect training pipelines from privacy attacks, including gradient leakage attacks, in which adversaries exploit shared gradients to reconstruct sensitive training data. A commonly adopted approach to mitigating such risks is the use of differential privacy [21], which introduces calibrated noise into gradients or parameter updates during training [110]. Prior work has shown that differential privacy can provide privacy guarantees while maintaining an acceptable model utility [26]. In addition, security teams must also protect trained models from backdoor attacks, in which adversaries insert hidden tokens to manipulate a model to behave in a particular way, for example, generating malicious output when a trigger token is invoked [21]. Recent studies suggest techniques such as gradient-based pruning, which removes model components that show low sensitivity to loss gradients which might potentially be backdoor triggers [95]. Accountability helps organizations demonstrate compliance with data protection laws and uphold ethical principles. Article 5(2) of GDPR clearly requires accountability, similar to the obligations outlined in PIPEDA Principle 1. To meet these principles, stakeholders should implement responsible AI and risk management frameworks [96], maintain data provenance records [111], verify consent for data related to children [90], and apply privacy and security protections against attacks [21]. In addition, DPIAs can help identify high-risk data categories [97]. Tools such as Datasheets for Datasets [49] and Model Cards can explain how the system handles data and uses models, improving transparency, trust, and ethical practices, particularly for applications designed for children. 3) Operation and Monitoring: During the operation and monitoring phase, LLM-based applications interact directly with users to provide real-time output. This phase introduces unique privacy and security challenges, particularly for chil- dren. Because interactions between children and LLM-based applications are live and dynamic, they involve ongoing data exchange and adaptive system responses, during which children can unintentionally reveal sensitive or personal information [89]. Such data may be stored in conversation histories and may be used for future model refinement or retraining [117]. Furthermore, in some jurisdictions, temporary storage for purposes such as safety monitoring or compliance with law enforcement requests is permitted, which can potentially increase the risk of unauthorized insider access. Another set of risks relates to information sharing activities with third parties. To reduce these risks, developers should apply the principles of Meaningful Consent, Transparency and Explainability, Purpose Limitation, and Security by Design. These safeguards should be complemented by continuous supervision and safety monitoring to ensure responses remain appropriate for children and to detect potential disclosures of distress or self-harm. There is a growing area of research in this direction [118]. Transparency and Explainability help build trust and support informed oversight in interactions between parents or guardians, their children, and LLM-based applications. Articles 12â13 of the GDPR and Section 312.4 of the COPPA Rule establish transparency requirements, calling for clear and accessible information about how personal data are processed, with par- ticular attention to age-appropriate communication. The system should therefore provide simple explanations on what kind of data it will collect during interactions with children, how these data are used, stored, and shared. In LLM-based applications, transparency might also involve explanations on how outputs are generated and how user inputs influence those outputs. Service providers can implement parental dashboards with clear and complete information on data flows to facilitate informed decision-making. In addition, service providers can implement visualizations and video-based explanations, as well as reduce the reliance on dense textual disclosures to support meaningful transparency, especially when communicating with children [91]. Furthermore, system developers should involve children in the design of transparency mechanisms. Participatory and co- design approaches that incorporate childrenâs preferences and perspectives can improve visibility into data processing using AI [91]. Recent work on childrenâs mental models of AI further suggests that it is useful to address common misconceptions and provide meaningful explanations of how AI decisions are made [92]. Meaningful Consent helps ensure parents or guardians have control over what happens to their childrenâs data once they start interacting with an LLM-based application. According to COPPA Rule 312.5 (Parental Consent), and GDPR Articles 6 and 8, parental or guardian consent must be verifiable. This means, service providers should maintain logs of parental or guardian consent. Consent should also be viewed as an ongoing process, reviewed whenever data use or system functionality changes, or when parents or guardians wish to make changes to what data about their children they want to share. As discussed above, this can be supported through a dynamic consent management framework, as proposed in [90]. Parental or guardian dashboards can help families understand what information is collected, adjust permissions, or withdraw consent [90]. This enables user autonomy while supporting compliance with regulatory requirements [83], [119]. Purpose Limitation ensures that user input during the Operation and Monitoring phase is used only for the immediate conversational context. This principle blocks secondary uses, such as behavioural profiling, analytics, or fine-tuning, unless there is a valid legal basis or explicit renewed consent. To implement these legal requirements, service providers must obtain parental or guardian consent discussed earlier. The purpose and subject tagging can then be used to ensure the data collected from children is only used for purposes for which parental or guardian consent was provided [90]. To avoid data leakage during interactions, systems can use ephemeral session memory designs to delete conversational context and temporary embeddings after each use [112]. Furthermore, LLM-based applications used by children should not collect user input for model training or sharing data with third parties, by default; such collection should be enabled only through parental or guardian controls. The only exception is when the sharing with third parties is necessary to provide the service [15]. Security by Design is important during live operations to protect against threats such as prompt injections, where malicious input bypasses LLM safety filters and manipulates model output, for example, to reveal sensitive information [21]; jailbreaks, which manipulates the applicationâs safety mechanisms to elicit prohibited responses [21]; and inference leakage, where the model inadvertently reveals sensitive or training-related information [120]. This principle is required under GDPR Article 32 and COPPA Rule Section 312.8. To maintain security during the Operation and Monitoring phase, security teams can implement input filtering to block harmful and adversarial prompts and perform runtime output filtering for unsafe content [21]. Earlier measures to reduce PII in training data can prevent it from being revealed during inference [21]. 4) Continuous Validation: The Continuous Validation phase involves the ongoing oversight, monitoring, and governance of an LLM-based application after its deployment [121]. This phase ensures that the system remains accurate and reliable, secure, and maintains ethical standards [47]. Unlike the Operation and Monitoring phase, which focuses on real-time user interactions, the Continuous Validation phase addresses long-term system integrity, mitigates biases, incorporates new knowledge and responds to emerging challenges [47]. This phase also involves verification of compliance with privacy and safety standards. Four principles are important during this phase: Accountability, User Rights, Meaningful Consent Mechanisms, and Security by Design. Accountability is necessary when systems undergo updates, re-evaluations, and ongoing monitoring. Articles 5(2) and 24 of the GDPR, and Principle 1 of PIPEDA require organizations to demonstrate their compliance with regulatory requirements. As discussed before, effective data and AI governance program as well as risk identification and mitigation can support organizations in achieving transparent, ethical, and auditable AI systems that align with regulatory requirements [96]. Additionally, DPIAs, along with regular audits and evaluations, help identify issues that arise from feature updates and changes [113]. Because iterative updates can introduce new risks and affect the way user rights are provided, maintaining traceability throughout the lifecycle is a key challenge. To reinforce transparency and trust, organizations should provide clear and accessible channels for regulatory inquiries, user complaints, and rights requests as required under GDPR Article 12 and supported by GDPR Recital 39. User Rights must be actively enforced during the continuous validation phase, as mandated by Section 312.6 of the COPPA Rule, Articles 15-17 of the GDPR, and Principle 9 of PIPEDA. These rights include accessing, correcting, or requesting deletion of personal data, especially critical for child-specific applications. To protect these rights, service providers can provide clear notices, pathways and interfaces support user engagement and understanding [114]. Such interfaces can allow parents or guardians to exercise these rights, such as requesting deletion of their childâs data, including conversation records, voice recordings, images, and other media, in a clear and actionable manner. Additionally, service providers can implement automated request workflows that log and time- stamp actions, creating a verifiable audit trail that supports regulatory compliance and strengthens trust [122]. Meaningful Consent Mechanisms must extend beyond initial onboarding to maintain compliance with COPPA Rule Section 312.5, GDPR Articles 6 and 8, and PIPEDA Principle 3. Con- sent must remain verifiable and informed, allowing parents or legal guardians to authorize, monitor, and revoke data collection or processing at any time. During continuous validation, this involves periodic revalidation, prompting guardians to review and reaffirm consent whenever new features, data uses, or policy changes arise. Security by Design is in accordance with COPPA Rule Section 312.8, GDPR Article 32, and PIPEDA Principle 7, which require controls to protect personal information from unauthorized access and modification, or inference leakage. For LLM-based applications designed for children, this involves implementing input and output filtering to prevent harmful or inappropriate content displayed to children. Adversarial testing can help identify and protect from prompt injection [115] and jailbreak attacks [21]. Additional measures, such as strong authentication, encrypted storage and transmission, and hash-based model integrity checks, enhance data protection and auditability [123]. Privacy-preserving techniques, including differential privacy, further protect sensitive data from being reconstructed in system responses [40], [124]. Periodic model audits and assessments should detect behavioural changes or data drift caused by authorized and unauthorized updates, providing ongoing verification of technical integrity and regulatory compliance [113]. C. Anthropomorphic and Nudging Risks In addition to the privacy risks associated with LLM-based applications previously discussed, it is important to highlight a new risk: that these applications may present themselves as human beings. This can potentially influence children to disclose more personal information than they would normally share. Therefore, it is crucial to design LLM-based applications in a way that does not unduly influence children. The United Nations Convention on the Rights of the Child (UNCRC) [125], the Age Appropriate Design Code by the UK Information Commissionerâs Office (ICO) [20], and previous work provide valuable guidance for achieving this objective. Regarding childrenâs rights, Article 17 of the UNCRC emphasizes a childâs right to access information that promotes their social, spiritual, and moral well-being, while also requiring the establishment of appropriate guidelines to protect children from harmful information [125]. Based on this definition, children have the right to know how service providers collect their data. Similarly, the AADC Transparency standard encourages developers to be open and transparent about their services, including how childrenâs personal data is collected, used, and shared, and to communicate this information in a manner appropriate for their age. The guidance further notes that transparency is linked to the principle of fairness: when service providers are unclear, opaque, or dishonest about how their services operate or how data is collected, the resulting data processing and use are likely to be unfair [20]. Anthropomorphic design may contravene both the right to information under the UNCRC and the AADC transparency guidance, because children may not recognize that they are dis- closing personal information when interacting with systems that present themselves as friendly or empathetic. Some researchers have therefore expressed concern that anthropomorphic design can function as a manipulative data-elicitation practice [32], [69]. These guidelines, taken together, suggest that service providers should be transparent about their data collection practices. Anthropomorphic effects raise concerns because they may (1) unduly influence childrenâs disclosure behaviours and (2) result in data collection and use that might not be sufficiently transparent. Reducing anthropomorphic effects is therefore a proactive measure and a design objective to protect privacy. Current recommendations suggest that children should be aware that they are communicating with an information system, not a human being [69]. For this, the systems need to explicitly identify itself as an AI system [32], [69]. Reminders can also be added throughout interactions. For instance, systems may explicitly inform children with messages such as, âI am a machine and do not possess an understanding of human emotionsâ [32]. Additional measures suggested in existing research to reduce anthropomorphic effects might involve preventive experiments in which developers identify whether certain agent dispositions or responses lead to more information disclosures and remove them before the system is released to actual users [70]. Nudging has been a recognized phenomenon in information systems, and, in fact, AADC explicitly recognizes and calls for avoiding nudging techniques for a child-centred AI [20]. Nudging, as discussed in Section I, refers to the design of online services that encourage children to provide more personal information than they would otherwise volunteer [20]. In the case of interactive agents, nudging can take the form of follow-up prompts such as âTell me moreâ or âWhat do you feel like right now?â, which can lead to more information disclosure. AADC standard guides against the use of any nudge techniques that undermine privacy. In addition, it encourages the use of pro-privacy nudges where appropriate and nudges that promote their health and well-being [20]. In practical cases involving LLM-based technologies, this can look like removing follow-up questions from system design, setting reminders that the information provided to systems might be accessed and read, and encouraging speaking to parents or guardians when certain information is shared. D. Case Study: Educational LLM Tutor for Children We present an LLM-based educational tutor for children under 13 to demonstrate how we can operationalize the proposed framework in a real-world application deployment (Figure 2). The application acts as a personalized homework assistant for math, language, arts, and science. Children interact through a conversational chatbot interface, powered by an LLM on the background. Data Collection Principles include Data Minimization, Purpose Limitation, and Meaningful Consent. Parents and guardians initiate the onboarding with verifiable parental consent (e.g., a microcharge on a credit card or an equivalent method such as ID verification), providing age-appropriate disclosures and a consent dashboard that supports granular choices made that can be revisited at a later date. The system minimizes and pseudonymizes inputs (input filtering), blocking off background identifiers such as cookies, device IDs, and precise location by default and only activating them with explicit, specific consent. Purpose tags are also attached to data at ingress, where training uses opt-in by default for childrenâs data. Model Training. Principles include Data Minimization, Security by Design, and Accountability. The system filters corpora gathered from child-tutor interactions and maintains exclusion lists for direct identifiers (e.g., names, addresses, etc.). The system uses differential privacy and strict access controls for any child-related fine-tuning [126]. The documentation records all training data lineage, lawful basis, and consent artifacts (in datasheets and model cards) to support audits and unlearning requests. Operation and Monitoring. Principles include Trans- parency/Explainability, Purpose Limitation, Security by Design, and User Rights. The chatbot identifies itself as an automated system, avoids human-like cues, and uses neutral language in its responses. Its dialogue policies are designed to prevent prompts that might elicit unnecessary personal disclosures. By default, its session memory is ephemeral, and conversation logs are minimized, encrypted, and retained only for explicitly stated safety or quality assurance purposes, under settings controlled by parents or guardians. Figure 2 illustrates the architecture of the LLM-based educational tutor, highlighting the alignment between the phases of the system lifecycle and the principles of privacy. The chatbot communicates with a child in an age-appropriate language. A parent or guardian can review the summary of the interactions. Rather than exposing full transcripts, the dashboard for parents or guardians presents aggregated insights such as discussion topics and safety alerts, providing oversight without undermining the childâs autonomy or trust [20]. Model retraining on childrenâs interactions is disabled by default and can only be enabled with an explicit parental or guardian consent. Runtime controls include input sanitization, adversarial prompt detection, output filtering, and secure data transmission and storage to prevent unauthorized access. Continuous Validation. The system performs ongoing moni- toring and evaluation of model behaviour, consent compliance, and privacy metrics, allowing detection of anomalies, drift, or privacy violations and supporting timely updates or interven- tions to maintain child safety and regulatory compliance. VII. DISCUSSION Given the rapid expansion of AI-based systems and childrenâs increasing engagement with them, governance mechanisms for child-facing AI systems must anticipate risks rather than respond to harms after they materialize. Our review shows that children may face multiple privacy risks that accumulate across the lifecycle of LLM-based applications. These include technical vulnerabilities, governance gaps, and design choices that may disproportionately affect children. Embedding privacy as a foundational design requirement can help mitigate these risks and ensure that AI systems are appropriate, lawful, and protective of childrenâs rights. To embed privacy protections into the design of LLM-based applications for children, we proposed an integrated framework based on the PbD approaches, that translates regulatory privacy principles from the GDPR, COPPA, and PIPEDA into actionable technical and organizational controls across the lifecycle of these systems. The framework maps regulatory requirements to the lifecycle stages of LLM-based applications, including Data Collection, Model Training, Operation and Monitoring, and Continuous Validation. This mapping identifies where privacy risks emerge and how designers and developers can mitigate them through design-time and operational controls informed by recent academic literature on privacy, security, and child-specific system design practices. In the remainder of this section, we examine how our pro- posed framework operationalizes PbD principles in childrenâs LLM-based applications and discuss the technical, regulatory, and practical implications of adopting a lifecycle-oriented privacy protection approach. We highlight both the strengths of this approach and the persistent challenges that remain when translating legal norms into deployable systems for child-facing AI. A. Alignment Between the Proposed Framework and PbD Our proposed framework operationalizes the seven PbD principles throughout the lifecycle of LLM-based applications for children: proactive risk management, privacy by default, privacy embedded into design, positive-sum functionality, end- to-end security, visibility and transparency, and respect for user privacy across the LLM lifecycle. Key technical measures include input filtering, subject-purpose-based data tagging, verifiable parental consent, adversarial training, differential privacy, ephemeral session memory, guardian dashboards, secure data transmission, age-appropriate explanations, and avoidance of human-like cue and nudging, among others. Specifically, our framework aligns practices with the first principle, proactive risk management, by implementing input Consent Form Child User Chatbot Interface (LLM Tutor) LLM (Model) Cloud / Server Guardian Consent Form Trigger Access Attempt Access Enabled Real-Time Interaction Guardian Dashboard User Input Personalized Response Interaction Log Data Retrieval Interaction Data Storage 1 1 2 34 2 7 7 456 5 3 The numbered labels represent embedded privacy principles: â Data Minimization ⥠Purpose Limitation âą Meaningful Consent Mechanisms ⣠Transparency & Explainability †Security by Design â„ User Rights ⊠Accountability Arrow Types: (dark teal) User Interaction (blue) Consent Flow (burnt orange) Data Handling (dashed crimson) Access Control & Oversight Fig. 2. Educational LLM Tutor Architecture Aligned to Lifecycle Stages and Privacy Principles filtering to minimize data collection and by proposing subject- and purpose-based tagging, which reduces the risk of over- collection and unauthorized uses. The framework further sup- ports privacy as a default principle by prohibiting the collection of user interactions for model training by default, unless parents or guardians explicitly enable it or provide consent. It supports positive-sum functionality by balancing privacy protections with system utility. Specifically, the framework shows how systems can still provide educational benefits without requiring unnecessary data collection, such as names or ages. The framework aligns with the principle of end-to-end security by in- corporating adversarial training, privacy-preserving techniques such as differential privacy, ephemeral session memory, and secure data storage and transmission, among other measures. The framework enables visibility and transparency through parental and guardian dashboards, age-appropriate explanations, and auditable system logs that enhance observability and accountability. Finally, the framework supports respect for user privacy throughout the LLM lifecycle through verifiable parental or guardian consent mechanisms, dynamic consent, and ongoing parental oversight. The framework also addresses child- specific considerations, such as age-appropriate explanations and minimizing nudging or human-like cues, further supporting ethical and legally compliant LLM deployment. B. Limitations Although this work demonstrates that PbD principles can be operationalized through regulatory and technical controls in LLM-based applications for children, significant challenges and gaps remain in the four key stages of the LLM lifecycle. These must be addressed to move from privacy-aware prototypes to scalable, legally compliant, and ethically robust deployments. C. Challenges in Implementing PbD Principles Across LLM- lifecycle Data Collection. Even with minimization policies and parental consent workflows in place, children may still disclose personal or sensitive details through open-ended interactions given they might not fully understand online privacy risks [99]. Additionally, it has been documented that children often access LLM-based applications to seek emotional support and companionship [83], which can reveal a lot of private details. Current input filtering and entity redaction techniques provide partial protections by removing explicit identifiers such as names, locations or phone numbers [89], however children may use indirect phrasing, speak of personal events without identifiers, which can still reveal details about them, their lifestyles, or associations. For example, a child can mention âthe park near grandmaâs houseâ or âmy school trip tomorrow,â which cannot always be captured by rule-based or keyword- driven filters. These limitations highlight the residual risk of identifiable data entering the system despite compliance mechanisms. To address this, future work should focus on adaptive filtering methods that combine natural language understanding, context-awareness, and child-specific linguistic patterns to detect disclosures. Importantly, such approaches must strike a balance between privacy protection and usability, ensuring that children are not discouraged from meaningful engagement with educational or assistive applications. Model Training. A persistent challenge is to mitigate model memorization, where rare or sensitive inputs, such as a childâs name, address, or personal story, become embedded in model parameters and risk being revealed during inference. This risk is heightened in child-specific applications because children may inadvertently share uniquely identifiable details in ways adults do not. While techniques such as differential privacy offer partial protection, empirical studies show that LLMs can still reproduce memorized phrases when subjected to targeted prompting or extraction attacks [127], [128]. Current defenses often struggle to balance utility with privacy, as strong noise addition can degrade the accuracy and general usability of the model [26], [126]. Emerging approaches, such as scalable machine unlearning methods [94] that selectively erase sensitive records without requiring full retraining, offer promising directions. Developing such strategies would directly reduce the long-term risks of exposing childrenâs personal information, while maintaining model quality for safe educational and developmental use. Operation and Monitoring. Providing meaningful trans- parency to children, parents and guardians remains difficult due to the probabilistic and opaque nature of LLMs. Simply logging responses or publishing model cards is insufficient, as these artifacts often do not capture how sensitive information is processed in real time [50], [113], [122]. The interfaces must therefore evolve to translate complex reasoning processes into age-appropriate explanations for children [92] and actionable oversight insights for parents and guardians. However, achiev- ing this balance requires advancing explainability methods beyond post-hoc interpretations towards accountability dash- boards that visualize various risks associated with a particular LLM-based tool [83]. Dynamic consent management further complicates this challenge: while static consent forms exist, mechanisms for guardians to update or revoke permissions in real time are underdeveloped, particularly for unstructured, conversational input streams, which characterize current AI technologies [129]. Consent, therefore, must be viewed as a living document [129]. Research into adaptive monitoring frameworks, coupled with lightweight privacy-preserving audit trails, is critical to operationalizing transparency and consent in practice [122], [130]. Finally, although ephemeral session memory can reduce privacy risk, it may conflict with safety, abuse prevention, and debugging needs, and therefore systems must clearly specify what is retained (if anything), for how long, and under what access model and purpose constraints. Continuous Validation. Maintaining privacy and safety protections after deployment is challenging due to both evolving threats and jurisdictional differences [21], [103]. Protective mechanisms implemented at launch can degrade over time if not reinforced through periodic testing, red teaming, and model audits [62], [113]. Moreover, cross-border use of LLM-based applications for children demands governance architectures that harmonize COPPA, GDPR, and PIPEDA obligations while adapting to local enforcement nuances [20], [131]. Finally, current engineering frameworks for AI are not well- suited to the developmental and cognitive needs of children. There is a critical need for children-specific interface patterns, consent workflows, and privacy risk models that integrate both technical and developmental protections [1]. Without such adaptations, PbD risks remaining a high-level aspiration rather than a fully operationalized standard in LLM-based ecosystems for children. D. Limitations of the Proposed Framework Although the proposed framework demonstrates how privacy and regulatory principles can be mapped onto the architecture of a child-specific LLM tutor, several limitations remain. First, the framework operates at the conceptual and architectural level. Translating these into production systems requires significant engineering effort, financial investment, and organizational commitment [103], [132]. In practice, smaller developers or educational institutions may struggle to implement the full set of recommendations, particularly those that require advanced cryptographic methods or dedicated compliance teams. Additionally, smaller developers may lack leverage against large vendors that control the infrastructure. Second, some of the technical protective mechanisms, such as differential privacy, machine unlearning, and tamper-evident logging, are still areas of active research [24], [27], [122]. Although promising in theory, their deployment at scale in multibillion-parameter models remains experimental and may introduce trade-offs in accuracy, usability, or system performance. This raises questions about whether current tech- nological maturity is sufficient to meet the high expectations set by regulatory frameworks like GDPR, COPPA, and PIPEDA [29], [133]. Finally, the legal and social contexts are dynamic. Regula- tions evolve, cultural expectations of childrenâs digital rights differ between jurisdictions, and parents can vary in their ability or willingness to engage with dashboards and consent mechanisms [83], [134]. The proposed model cannot fully capture these contextual nuances, and therefore, its effectiveness will depend on ongoing adaptation, stakeholder participation, and iterative design informed by real-world use. Collectively, these limitations indicate that the proposed framework should be regarded as a foundational basis for on- going development. It is the collaboration among technologists, educators, regulators, and families, which will be essential to refine and operationalize these principles in practice. Nonetheless, the proposed framework represents a mean- ingful step toward designing safer, more secure, and more compliant LLM-based tools for children. Despite the unprece- dented exposure to these tools during the formative years and their unique vulnerabilities, children remain underrepre- sented in discussions of AI risks. Recent policy developments, particularly in the United States and increasingly in the United Kingdom, emphasize age verification as a primary mechanism for protecting childrenâs online privacy and safety [18], [135]. While these measures may reduce underage access to certain applications, they risk overemphasizing verification at the expense of broader privacy-preserving mechanisms. Combining age verification with additional measures, such as those proposed in our framework, including data minimization, dynamic consent, security by design, and protections throughout the data lifecycle, offers a more balanced approach. Together, these mechanisms help build trust with parents and guardians, who need assurance that their childrenâs data is handled responsibly [130], [136], and support children in exploring and interacting with technology while minimizing risks to their privacy. VIII. CONCLUSION In this paper, we examined how privacy and data protection principles derived from COPPA, GDPR, and PIPEDA can be operationalized in the design and deployment of LLM-based applications for children. Through comparative legal and tech- nical analysis, we articulated seven core regulatory principles, Data Minimization, Purpose Limitation, Meaningful Consent Mechanisms, Transparency and Explainability, Security by Design, Accountability, and User Rights, as fundamental for the development of ethically aligned and legally compliant LLM systems [14]â[16]. We mapped these principles to four stages of the LLM lifecycle: Data Collection, Model Training, Operation and Monitoring, and Continuous Validation. Each stage presents unique privacy risks for children, requiring technical and organizational controls to support compliance and protect child data. To address these risks, we proposed a PbD-aligned privacy engineering framework informed by guidance from data protection authorities and childrenâs rights organizations. Some recommended protective mechanisms include local or ephemeral data processing, session-based memory isolation, real-time input filtering, and guardian-based dashboards to enhance transparency and oversight. Where limited retention is necessary for safety monitoring, abuse prevention, or incident response, retained data should be minimized, time-bounded, access-controlled, and purpose-separated from model training or unrelated analytics. Security-by-design controls in this context are also directly tied to child privacy outcomes, because failures such as prompt injection, jailbreaks, poisoning, or backdoors can lead to disclosure of sensitive child data, bypass of parental controls, or unsafe elicitation of unnecessary personal informa- tion. The proposed framework demonstrates how abstract legal obligations can be translated into enforceable technical controls, accessible interfaces, and auditable governance, supporting both regulatory compliance and the cultivation of trust in LLM- based applications for children. Using an educational LLM tutor case study, we illustrated how these principles can be operationalized in real-world contexts. By operationalizing these privacy principles into actionable design practices, this work provides a roadmap for supporting compliance with legal requirements in LLM-based applications for children, along with enhancing safety, privacy, and well-being of children. While controls mentioned in our framework offer improved privacy protection, persistent vulnerabilities remain, including implicit profiling, model memorization of sensitive disclosures, and insufficient transparency for parents or guardians [10], [59], [131]. Additionally, the proposed framework remains conceptual. Future research should focus on advancing scalable technical solutions that extend beyond conceptual frameworks. Priority directions include improving robustness against mem- orization and extraction attacks, developing explainability methods tailored to children and guardians, enabling dynamic consent management through mechanisms such as verifiable credentials or smart contracts [97], [127], [137], and supporting cross-border regulatory interoperability for globally deployed LLM systems. REFERENCES [1]UNICEF Office of Global Insight and Policy, âPolicy guidance on AI for children (version 2.0),â UNICEF, Tech. Rep., 2021, accessed: 2025-10-02. [Online]. Available: https://w.unicef.org/innocenti/ media/1341/file/UNICEF-Global-Insight-policy-guidance-AI-children- 2.0-2021.pdf [2]H. Goel and G. Chaudhary, âSecuring the digital footprints of minors: Privacy implications of ai,â Balkan Soc. Sci. Rev., vol. 23, p. 235, 2024. [3]C. Zhang, X. Liu, K. Ziska, S. Jeon, C.-L. Yu, and Y. Xu, âMathemyths: leveraging large language models to teach mathematical language through child-AI co-creative storytelling,â in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 2024, p. 1â23. [4]I. Carmo, P. Costa, and P. Santana, âBoosting childrenâs reading motivation with llm-generated story crossovers,â in 2024 International Conference on Graphics and Interaction (ICGI). IEEE, 2024, p. 1â8. [5]S. Sun, Z. Wang, J. Zhang, X. Zhao, and X. Qiao, âStorychat: An interactive storytelling system for engaging with agent characters in a story,â in 2024 17th International Symposium on Computational Intelligence and Design (ISCID). IEEE, 2024, p. 213â217. [6]H. Kibriya, W. Z. Khan, A. Siddiqa, and M. K. Khan, âPrivacy issues in large language models: a survey,â Computers and Electrical Engineering, vol. 120, p. 109698, 2024. [7]I. Barbera, âAI privacy risks & mitigations â large language models(llms),âEuropeanDataProtectionBoard(EDPB), Technical Report, 2025, accessed: 2025-10-22. [Online]. Available: https://w.edpb.europa.eu/system/files/2025- 04/ai- privacy- risks- and-mitigations-in-llms.pdf [8] S. A. Gelman, S. E. Nancekivell, Y. Lee, and F. Schaub, âChildrenâs understanding of digital tracking and digital privacy,â Handbook of Children and Screens, p. 643, 2024. [9] N. Kurian, ââno, alexa, no!â: designing child-safe AI and protecting children from the risks of the âempathy gapâ in large language models,â Learning, Media and Technology, p. 1â14, 2024. [10]N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel, âExtracting training data from large language models,â in Proceedings of the USENIX Security Symposium, 2021, p. 2633â2650. [11]N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, âQuantifying memorization across neural language models,â in The Eleventh International Conference on Learning Representations, 2022. [12](2025) Childrenâs advertising review unit â decision: Buddy AI. Accessed: 2025-10-24. [Online]. Available: https://bbbprograms.org/ media/newsroom/decisions/buddy-ai [13] J. E. Cohen, Between Truth and Power: The Legal Constructions of Informational Capitalism. Oxford University Press, 2019. [14]European Parliament and the Council of the European Union, âRegulation (EU) 2016/679 of the european parliament and of the council on the protection of natural persons,â Official Journal of the European Union, Tech. Rep., 2016, accessed: 2025-07-22. [Online]. Available: https://eur-lex.europa.eu/eli/reg/2016/679/oj [15]Federal Trade Commission, âChildrenâs online privacy protection rule (COPPA),â Federal Trade Commission (FTC), 2020. [Online]. Available: https://w.ftc.gov/legal- library/browse/rules/childrens- online-privacy-protection-rule-coppa [16]Office of the Privacy Commissioner of Canada. (2000) The personal information protection and electronic documents act (pipeda). Accessed: 2025-10-24. [Online]. Available: https://laws-lois.justice.gc.ca/eng/acts/ p-8.6/page-1.html [17]European Union, âGdpr recital 38 - childrenâs personal data,â 2016, accessed: 2025-12-08. [Online]. Available: https://gdpr-info.eu/recitals/ no-38 [18] O.ofthePrivacyCommissionerofCanada,âExploring thechildrenâsprivacycode,âOfficeofthePrivacy Commissioner of Canada, Tech. Rep., 2023, accessed: 2025-12-08. [Online]. Available: https://w.priv.gc.ca/en/about-the-opc/what-we- do/consultations/completed-consultations/consultation-children-code/ expl_children-code/ [19] European Union, âGdpr article 8 - conditions applicable to childâs consent in relation to information society services,â https://gdpr-info.eu/ art-8-gdpr/, 2016, accessed: 2025-12-08. [20]UK Information Commissionerâs Office, âAge appropriate design: A code of practice for online services,â Information Commissionerâs Office, United Kingdom, Tech. Rep., 2020, accessed: 2025- 10-02. [Online]. Available: https://ico.org.uk/for- organisations/uk- gdpr- guidance- and- resources/childrens- information/childrens- code- guidance- and- resources/age- appropriate- design- a- code- of- practice- for-online-services/ [21]B. C. Das, M. H. Amini, and Y. Wu, âSecurity and privacy challenges of large language models: A survey,â ACM Computing Surveys, vol. 57, no. 6, p. 1â39, 2025. [22]K. Chen, X. Zhou, Y. Lin, S. Feng, L. Shen, and P. Wu, âA survey on privacy risks and protection in large language models,â Journal of King Saud University Computer and Information Sciences, vol. 37, no. 7, p. 163, 2025. [23]Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, âA survey on large language model (llm) security and privacy: The good, the bad, and the ugly,â High-Confidence Computing, vol. 4, no. 2, p. 100211, 2024. [24] A. Ginart, M. Guan, G. Valiant, and J. Y. Zou, âMaking AI forget you: Data deletion in machine learning,â Advances in Neural Information Processing Systems, vol. 32, 2019. [25]L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot, âMachine unlearning,â in 2021 IEEE Symposium on Security and Privacy (SP).IEEE, 2021, p. 141â159. [26]M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, âDeep learning with differential privacy,â in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 2016, p. 308â318. [27]D. Yu, E. Bagdasaryan, and V. Shmatikov, âDifferentially private fine- tuning of language models,â in International Conference on Learning Representations (ICLR), 2021. [28]W. Kuang, B. Qian, Z. Li, D. Chen, D. Gao, X. Pan, Y. Xie, Y. Li, B. Ding, and J. Zhou, âFederatedscope-llm: A comprehensive package for fine-tuning large language models in federated learning,â in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, p. 5260â5271. [29]G. Feretzakis, E. Vagena, K. Kalodanis, P. Peristera, D. Kalles, and A. Anastasiou, âGdpr and large language models: Technical and legal obstacles,â Future Internet, vol. 17, no. 4, p. 151, 2025. [30] D. Zhang, P. Finckenberg-Broman, T. Hoang, S. Pan, Z. Xing, M. Sta- ples, and X. Xu, âRight to be forgotten in the era of large language models: Implications, challenges, and solutions,â AI and Ethics, vol. 5, no. 3, p. 2445â2454, 2025. [31]H. Al Breiki and Q. H. Mahmoud, âA framework for integrating privacy by design into generative AI applications,â in Proceedings of the AAAI Symposium Series, vol. 6, no. 1, 2025, p. 2â9. [32] N. Kurian, âDevelopmentally aligned ai: a framework for translating the science of child development into AI design,â AI, Brain and Child, vol. 1, no. 1, p. 1â13, 2025. [33]A. Cavoukian, âPrivacy by design: The 7 foundational principles, implementation and mapping of fair information practices,â Information and Privacy Commissioner of Ontario, Canada, Tech. Rep., 2011. [Online]. Available: https://w.privacysecurityacademy.com/wp- content/uploads/2020/08/PbD-Principles-and-Mapping.pdf [34]â,âOperationalizingprivacybydesign:Aguideto implementing strong privacy practices,â 2012. [Online]. Available: https://gpsbydesigncentre.com/wp- content/uploads/2021/08/Doc- 5- Operationalizing-pbd-guide.pdf [35]HillNotes. (2021) Privacy by design: Origin and purpose. Accessed: 2025-06-07. [Online]. Available: https://hillnotes.ca/2021/12/09/privacy- by-design-origin-and-purpose/ [36]P. Kodakandla, âUnified data governance: Embedding privacy by design into AI model pipelines,â International Journal of Novel Research and Development, vol. 9, no. 10, p. 1â5, 2024. [37]M. Barati, G. Theodorakopoulos, and O. Rana, âAutomating GDPR compliance verification for cloud-hosted services,â in International Symposium on Networks, Computers and Communications (ISNCC). IEEE, 2020, p. 1â6. [38]M. Barati, K. Adu-Duodu, O. Rana, G. S. Aujla, and R. Ranjan, âCompliance checking of cloud providers: Design and implementation,â Distributed Ledger Technologies: Research and Practice, vol. 2, no. 2, p. 1â20, 2023. [39] California Legislature, âCalifornia consumer privacy act of 2018 (as amended by the california privacy rights act of 2020),â Cal. Civ. Code § 1798.100â1798.199, 2018. [Online]. Available: https: //leginfo.legislature.ca.gov/faces/codes_displayText.xhtml?lawCode= CIV&division=3.&part=4.&title=1.81.5 [40]M. Miranda, E. S. Ruzzetti, A. Santilli, F. M. Zanzotto, S. BratiĂšres, and E. RodolĂ , âPreserving privacy in large language models: A survey on current threats and solutions,â arXiv preprint arXiv:2408.05212, 2024. [41]Y. Liu, T. Han, S. Ma, J. Zhang, Y. Yang, J. Tian, H. He, A. Li, M. He, Z. Liu et al., âSummary of chatgpt-related research and perspective towards the future of large language models,â Meta-radiology, vol. 1, no. 2, p. 100017, 2023. [42]A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ć. Kaiser, and I. Polosukhin, âAttention is all you need,â Advances in Neural Information Processing Systems, vol. 30, 2017. [43]H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, âA comprehensive overview of large language models,â ACM Transactions on Intelligent Systems and Technology, vol. 16, no. 5, p. 1â72, 2025. [44] X. Junlin, Z. Chen, R. Zhang, and G. Li, âLarge multimodal agents: a survey,â Visual Intelligence, vol. 3, no. 1, p. 24, 2025. [45]W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al., âA survey of large language models,â arXiv preprint arXiv:2303.18223, vol. 1, no. 2, 2023. [46]L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima et al., âThe pile: An 800gb dataset of diverse text for language modeling,â arXiv preprint arXiv:2101.00027, 2020. [47] IEEE Computer Society. (2024) Large language model lifecycle: An examination of the development, deployment, and maintenance of llms. Accessed: 2025-10-24. [Online]. Available: https://w.computer.org/ publications/tech-news/trends/large-language-model-lifecycle/ [48] M. Gupta, C. Akiri, K. Aryal, E. Parker, and L. Praharaj, âFrom chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy,â IEEE access, vol. 11, p. 80 218â80 245, 2023. [49] T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. DaumĂ© I et al., âDatasheets for datasets,â Communications of the ACM, vol. 64, no. 12, p. 86â92, 2021. [50]M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, âModel cards for model reporting,â in Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* â19).New York, NY, USA: Association for Computing Machinery, 2019, p. 220â229. [51] Ofcom, âOnline nation 2024 report,â 2024, accessed: 2025- 10-02. [Online]. Available: https://w.ofcom.org.uk/siteassets/ resources/documents/research-and-data/online-research/online-nation/ 2024/online-nation-2024-report.pdf?v=386238 [52]â,âChildrenandparents:Mediauseandattitudes report 2024,â 2024, accessed: 2025-06-07. [Online]. Available: https://w.ofcom.org.uk/research-and-data/media-literacy-research/ childrens/children-and-parents-media-use-and-attitudes-report-2024 [53]M. Madden, A. Calvin, A. Hasse, and A. Lenhart, âThe dawn of the AI era: Teens, parents, and the adoption of generative AI at home and school,â Common Sense Media, San Francisco, CA, Tech. Rep., 2024. [Online]. Available: https: //w.commonsensemedia.org/sites/default/files/research/report/2024- the-dawn-of-the-ai-era_final-release-for-web.pdf [54]OpenAI, âIntroducing parental controls,â 2025, accessed: 2025-10-02. [Online]. Available: https://openai.com [55]J. Jiao, S. Afroogh, K. Chen, A. Murali, D. Atkinson, and A. Dhurandhar, âSafe-child-llm: A developmental benchmark for evaluating llm safety in child-AI interactions,â arXiv preprint arXiv:2506.13510, 2025. [56]M. Eira, A. Rasouli, and V. Charisi, âParentsâ perceptions about the use of generative AI systems by adolescents,â in Proceedings of the 24th Interaction Design and Children. ACM, 2025, p. 927â931. [57] P. Wisniewski, H. Jia, H. Xu, M. B. Rosson, and J. M. Carroll, â"preventative" vs. "reactive": How parental mediation influences teensâ social media privacy behaviors,â in Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing, 2015, p. 302â316. [58] Z. Chen, L. Xie, L. Zhang, B. Li, Y. Wang, W. Wang, Y. Wang, and J. Chen, âBackdoor attacks and defenses in feature-partitioned collaborative learning,â arXiv preprint arXiv:2202.08313, 2022. [Online]. Available: https://arxiv.org/abs/2202.08313 [59]N. Mireshghallah, H. Kim, X. Zhou, Y. Tsvetkov, M. Sap, R. Shokri, and Y. Choi, âCan llms keep a secret? testing privacy implications of language models via contextual integrity theory,â arXiv preprint arXiv:2310.17884, 2023. [60]R. Behnia, M. R. Ebrahimi, J. Pacheco, and B. Padmanabhan, âEw- tune: A framework for privately fine-tuning large language models with differential privacy,â in 2022 IEEE International Conference on Data Mining Workshops (ICDMW). IEEE, 2022, p. 560â566. [61]O. Shvetsova et al., âDesigning an intelligent filter for safe and responsible deployment of large language models,â Applied Sciences, vol. 15, no. 13, p. 7298, 2025, accessed: 2025-10-28. [Online]. Available: https://w.mdpi.com/2076-3417/15/13/7298 [62]X. Qi, B. Wei, N. Carlini, Y. Huang, T. Xie, L. He, M. Jagielski, M. Nasr, P. Mittal, and P. Henderson, âOn evaluating the durability of safeguards for open-weight llms,â arXiv preprint arXiv:2412.07097, 2024. [63]Y. Sens, âTowards a privacy-by-design framework for ml-enabled systems,â in 2025 IEEE/ACM 4th International Conference on AI EngineeringâSoftware Engineering for AI (CAIN).IEEE, 2025, p. 244â246. [64]N. Oh, G. D. Vrakas, S. J. Brooke, S. Moriniere, and T. Duke, âPetlp: A privacy-by-design pipeline for social media data in AI research,â in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 8, no. 2, 2025, p. 1926â1938. [65]F. Tanaka, A. Cicourel, and J. R. Movellan, âSocialization between toddlers and robots at an early childhood education center,â Proceedings of the National Academy of Sciences, vol. 104, no. 46, p. 17 954â17 958, 2007. [66]V. Andries and J. Robertson, âAlexa doesnât have that many feelings: Childrenâs understanding of ai through interactions with smart speakers in their homes,â Computers and Education: Artificial Intelligence, vol. 5, p. 100176, 2023. [67]K. Darling, ââwhoâs johnny?âanthropomorphic framing in human-robot interaction, integration, and policy,â Anthropomorphic framing in human- robot interaction, integration, and policy (March 23, 2015). Robot Ethics, vol. 2, 2015. [68] N. I. Abbasi, M. Spitale, J. Anderson, T. Ford, P. B. Jones, and H. Gunes, âCan robots help in the evaluation of mental wellbeing in children? an empirical study,â in 2022 31st IEEE international conference on robot and human interactive communication (RO-MAN).IEEE, 2022, p. 1459â1466. [69] L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh et al., âEthical and social risks of harm from language models,â arXiv preprint arXiv:2112.04359, 2021. [70] C. Akbulut, L. Weidinger, A. Manzini, I. Gabriel, and V. Rieser, âAll too human? mapping and mitigating the risk from anthropomorphic ai,â in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 7, no. 1, 2024, p. 13â26. [71]The Verge, âMeta confirms public facebook and instagram posts used to train ai,â The Verge, 2024. [Online]. Available: https: //w.theverge.com/2024/06/meta-facebook-instagram-ai-training [72](2023)Amazonusesalexachilddatatotunevoice algorithm.Accessed:2025-10-24.[Online].Available:https: / / w.aiaaic.org / aiaaic - repository / ai - algorithmic - and - automation - incidents/amazon-uses-alexa-child-data-to-tune-voice-algorithm [73]Character.AI. (2024) Sharing more about the recent incident. Accessed: 2025-10-24. [Online]. Available: https://blog.character.ai/sharing-more- about-the-recent-incident/ [74]A. OâDaffer, W. Liu, and C. S. Bloss, âSchool-based online surveillance of youth: Systematic search and content analysis of surveillance company websites,â Journal of Medical Internet Research, vol. 27, p. e71998, 2025. [75]K. A. Farkhodovich and R. D. Salixovna, âExploring language learning apps: A focus on duolingo, lingo kids, and Buddy.AI for enhanced english acquisition,â Hamkor konferensiyalar, vol. 1, no. 6, p. 94â97, 2024. [76]J. Taylor, âMetaâs AI is scraping usersâ photos and posts. Europeans can opt out, but Australians cannot,â The Guardian, 2024, accessed: 2025- 10-24. [Online]. Available: https://w.theguardian.com/technology/ article/2024/sep/11/meta- ai- post- scraping- security- opt- out- privacy- laws [77](2024) Facebook scrapes photos of kids from australian user profiles to train its ai. Accessed: 2025-10-24. [Online]. Available: https://w.malwarebytes.com/blog/news/2024/09/facebook-scrapes- photos-of-kids-from-australian-user-profiles-to-train-its-ai [78]R. Beaumont and LAION, âLAION-5B: A new era of open large-scale multi-modal datasets,â 2022, accessed: 2025-10-24. [Online]. Available: https://laion.ai/blog/laion-5b/ [79]J. Taylor. (2024) Australian childrenâs images found in AI training dataset used by stability AI and midjourney. Accessed: 2025-10-24. [Online]. Available: https://w.theguardian.com/technology/article/ 2024/jul/03/australian-children-used-ai-data-stability-midjourney [80]StanfordInternetObservatory,âMachinelearningtraining datasetscontainchildsexualabusematerial,âStanford University,Tech.Rep.,2023,accessed:2025-10-24.[On- line]. Available: https://stacks.stanford.edu/file/druid:kh752sm9123/ ml_training_data_csam_report-2023-12-23.pdf [81]C. Caetano, G. O. d. Santos, C. Petrucci, A. Barros, C. Laranjeira, L. S. F. Ribeiro, J. F. de Mendonça, J. A. dos Santos, and S. Avila, âNeglected risks: The disturbing reality of childrenâs images in datasets and the urgent call for accountability,â in Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, 2025, p. 2542â2553. [82]K. Kireev, A.-M. Cre ̧tu, R. Meier, S. A. Bargal, E. Redmiles, and C. Troncoso, âA manually annotated image-caption dataset for detecting children in the wild,â arXiv preprint arXiv:2506.10117, 2025. [83]Y. Yu, T. Sharma, M. Hu, J. Wang, and Y. Wang, âExploring parent-child perceptions on safety in generative ai: concerns, mitigation strategies, and design implications,â in 2025 IEEE Symposium on Security and Privacy (SP). IEEE, 2025, p. 2735â2752. [84]Character.AI. (2024) Privacy policy. Accessed: 2025-10-24. [Online]. Available: https://character.ai/privacy [85] Gaggle.Net, Inc., âGaggle safety management,â 2025, accessed: 2025- 10-24. [Online]. Available: https://w.gaggle.net/safety-management/ [86]The Detroit News. (2025) Schools use AI to monitor kids, hoping to prevent violence. an investigation found security risks. Accessed: 2025- 10-24. [Online]. Available: https://w.detroitnews.com/story/news/ nation/2025/03/12/schools-use-ai-to-monitor-kids-hoping-to-prevent- violence-an-investigation-found-security-risks/82307942007/ [87]Associated Press and Education Reporting Collaborative, âSchools use ai to monitor kids, hoping to prevent violence. our investigation found security risks,â 2025, accessed: 2025-10-24. [Online]. Available: https://apnews.com/article/25a3946727397951fd42324139aaf70f [88]M. A. Garcia, D. Culbreth, S. P. Foxx, F. Martin, and W. Wang, âKeeping students safe: the experiences of educators implementing monitoring applications in k-12 settings,â Education and Information Technologies, p. 1â31, 2025. [89] I. C. Ngong, S. R. Kadhe, H. Wang, K. Murugesan, J. D. Weisz, A. Dhurandhar, and K. Natesan Ramamurthy, âProtecting users from themselves: Safeguarding contextual privacy in interactions with con- versational agents,â in Findings of the Association for Computational Linguistics.Association for Computational Linguistics, 2025, p. 26 196â26 220. [90]S. Tokas and O. Owe, âA formal framework for consent management,â in International Conference on Formal Techniques for Distributed Objects, Components, and Systems. Springer, 2020, p. 169â186. [91]I. Milkaite and E. Lievens, âChild-friendly transparency of data processing in the eu: from legal requirements to platform policies,â Journal of Children and Media, vol. 14, no. 1, p. 5â21, 2020. [92]A. Dangol, R. Wolfe, R. Zhao, J. Kim, T. Ramanan, K. Davis, and J. A. Kientz, âChildrenâs mental models of AI reasoning: Implications for AI literacy education,â arXiv preprint arXiv:2505.16031, 2025. [93]G. Feretzakis, K. Papaspyridis, A. Gkoulalas-Divanis, and V. S. Verykios, âPrivacy-preserving techniques in generative AI and large language models: A narrative review,â Information, vol. 15, no. 11, p. 697, 2024, accessed: 2025-10-28. [Online]. Available: https://w.mdpi.com/2078-2489/15/11/697 [94]A. Blanco-Justicia, N. Jebreel, B. Manzanares-Salor, D. SĂĄnchez, J. Domingo-Ferrer, G. Collell, and K. Eeik Tan, âDigital forgetting in large language models: A survey of unlearning methods,â Artificial Intelligence Review, vol. 58, no. 3, p. 90, 2025. [95]S. Chapagain, S. M. Hamdi, and S. F. Boubrahimi, âPruning strategies for backdoor defense in llms,â in Proceedings of the 34th ACM International Conference on Information and Knowledge Management, 2025, p. 4633â4638. [96]B. Xia, Q. Lu, L. Zhu, S. U. Lee, Y. Liu, and Z. Xing, âTowards a responsible ai metrics catalogue: A collection of metrics for ai accountability,â in Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI, 2024, p. 100â111. [97]D. Kloza, J. Ausloos, and P. Dewitte, âTowards a method for data protection impact assessment: From legal requirements to technical implementation,â in Privacy Technologies and Policy (APF). Springer, 2019, p. 43â65. [98]CNIL.(2024)Carryingoutadataprotectionimpact assessment if necessary. Accessed: 2025-06-07. [Online]. Available: https://w.cnil.fr/en/carrying- out- protection- impact- assessment- if- necessary [99]J. Zhao, G. Wang, C. Dally, P. Slovak, J. Edbrooke-Childs, M. Van Kleek, and N. Shadbolt, ââI make up a silly nameâ understanding childrenâs perception of privacy risks online,â in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019, p. 1â13. [100]J. Irwin, A. Dharamshi, and N. Zon, âChildrenâs privacy in the age of artificial intelligence,â Canadian Standards Association, 2021. [Online]. Available: https://w.csagroup.org/wp-content/uploads/CSA-Group- Research-Children_s-Privacy-in-the-Age-of-Artificial-Intelligence.pdf [101]European Data Protection Board, âGuidelines 05/2020 on consent under regulation 2016/679,â 2020, accessed: 2025-06-07. [Online]. Available: https://w.edpb.europa.eu/our-work-tools/our-documents/ guidelines/guidelines-052020-consent-under-regulation-2016679_en [102]S. Wachter and B. Mittelstadt, âA right to reasonable inferences: Re-thinking data protection law in the age of big data and AI,â Columbia Business Law Review, no. 2, p. 494â620, 2020. [Online]. Available: https://ssrn.com/abstract=3248829 [103]International Organization for Standardization, âISO/IEC 5338:2023 Artificial Intelligence â Lifecycle-oriented Requirements for LLM Applications,â ISO/IEC, International Standard ISO/IEC 5338:2023, 2023, accessed: 2025-06-07. [Online]. Available: https://w.iso.org/ standard/81118.html [104]S. Spiekermann and L. F. Cranor, âEngineering privacy,â IEEE Trans- actions on Software Engineering, vol. 35, no. 1, p. 67â82, 2009. [105]R. Staab, N. Jovanovic, K. Mai, P. Ganesh, M. Vechev, F. Fioretto, and M. Jagielski, âSok: Data minimization in machine learning,â arXiv preprint, 2025. [Online]. Available: https://arxiv.org/abs/2508.10836 [106]S. Pasch and M. C. Cha, âBalancing privacy and utility in personal llm writing tasks: An automated pipeline for evaluating anonymizations,â in Proceedings of the Sixth Workshop on Privacy in Natural Language Processing, 2025, p. 32â41. [107] M. Kuo, J. Zhang, J. Zhang, M. Tang, L. DiValentin, A. Ding, J. Sun, W. Chen, A. Hass, T. Chen et al., âProactive privacy amnesia for large language models: Safeguarding pii with negligible impact on model utility,â arXiv preprint arXiv:2502.17591, 2025. [108]M. Ashouri-Talouki, N. Kahani, and M. Barati, âPrivacy-preserving attribute-based access control with non-monotonic access structure,â in 2023 7th Cyber Security in Networking Conference (CSNet). IEEE, 2023, p. 32â38. [109] M. Ashouri-Talouki, N. Kahani, M. Barati, and Z. Abedini, âA revocable attribute-based access control with non-monotonic access structure,â Annals of Telecommunications, vol. 79, no. 11, p. 833â842, 2024. [110]E. Elabd, âDynamic differential privacy technique for deep learning models,â Scientific Reports, vol. 15, no. 1, p. 39353, 2025. [111]S. Sultan and C. D. Jensen, âEnsuring purpose limitation in large- scale infrastructures with provenance-enabled access control,â Sensors, vol. 21, no. 9, p. 3041, 2021. [112] Q. Zhang, E. E. Wang, J. Li, and X. Wang, âBurn-after-use for preventing data leakage through a secure multi-tenant architecture in enterprise llm,â arXiv preprint arXiv:2601.06627, 2026. [113]I. D. Raji, A. Smart, P. B. White, M. Mitchell, and T. Gebru, âOutsider oversight: Designing a third party audit ecosystem for AI governance,â Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p. 239â252, 2022. [114] F. Schaub, R. Balebako, A. L. Durity, and L. F. Cranor, âA design space for effective privacy notices,â in Eleventh symposium on usable privacy and security (SOUPS 2015), 2015, p. 1â17. [115]Y. Liu, G. Deng, Y. Li, K. Wang, Z. Wang, X. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng et al., âPrompt injection attack against llm-integrated applications,â arXiv preprint arXiv:2306.05499, 2023. [116]M. I. Khalid, M. Ahmed, and J. Kim, âEnhancing data protection in dynamic consent management systems: formalizing privacy and security definitions with differential privacy, decentralization, and zero- knowledge proofs,â Sensors, vol. 23, no. 17, p. 7604, 2023. [117]S. Zanella-BĂ©guelin, L. Wutschitz, S. Tople, V. RĂŒhle, A. Paverd, O. Ohrimenko, B. Köpf, and M. Brockschmidt, âAnalyzing information leakage of updates to natural language models,â in Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, 2020, p. 363â375. [118]G. Holmes, B. Tang, S. Gupta, S. Venkatesh, H. Christensen, and A. Whitton, âApplications of large language models in the field of suicide prevention: Scoping review,â Journal of Medical Internet Research, vol. 27, p. e63126, 2025. [119]Canadian Institute for Advanced Research (CIFAR), âResponsible AI and children: Insights, implications, and best practices,â CIFAR, Canada, Tech. Rep., 2024. [Online]. Available: https://cifar.ca/wp- content/ uploads/2024/04/CIFAR-Responsible-AI-and-Children-EN_Final.pdf [120]M. Duan, A. Suri, N. Mireshghallah, S. Min, W. Shi, L. Zettlemoyer, Y. Tsvetkov, Y. Choi, D. Evans, and H. Hajishirzi, âDo membership inference attacks work on large language models?â arXiv preprint arXiv:2402.07841, 2024. [121]L. Myllyaho, M. Raatikainen, T. MĂ€nnistö, T. Mikkonen, and J. K. Nurminen, âSystematic literature review of validation methods for ai systems,â arXiv preprint arXiv:2107.12190, 2021. [122]R. Samavi and M. P. Consens, âPublishing privacy logs to facilitate transparency and accountability,â Journal of Web Semantics, vol. 50, p. 1â20, 2018. [123]T. Shevlane, âStructured access: an emerging paradigm for safe AI deployment,â arXiv preprint arXiv:2201.05159, 2022. [124]S. Liu and W. Ding, âArtificial intelligence for children: Unicefâs policy guidance and beyond,â Children & Society, vol. 39, no. 1, p. 374â382, 2025. [125]United Nations, âConvention on the rights of the child,â https:// w.ohchr.org/en/instruments- mechanisms/instruments/convention- rights-child, 1989, accessed: 2025-06-07. [126]C. Dwork, A. Roth et al., âThe algorithmic foundations of differential privacy,â Foundations and trendsÂź in theoretical computer science, vol. 9, no. 3â4, p. 211â407, 2014. [127]A. M. Kassem, O. Mahmoud, and S. Saad, âPreserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models,â in The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [128]Z. Zhou, J. Xiang, C. Chen, and S. Su, âQuantifying and analyzing entity-level memorization in large language models,â in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, 2024, p. 19 741â19 749. [129]A. Skulmowski, âInformed consent in educational AI research needs to be transparent, flexible, and dynamic,â Mind, Brain, and Education, vol. 19, no. 1, p. 32â36, 2025. [130]S. Pearson, âAccountability frameworks for personal data protectionâa research direction,â Computer Law & Security Review, vol. 29, no. 3, p. 232â240, 2013. [131]S. Neel and P. Chang, âPrivacy issues in large language models: A survey,â arXiv preprint arXiv:2312.06717, 2023. [Online]. Available: https://arxiv.org/abs/2312.06717 [132]J. Diaz-De-Arcaya, J. LĂłpez-De-Armentia, R. Miñón, I. L. Ojanguren, and A. I. Torre-Bastida, âLarge language model operations (llmops): Definition, challenges, and lifecycle management,â in 2024 9th Interna- tional Conference on Smart and Sustainable Technologies (SpliTech). IEEE, 2024, p. 1â4. [133] S. Egelman, âInforming future privacy enforcement by examining 20+ years of coppa,â Harv. JL & Tech., vol. 37, p. 1387, 2023. [134] S. Livingstone and E. Lievens, âA child rights approach to childrenâs data protection,â International Journal of Childrenâs Rights, vol. 29, no. 2, p. 261â282, 2021. [135] Irish Data Protection Commission, âFundamentals for a child- oriented approach to data processing,â Data Protection Commission, Ireland, Dublin, Ireland, Tech. Rep., 2021, accessed: 2025-06- 07. [Online]. Available: https://w.dataprotection.ie/sites/default/ files/uploads/2021-12/Fundamentals%20for%20a%20Child-Oriented% 20Approach%20to%20Data%20Processing_FINAL_EN.pdf [136]D. Long and J. Harrison, âTrust and transparency in neural language models,â Communications of the ACM, vol. 66, no. 4, p. 34â42, 2023. [137]C. Uddagiri and B. V. Isunuri, âEthical and privacy challenges of generative ai,â in Generative AI: Current Trends and Applications. Springer, 2024, p. 219â244.