Paper deep dive
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization
Mazene Ameur, Abdelkader Mekrache, Bouziane Brik, Adlen Ksentini
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 4:20:15 AM
Summary
This paper presents a comprehensive tutorial and survey on LLM-powered Agentic AI for 5G and 6G Next-Generation Networks (NGNs). It bridges the gap between Agentic AI and telecommunications by formalizing control, management, and AI-native planes, mapping agentic capabilities to network control surfaces, and addressing standardization, protocols (MCP, A2A), and evaluation benchmarks. The work highlights the shift from rule-based automation to autonomous, goal-driven intelligence in network management.
Entities (14)
Relation Signals (10)
Large Language Models → powers → Agentic AI
confidence 97% · Agentic Artificial Intelligence (AI), enabled by Large Language Models
Agentic AI → enables → Next-Generation Networks
confidence 95% · Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs).
3GPP → standardizes → 5G
confidence 92% · leading standardization bodies such as 3rd Generation Partnership Project (3GPP)... consistently highlight the need for cohesive frameworks
ETSI → standardizes → Zero-Touch Service Management
confidence 90% · ZSM, standardized by ETSI
Agentic AI → supports → Intent-Based Networking
confidence 90% · Next-generation IBN therefore moves toward natural-language interfaces enabled by Agentic AI
Agentic AI → supports → Zero-Touch Service Management
confidence 90% · ZSM, standardized by ETSI, targets fully autonomous network operation... The reasoning requirement of the third stage explains the recent surge of interest in Agentic AI for ZSM
Agentic AI → uses → Model Context Protocol
confidence 88% · The survey analyzes emerging Agentic AI protocol frameworks, including MCP and A2A Communication
Agentic AI → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundations of agentic systems: reasoning, planning, tool use, multi-agent coordination, and evaluation. Part II maps agentic capabilities onto 5G/6G control surfaces, standardization, and major 6G initiatives. Finally, it identifies open challenges shaping autonomous telecommunications.
Tags
Links
- Source: https://arxiv.org/abs/2607.16066v1
- Canonical: https://arxiv.org/abs/2607.16066v1
Trouble viewing inline? Open PDF directly →
Full Text
153,210 characters extracted from source content.
Expand or collapse full text
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization Mazene Ameur mazene.ameur@eurecom.fr EURECOMCommunication Systems DepartmentSophia AntipolisFrance , Abdelkader Mekrache abdelkader@simula.no Simula MetropolitanCRNAOsloNorway , Bouziane Brik bbrik@sharjah.ac.ae University of SharjahComputer Science DepartmentSharjahUAE and Adlen Ksentini adlen.ksentini@eurecom.fr EURECOMCommunication Systems DepartmentSophia AntipolisFrance Abstract. Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundations of agentic systems: reasoning, planning, tool use, multi-agent coordination, and evaluation. Part I maps agentic capabilities onto 5G/6G control surfaces, standardization, and major 6G initiatives. Finally, it identifies open challenges shaping autonomous telecommunications. Agentic AI, Large Language Models, 5G, 6G, Next-Generation Networks, Network Management, AI-native Networks †copyright: none†ccs: General and reference Surveys and overviews†ccs: Networks Network design principles†ccs: Computing methodologies Artificial intelligence 1. Introduction 1.1. Context and Motivation The telecommunications sector is at a critical inflection point shaped by the maturation of Fifth Generation (5G) networks and the early architectural formulation of Sixth Generation (6G) systems (Andrews et al., 2014; Saad et al., 2020). Modern network infrastructures are increasingly defined by extreme heterogeneity, ultra-dense deployments, and stringent performance requirements across latency, reliability, and Quality of Service (QoS) (Ameur et al., 2025d, b). As the vision for beyond-5G and 6G networks matures, the community increasingly converges on architectures that demand unprecedented levels of automation, closed-loop intelligence, and operational autonomy across heterogeneous infrastructures (Tariq and others, 2020). Within this trajectory, leading standardization bodies such as 3rd Generation Partnership Project (3GPP111https://w.3gpp.org/), European Telecommunications Standards Institute (ETSI222https://w.etsi.org/), and TeleManagement (TM) Forum333https://w.tmforum.org/, together with the global research community, consistently highlight the need for cohesive frameworks that translate high-level service expectations into actionable, verifiable, and continuously optimized network behaviors at scale (Leivadeas and Falkner, 2023). In this context, Zero-Touch network and Service Management (ZSM) (ETSI, 2023) and Intent-Based Networking (IBN) emerge as pivotal enablers, not as isolated features but as integral pillars that support the end-to-end autonomy envisioned for Next-Generation Networks (NGNs) (Mekrache et al., 2024). Against this backdrop, the emergence of Agentic Artificial Intelligence (AI), enabled by recent advances in Large Language Models (LLMs) and reasoning-centric architectures, marks a fundamental transition from reactive, rule-based automation to proactive, goal-driven intelligence (Murugesan, 2025). Unlike traditional AI solutions deployed as isolated optimization components, agentic systems exhibit autonomy, contextual awareness, and coordinated behavior, allowing them to reason over network states, translate high-level intents into actionable policies, and operate across heterogeneous infrastructure layers (Pati, 2025). This capability enables continuous adaptation, self-directed decision-making, and scalable intelligence at the network level, effectively redefining the operational paradigm of NGNs (Chen et al., 2024). 1.2. Related Surveys, Gaps, Scope and Objectives Selection Process. The body of related literature considered in this survey was collected through systematic searches on Google Scholar, complemented by a targeted examination of existing surveys with narrower scopes. The search strategy employed representative keywords such as “Agentic AI,” “LLM agents,” and “Large Language Model agents,” augmented with networking-specific terms including “Next Generation Network,” “5G,” and “6G networks,” to ensure comprehensive coverage of potentially relevant studies. From the initial pool of collected publications, a rigorous pre-selection process was conducted to exclude out-of-scope studies, non-survey papers, and sources lacking sufficient academic rigor. This process resulted in a refined and high-quality set of surveys that are closely aligned with the objectives of this work. A detailed comparative analysis of this curated set enabled the identification of key limitations, unresolved challenges, and underexplored research directions in the existing literature. These observations directly motivated the development of the present survey, which aims to address a critical gap in the intersection of Agentic AI and NGNs, as discussed in this subsection. Table 1. Summary of related surveys on LLM-based and Agentic AI for networks. Ref. Key Contributions LLMs Agentic Concepts Agentic Protocols Agentic Evaluation Agentic Benchmarks Next Gen. Networks Standards & Projects (Jiang and others, 2025) LLM survey for 6G: architectures, communication applications, agentic concepts. ◐ ◐ ✗ ✓ ✗ ✓ ✗ (Kong and others, 2025) Agent communication security: lifecycles, MCP/A2A protocols, defense mechanisms. ✓ ✓ ✓ ◐ ✗ ✗ ✗ (Long and others, 2025) LLM-enabled network ops: fault diagnosis, causal inference, optimization. ✓ ✗ ✗ ✗ ✗ ✓ ✗ (Pati, 2025) Agentic AI survey: autonomy, memory, goal-driven reasoning, ethics. ◐ ✓ ✗ ✗ ✗ ◐ ✗ (Saad and others, 2025) AGI-native wireless: cognitive telecom brain for autonomous network control. ◐ ◐ ✗ ✗ ◐ ✓ ✗ (Celik and Eltawil, 2024) GenAI survey for 6G: model families, learning paradigms, security applications. ◐ ✗ ✗ ✗ ✗ ✓ ✗ (Zhao and others, 2025) Edge General Intelligence: distributed agents, world models, foundation models. ◐ ✓ ✗ ◐ ✗ ✓ ✗ (Chen et al., 2024) AI agents in 6G: use cases, collaborative architectures, system challenges. ◐ ✓ ✗ ✗ ✗ ✓ ✗ (Mohammadi et al., 2025) LLM agent evaluation: taxonomies for objectives, processes, reliability. ◐ ✓ ✗ ✓ ✓ ✗ ✗ (Zhu and others, 2025) LLM chatbots vs. agents: evaluation benchmarks and assessment dimensions. ✓ ✓ ✗ ✓ ✓ ✗ ✗ (Jiang et al., 2025) LLM and Agentic AI for 6G: design principles, multi-agent frameworks. ✓ ✓ ✓ ✓ ✗ ✓ ✗ (Masterman et al., 2024) AI agent study: single- vs. multi-agent designs, coordination mechanisms. ◐ ✓ ✗ ✓ ✓ ✗ ✗ (Sapkota et al., 2026) AI Agents vs. Agentic AI: taxonomies, hallucination, coordination failures. ◐ ✓ ✗ ✓ ◐ ✗ ✗ (Ferrag et al., 2025) LLM/agent benchmarks: evaluation frameworks, protocols, security vulnerabilities. ◐ ✓ ✓ ✓ ✓ ◐ ✗ (Zhang and others, 2026) Agentic AI for edge networks: agentification frameworks, edge intelligence. ◐ ✓ ◐ ◐ ◐ ◐ ✗ Ours Full survey: Agentic AI for NGNs, protocols, benchmarks, standardization. ✓ ✓ ✓ ✓ ✓ ✓ ✓ Legend: ✓= Addressed ◐= Partially Addressed ✗ = Not Addressed Related Surveys. The intersection of LLMs and 5G/6G Networks has only recently attracted survey-level attention. Existing works primarily examine isolated aspects of AI-enabled networking and reflect an early transition from learning-based automation toward foundation model-driven designs. As summarized in Table 1, the available surveys remain limited in number and fragmented in scope, with gaps in agentic depth, architectural coverage, protocol realization, evaluation rigor, and alignment with standardization efforts. This fragmentation prevents a unified system-level view of LLM-powered Agentic AI for 5G evolution and emerging 6G networks. A first group of surveys examines the adoption of LLMs and generative AI in 6G architectures (Jiang and others, 2025; Long and others, 2025; Celik and Eltawil, 2024). These works discuss model families, learning paradigms, and application domains spanning communications, sensing, and network management. They highlight the role of foundation models in fault diagnosis, monitoring, causal inference, and optimization, and propose high-level visions for AI-native networks. However, intelligence is largely treated as centralized inference rather than autonomous agency. Agentic reasoning, coordination mechanisms, and protocol-level integration receive limited attention. As a result, closed-loop autonomy, multi-agent orchestration, and lifecycle management remain insufficiently addressed. A second cluster of studies focuses on conceptual and theoretical foundations of Agentic AI, including formalizations of autonomy, memory, goal-driven reasoning, and agent taxonomies (Pati, 2025; Zhu and others, 2025; Sapkota et al., 2026). Complementary works develop evaluation frameworks and taxonomies for LLM agents, distinguishing objectives, processes, and assessment dimensions (Mohammadi et al., 2025; Zhu and others, 2025). While these contributions provide essential conceptual clarity and methodological structure, they are largely domain agnostic. They do not map agentic constructs to telecom network management, control planes, or orchestration layers, nor do they consider telecom-specific Key Performance Indicators (KPIs), latency constraints, reliability requirements, or regulatory obligations. A third category of surveys focuses on security and protocol aspects of agentic systems, addressing agent communication, multi-agent coordination, emerging Model Context Protocol (MCP) and Agent-to-Agent (A2A) paradigms, and vulnerabilities such as prompt injection and trust violations (Kong and others, 2025; Ferrag et al., 2025). While these works explicitly consider inter-agent interactions, they remain largely detached from telecom architectures and operational constraints. Integration with carrier-grade control frameworks, network management, and real-time radio access processes is not examined. Related analytical studies on single-agent and multi-agent designs explore coordination trade-offs (Masterman et al., 2024), but do not embed these mechanisms within telecom stacks or end-to-end network architectures. Visionary and forward-looking contributions propose Artificial General Intelligence (AGI) native wireless systems, cognitive telecom brains, and fully autonomous network control (Saad and others, 2025; Chen et al., 2024; Zhang and others, 2026). Tutorials on LLM-centric design and multi-agent frameworks for 6G attempt to bridge Agentic AI and telecom systems (Zhao and others, 2025; Jiang et al., 2025). While these works are valuable in articulating long-term trajectories and high-level design principles, they remain largely qualitative, with limited treatment of evaluation benchmarks and standardization pathways. Their speculative nature and lack of technical details reduce their utility for near-term engineering and deployment. Current Research Gaps & Motivation. Existing surveys either focus on LLM-enabled networking without agentic depth (Jiang and others, 2025; Long and others, 2025; Celik and Eltawil, 2024), or develop agentic theory and evaluation frameworks without grounding in telecom architectures and standards (Pati, 2025; Mohammadi et al., 2025; Zhu and others, 2025; Sapkota et al., 2026), or analyze security and protocols without domain-specific integration (Kong and others, 2025; Ferrag et al., 2025). Visionary works outline AGI native networks but lack actionable detail (Saad and others, 2025; Chen et al., 2024), while tutorials and analytical studies provide partial bridges that remain insufficiently comprehensive (Jiang et al., 2025; Masterman et al., 2024). To the best of our knowledge, no prior work jointly and systematically addresses LLMs, agentic concepts, agentic protocols, evaluation methodologies, benchmarks, telecom network integration, and standardization initiatives in a unified study. This gap motivates the need for a holistic survey that positions LLM-powered Agentic AI as a first-class architectural paradigm for next-generation 5G and 6G network management, spanning protocol design, autonomous control, performance evaluation, and alignment with ongoing industry and standardization efforts. Our survey addresses these deficiencies by offering the first comprehensive and tutorial-style study of LLM-powered Agentic AI with explicit emphasis on next-generation telecommunications. 1.3. Survey Contributions The key contributions of this survey can be summarized as follows: • Comprehensive technical treatment of Agentic AI for 5G/6G Networks: This survey adopts a tutorial-first and review-driven methodology that establishes a unified conceptual foundation before delivering a concise yet rigorous technical survey of Agentic AI in the context of 5G and 6G network management. • Cross-domain tutorial perspective: The survey is structured to be accessible to both AI and telecommunications communities, providing network researchers with essential background on LLMs, agentic reasoning, and autonomy, while offering AI researchers a clear overview of networking fundamentals, thereby fostering cross-disciplinary understanding. • Review of emerging agentic protocols and evaluation benchmarks: The survey analyzes emerging Agentic AI protocol frameworks, including MCP and A2A Communication, from a telecommunications perspective, and reviews relevant evaluation and benchmarking approaches for agentic networking systems. • Unified taxonomy of LLM-based Agentic AI applications in telecom: A compact taxonomy is introduced to categorize Agentic AI applications across the telecom ecosystem, covering infrastructure, management, and AI-native control domains, including radio, core, transport, and edge environments, as well as ZSM, IBN, Explainable AI (XAI), and AI Operations (AIOps). • Standardization, industry relevance, and future directions: The survey summarizes ongoing standardization and industry initiatives related to AI-native and autonomous networks, and identifies key challenges and research directions for the design and deployment of next-generation agentic networking systems. Figure 1. Survey Structure. 1.4. Survey Organization Fig. 1 depicts the structure of this paper, which combines a tutorial foundation with a state-of-the-art survey. Section 2 scopes the architectural foundations of NGNs across the radio access, transport, core, and edge/cloud domains, together with the IBN and ZSM management paradigms. Section 3 covers core concepts, architectural patterns, inter-agent communication, and evaluation methodologies relevant to networking. Section 4 surveys Agentic AI applications across telecommunication layers, spanning infrastructure intelligence, management and orchestration, and AI-native control. Section 5 reviews current standardization activities and major 6G research initiatives and their convergence toward agentic design principles. Section 6 discusses open challenges and future research directions, and Section 7 concludes the survey. 2. Next-Generation Networks: An Overview While prior generations from 1G through 5G progressively introduced digital transmission, mobile broadband, and softwarization, 6G is expected to be distinguished by the pervasive integration of intelligence across all architectural layers (Giordani et al., 2020). As in 5G, the 6G control plane will continue to host a subset of management-related functionalities, increasingly augmented by AI-driven mechanisms for tasks such as mobility management, resource allocation, and network optimization. In parallel, the management plane, encompassing paradigms such as ZSM and IBN, provides higher-level orchestration, automation, and policy enforcement capabilities. These management functionalities inherently span both planes, reflecting their cross-layer nature. Therefore, this section focuses on the architectural components most relevant to the deployment of Agentic AI for 6G network management. As illustrated in Fig. 2, we organize the discussion around three conceptual layers: the control plane, the management plane, and a vertical AI-native plane that intersects both. This abstraction aligns with the AI-native vision advocated in recent standardization efforts (Shehzad et al., 2022). Figure 2. Overview of AI-native 6G technologies. 2.1. Control plane In next-generation networks, the control plane is decomposed across multiple technological domains, namely the Radio Access Network (RAN), Transport Network (TN), Core Network (CN), and Edge/Cloud infrastructure. Each of these domains adopts the principle of Control and User Plane Separation (CUPS), enabling independent scaling, modularization, and programmable control of Network Functions (NFs) (Choi and others, 2022). In this work, we focus on the control plane, as it increasingly hosts AI-driven functionalities supporting a wide range of management tasks. In the RAN, the O-RAN initiative has redefined the control plane through the RAN Intelligent Controller (RIC), which exposes two programmable surfaces: the Near-Real-Time (Near-RT) RIC for sub-second control loops via xApps, and the Non-Real-Time (Non-RT) RIC for policy-driven loops via rApps (Azariah and others, 2024). These constructs allow third-party agents to manipulate radio resource allocation, beamforming, and handover decisions, which is essential for embedding agentic reasoning at the edge of the network. On the data plane side, 6G is expected to extend this programmability to mmWave, Terahertz bands, and cell-free massive MIMO architectures (Gupta and Jha, 2015). The TN adopts Software-Defined Networking (SDN) to centralize routing decisions in a controller that maintains a global topology view, enabling dynamic path computation, slicing, and load balancing (Long et al., 2022). Mainstream controllers such as ONOS444https://github.com/opennetworkinglab/onos/, OpenDaylight555https://w.opendaylight.org, and the 6G-oriented TeraFlowSDN666https://tfs.etsi.org expose northbound APIs through which AI agents can inject routing intents or trigger reconfiguration based on predicted traffic patterns (Prabha et al., 2022). Performance-critical capabilities such as Segment Routing (SRv6) and Time-Sensitive Networking (TSN) provide the deterministic substrate over which agentic decisions are enforced. The CN follows a Service-Based Architecture (SBA) in which NFs communicate through standardized HTTP/2 APIs, providing a natural integration point for AI agents (Husain et al., 2022). Key NFs include the Access and Mobility Management Function (AMF), Session Management Function (SMF), and the User Plane Function (UPF) accelerated through eBPF/XDP for kernel-level packet processing (do Amaral et al., 2021). Of particular relevance to Agentic AI are the Network Data Analytics Function (NWDAF), which supplies analytics events such as mobility prediction or abnormal traffic detection (Mekrache et al., 2023; Ameur et al., 2026b), and the Network Exposure Function (NEF), which provides a secure boundary through which external agents can interact with CN services (Fragkos et al., 2022). The Edge/Cloud domain hosts vertical applications and is governed by container orchestration platforms such as Kubernetes777https://kubernetes.io and OpenShift888https://github.com/openshift, which expose declarative APIs for workload placement, scaling, and lifecycle management (Younis and others, 2024). This declarative interface is well-suited for Agentic AI, since high-level placement decisions, including the dynamic migration of latency-sensitive workloads between centralized clouds and edge sites, can be expressed as desired-state manifests that the orchestrator reconciles (Kitanov and others, 2024). 2.2. Management plane The management plane orchestrates services end-to-end across the four infrastructure domains. Given the scale and complexity of 6G, manual management is impractical, motivating two complementary paradigms that are increasingly addressed through Agentic AI and constitute the focus of this survey: IBN and ZSM. These paradigms are jointly shaped by the IETF999https://w.ietf.org, TM Forum, 3GPP, and ETSI, and although they overlap in scope, we describe them separately for clarity. 2.2.1. Intent-Based Networking IBN abstracts network governance to high-level declarative goals, called intents, that specify the desired outcome rather than the configuration steps required to achieve it (Leivadeas and Falkner, 2023). The IBN lifecycle consists of five canonical stages: (i) profiling, in which the requirement is captured; (i) translation, in which the intent is decomposed into low-level configurations targeting the RAN, TN, CN, and Edge/Cloud; (i) resolution, in which conflicts between competing resource requests are reconciled; (iv) activation, in which the configurations are deployed; and (v) assurance, in which compliance is continuously monitored. A representative intent, such as instantiating an end-to-end service composed of a gNB, a CN slice, and an edge application with 1 Gbps guaranteed throughput and ultra-low latency, illustrates the cross-domain coordination required. Current standards encode intents in structured formats such as JSON or YAML, yet these still impose data-model expertise on operators. Next-generation IBN therefore moves toward natural-language interfaces enabled by Agentic AI, allowing operators to express intents conversationally and delegate the technical decomposition to autonomous agents (Mekrache et al., 2024). 2.2.2. Zero-touch Service Management ZSM, standardized by ETSI, targets fully autonomous network operation in which detection, prediction, and remediation occur without human intervention (Coronado and others, 2022; ETSI ISG ZSM, 2024). When combined with IBN, ZSM operationalizes the assurance stage by enforcing intent compliance throughout the service lifecycle. The closed automation loop comprises three stages: (i) anomaly detection and prediction, typically driven by time-series AI models that flag degradations before they impact users; (i) Root Cause Analysis (RCA), increasingly performed with explainable AI to expose the causal chain behind a fault (Dwivedi and others, 2023); and (i) anomaly resolution, which requires planning and reasoning to select corrective actions. The reasoning requirement of the third stage explains the recent surge of interest in Agentic AI for ZSM, since LLM-based agents can synthesize multi-step remediation plans grounded in current network state. 2.3. AI-native plane The AI-native plane is a vertical pillar that supplies intelligence to both the infrastructure and the management layers. In the control plane, it manifests as xApps and rApps in the RIC for radio control, AI-augmented SDN routing, NWDAF-driven analytics in the CN, and predictive workload placement in the Edge/Cloud. In the management plane, it underpins natural-language intent translation, conflict resolution, and the closed-loop assurance enabled by ZSM. Standards bodies including TM Forum and ETSI have articulated this AI-native vision while emphasizing that explainability is a prerequisite for autonomous decision-making (Wu and others, 2021). We organize the AI-native layer around three pillars: Data Operations, AI Operations (AIOps), and Explainability. Data Operations. High-quality data is the prerequisite of any AI-driven function, and the AI-native layer must collect, store, and preprocess telemetry from across all domains (Chochliouros and others, 2025). Domain-specific telemetry pipelines include monitoring xApps in the RAN (Santos and others, 2025), In-band Network Telemetry (INT) in the TN (Tan and others, 2021), NWDAF events and 3GPP Event Exposure APIs in the CN (Mekrache et al., 2023), and Prometheus-based metrics in the Edge/Cloud (Ksentini and others, 2025). Beyond raw telemetry, human feedback is becoming a first-class data source for IBN, enabling continuous retraining of intent translation models to improve user satisfaction. AI Operations. AIOps cover the lifecycle of AI models, from offline training on historical datasets to online adaptation under dynamic conditions (Kukliński et al., 2025). Federated Learning is increasingly adopted to train models locally at the edge while preserving data privacy and capturing global patterns (Yang et al., 2022). On the inference side, the millisecond budget of 6G control loops elevates the importance of optimization techniques such as pruning, quantization, and Knowledge Distillation, the last allowing a compact student model to approximate a larger teacher with minimal accuracy loss (Ameur et al., 2025a; Fang and others, 2026). Explainability. Black-box AI is incompatible with the accountability requirements of operational networks and with regulatory frameworks such as the EU AI Act. Explainable AI techniques therefore play a central role in the AI-native layer: SHAP (Antwarg et al., 2021) and LIME (Zafar and Khan, 2021) attribute decisions to specific input features such as traffic load or latency, while Agentic AI itself is increasingly leveraged to produce natural-language explanations of complex network states (Mekrache and others, 2025), closing the loop between automation and human oversight. 3. Agentic AI Foundations This section reviews and analyzes recent advancements in LLM-based Agentic AI from a telecom perspective, with emphasis on their implications for future 5G/6G Networks. 3.1. From LLMs to Agentic AI The progression from LLMs to Agentic AI represents a shift from static, prompt-driven predictors to systems capable of autonomous operation within complex and evolving environments (Celik and Eltawil, 2024; Ferrag et al., 2025). As shown in Fig. 3, this shift follows a clear architectural trajectory that begins with transformer-based and foundation models, progresses through reasoning-optimized LLMs, and culminates in agentic systems that integrate language understanding with decision-making and control. While modern LLMs exhibit strong generalization, reasoning, and multimodal understanding, these abilities alone are insufficient for domains that require sustained interaction, adaptive decision processes, and operational accountability (Chen et al., 2024). NGNs exemplify such domains, where tasks involve continuous monitoring, intent-driven optimization, coordinated control across layers, and robust handling of uncertainty (Zhao and others, 2025). Agentic AI emerges as an architectural response that integrates LLMs with memory, planning, tool access, and alignment mechanisms to support reliable and context aware autonomy (Zhao and others, 2025; Zhu and others, 2025). Figure 3. LLM-based Agentic AI paradigm Evolution. 3.2. LLM Adaptation Strategies for Agentic Workflows Adaptation strategies transform general-purpose LLMs into reliable, context-aware components for intelligent network management. Formally, let fθ:→f_θ:X denote a pre-trained LLM with parameters θ∈ℝdθ ^d. Adaptation seeks a task-specialized predictor fθ′(⋅;c)f_θ (·\,;\,c), where θ′θ denotes (possibly updated) parameters and c an in-context conditioning signal. Two complementary paradigms are employed: Fine-Tuning (FT), which modifies θ→θ′θ→θ via optimization on task data, and In-Context Learning (ICL) (Dong and others, 2024b), which fixes θ′=θ =θ and adapts behavior through c alone. Table 2 summarizes both paradigms, their core mechanisms, and telecom alignment. Table 2. Taxonomy of LLM Adaptation Strategies and their Telecom Alignment. Paradigm Technique Core Mechanism Telecom Alignment Fine-Tuning SFT (Dong and others, 2024a) Cross-entropy minimization on labeled pairs D Fault diagnosis, traffic classification, KPI prediction PEFT / LoRA (Xu et al., 2023; Hu et al., 2022) Low-rank update ΔW=BA W=BA; optimizes ϕφ, |ϕ|≪d|φ| d Efficient adaptation at edge/RAN with limited compute RL / RLHF / GRPO (DeepSeek-AI and others, 2026) KL-regularized reward maximization: [rψ]−βKLE[r_ψ]-β\,KL QoS/latency alignment; policy-compliant decision-making Knowledge Distillation (Xu and others, 2024; Ameur et al., 2025a) minKL(pθT∥pθS) \,KL(p_ _T\|p_ _S) from teacher to student Compact models for base stations and edge inference ReFT (Wu and others, 2024b) Intervention Φϕ _φ on hidden states h(ℓ)h^( ); |ϕ|≪d|φ| d Traffic dynamics and anomaly signature modeling In-Context Learning Zero-Shot (Kojima et al., 2022) c=ℐc=I; no examples or retraining Intent-based management; NL network configuration queries Few-Shot (Ye and Durrett, 2022) c=ℐ,(xj,yj)j=1kc=\I,(x_j,y_j)_j=1^k\; small-k exemplars Alarm/log-guided classification and summarization CoT / ToT (Nguyen and Xu, 2025) Marginalizes over latent steps z=(z1,…,zT)z=(z_1,…,z_T) Root cause analysis; multi-step troubleshooting workflows RAG (Lewis et al., 2020) Top-k retrieval ℛ(x)R(x) from corpus K Standards/spec grounding; up-to-date operational knowledge Tool / API Integration (Li and others, 2023) Action a=(τ,u)a=(τ,u); generation conditioned on o=τ(u)o=τ(u) OSS/BSS and SDN closed-loop control; real-time monitoring Fine-Tuning methods (Parthasarathy et al., 2024) optimize over a labeled telecom corpus =(xi,yi)D=\(x_i,y_i)\. Supervised Fine-Tuning (SFT) (Dong and others, 2024a) minimizes token-level cross-entropy: θ∗=argminθ1N∑iℒ(fθ(xi),yi)θ = _θ 1N _iL(f_θ(x_i),y_i), specializing the model to domain tasks such as fault diagnosis and traffic classification. Parameter-Efficient Fine-Tuning (PEFT) (Xu et al., 2023), including Low-Rank Adaptation (LoRA) (Hu et al., 2022), which decomposes weight updates as ΔW=BA W=BA with rank r≪min(m,n)r (m,n) that restricts optimization to a low-dimensional subspace ϕ∈ℝkφ ^k, k≪dk d, enabling deployment at edge and RAN nodes with limited compute. Reinforcement and reward-based methods (DeepSeek-AI and others, 2026) (e.g., Reinforcement Learning Human Feedback (RLHF), Group Relative Policy Optimization (GRPO)) treat the LLM as a stochastic policy πθ(y∣x) _θ(y x) and maximize expected reward under a KL-regularized objective [rψ(x,y)]−βKL(πθ∥πref)E[r_ψ(x,y)]-β\,KL( _θ\| _ref), aligning outputs with QoS and operational constraints. Knowledge Distillation (Xu and others, 2024; Ameur et al., 2025a) minimizes KL(pθT∥pθS)KL(p_ _T\|p_ _S) to transfer capabilities from a large teacher to a compact student suitable for base-station or edge deployment. Representation Fine-Tuning (ReFT) (Wu and others, 2024b) instead learns an intervention Φϕ:h(ℓ)↦h~(ℓ) _φ:h^( ) h^( ) on hidden states, capturing telecom-specific patterns (e.g., traffic dynamics, anomaly signatures) without full weight updates. In-Context Learning fixes θ and adapts f through the conditioning context c, yielding y^=argmaxypθ(y∣c,x) y= _yp_θ(y c,x) (Dong and others, 2024b). Zero-shot prompting (Kojima et al., 2022) sets c=ℐc=I (a task instruction), enabling intent-driven network queries without retraining. Few-shot prompting (Ye and Durrett, 2022) augments c with k telecom examples (xj,yj)j=1k\(x_j,y_j)\_j=1^k, improving performance on classification, summarization, and diagnosis. Chain-/Tree-of-Thought (CoT/ToT) (Nguyen and Xu, 2025) introduces latent reasoning steps z=(z1,…,zT)z=(z_1,…,z_T), marginalizing pθ(y∣c,x)=∑zpθ(y∣z,c,x)pθ(z∣c,x)p_θ(y c,x)= _zp_θ(y z,c,x)\,p_θ(z c,x) over sequential or branching trajectories for structured telecom workflows (e.g., root cause analysis). Retrieval-Augmented Generation (RAG) (Lewis et al., 2020) conditions generation on top-k documents retrieved from a telecom knowledge corpus K including standards, specifications, operational logs, via a retriever ℛ:→kR:X ^k, ensuring factual grounding without parameter modification. Finally, Tool/API Integration (Li and others, 2023) exposes a set =τ1,…,τMT=\ _1,…, _M\ of telecom system interfaces (e.g. Operations Support System (OSS) and Business Support System (BSS), SDN controllers), letting the model select action a=(τ,u)a=(τ,u) and condition subsequent generation on the observation o=τ(u)o=τ(u), enabling closed-loop automation. 3.3. Agentic AI Core Concepts Agentic AI introduces a set of architectural abstractions that extend LLMs from passive inference engines to autonomous decision-making systems (Kong and others, 2025). As illustrated in Fig. 4, the core components include the agent powered by LLMs, memory systems, planning mechanisms, tool interfaces, and action execution pathways (Sapkota et al., 2026). Short-term and long-term memory support contextual grounding and continuity across tasks. Planning modules provide structured reasoning capabilities, including reflection, self-critiquing, and task decomposition. Tool interfaces enable the agent to invoke analytical and operational functions such as anomaly detection, load prediction, and traffic estimation. The resulting actions are executed through standardized interfaces spanning the CN, RAN, and cloud/edge domains (Pati, 2025; Jiang et al., 2025). Collectively, these components form a modular and interoperable foundation that governs the autonomy, adaptability, and reliability of agentic systems operating in NGNs. Formally, an agentic system can be represented as a tuple: (1) =⟨,,c,ℳ,,πθ,⟩,A= ,O,A_c,M,T, _θ,G , where S is the environment state space (e.g., CN/RAN/edge conditions), O the observation space, cA_c the action space, ℳM the memory system, T the set of available tools, πθ _θ the LLM-based policy, and G the goal specification. The following subsections delve into the details of these concepts. Figure 4. Telecom-oriented Agentic AI Framework Overview. 3.3.1. Agents An agent is an autonomous computational system that interprets observations, maintains state, makes decisions, and performs actions to achieve defined goals (Kong and others, 2025; OpenAI, 2026). In LLM-based settings, agents extend pretrained models with planning, tool use, and persistent memory, enabling them to operate through iterative control loops rather than single-step inference. This differentiates agentic systems from conventional conversational LLMs, which lack self-directed action and sustained contextual reasoning (Pati, 2025; Jiang et al., 2025). Formally, at each decision step t, the agent receives an observation ot∈o_t , updates its internal memory state mt∈ℳm_t , and selects an action at∈ca_t _c through an LLM-parameterized policy: (2) at∼πθ(at|ot,mt,),mt+1=(mt,ot,at),a_t _θ\! (a_t\, |\,o_t,\,m_t,\,G ), m_t+1=U(m_t,o_t,a_t), where U denotes the memory update operator. The agent seeks to maximize a cumulative goal-aligned utility (3) J(πθ)=πθ[∑t=0Hγtℛ(st,at;)],J( _θ)=E_ _θ\! [ _t=0^Hγ^t\,R(s_t,a_t\,;\,G) ], with horizon H, discount factor γ∈(0,1]γ∈(0,1], and goal-conditioned reward ℛR capturing objectives such as QoS satisfaction or SLA compliance. Leading industry bodies like OpenAI101010https://openai.com/fr-FR/ and Anthropic111111https://w.anthropic.com/ emphasize that agents can dynamically structure their workflows, select tools, and adapt strategies in response to changing conditions (OpenAI, 2026; Anthropic, 2026). Agent architectures span reactive, deliberative, and hybrid forms, allowing a balance between rapid response and long-horizon reasoning (Pati, 2025). 3.3.2. Memory Memory is a foundational component of agentic intelligence, enabling temporal continuity, contextual reasoning, and adaptive decision-making across dynamic telecom environments. Native LLMs are constrained by fixed context windows and stateless inference, whereas agentic systems introduce hierarchical memory layers to persist and structure information across time and control loops. Formally, the memory system is decomposed as (Pati, 2025; OpenAI, 2026): (4) ℳ=ℳST∪ℳLT,ℳLT=ℳep∪ℳsem∪ℳproc,M=M_ST _LT, _LT=M_ep _sem _proc, where ℳSTM_ST and ℳLTM_LT denote short- and long-term memory, and ℳep,ℳsem,ℳprocM_ep,M_sem,M_proc denote episodic, semantic, and procedural components. Short-term memory (STM) corresponds to transient, high-resolution context maintained within or alongside the prompt window. It evolves as a bounded sliding buffer: (5) ℳST(t)=(oτ,aτ,rτ)τ=t−W+1t,|ℳST(t)|≤W,M_ST^(t)=\(o_τ,a_τ,r_τ)\_τ=t-W+1^t, |M_ST^(t)|≤ W, where W is the effective context capacity. In 5G/6G management, STM captures real-time telemetry streams (e.g., KPIs, alarms, channel conditions, traffic load) and intermediate reasoning states during tasks such as fault localization or resource allocation. It enables fine-grained correlation across recent events, supports multi-step reasoning (e.g., root cause analysis pipelines), and facilitates rapid reaction in near-real-time control loops (e.g., RAN scheduling, slicing adaptation). However, STM is inherently limited in capacity and temporal scope, requiring efficient summarization and filtering mechanisms (OpenAI, 2026; Anthropic, 2026). Long-term memory (LTM) provides persistent storage of structured and unstructured knowledge across sessions and time horizons (OpenAI, 2026; Anthropic, 2026). It is typically realized as an indexed key–value store ℳLT=(ki,vi)i=1NM_LT=\(k_i,v_i)\_i=1^N, where ki∈ℝdek_i ^d_e is an embedding key and viv_i the stored content. Retrieval given a query q is performed via a similarity-based selection (Pati, 2025): (6) ℛ(q;ℳLT)=Top-k(ki,vi)∈ℳLTsim(ϕ(q),ki),R(q;M_LT)= (k_i,v_i) _LTTop-k\;sim\! (φ(q),\,k_i ), with ϕ(⋅)φ(·) an embedding function and sim(⋅,⋅)sim(·,·) a similarity metric (e.g., cosine). In telecom settings, LTM includes historical network states, configuration changes, policy rules, learned traffic patterns, and prior optimization outcomes. It is typically implemented via external vector databases, knowledge graphs, or time-series repositories, enabling retrieval-augmented reasoning and continual learning. LTM supports trend analysis, seasonal pattern recognition, and predictive modeling (e.g., congestion forecasting, anomaly anticipation), which are critical for proactive and closed-loop network management in 6G systems. Specialized memory forms further enhance functionality. Episodic memory ℳep=ej=(sj,aj,rj,sj′)M_ep=\e_j=(s_j,a_j,r_j,s_j )\ stores past operational events (e.g., incident-response traces), enabling experience replay and case-based reasoning. Semantic memory ℳsemM_sem encodes domain knowledge such as standards, protocols, and network models (often as a knowledge graph sem=(,ℰ)G_sem=(V,E)), supporting consistent and interpretable decisions. Procedural memory ℳproc=πj:→cM_proc=\ _j:S _c\ captures learned action policies and workflows (e.g., automated remediation sequences), enabling execution of complex control strategies. The integration of STM and LTM is particularly critical in telecom agentic systems, where decisions must balance immediate network conditions with long-term objectives such as QoS guarantees, energy efficiency, and SLA compliance (OpenAI, 2026; Anthropic, 2026). 3.3.3. Tools Tool use is a defining mechanism that expands the operational scope of Agentic AI by enabling interaction with external systems beyond language generation. Through interfaces to Application Programming Interfaces (APIs), controllers, simulators, and analytical engines, agents can retrieve measurements, perform computations, and execute control actions (Qin and others, 2023). Formally, let =τ1,…,τMT=\ _1,…, _M\ be the set of available tools, where each tool τi:i→i _i:U_i _i maps an argument space iU_i to an observation space iO_i. At step t, the agent jointly selects a tool and its invocation arguments (Pati, 2025): (7) (τt,ut)∼πθ(⋅|ot,mt,),ottool=τt(ut),( _t,u_t) _θ\! (·\, |\,o_t,m_t,G ), o_t^tool= _t(u_t), and integrates the resulting observation ottool∈τto_t^tool _ _t into subsequent reasoning. Effective tool use requires accurate selection, invocation, and interpretation of outputs, providing grounding between internal reasoning and external environments (Anthropic, 2026). In communication networks, tool-enabled agents support telemetry analysis, diagnostics, configuration management, and policy evaluation, allowing agents to translate high-level intents into actionable network operations. 3.3.4. Planning Planning is a core capability of Agentic AI that enables the translation of high-level intents into structured and executable action sequences. Given a goal G and current context ctc_t, the planner produces a plan: (8) P=(p1,p2,…,pK)∼πθplan(P∣,ct),P=(p_1,p_2,…,p_K) _θ^plan(P ,c_t), where each subtask pk=(opk,prek,postk)p_k=(op_k,pre_k,post_k) specifies an operation together with its preconditions and postconditions, subject to a partial order ≺ encoding dependencies (pi≺pjp_i p_j if pjp_j depends on pip_i). It decomposes complex network objectives into ordered tasks while accounting for dependencies, constraints, and cross-domain interactions. Planning operates in close coordination with memory and tool interfaces to adapt to real-time network conditions and reuse prior knowledge. Iterative refinement through feedback and self-critique supports closed-loop control and continuous adjustment, formalized as: (9) P(i+1)=πθplan(P(i),ℱ(P(i),otexec)),P^(i+1)= _θ^plan\! (P^(i),\,F(P^(i),o_t^exec) ), where ℱF is a feedback/critique operator evaluating execution outcomes otexeco_t^exec, and the iteration terminates when the goal predicate (ct)G(c_t) is satisfied or a budget is exhausted (Sapkota et al., 2026; OpenAI, 2026). These capabilities make planning a foundational element for intent-driven and autonomous operation in AI-native 5G and 6G networks. 3.4. Agentic AI Architectural Patterns Agentic AI architectural patterns define modular design abstractions for orchestrating reasoning, planning, memory, and tool interaction within LLM-based systems. These patterns formalize how agents interact with external environments and internal cognitive modules, enabling structured autonomy, improved controllability, and scalable deployment in NGNs (Jaggavarapu, 2025). They further provide a principled basis for balancing latency, reliability, and adaptability under telecom-grade operational constraints. The key trade-offs across different agentic AI architectural patterns in telecom scenarios are summarized in Table 3. The following patterns represent widely adopted agentic designs in complex networked systems. Table 3. Trade-offs of Agentic AI Architectural Patterns in Telecom Scenarios. Arch. Strengths Limitations Best-Fit Telecom Scenarios ReAct Low-latency iterative reasoning with tool grounding; strong for telemetry-driven decision loops. Limited long-horizon planning; sensitive to noisy observations. Real-time fault detection, KPI-driven monitoring, RAN anomaly diagnosis. AutoGPT Strong long-horizon autonomy; supports hierarchical task decomposition and self-refinement. High computational overhead; potential error accumulation without supervision. Intent-based orchestration, automated troubleshooting, lifecycle management. Agentic RAG High factual grounding via adaptive retrieval; improves compliance and knowledge freshness. Retrieval latency; dependency on external knowledge quality. Standards compliance checks, policy validation, root cause analysis. Multi-Agent Systems High scalability and parallelism; robust to domain decomposition and partial failures. Coordination complexity; communication overhead; non-trivial convergence. Network slicing orchestration, cross-domain optimization, large-scale SON control. 3.4.1. ReAct (Reason + Act) ReAct interleaves explicit reasoning traces with tool execution, enabling stepwise grounding of decisions in external observations (Yao et al., 2023). This tight coupling between inference and action supports iterative telemetry acquisition and corrective control, making it suitable for closed-loop network monitoring and online diagnosis in dynamic RAN and core environments. 3.4.2. AutoGPT AutoGPT implements a fully autonomous loop combining task decomposition, persistent memory, and iterative self-refinement (Qian et al., 2023). It supports long-horizon execution where objectives are recursively decomposed into subtasks, executed via tools, and updated based on feedback signals. This makes it relevant for autonomous operations such as intent-driven orchestration, multi-stage troubleshooting, and optimization workflows across distributed telecom domains. 3.4.3. Agentic RAG Agentic RAG extends retrieval-augmented generation by embedding retrieval decisions within the agent’s policy loop. The system dynamically determines retrieval timing, query formulation, and source ranking, followed by verification-aware integration of external knowledge (Ameur et al., 2025c). This improves factual grounding and contextual precision in telecom scenarios involving standards compliance, fault localization, and performance auditing. 3.4.4. Multi-Agent Systems Multi-Agent Systems (MAS) distribute reasoning and action across specialized agents operating under coordination protocols. Each agent may focus on a functional plane (e.g., radio optimization, transport control, service assurance), enabling parallelism, fault isolation, and scalable decision-making (Gupta et al., 2023). This paradigm is particularly effective for large-scale NGNs requiring cross-domain orchestration, resilience under partial failure, and concurrent optimization of heterogeneous network slices. 3.5. Agentic AI Protocols The rapid proliferation of Agentic AI systems has exposed challenges in interoperability, scalability, and standardization, as early solutions relied on proprietary frameworks and ad-hoc integrations. This fragmentation hinders reproducibility and cross-system coordination, particularly in large-scale domains such as telecommunications. In response, emerging agentic protocols standardize context access, tool invocation, intent exchange, and multi-agent collaboration, forming a foundation for consistent interaction and semantic alignment across distributed systems (Kong and others, 2025; Ehtesham et al., 2025). Table 4. Summary of Agentic AI protocols and their relevance to Future 5G/6G Networks. Protocol Primary Objective Key Mechanisms 6G Integration Point Transport Constraint MCP Standardize LLM-based agent access to external tools and data sources. JSON-RPC 2.0 over stdio or Streamable HTTP/SSE; tools/call with tool name and typed arguments. Non-RT RIC, SMO, NWDAF, OSS/BSS interfaces. stdio (local), HTTP/1.1 + (remote, TLS) 50–200 ms per tool chain on loaded edge. A2A Interoperable coordination among heterogeneous LLM-based agents. Capability cards, task lifecycle objects, and streaming; JSON-RPC or gRPC transport. Non-RT and near-RT RIC; RAN optimization, slicing, anomaly detection. HTTP+SSE, gRPC (TLS) gRPC serialization overhead in high-frequency small-message exchanges. ANP Open-internet agent discovery and decentralized collaboration. Decentralized identity, structured capability descriptions, web-standard discovery endpoints. Edge and core orchestration; hierarchical and flat agent topologies. HTTPS, secured P2P Identity resolution latency; stale routing during rapid topology changes. ACP Interoperable messaging between LLM agents and legacy network interfaces. Multipart typed payloads over REST; synchronous and asynchronous patterns with session management. Full RAN–core–edge stack; legacy NF and OSS/BSS interoperability. HTTP (TLS) Polling overhead for async flows; no native push without streaming extension. AP2 Secure verifiable agent-initiated transactions. Cryptographically signed mandate chain (intent, cart, payment) as verifiable credentials. Business and service layers; multi-operator leasing and spectrum trading. HTTPS (TLS) Signing latency (1–2 ms/mandate) confines use to non-RT economic loops. 3.5.1. Model Context Protocol (MCP) MCP was introduced by Anthropic in November 2024 to standardize how LLM-based applications communicate with external tools, data sources, and services through a formal client-server architecture (Ehtesham et al., 2025). At the wire level, MCP messages are serialized as JavaScript Object Notation/Remote Procedure Call (JSON-RPC) 2.0 objects. The protocol defines two standard transport mechanisms: stdio, where JSON-RPC messages are exchanged over standard input/output streams, and Streamable HTTP, where client-to-server messages are sent as HTTP POST requests and server-to-client streaming is delivered via Server-Sent Events (SSE). Tool invocation is performed via a tools/call request, whose params object carries a name field identifying the target tool and an arguments field containing the typed key-value input map validated against the tool’s inputSchema. In future 5G/6G deployments, MCP has been adopted to interface LLM agents with NWDAF analytics endpoints, O-RAN xApp registries, and network configuration databases (Li et al., 2025a; Ameur et al., 2026a), primarily at the non-RT RIC and management plane, where its latency profile is compatible with control timescales. 3.5.2. Agent-to-Agent Communication (A2A) Announced by Google in April 2025, A2A defines structured coordination among heterogeneous LLM-based agents (Ehtesham et al., 2025). Messages are organized around four core objects: an Agent Card advertising capabilities, a Task with a defined lifecycle, an Artifact carrying output, and a Message as the atomic communication unit. Transport bindings include JSON-RPC 2.0 over HTTP with SSE streaming, gRPC (Protocol Buffers v3), and HTTP/REST. In 6G contexts, A2A supports decentralized collaboration for RAN optimization, slicing, and anomaly detection, with gRPC offering lower serialization overhead than JSON-RPC for high-frequency inter-agent exchanges (Kong and others, 2025). 3.5.3. Agent Network Protocol (ANP) ANP models agents as networked entities that advertise capabilities and discover peers on the open internet (Agent Network Protocol Contributors, 2024). Built on Decentralized Identifiers (DIDs), Verifiable and Credentials, it defines a three-layer architecture covering identity, meta-protocol negotiation, and application interactions. Agent discovery relies on the well-known URI path, returning a JSON document listing public agent descriptions under a domain. For 5G/6G networks, ANP supports deployment across RAN, core, and edge domains; however, decentralized identity resolution introduces latency that can cause stale routing during rapid topology changes, such as handovers, a failure mode requiring bounded-latency extensions not yet addressed in the literature. 3.5.4. Agent Communication Protocol (ACP) Developed by IBM Research and released in May 2025 under Linux Foundation governance, ACP defines a RESTful HTTP interface for LLM-driven agent interoperability (IBM BeeAI, 2024). Messages are Media types multipart payloads over HTTP, supporting synchronous and asynchronous patterns with structured session management and authentication via Role-Based Access Control (RBAC). In NGNs, ACP facilitates intent exchange between LLM agents and legacy NF management interfaces; its framework-agnostic REST transport and offline agent packaging suit heterogeneous RAN–core–edge deployments without requiring persistent connections. 3.5.5. Agent Payments Protocol (AP2) Announced by Google in September 2025, AP2 standardizes secure financial interactions among autonomous agents via cryptographically signed mandates structured as Verifiable Credentials over HTTPS with Transport Layer Security (TLS) (Parikh and Surapaneni, 2025). Three mandate types, Intent, Cart, and Payment, are signed using encryption protocols, forming a non-repudiable audit chain from user delegation to settlement. From a telecom perspective, AP2 supports multi-operator resource leasing, spectrum trading, and autonomous service procurement (Barros, 2025), though its cryptographic overhead confines it to non-RT business-layer transactions rather than sub-second RAN-tier resource negotiation. 3.5.6. Integration of Agentic Protocols in 6G Networks Table 4 maps each protocol to its target 6G integration point. MCP aligns with the non-RT RIC and management plane, where its tool-invocation model interfaces directly with NWDAF, OSS/BSS, and O-RAN SMO. A2A operates across the non-RT and near-RT RIC tiers, enabling inter-agent coordination for RAN optimization and slicing via JSON-RPC or gRPC transport. ANP targets edge and core orchestration, where web-standard discovery and decentralized identity support large agent collectives but introduce resolution latency under mobility. ACP bridges agentic systems with legacy NF management interfaces across the full RAN–core–edge stack via a framework-agnostic REST interface. AP2 is confined to the business and service layers, supporting autonomous economic transactions through verifiable signed mandates in multi-operator ecosystems. A cross-cutting constraint applies to all five protocols: JSON serialization introduces per-message payload overhead of 0.5–several Kilobyte, making migration to binary encoding and persistent connection reuse a prerequisite for any protocol targeting near-RT or RT 6G control tiers. 3.6. Evaluation Methods Agentic AI systems introduce capabilities such as tool use, planning, multi-step reasoning, and autonomous decision-making, which extend beyond traditional language modeling benchmarks. Accordingly, evaluation must capture not only output correctness but also behavioral reliability, task efficacy, safety, and robustness in dynamic operational environments (Bandi et al., 2025; Farooq et al., 2025). We organize existing evaluation approaches into a compact taxonomy of metrics, while directing readers to prior works for detailed formulations due to space constraints. 3.6.1. Task-Level Performance Metrics This category captures an agent’s ability to complete complex, multi-step workflows involving planning, tool invocation, and interaction with external systems. Representative metrics include task success rate, execution efficiency, error propagation, and constraint adherence. These measures are particularly relevant for telecom scenarios such as fault resolution, resource optimization, and closed-loop automation, with detailed treatments provided in (Bandi et al., 2025). 3.6.2. Tool Interaction and Reliability Metrics Given the reliance of Agentic AI on external tools, APIs, and data sources, this category evaluates tool call correctness, argument validity, latency sensitivity, and recovery from failures. In network management contexts, such metrics reflect robustness in interfacing with telemetry systems, configuration platforms, and monitoring services. Comprehensive definitions and evaluation protocols are discussed in (Farooq et al., 2025). 3.6.3. Efficiency and System-Level Metrics This category addresses the computational and operational overhead introduced by Agentic AI, including resource consumption, inference latency, and tool invocation cost. It further considers trade-offs between autonomy and system load, which are critical in telecom environments with strict real-time and energy constraints. Detailed analyses can be found in (Sapkota et al., 2026; Bandi et al., 2025; Farooq et al., 2025). 4. Agentic AI for 5G/6G Networks: A Survey In this section, we review recent literature exploring LLM-based Agentic AI approaches for future networks, following the architectural framework established in Section 2. We organize prior work according to the three primary 6G layers illustrated in Fig. 2: (i) the control plane, comprising the RAN, CN, TN, and the Edge/Cloud continuum; (i) the management plane, which focuses on frameworks enabling autonomous 6G operations such as IBN and ZSM; and (i) the AI-native plane, which encompasses the full lifecycle of agentic intelligence, including data operations, AI operations (training, inference, and orchestration), and AI explainability. Under this taxonomy, we analyze how Agentic AI optimizes 6G infrastructure resources, enables high-level closed-loop automation, and defines systematic methodologies for deploying intelligent and interpretable models across the network fabric. Collectively, these works illustrate a paradigm shift toward AI-native networking, where intelligence is no longer an overlay but a first-class architectural component of 6G systems. A structured summary of the surveyed contributions is provided in Table 5. Table 5. Summary of LLM-based Agentic AI for 5G/6G Networks. Layer Domain Task Focus References Control plane RAN Resource Allocation Agentic xApps/rApps for O-RAN control and RRM (Elkael and others, 2025; Bao et al., 2025; Salan and others, 2025; Kamatani and others, 2025; Feng and others, 2025) Security Agentic and LLM-based threat detection and security compliance in O-RAN (Chatzimiltis et al., 2025; Moore et al., 2025; Wen and others, 2024) Configuration Agentic observability platforms and conflict evaluation for xApps (Chatzistefanidis and others, 2025; Sharma et al., 2025; Maxenti and others, 2025) Intent-Based Control LLM-as-operator and hierarchical RAN control (Giwa et al., 2025; Elkael and others, 2025; Bao et al., 2025; Salan and others, 2025) Simulation Multi-agent LLM-driven RAN simulation and evaluation (Rezazadeh and others, 2024; Hu and others, 2025) Marketplace Symbiotic/agentic marketplaces and safe coordination (Chatzistefanidis and Nikaein, 2025; Zhang, 2025) TN Configuration LLM-driven SDN/SDM configuration automation (Wang et al., 2024; Neupane and others, 2025) Security LLM-based intrusion and DDoS detection in SDN/SD-WAN (Swileh and Zhang, 2025; Lodh et al., 2025; Zhang and others, 2023; Cao et al., 2019) Traffic Engineering LLM-aware load balancing and graph-based dynamic networking (Li et al., 2025b; Sun and others, 2024) CN Intent Management LLM-assisted intent extraction and semantic routing in 5G/6G (Manias et al., 2024b, a; Rodriguez-Navas et al., 2024; Kan et al., 2024) Agentic Core Control Mission-oriented and autonomous 6G core orchestration (Tong and others, 2025; Li and others, 2025; Yu et al., 2025) LLM-Aware Slicing Core/edge slicing tailored to LLM compute requirements (Liu et al., 2024; Ameur et al., 2024) Edge/Cloud Service Placement Agentic edge for AI services and semantic-aware offloading (Tang and others, 2025; Feng and others, 2025; Li et al., 2025a; Qian and others, 2024) Resource Allocation Orchestration for AI-native and containerized applications (Tang and others, 2025; Feng and others, 2025; Kalafatidis and others, 2025; Li et al., 2025a) Security Intrusion/threat detection for container environments (Kalafatidis and others, 2025; Rigaki et al., 2024) Agent Economy AI-native APIs enabling telco agent economies (Barros, 2025; Zhang, 2025) Management plane IBN Translation Intent parsing and translation to network policies (Mekrache et al., 2024; Mekrache and Ksentini, 2024; Mekrache et al., 2025c, a; Hossain and Aljoby, 2025; Angi et al., 2025; Lira et al., 2024; Ifland and others, 2024) Resolution Intent conflict resolution and explainable justifications (Salmi and others, 2025; Ali et al., 2023) Assurance End-to-end monitoring and SLA verification (Tang and others, 2024; Mekrache and others, 2025; Abbas and others, 2025) ZSM Service Automation Zero-touch instantiation across TN/NTN segments (Abbas and others, 2025; Mekrache and others, 2025) Closed-Loop Control Multi-agent LLM architectures for self-organizing networks (Qayyum and others, 2025; Qu and others, 2025; Gemayel and Mokh, 2025; Shah and Shen, 2025) AI-native plane Data Ops. Data/Telemetry NWDAF-style collection and log semanticization (Quadrini et al., 2023; Rodriguez-Navas et al., 2024; Kamatani and others, 2025; Wu and others, 2024a; Chochliouros and others, 2025) Datasets Domain-specific QA and evaluation datasets for traffic data (Said and others, 2024; Wu and others, 2024a) AI Ops. LLM Adaptation Fine-tuning LLMs (Mobile-LLaMA) and context routing (Said and others, 2024; Kan et al., 2024; Wu and others, 2024a; Li et al., 2025a) Orchestration Multi-agent design patterns for agentic RAN/Core/SON (Mekrache et al., 2025c, a; Qayyum and others, 2025; Qu and others, 2025; Elkael and others, 2025; Li and others, 2025; Barros, 2025; Ameur et al., 2025c) Architecture Blueprints for AI-native 6G and IBN-ZTSA integration (Salmi and others, 2025; Abbas and others, 2025; Feng and others, 2025; Boutouchent and others, 2025) Optimization DRL-based task scheduling and inference optimization (Mekrache et al., 2025b; He et al., 2024) Evaluation Benchmarking frameworks for LLMs and agents (Maatouk et al., 2025; Gupta et al., 2025; Colle et al., 2025; Ferrag et al., 2026; Nikbakht et al., 2024; Wu et al., 2025a) Explainability & Compliance XAI Causal inference and XAI for trustworthy control (Sharma et al., 2025; Mekrache and others, 2025; Wen and others, 2024; Lodh et al., 2025) Compliance Collaborative Agents for MLOps compliance with AI/Telecom regulation (Ameur et al., 2025c) 4.1. Control plane Agentic AI approaches are being extensively explored across all control plane components, spanning the RAN to the Edge/Cloud continuum. These works address critical challenges such as resource allocation, automated configuration, and security within each network domain. In the following, we review recent literature and developments specific to each of these areas. 4.1.1. Radio Access Network In the RAN domain, AgentRAN introduces an AI-native, O-RAN-aligned agentic architecture in which LLM-powered agents interpret natural-language intents and orchestrate hierarchical control loops across rApps, xApps, and dApps for tasks such as scheduling and power control (Elkael and others, 2025). This is complemented by ALLSTaR, which focuses on automated scheduler generation by synthesizing Medium Access Control (MAC) schedulers from intent-like specifications, and LLM-hRIC, which proposes a hierarchical RAN intelligent control framework (Bao et al., 2025). Furthermore, Giwa et al. explore LLM-based assistants for intent-driven RAN management (Giwa et al., 2025). Other proposals, such as RAG-empowered Radio Resource Management (RRM), LLM-xApp, and LLM-5GMAC, demonstrate how agentic reasoning improves spectral efficiency, observability, and troubleshooting in O-RAN deployments (Salan and others, 2025; Wu et al., 2025b; Kamatani and others, 2025). Beyond classical RRM, Agentic AI is increasingly explored for edge-native coordination and network assurance. Frameworks for Internet-of-Vehicles and semantic-aware 6G edge networks leverage agents to coordinate resource allocation across distributed nodes (Feng and others, 2025), while MAS handle traffic management and slice assurance in dense deployments (Salama and others, 2025). On the observability side, MX-AI introduces an agentic control platform for AI-RAN, and AutoRAN advocates zero-touch RAN operation through learning-driven automation (Chatzistefanidis and others, 2025; Maxenti and others, 2025). To address operational conflicts, recent work combines explainable machine learning and causal inference to diagnose and resolve conflicting control logic across xApps (Sharma et al., 2025). Security and trustworthiness remain central challenges in agentic RAN design. Recent works include security-compliance agents, threat-mitigation frameworks such as MobiLLM, and secure slicing xApps for Open RAN (Chatzimiltis et al., 2025; Moore et al., 2025; Wen and others, 2024). To ensure safe agent interactions in open marketplaces, Zhang et al. propose frameworks for secure multi-agent coordination (Zhang, 2025). Additionally, reflection-driven self-optimization and symbiotic agent models are proposed to enhance trustworthiness in AGI-driven RANs (Hu and others, 2025; Chatzistefanidis and Nikaein, 2025). Finally, simulation platforms such as GenoNet enable multi-agent RAN experimentation using ns-3, while Tele-LLM-Hub and NetMCP provide operator assistants and network-aware context protocols for RAN controllers (Rezazadeh and others, 2024; Shah and Shen, 2025; Li et al., 2025a). 4.1.2. Transport Network In the TN, agentic controllers are primarily deployed for configuration automation, security enforcement, and traffic engineering. LLM-enabled full-stack configuration frameworks for Space Division Multiplexing (SDM) networks and LLM-driven SDN controllers for optical transport translate high-level intents into device-level configurations, significantly simplifying the management of complex IP and optical infrastructures (Wang et al., 2024). NetPrompt extends this paradigm by synthesizing SDN policies from operator directives, while graph-based LLM frameworks abstract complex path computation and traffic engineering decisions (Neupane and others, 2025; Sun and others, 2024). Security in TNs is addressed through LLM-based intrusion and anomaly detection in SDN and Software-Defined Wide Area Network (SD-WAN) environments. This includes lightweight fine-tuning for explainable intrusion detection, zero-training proactive Distributed Denial of Service (DDoS) mitigation using port-level monitoring, and malicious packet recognition (Lodh et al., 2025; Swileh and Zhang, 2025). Earlier agent-based models laid the foundation for network service prediction and resource scheduling, while contemporary LLM-driven approaches focus on intelligent SD-WAN maintenance, including alarm correlation and root-cause analysis (Cao et al., 2019; Zhang and others, 2023). Finally, agentic load-balancing mechanisms leverage LLM-based reasoning to optimize flow distribution across dynamic transport topologies (Li et al., 2025b). 4.1.3. Core Network In the CN, semantic and agentic control frameworks are being developed to realize intent-based and mission-oriented operations. Semantic routing mechanisms leverage LLMs to map natural-language intents into declarative policies that steer traffic through appropriate NF chains (Manias et al., 2024b). This direction is further supported by research on LLM-based intent extraction and trace analytics for proactive core monitoring (Manias et al., 2024a; Rodriguez-Navas et al., 2024). Mobile-LLaMA extends this capability through instruction fine-tuning of open-source models, enabling precise analysis of GTP protocols, control-plane messaging, and KPIs for rapid troubleshooting (Kan et al., 2024). Agentic AI has also been proposed to redefine core orchestration. Frameworks such as A-Core and Agentic-AI Core introduce mission-oriented agentic controllers that coordinate core functions, while autonomous cognitive architectures rely on collaborative agents to manage mobility, session continuity, and policy enforcement (Tong and others, 2025; Li and others, 2025; Yu et al., 2025). Moreover, LLM-Slice introduces dedicated network slicing for LLM traffic by jointly considering compute and network requirements, directly coupling agentic application needs with core resource provisioning (Liu et al., 2024; Ameur et al., 2024). 4.1.4. Edge/Cloud At the Edge/Cloud layer, Agentic AI primarily targets AI-service hosting, resource management, and cloud-native security. End-to-end edge AI provisioning frameworks orchestrate AI workloads across RAN and edge segments in 6G O-RAN environments, while semantic-aware edge networks dynamically determine optimal model placement and task offloading under latency and energy constraints (Tang and others, 2025; Feng and others, 2025). NetMCP facilitates this integration by introducing a network-aware MCP that enables LLMs to discover and interact with network APIs for fine-grained resource control (Li et al., 2025a). Large-scale deployments, such as Alibaba HPN, further demonstrate data-center networks optimized for LLM training and inference traffic patterns (Qian and others, 2024). Security in distributed Edge/Cloud environments is increasingly handled by specialized agents. Recent works propose LLM-enhanced intrusion detection for containerized applications using multi-tier IDS architectures for SDN and Kubernetes (Kalafatidis and others, 2025). Similarly, Hackphyr explores locally fine-tuned LLM agents for secure network operations in edge and on-premise settings (Rigaki et al., 2024). Finally, the emergence of a “telco agent economy” is supported by AI-native network APIs that enable external agents to safely invoke, compose, and monetize network services over programmable Edge/Cloud infrastructures (Barros, 2025; Zhang, 2025). Synthesis and Takeaway Lessons. The surveyed literature collectively evidences a transition toward intent-driven, agentic control across RAN, transport, core, and edge domains, yet converges on a consistent set of unresolved system-level constraints. The main conclusions are as follows: (i) there is a fundamental latency–accuracy trade-off, as LLM-based reasoning cannot satisfy sub-10 ms control-loop requirements, which necessitates hierarchical decomposition across non-RT and near-RT planes with model compression or speculative execution, albeit introducing additional and insufficiently quantified E2/A1/O1 signaling overhead; (i) intent-to-policy translation is fragile under topology changes and KPI variability, requiring continuous re-grounding mechanisms closely integrated with real-time observability; (i) embedding agents within the managed infrastructure creates circular resource dependencies that complicate system stability and resource allocation. 4.2. Management plane The management plane encompasses Agentic AI frameworks that operate above the raw infrastructure, with a primary focus on IBN and ZSM. In this layer, LLMs and MAS are leveraged to close the loop between high-level operator goals and low-level network configurations, as well as to automate service lifecycle operations across heterogeneous domains. 4.2.1. Intent-Based Networking A primary group of works focuses on the intent translation problem, which consists of mapping natural-language goals into machine-executable policies. Mekrache et al. propose several LLM-centric IBN pipelines in which intents are parsed, semantically structured, and mapped to network actions, with a particular emphasis on next-generation softwarized networks, Operations Support Systems (OSS), and Business Support Systems (BSS) integration (Mekrache et al., 2024; Mekrache and Ksentini, 2024; Mekrache et al., 2025c, a). Complementarily, NetIntent introduces an end-to-end SDN framework and the IBNBench dataset for benchmarking LLMs on intent translation and flow-conflict detection while coordinating agents across ONOS controllers (Hossain and Aljoby, 2025). At the device level, LLNeT employs a Small Language Model (SLM) to directly instruct softwarized network elements, whereas Genet provides multimodal interfaces that translate high-level descriptions into low-level configurations, reducing operational errors (Angi et al., 2025; Ifland and others, 2024; Lira et al., 2024). Beyond translation, recent literature addresses intent conflict resolution. AI-native O-RAN architectures for 6G incorporate intent-aware control mechanisms to reconcile competing objectives, such as throughput versus energy efficiency (Salmi and others, 2025). Similarly, LLM-assisted Deep Reinforcement Learning (DRL) strategies have been proposed for anti-jamming, where agents iteratively refine policies under adversarial conditions (Ali et al., 2023). A third stream of work focuses on intent assurance and monitoring. LLM-assisted health management architectures perform semantic analysis of telemetry to verify whether deployed configurations satisfy the original intents (Tang and others, 2024). Mekrache et al. further combine XAI and LLMs to provide human-understandable justifications for automated decisions across the IBN pipeline, thereby increasing trust in closed-loop adaptations (Mekrache and others, 2025). This vision is extended by IBN-ZTSA frameworks, which integrate intent-based logic with zero-touch automation for continuous SLA verification in hybrid Terrestrial and Non-Terrestrial Networks (TN/NTN) (Abbas and others, 2025). 4.2.2. Zero-touch network and Service Management In the ZSM domain, Agentic AI is utilized to realize self-organizing, self-healing, and self-optimizing networks across the RAN, CN, TN, and Edge/Cloud. LaMA-SON introduces an LLM-driven multi-agent architecture for intelligent Self-Organizing Network (SON) management, in which specialized agents handle traffic management, QoS optimization, and security threat detection using role-specific reasoning (Qayyum and others, 2025). This multi-agent paradigm is further extended by Qu et al., who propose a dual-loop edge–terminal collaboration framework for 6G that coordinates agents at both the network edge and user terminals to jointly optimize radio and computing resources (Qu and others, 2025). At the orchestration plane, LLM-based MAS automates the lifecycle of Virtualized Network Functions (VNFs) by mapping service descriptions into NF graphs and triggering scaling operations when needed (Gemayel and Mokh, 2025). Tele-LLM-Hub generalizes this concept by providing a context-aware platform where agents collaborate via the TeleMCP protocol to support fault management and performance analysis across domains (Shah and Shen, 2025). Finally, zero-touch service automation is explicitly addressed in the IBN-ZTSA framework, which couples intent-based pipelines with ZSM to automate service instantiation and scaling across TN/NTN segments (Abbas and others, 2025). Complementing this, Mekrache et al. demonstrate that combining XAI and LLMs embeds trust and interpretability into the ZSM loop, enabling autonomous yet auditable scaling and anomaly resolution (Mekrache and others, 2025). Table 6. Comparison of Agentic AI Approaches in the Management plane. Work Agentic Arch. LLM Memory Tools Domain KPI (Mekrache et al., 2024) Single LLM (ReAct) Ollama (Mistral:7B, Llama:13B) Ext. KB + Prompt Eng. OSS API Cloud-Edge/RAN Avg. score, Execution time (Mekrache et al., 2025c) MAS (Hierarchical) Finetuned LLM + GPT-4 Ext. KB + Prompt Eng. OSS API Cloud-Edge/RAN/Core BERT, Cosine, Exact Match (Hossain and Aljoby, 2025) MAS-Hybrid (LLM/Non-LLM) Ollama (Qwen, Phi, Llama, etc.) Prompt Eng. (short-term) SDN (ONOS) SDN Accuracy, Execution time (Angi et al., 2025) Single LLM (ReAct) Gemini-1-Pro, Phi3-mini, Llama3:8B Prompt Eng. (short-term) SDN (Ryu) SDN Latency, Accuracy, Energy, Token, Cost (Ifland and others, 2024) Single Multimodal LLM GPT-4-Vision Prompt Eng. (short-term) Network Topology Net. Configuration Human/LLM Correlation (Lira et al., 2024) Single LLM (ReAct) Ollama (zephyr:7b) Prompt Eng. (short-term) Network Topology Net. Configuration Accuracy, Execution time (Salmi and others, 2025) MAS (Hierarchical) Not Mentioned Ext. Memory + Prompt Eng. O-RAN xApps/rApps RAN No Evaluation (Ali et al., 2023) MAS-Hybrid (DRL-LLM) Ollama (Falcon) Prompt Eng. (short-term) RL environment RAN Policy reward signal (Ameur et al., 2025c) MAS (Hierarchical) Ollama (Mistral, Qwen, deepseek-r1, etc.) Agentic RAG (long+short) Regulation API, Web Search Core/RAN/Edge/ Compliance Resp. time, Score, LLM-as-Judge (Qayyum and others, 2025) MAS (Custom) Ollama (Llama, Qwen, etc.) Cloud memory (centralized) Network management Self-Org. Networks Accuracy, Score rating (Qu and others, 2025) MAS (Distributed) MiniCPM-V2.6 Ext. Memory + Prompt Eng. Map, Meteorology APIs Urban Emergency Success rate, Avg. delay (Gemayel and Mokh, 2025) MAS (CoT) GPT-2, RoBERTa, etc. RAG rApp integration RAN Accuracy (Ameur et al., 2026a) MAS (Hierarchical) Ollama (gpt-oss, gemma3, deepseek-r1, etc.) Agentic RAG + MCP MCP (KB + NWDAF) Core (NWDAF) Task/Tool accuracy, LLM-as-Judge Agentic Capabilities for the Management plane To ensure both depth and clarity, we deliberately chose to expand our focus to the management layer, which represents the most immediate and impactful domain for integrating LLM-based agentic frameworks. This layer inherently aligns with the strengths of Agentic AI, including decision-making, orchestration, policy adaptation, and closed-loop automation across complex and dynamic network environments. Moreover, this focus is strongly motivated by the fact that contemporary telecom standardization efforts, such as those led by the ETSI within the ETSI GR ENI framework, explicitly emphasize intelligence and automation at the management and orchestration layers. These initiatives highlight the centrality of this layer in enabling cognitive network operations and validating AI-driven control loops. Table 6 synthesizes representative LLM-based agentic frameworks for NGN management along the axes of architectural paradigm, model choice, tool ecosystem, memory design, and reported KPIs. The table reveals a clear progression from monolithic LLM agents toward distributed, hierarchical, and hybrid MAS, driven by the scalability, robustness, and control-loop requirements of telecom deployments. Early works such as (Mekrache et al., 2024) and (Lira et al., 2024) adopt a single-agent ReAct paradigm in which reasoning and acting are co-located within one LLM instance, typically a GPT-4-class model selected for its zero-shot reasoning and tool-use capabilities. Their evaluation centers on semantic accuracy, intent translation fidelity, and execution latency. While these models deliver strong reasoning precision, they are constrained by bounded context windows, the absence of persistent memory, and prompt sensitivity, which limit reproducibility and robustness in multi-domain scenarios. Hierarchical MAS frameworks, including (Mekrache et al., 2025c), (Ameur et al., 2025c), and (Ameur et al., 2026a), mark a substantive architectural shift. They deploy specialized agents backed by heterogeneous LLMs, combining fine-tuned domain-specific models with high-capacity proprietary ones to match reasoning depth and latency to task requirements. Evaluation moves beyond isolated accuracy toward system-level indicators such as SLA compliance, orchestration latency, and cross-domain coordination efficiency. Agentic RAG and hybrid memory mechanisms strengthen contextual grounding across iterative control loops, and the emerging LLM-as-a-judge paradigm enables meta-evaluation that partially compensates for the lack of standardized telecom benchmarks. Hybrid architectures, exemplified by (Hossain and Aljoby, 2025) and (Ali et al., 2023), separate high-level reasoning from low-level control execution. LLMs handle intent interpretation, policy abstraction, and decision guidance, while deterministic controllers such as SDN frameworks or DRL agents perform real-time optimization. This decoupling improves convergence time, policy optimality, and stability in dynamic RAN environments. The LLM–DRL coupling is particularly illustrative: the LLM constrains the policy search space while DRL provides adaptability and performance guarantees under stochastic conditions (Ameur et al., 2024), addressing the lack of formal convergence properties in purely generative approaches. Two transversal dimensions further differentiate these systems. First, tool ecosystems have evolved from narrow OSS/BSS and SDN bindings toward broader integration with O-RAN rApps/xApps, NWDAF analytics, and external data sources, with standardized protocols such as MCP (Ameur et al., 2026a) enabling structured and interoperable tool invocation. This repositions the LLM as an orchestrator of tool-augmented cognition rather than a standalone decision engine, at the cost of new overheads in tool invocation latency and interoperability. Second, memory design has progressed from stateless prompt-based context toward multi-layered architectures combining short-term conversational state with long-term knowledge persistence via agentic RAG or centralized repositories. These designs improve long-horizon task performance but introduce consistency, synchronization, and retrieval-efficiency challenges, particularly in distributed deployments such as (Qu and others, 2025). Synthesis and Takeaway Lessons. At the management plane, Agentic AI strengthens the coupling between high-level intents and automated execution across IBN and ZSM paradigms, but several limitations persist. The main conclusions are as follows: (i) intent translation fidelity degrades under ambiguity, multi-domain interactions, and conflicting policies, due to context limitations and the absence of consistency-preserving mechanisms; (i) hierarchical multi-agent systems improve SLA alignment and cross-domain coordination, but introduce orchestration and inter-agent communication overhead that is rarely benchmarked against ZSM latency constraints; (i) hybrid LLM–DRL approaches accelerate convergence, yet lack formal stability guarantees under stochastic dynamics; (iv) tool ecosystems remain heterogeneous across OSS/BSS, SDN, O-RAN, and NWDAF, with invocation latency and interoperability costs largely unreported, while memory architectures trade temporal coherence against synchronization and retrieval overhead; (v) the tight integration of semantic intent handling with autonomous orchestration creates compounded failure modes, where misinterpretations propagate across control loops in the absence of explicit validation checkpoints. 4.3. AI-native plane The AI-native plane encompasses the full lifecycle of agentic intelligence within the network, spanning data collection and semanticization (Data Operations), model adaptation and multi-agent orchestration (AI Operations), and mechanisms for ensuring trustworthy, interpretable deployments (Explainability). This holistic perspective positions Agentic AI as a native architectural component of the 6G ecosystem rather than an external add-on. 4.3.1. Data Operations Several works focus on the data plane that feeds Agentic AI. Quadrini et al. demonstrate how the 3GPP NWDAF can be used to collect real 5G core traffic for analytics and downstream AI tasks, providing a practical blueprint for data acquisition in operational networks (Quadrini et al., 2023). Building on this, Rodriguez-Navas et al. propose LLM-assisted trace analytics for next-generation mobile cores, in which traces are ingested, semantically transformed, and queried via LLMs to support troubleshooting and performance analysis (Rodriguez-Navas et al., 2024). At the RAN side, Kamatani et al. leverage LLMs to analyze MAC-layer logs in O-RAN split 7.2, semantically summarizing log sequences and identifying performance bottlenecks that are difficult to capture with rule-based tools (Kamatani and others, 2025). NetLLM generalizes this idea through a multimodal encoder and networking heads that transform heterogeneous inputs such as time series, logs, and topologies, into unified token-like embeddings, enabling a single model to handle diverse networking tasks (Wu and others, 2024a). Chochliouros et al. further emphasize the importance of telecom-specific dataset design and data quality for LLM training in AI-native networks (Chochliouros and others, 2025). On the evaluation side, TrafficNetQA introduces question-answering datasets tailored to traffic network files, enabling systematic assessment of an LLM’s ability to interpret spatial-temporal traffic data (Kwon et al., 2025). Similarly, Abderrahmane et al. construct 5G-focused datasets for tasks such as KPI forecasting and anomaly detection, bridging the gap between generic benchmarks and telecom-specific distributions (Said and others, 2024). 4.3.2. AI Operations AI Operations (AIOps) focus on how models and agents are adapted, orchestrated, and optimized within 6G systems. Mobile-LLaMA applies instruction fine-tuning to open-source LLMs so they can interpret 5G logs and KPIs, demonstrating that domain adaptation significantly improves alarm classification and root-cause analysis (Kan et al., 2024). NetLLM proposes a Low-Rank Networking Adaptation (D-LRNA) scheme that enables a base LLM to solve a wide range of networking tasks with modest overhead (Wu and others, 2024a). NetMCP further extends this capability by defining a network-aware MCP and the SONAR routing algorithm, which jointly consider semantic similarity and real-time QoS when invoking external tools (Li et al., 2025a). Multi-agent orchestration is studied at both the management plane and within network architectures. Qayyum et al. design LLM-driven multi-agent SON frameworks where specialized agents handle traffic engineering and fault management (Qayyum and others, 2025). Qu et al. propose a dual-loop edge–terminal system in which terminal-side and edge-side agents collaborate to optimize radio and compute resources (Qu and others, 2025). Architecturally, AgentRAN and Agentic-AI Core embed hierarchies of agents directly into RAN and core designs, enabling mission-oriented control loops (Elkael and others, 2025; Li and others, 2025). Barros introduces AI-native network APIs for a “telco agent economy,” defining how external agents can safely monetize network capabilities (Barros, 2025). Several works advocate fully AI-native 6G architectures. AI-native O-RAN designs assume AI components as first-class citizens in both control and data planes (Salmi and others, 2025). IBN-ZTSA extends this philosophy to end-to-end service automation across TN/NTN domains (Abbas and others, 2025). Feng et al. propose 6G native-AI edge networks where semantic intelligence is tightly coupled with task-oriented communication (Feng and others, 2025), while Boutouchent et al. outline high-level principles for placing AI functions across the RAN–core–cloud continuum (Boutouchent and others, 2025). Runtime optimization is also a key research direction. Mekrache et al. introduce DRL-based schedulers for LLM workloads that dynamically allocate GPU and network resources (Mekrache et al., 2025b), while Fang et al. study DRL-controlled inference pipelines that decide between edge and cloud execution (He et al., 2024). In parallel, a growing ecosystem of benchmarks (summarized in Table 7), TeleQnA, TeleTables, TelAgentBench, MMTelCo, TeleMath, TSpec-LLM, α 3-Bench, and TeleYAML, enables systematic evaluation of telecom-focused LLMs and agents (Maatouk et al., 2025; Gupta et al., 2025; Colle et al., 2025; Ferrag et al., 2026; Nikbakht et al., 2024; Wu et al., 2025a; GSMA, 2026). Table 7. Summary of Telecom-Oriented Agentic AI Benchmarks. Benchmark Core Task Metric Benchmark Focus TeleQnA Telecom Knowledge QA Exact-Match Accuracy Multiple-choice questions covering telecom standards, terminology, and domain knowledge. TeleTables Table Interpretation Exact-Match / MCQ Accuracy Evaluation of LLM ability to understand and reason over telecom standard tables. TelAgentBench Agentic Telecom Eval. Multi-capability Metrics Benchmark for agentic capabilities (reasoning, planning, tool use, RAG, instruction following) in telecom tasks. M-Telco Multimodal Telecom Tasks Mixed (MCQ, retrieval) Suite of multimodal benchmarks involving text/image tasks for telecom use cases. TeleMath Mathematical Reasoning Exact-Match Telecom-specific math problem solving and numerical reasoning. TSpec-LLM Standards Comprehension Exact-Match / Retrieval Comprehensive dataset for LLM interpretation of 3GPP technical specifications. TeleYAML Intent-to-Config. Gen. Graded Score Generation of structured telecom config from natural language intent (GSMA Open-Telco task). α 3-Bench Conversational UAV Autonomy Composite Score Benchmarking safe, network-aware, and resource-efficient LLM agents in 6G-enabled aerial systems. 4.3.3. Explainability & Compliance Explainability is essential to ensure that Agentic AI decisions remain transparent, auditable, and trustworthy. In the RAN, Sharma et al. combine explainable ML and causal inference to diagnose and mitigate xApp conflicts (Sharma et al., 2025). 6G-xSec introduces an explainable edge security framework for Open RAN that clarifies why specific flows are flagged as malicious (Wen and others, 2024). At the management level, Mekrache et al. advocate the joint use of XAI and LLMs to support human-in-the-loop validation within ZSM (Mekrache and others, 2025). In transport security, Lodh et al. show that feature-attribution explanations can help operators calibrate trust in LLM-based SDN intrusion detection systems (Lodh et al., 2025). Beyond explainability, regulatory compliance is an emerging concern: EUROCOMPLY proposes an agentic MLOps framework that automates compliance verification against AI and telecom regulations, embedding policy-aware agents directly into the model lifecycle (Ameur et al., 2025c). Synthesis and Takeaway Lessons. At the AI-native plane, the viability of agentic 6G systems is constrained by tightly coupled challenges across data pipelines, model orchestration, and runtime deployment. The main conclusions are as follows: (i) telecom-specific DataOps pipelines emphasize semanticization and quality control, but the lack of standardized interfaces between NWDAF outputs and LLM-ready representations leads to ad hoc preprocessing, limiting reproducibility and benchmark portability; (i) in AI Operations, multi-agent decomposition, model adaptation, and DRL-based scheduling enable functional scalability, yet inter-agent communication overhead, consensus latency, and tool-routing delays remain unquantified relative to ZSM and SON timing constraints, while inference optimization is not aligned with network control objectives, leaving end-to-end latency budgets unmet; (i) AI-native architectures distributing inference across the RAN–core–cloud continuum introduce significant context migration and session state transfer costs that scale with model size and interaction history; (iv) explainability is treated as a post-hoc process and does not provide corrective feedback into agent decision loops, limiting its utility for real-time control; (v) despite emerging telecom-specific benchmarks, evaluation remains fragmented and lacks standardization for consistent performance assessment. 5. Standardization and Projects This section reviews key standardization activities and collaborative research projects that explore the integration of LLM-based Agentic AI into NGN architectures. 5.1. Standardization Efforts Standardization plays a foundational role in shaping the evolution of the telecommunications industry by translating emerging research concepts into interoperable, scalable, and deployable technologies. By defining common architectures, interfaces, and operational principles, standards enable multi-vendor adoption of AI-driven automation with predictable behavior and trust. As the industry moves toward 6G, Standards Development Organizations (SDOs) are formalizing LLM-based agents as native architectural components of network control, management, and orchestration, as summarized in Table 8. Table 8. Key Standardization and Industry Efforts Relevant to LLM-Based Agentic AI in 6G Networks. Org. Scope and Key Docs Agentic AI Technical Focus 6G Integration Point Open Technical Gap 3GPP 6G service/architecture; TR 22.870, SA2 AI-native studies LLM agents as native 6G NFs; intent-aware SBA interfaces; multi-agent service coordination across network domains. 6G Core SBA (N-series interfaces); NRF-based NF registration extended to agent identities. LLM agent identity and OAuth2 token binding for ephemeral agent instances not yet specified. ETSI ZSM ISG (GS ZSM 020/022); ENI ISG (GR 051, GR 055, GR 056, GS 059, GR 060–062); NFV ISG Closed-loop intent realization via ZSM management domains; ENI GR 056 multi-agent coordination for core networks; GS 059 agent interface and protocol specification (telecom-native agentic protocol); GR 062 security aspects; NFV MANO intent interfaces. E2E management fabric; cross-domain intent interfaces; MANO orchestration workflows; GS 059 agent interaction primitives for 5G/6G core. No binding latency requirements for closed-loop intent realization; GS 059 lacks wire-level serialization format, leaving 3GPP N-series and TM Forum API interoperability implementation-dependent. TM Forum Autonomous Networks program; TMF Open APIs (TMF620/641/688); ODA Agent-driven OSS/BSS workflows; LLM intent decomposition over Open APIs; model-agnostic A2A protocol for inter-agent coordination; ODA AI/data capabilities. OSS/BSS northbound; TMF Open API tool invocation via MCP-compatible function-call schemas. A2A protocol lacks binding serialization format; wire-level interoperability with 3GPP/ETSI intent interfaces underspecified. ITU-T IMT-2030 (Focus Group AI-native Networks, SG13); X.700-series alignment AI agents as perception-reasoning-decision-coordination functional entities; LLMs for intent reasoning over X.700-series management interfaces; agent lifecycle management. AI-native management plane; supervisory agent trust frameworks with configurable autonomy thresholds. Agent lifecycle procedures for stateful LLM sessions not yet harmonized with 3GPP NF lifecycle management. O-RAN Alliance NG-RG 6G studies; O-RAN WG2 (non-RT RIC/A1); WG3 (near-RT RIC/E2) LLM agents as rApps (non-RT RIC, A1 intent policies) and xApps (near-RT RIC, E2 control, 10–100 ms loop); intent mediation, policy planning, adaptive RAN control. Non-RT RIC (rApp, A1/O1 interfaces); near-RT RIC (xApp, E2 interface); open fronthaul for edge inference. No certification path for distilled LLM xApps at near-RT RIC; latency-accuracy trade-off at E2 boundary unresolved. IETF IETF 124 Proceedings and draft_agentic_usecases Protocol-level agent capability discovery and context sharing; YANG model extensions for agent-readable network state; delegated OAuth 2.0 credentials for agent authorization. Northbound RESTCONF/NETCONF interfaces; multi-operator agent identity federation. Delegated credential binding for ephemeral LLM agents across multi-operator domains lacks standardized solution. GSMA Open-Telco LLM Benchmarks; TelecomGPT; Foundry pilots Quantitative benchmarking of LLM/agent capabilities on 3GPP spec comprehension, KPI-driven configuration, fault diagnosis; end-to-end agent workflow evaluation covering latency, energy, and API overhead. Operator validation layer bridging research benchmarks and 3GPP/ETSI deployment requirements. Benchmark KPI sets incompatible with 3GPP/ETSI metrics; unified cross-SDO evaluation framework absent. At the service and core architecture layer, 3GPP TR 22.870 envisions AI agents as first-class entities within the 6G SBA, interpreting user intents and coordinating services across domains through intent-aware NF interfaces (3GPP, 2024; 3GPP TSG SA WG2, 2024). The International Telecommunication Union Telecommunication Standardization Sector (ITU-T) Focus Group on AI-native Networks (SG13) complements this by specifying perception–reasoning–decision–coordination roles for agents and aligning LLM-driven intent reasoning with X.700-series management interfaces, while mandating human-in-the-loop validation above configurable autonomy thresholds (ITU, 2025b, a). Management and orchestration frameworks are addressed by ETSI, whose ZSM ISG defines closed-loop, cross-domain intent interfaces (ETSI ISG ZSM, 2024; ETSI, 2023) and whose ENI ISG has produced the most mature telecom-native agentic protocol stack to date, spanning multi-agent coordination (GR 056), the agent interface specification (GS 059), and security aspects of AI agent-based cores (GR 062), with complementary integration into NFV-MANO workflows (ETSI, 2025a, b). TM Forum operationalizes these principles in OSS/BSS through its five-level Autonomous Networks model, in which agents invoke TMF620/641/688 Open APIs as tools via MCP-compatible function-call schemas and coordinate through a model-agnostic A2A protocol (TM Forum, 2025a, b; TM Forum Catalyst Project, 2024). At the RAN, the O-RAN Alliance maps LLM agents onto its disaggregated RIC architecture, deploying them as rApps consuming A1 policy intents at the non-RT tier and, where latency permits, as distilled xApps acting on E2 indications within the 10–100 ms control loop (O-RAN Alliance, 2025b, a, Technical report; Chatzistefanidis and others, 2025). Internet Engineering Task Force (IETF) provides the underlying protocol substrate through use-case-driven requirements for agent capability discovery and context sharing, alongside YANG/RESTCONF extensions and OAuth 2.0-based delegated credentials for agent authorization (Internet Engineering Task Force (IETF), 2024a, b). Finally, Global System for Mobile communications (GSMA) anchors these efforts empirically: its Open-Telco LLM Benchmarks quantify the performance gap between general-purpose and telecom fine-tuned models (e.g., TelecomGPT) on 3GPP comprehension and KPI-driven configuration tasks, and extend evaluation to end-to-end agent workflows covering latency, energy, and API overhead (GSMA and Hugging Face, 2025; GSMA Foundry and Khalifa University, 2025; GSMA and others, 2025). Across all SDOs, three open gaps recur and are detailed per-organization in Table 8: (i) identity and credential binding for ephemeral LLM agents, unresolved in both 3GPP NRF/OAuth2 procedures and IETF delegated-credential schemes; (i) wire-level interoperability between ETSI GS 059, TM Forum A2A, and 3GPP N-series interfaces, which currently lack a shared serialization format; and (i) the latency–accuracy trade-off for agent placement, most acute at the O-RAN near-RT RIC boundary and compounded by the absence of unified cross-SDO benchmarking KPIs. 5.2. 6G Project Initiatives Beyond formal standardization, large-scale 6G research initiatives validate AI-native architectural concepts through experimentation and prototyping, with LLM-based Agentic AI emerging as a recurring design primitive for translating service objectives into coordinated cross-domain actions. Table 9 summarizes the representative initiatives discussed below, detailing their agentic architectures, integration points, and current limitations. Within the European SNS-JU programme, four complementary projects explore distinct layers of the agentic stack. SUNRISE-6G deploys LLM agents as intent mediators that translate natural-language service requests into platform-specific orchestration primitives over TMF-aligned northbound APIs, establishing Agentic AI as a federation layer across heterogeneous 6G testbeds (SUNRISE-6G Consortium, 2024). 6G-INTENSE advances this toward full Intent-Based Networking (IBN) lifecycle management, orchestrating a multi-agent pipeline that decomposes intents into domain sub-intents, verifies feasibility against real-time network state, enforces policies, and closes the assurance loop through KPI monitoring aligned with ETSI ZSM and TM Forum TMF921 interfaces (6G-INTENSE Consortium, 2021). 6G-DALI extends agentic control into the data and model-lifecycle plane, using LLM agents to orchestrate DataOps and MLOps workflows, including dataset generation, preprocessing, training scheduling, and reproducibility, across NWDAF-compatible nodes (6G-DALI Consortium, ). FLECON-6G complements these with a Network Digital Twin (NDT) substrate over which agents query twin state to compute cross-domain control updates, with explainability enforced through causal justification logs traceable to originating intents (FLECON-6G Consortium, 2024). Beyond Europe, the U.S. DoD-sponsored OPEN6G initiative targets the RAN tier through its AgentRAN hierarchical multi-agent system, deploying LLM rApps that consume A1 intent policies at the non-RT RIC and lightweight inference xApps that act on E2 indications within the 10–100 ms control loop, with digital-twin validation gating live RAN commitment (Open6G OTIC, ). Collectively, these projects converge on three common limitations detailed per-project in Table 9: the absence of formal schema alignment and convergence guarantees for intent grounding under conflicting or concurrent requests, the lack of quantified end-to-end latency characterization for closed-loop agentic workflows, and the missing conformance and certification paths, most acute for LLM xApps operating under sub-10 ms E2 deadlines, required to transition prototype results into standardized deployment. Table 9. Summary of 6G Project Initiatives Leveraging LLM-Based Agentic AI. Project Timeline Funding Agentic Arch. Agentic Contribution Integration Point Limitation SUNRISE-6G 2024–2027 EU SNS-JU Single/multi-agent intent mediators over federated testbeds LLM agents parse natural-language intents and translate them into platform-specific orchestration primitives via northbound APIs; enables cross-platform service coordination. TMF-aligned northbound APIs; federated testbed orchestration layer. Intent grounding via prompt engineering; no formal schema alignment across heterogeneous platforms. 6G-INTENSE 2024–2027 EU SNS-JU Multi-agent IBN pipeline (spec, decompose, enforce, assure) Full intent lifecycle management: LLM agents decompose intents into domain sub-intents, assess real-time feasibility against network state, enforce policies, and close the assurance loop via KPI monitoring. ETSI ZSM cross-domain fabric; TM Forum TMF921 intent interfaces. No formal convergence guarantees under simultaneous conflicting intents; SLA compliance under KPI degradation unverified. 6G-DALI 2025–2028 EU SNS-JU LLM agents as AI lifecycle managers over distributed pipelines Agents interpret natural-language intents to trigger DataOps and MLOps workflows: dataset generation, preprocessing, model training scheduling, and experiment reproducibility across NWDAF-compatible nodes. NWDAF data pipelines; distributed compute schedulers; model registries. End-to-end MLOps loop latency unquantified; consistency under concurrent agent requests on shared infrastructure undefined. FLECON-6G 2025–2028 EU SNS-JU Agentic AI over multi-layer Network Digital Twins Agents query digital twin state to compute cross-domain control updates and trigger automated workflows; explainability constraints require causal justification logs traceable to originating intents. Multi-layer digital twin control plane; cross-domain intent decomposition with explainable decision logging. Twin-reality divergence under anomalies propagates to incorrect agent actions; sync latency under high query rates uncharacterized. OPEN6G 2022–Ongoing U.S. DoD (IB5G) Hierarchical MAS (AgentRAN) across non-RT and near-RT RIC tiers LLM rApps consume A1 intent policies at non-RT RIC; lightweight inference xApps act on E2 indications within 10–100 ms loop; digital twin validates actions before live RAN commitment. O-RAN RIC stack; A1/E2 interfaces; programmable Open RAN testbed with digital twin validation. No O-RAN conformance procedure for LLM xApps; distilled model accuracy under real traffic at <<10 ms E2 deadline undemonstrated. 6. Open Challenges and Future Perspectives While Agentic AI promises to revolutionize telecom network management, it also introduces significant new challenges that must be addressed for 6G. The convergence of LLM capabilities with telecom requirements creates a unique research frontier where scalability, real-time responsiveness, interpretability, security, and regulatory considerations intersect in unprecedented ways (Jiang et al., 2025). In this section, we discuss the major challenges and open research questions of Agentic AI in telecom networks 5G/6G. 6.1. Lack of Deterministic Guarantees: Stochasticity, Hallucination, and Cascading Misinformation LLMs are inherently stochastic systems whose outputs are sampled from a learned probability distribution over token sequences conditioned on the input context, rather than derived from formally specified operational semantics. This stochasticity is the generative source of the hallucination phenomenon, wherein a model produces syntactically coherent but factually or semantically incorrect outputs with a non-zero and generally unquantifiable probability (Huang et al., 2025). In network management, this property is structurally incompatible with the deterministic guarantees required by telecommunications control planes. A hallucinated root-cause attribution, such as misclassifying a radio-frequency interference event as a hardware failure, can trigger erroneous reconfiguration of live gNodeB parameters with cascading effects across co-scheduled user equipment. The problem is compounded in multi-agent pipelines, where the stochastic output of one LLM agent serves as the grounding context for subsequent agents, enabling what the literature terms cascading hallucinations: misinformation is reinforced through memory retrieval, tool invocation, or inter-agent communication and amplified across multiple decision steps before any corrective mechanism intervenes. Future Perspectives. Addressing this challenge requires research into formal uncertainty quantification methods for LLM agents operating in closed-loop network control settings, including Bayesian approaches that produce calibrated confidence estimates alongside network management decisions. Hybrid neuro-symbolic architectures that enforce 3GPP-specified operational constraints as hard logical invariants, rather than as soft learned preferences, offer a promising path toward determinism-compatible agentic control. 6.2. Computational Cost and Energy Footprint The deployment of LLM-based agents as components of network management systems introduces a resource consumption profile that conflicts with the sustainability objectives of 6G, which targets an order-of-magnitude improvement in energy efficiency relative to 5G. State-of-the-art LLMs in the 7B to 70B parameter range require 14 GB to 140 GB of GPU memory for inference, a resource profile that directly conflicts with the 6G target of an order-of-magnitude improvement in energy efficiency relative to 5G. Compression techniques, including 4-bits quantization, structured pruning, and knowledge distillation reduce the memory footprint by factors of four to eight but introduce accuracy degradation on domain-specific reasoning tasks that remains uncharacterized for telecom applications (Jain et al., 2025). In multi-agent deployments spanning Non-RT and Near-RT RIC layers, NWDAF, and distributed edge orchestrators, aggregate inference energy budgets dominate total system cost: in hybrid LLM-MARL architectures, LLM inference accounts for the majority of system energy expenditure even when invoked in fewer decision cycles. Future Perspectives. Purpose-built Small Telecom Language Models targeting 1B to 3B parameters via domain-adaptive fine-tuning, energy-aware agent scheduling policies that gate LLM inference based on decision complexity, and the integration of inference costs as explicit constraints within network resource optimization frameworks are priority directions for future work. 6.3. Real-Time Inference Latency vs. Network Control Time Scales The temporal architecture of 5G/6G control loops imposes strict latency bounds that current LLM inference pipelines cannot satisfy at the functional layers where autonomous control is most consequential. O-RAN Near-RT RIC control loops operate in the 10 ms to 1 s range, whereas LLM autoregressive inference latency is highly dependent on model size, batch size, decoding strategy, and hardware configuration. For instance, large-scale models (e.g., 30B–70B parameters) running with batch size 1 and standard autoregressive decoding (greedy or top-k) on a single high-end GPU (e.g., A100/H100) typically exhibit end-to-end response times on the order of 200 ms to several seconds per query, with per-token latencies in the tens of milliseconds. Smaller models (e.g., 7B–13B) or quantized variants can reduce latency, particularly under larger batch sizes or with aggressive decoding approximations, but this introduces accuracy degradation and still struggles to satisfy strict real-time constraints when accounting for end-to-end pipeline overheads (prompt construction, tool invocation, and network I/O) (Pellejero et al., 2025). Future Perspectives. Hierarchical control architectures separating LLM strategic planning on second-to-minute horizons from lightweight reactive policies executing on millisecond time scales, combined with formal stability and convergence analysis of LLM-in-the-loop RAN control systems, currently absent from the literature, constitute the most urgent research priorities for this challenge. 6.4. Security Vulnerabilities in Multi-Agent Systems Multi-agent LLM architectures introduce AI-to-AI privilege escalation as a structurally novel vulnerability: an agent that resists direct prompt injection will nonetheless execute identical malicious payloads originating from a peer agent, because current safety training addresses human-to-AI rather than AI-to-AI boundaries. Empirically, 82.4% of evaluated models are compromised through inter-agent communication, compared to 52.9% via RAG backdoor attacks and 41.2% via direct injection (Lupinacci et al., 2025). In O-RAN, the openness of A1, E2, and O1 interfaces constitutes an ingress vector for adversarial content into xApp input streams, while in the 5G core, a compromised agent interacting with the NEF or Policy Control Function (PCF) can exfiltrate subscriber data or inject policy rules affecting multiple shared-slice tenants simultaneously. Future Perspectives. Zero-trust security architectures for multi-LLM network management systems, enforcing mutual cryptographic authentication of agent identities and continuous behavioral attestation, represent a necessary and underspecified research direction. The development of guardian agent architectures that monitor the chain-of-thought and tool invocation sequences of primary management agents in real time, intervening before policy-violating actions reach NF APIs, merits investigation. 6.5. Lack of Telecom-Specific Agentic AI Benchmarks Although some benchmarks for LLM agents exist (listed in Table 7), they remain limited in scope and do not adequately capture the complexity of telecom network environments. Current benchmarks focus primarily on general reasoning or software tasks, with limited representation of network dynamics, multi-domain interactions, and real-time constraints. The shortage of realistic datasets and evaluation frameworks hinders the ability to assess the performance, robustness, and scalability of agentic AI systems in telecom contexts. Future Perspectives. Open benchmark suites covering heterogeneous multi-domain scenarios spanning RAN optimization, core orchestration, fault diagnosis, and intent translation, instantiated on open emulation platforms such as Open5GS, OpenAirInterface, and Free5GC, and incorporating standardized adversarial stress tests referenced to standards like 3GPP and ETSI, are a prerequisite for rigorous scientific progress and operator confidence in agentic network management systems. 6.6. Agentic Explainability, Governance, and Regulatory Compliance The EU AI Act classifies autonomous systems operating in critical infrastructure as high-risk, mandating transparency, auditability, and human oversight requirements that current LLM-based agentic architectures are structurally unable to satisfy. The autoregressive attention mechanism underlying LLMs does not produce decision traces interpretable in terms of 3GPP parameter bounds, interference constraints, or operator policy rules, and post-hoc explanation methods, rendering them insufficient for regulatory audit purposes (Ameur et al., 2025c). In multi-operator and cross-border 6G deployments, agentic decisions mediated by third-party-hosted LLMs raise data sovereignty concerns that prevent operators from exposing the full network state required for effective agent operation without violating national data residency regulations, a structural tension addressed by the Sovereign AI paradigm, which advocates embedding LLM inference within operator-controlled infrastructure under nationally governed AI frameworks, though its implementation within Near-RT and Non-RT RIC components at production scale has not yet been demonstrated. Future Perspectives. Future work should prioritize explainable agentic architectures that generate structured, standards-referenced decision logs mapping each management action to specific 3GPP specification clauses, KPI thresholds, and operator policy rules, enabling post-hoc audit trails compatible with regulatory requirements. Governance frameworks incorporating human-in-the-loop validation gates for high-impact control actions, policy-constrained inferencing, and zero-trust model update pipelines are essential operational safeguards. Active collaboration between AI researchers, network operators, 3GPP, ETSI, and national telecommunications regulators will be necessary to establish conformance testing regimes and certification procedures for AI-native NFs prior to large-scale 6G deployment. 7. Conclusion In this paper, we provide a comprehensive, tutorial-based survey of Agentic AI empowered by LLMs and its profound implications for the evolution of 5G/6G Networks. We articulate the progression from early predictive language models to autonomous and collaborative agentic systems, and rigorously examine their integration across 5G infrastructures and emerging visions for 6G. By coherently synthesizing advances in agent-centric intelligence with contemporary network architectures, protocol frameworks, and standardization initiatives, this survey elucidates how Agentic AI enables intent-driven operation, ZSM, and adaptive network orchestration at scale. Beyond consolidating the state of the art, we identify key research frontiers and architectural imperatives, positioning Agentic AI as a foundational pillar for intelligent, resilient, and self-evolving future networks. This work aspires to inform, inspire, and guide both academic inquiry and industrial innovation at the intersection of AI and telecommunications. Acknowledgments This work is supported by the European Union Horizon Program under the 6G-INTENSE project (Grant No. 101139266) and the FLECON-6G project (Grant No. 101192462). References 3GPP TSG SA WG2 (2024) System architecture evolution for AI-native and intent-based 6g networks. Technical report 3rd Generation Partnership Project (3GPP). Note: SA2 Contributions on 6G Architecture Cited by: §5.1. 3GPP (2024) TR 22.870: study on 6g use cases and service requirements. Technical report Technical Report Release 20, 3rd Generation Partnership Project (3GPP), Services and System Aspects Group SA1. Cited by: §5.1. [3] 6G-DALI Consortium 6G-DALI: 6g data and ML operations automation via an end-to-end AI framework. Note: https://6gdali.eu/Accessed: 11 December 2025 Cited by: §5.2. 6G-INTENSE Consortium (2021) 6G-INTENSE: intent-driven native AI architecture supporting compute-network abstraction and sensing at the deep edge. Note: https://6g-intense.eu/Accessed: 11 December 2025 Cited by: §5.2. K. Abbas et al. (2025) IBN-ztsa: ai-ibn for zero touch service automation of b5g terrestrial and non-terrestrial networks. IEEE Communications Standards Magazine. Cited by: §4.2.1, §4.2.2, §4.3.2, Table 5, Table 5, Table 5. Agent Network Protocol Contributors (2024) Agent network protocol (ANP). Note: https://github.com/agent-network-protocol/AgentNetworkProtocolAccessed: 30 April 2025 Cited by: §3.5.3. A. S. Ali, D. M. Manias, A. Shami, and S. Muhaidat (2023) Leveraging large language models for drl-based anti-jamming strategies in zero touch networks. External Links: 2308.09376 Cited by: §4.2.1, §4.2, Table 5, Table 6. M. Ameur, A. Bradai, and N. Lagraa (2025a) Exploring teacher-student learning with multi-agent DRL for QoS routing in SDN. In Proceedings of ICC 2025 - IEEE International Conference on Communications, Montreal, QC, Canada, p. 5933–5938. External Links: Document Cited by: §2.3, §3.2, Table 2. M. Ameur, B. Brik, and A. Ksentini (2024) Leveraging llms to explain drl decisions for transparent 6g network slicing. In Proceedings of the 2024 IEEE 10th International Conference on Network Softwarization (NetSoft), Saint Louis, MO, USA, p. 204–212. External Links: Document Cited by: §4.1.3, §4.2, Table 5. M. Ameur, B. Brik, and A. Ksentini (2025b) Adapt but do not forget: towards enhancing drift handling in 6g networks. In Proceedings of the 2025 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN), Barcelona, Spain, p. 1–6. External Links: Document Cited by: §1.1. M. Ameur, B. Brik, and A. Ksentini (2025c) EUROCOMPLY: enabling zero-touch AI compliance auditing via LLM-based agentic AI. IEEE Communications Magazine. External Links: Document Cited by: §3.4.3, §4.2, §4.3.3, Table 5, Table 5, Table 6, §6.6. M. Ameur, B. Brik, and A. Ksentini (2026a) Agentic-nwdaf: enabling intent-driven agentic intelligence for autonomous 6g network analytics. In Proceedings of the IEEE International Conference on Communications (ICC), Glasgow, Scotland, UK. Cited by: §3.5.1, §4.2, §4.2, Table 6. M. Ameur, B. Brik, and A. Ksentini (2025d) Dual self-attention is what you need for model drift detection in 6g networks. IEEE Transactions on Machine Learning in Communications and Networking 3 (), p. 690–709. External Links: Document Cited by: §1.1. M. Ameur, B. Brik, and A. Ksentini (2026b) When mlops meets nwdaf to enable autonomous next generation network analytics. In ICMLCN 2026, IEEE International Conference on Machine Learning in Communications and Networking, Cited by: §2.1. J. G. Andrews, S. Buzzi, W. Choi, S. V. Hanly, A. Lozano, A. C. K. Soong, and J. C. Zhang (2014) What will 5g be?. IEEE Journal on Selected Areas in Communications 32 (6), p. 1065–1082. External Links: Document Cited by: §1.1. A. Angi, A. Sacco, and G. Marchetto (2025) LLNeT: an intent-driven approach to instructing softwarized network devices using a small language model. IEEE Transactions on Network and Service Management. Cited by: §4.2.1, Table 5, Table 6. Anthropic (2026) Building effective AI agents. Note: https://w.anthropic.com/engineering/building-effective-agentsAccessed: 20 January 2026 Cited by: §3.3.1, §3.3.2, §3.3.2, §3.3.2, §3.3.3. L. Antwarg, R. M. Miller, B. Shapira, and L. Rokach (2021) Explaining anomalies detected by autoencoders using shapley additive explanations. Expert Systems with Applications 186, p. 115736. Cited by: §2.3. W. Azariah et al. (2024) A survey on open radio access networks: challenges, research directions, and open source approaches. Sensors 24 (3), p. 1038. Cited by: §2.1. A. Bandi, B. Kongari, R. Naguru, S. Pasnoor, and S. V. Vilipala (2025) The rise of agentic AI: a review of definitions, frameworks, architectures, applications, evaluation metrics, and challenges. Future Internet 17 (9), p. 404. Cited by: §3.6.1, §3.6.3, §3.6. L. Bao, S. Yun, J. Lee, and T. Q. S. Quek (2025) LLM-hRIC: LLM-empowered hierarchical RAN intelligent control for O-RAN. arXiv preprint arXiv:2504.18062. Cited by: §4.1.1, Table 5, Table 5. S. Barros (2025) AI-native network APIs: a telco framework for the agent economy. SSRN (5223311). Cited by: §3.5.5, §4.1.4, §4.3.2, Table 5, Table 5. A. Boutouchent et al. (2025) 6G-intense: intent-driven native artificial intelligence architecture supporting network-compute abstraction and sensing at the deep edge. IEEE Vehicular Technology Magazine. Cited by: §4.3.2, Table 5. Y. Cao, R. Wang, C. Min, and A. Barnawi (2019) AI agent in software-defined network: agent-based network service prediction and wireless resource scheduling optimization. IEEE Internet of Things Journal 7 (7), p. 5816–5826. Cited by: §4.1.2, Table 5. A. Celik and A. M. Eltawil (2024) At the dawn of generative AI era: a tutorial-cum-survey on new frontiers in 6g wireless intelligence. IEEE Open Journal of the Communications Society 5, p. 2433–2489. Cited by: §1.2, §1.2, Table 1, §3.1. S. Chatzimiltis, M. B. Mashhadi, M. Shojafar, M. Debbah, and R. Tafazolli (2025) Agentic AI for 6g: a new paradigm for autonomous RAN security compliance. arXiv preprint arXiv:2512.12400. Cited by: §4.1.1, Table 5. I. Chatzistefanidis et al. (2025) MX-AI: agentic observability and control platform for open and AI-RAN. arXiv preprint arXiv:cs.NI. Cited by: §4.1.1, Table 5, §5.1. I. Chatzistefanidis and N. Nikaein (2025) Symbiotic agents: a novel paradigm for trustworthy AGI-driven networks. Computer Networks, p. 111749. Cited by: §4.1.1, Table 5. Z. Chen, Q. Sun, N. Li, X. Li, Y. Wang, and I. Chih-Lin (2024) Enabling mobile AI agent in 6g era: architecture and key technologies. IEEE Network 38 (5), p. 66–75. Cited by: §1.1, §1.2, §1.2, Table 1, §3.1. I. P. Chochliouros et al. (2025) Developing a 6g data and ML operations automation via an end-to-end AI framework: the 6G-DALI context. In Proceedings of the IFIP International Conference on Artificial Intelligence Applications and Innovations, p. 129–144. Cited by: §2.3, §4.3.1, Table 5. J. Choi et al. (2022) RAN-CN converged control-plane for 6g cellular networks. In Proceedings of GLOBECOM 2022 - 2022 IEEE Global Communications Conference, p. 1253–1258. Cited by: §2.1. V. Colle, M. Sana, N. Piovesan, A. D. Domenico, F. Ayed, and M. Debbah (2025) TeleMath: a benchmark for large language models in telecom mathematical problem solving. External Links: 2506.10674 Cited by: §4.3.2, Table 5. E. Coronado et al. (2022) Zero touch management: a survey of network automation solutions for 5g and 6g networks. IEEE Communications Surveys & Tutorials 24 (4), p. 2535–2578. Cited by: §2.2.2. DeepSeek-AI et al. (2026) DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:cs.CL. Cited by: §3.2, Table 2. T. A. N. do Amaral, R. V. Rosa, D. F. C. Moura, and C. E. Rothenberg (2021) An in-kernel solution based on XDP for 5g UPF: design, prototype and performance evaluation. In Proceedings of the 17th International Conference on Network and Service Management (CNSM), p. 146–152. Cited by: §2.1. G. Dong et al. (2024a) How abilities in large language models are affected by supervised fine-tuning data composition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 177–198. Cited by: §3.2, Table 2. Q. Dong et al. (2024b) A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, p. 1107–1128. Cited by: §3.2, §3.2. R. Dwivedi et al. (2023) Explainable AI (XAI): core ideas, techniques, and solutions. ACM Computing Surveys 55 (9), p. 1–33. Cited by: §2.2.2. A. Ehtesham, A. Singh, G. K. Gupta, and S. Kumar (2025) A survey of agent interoperability protocols: model context protocol (MCP), agent communication protocol (ACP), agent-to-agent protocol (A2A), and agent network protocol (ANP). arXiv preprint arXiv:cs.AI. Cited by: §3.5.1, §3.5.2, §3.5. M. Elkael et al. (2025) AgentRAN: an agentic AI architecture for autonomous control of open 6g networks. arXiv preprint arXiv:2508.17778. Cited by: §4.1.1, §4.3.2, Table 5, Table 5, Table 5. ETSI ISG ZSM (2024) Zero-touch network and service management (ZSM); intent-driven autonomous networks. Technical report Technical Report GR ZSM 011 V2.1.1, European Telecommunications Standards Institute (ETSI). Cited by: §2.2.2, §5.1. ETSI (2023) ETSI GR ZSM 009-3 V1.1.1: closed-loop automation (part 3). Technical report European Telecommunications Standards Institute (ETSI). Cited by: §1.1, §5.1. ETSI (2025a) ETSI GR ENI 051 V4.1.1: study on AI agents based next-generation network slicing. Technical report European Telecommunications Standards Institute (ETSI). Cited by: §5.1. ETSI (2025b) ETSI GR ENI 055 V4.1.1: use cases and requirements for AI agents based core network. Technical report European Telecommunications Standards Institute (ETSI). Cited by: §5.1. L. Fang et al. (2026) Knowledge distillation and dataset distillation of large language models: emerging trends, challenges, and future directions. Artificial Intelligence Review 59 (1), p. 17. Cited by: §2.3. A. Farooq, S. Raza, N. Karim, H. Iqbal, A. V. Vasilakos, and C. Emmanouilidis (2025) Evaluating and regulating agentic AI: a study of benchmarks, metrics, and regulation. TechRxiv. Cited by: §3.6.2, §3.6.3, §3.6. C. Feng et al. (2025) Towards 6g native-AI edge networks: a semantic-aware and agentic intelligence paradigm. arXiv preprint arXiv:2512.04405. Cited by: §4.1.1, §4.1.4, §4.3.2, Table 5, Table 5, Table 5, Table 5. M. A. Ferrag, A. Lakas, and M. Debbah (2026) α3α^3-Bench: a unified benchmark of safety, robustness, and efficiency for llm-based uav agents over 6g networks. External Links: 2601.03281 Cited by: §4.3.2, Table 5. M. A. Ferrag, N. Tihanyi, and M. Debbah (2025) From LLM reasoning to autonomous AI agents: a comprehensive review. arXiv preprint arXiv:cs.AI. Cited by: §1.2, §1.2, Table 1, §3.1. FLECON-6G Consortium (2024) FLECON-6G: flexible open architecture and AI-driven enabling technologies for a novel 6g connectivity platform. Note: https://flecon6g.eu/Accessed: 11 December 2025 Cited by: §5.2. D. Fragkos, G. Makropoulos, A. Gogos, H. Koumaras, and A. Kaloxylos (2022) NEFSim: an open experimentation framework utilizing 3GPP’s exposure services. In Proceedings of the 2022 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit), p. 303–308. Cited by: §2.1. J. Gemayel and A. Mokh (2025) Network function orchestration with llm based multi-agent system. In Proceedings of the 2025 IEEE International Conference on Communications Workshops (ICC Workshops), p. 262–267. Cited by: §4.2.2, Table 5, Table 6. M. Giordani, M. Polese, M. Mezzavilla, S. Rangan, and M. Zorzi (2020) Toward 6g networks: use cases and technologies. IEEE Communications Magazine 58 (3), p. 55–61. Cited by: §2. O. Giwa, M. Adewole, T. Awodumila, and P. Aderinto (2025) The LLM as a network operator: a vision for generative AI in the 6g radio access network. arXiv preprint arXiv:cs.NI. Cited by: §4.1.1, Table 5. GSMA Foundry and Khalifa University (2025) TelecomGPT initiative announcement. Note: GSMA Foundry Cited by: §5.1. GSMA and Hugging Face (2025) OTELLM benchmark technical report. Technical report GSMA. Cited by: §5.1. GSMA et al. (2025) AI telco troubleshooting challenge. Note: GSMA Cited by: §5.1. GSMA (2026) GSMA open-telco llm benchmarks 2.0: the first dedicated llm evaluation for telecoms. Note: https://huggingface.co/blog/otellm/gsma-benchmarks-02Accessed: 2026-01-15 Cited by: §4.3.2. A. Gupta, S. Karamcheti, R. Krishna, and C. Manning (2023) Multi-agent collaboration: harnessing the power of LLMs. arXiv preprint arXiv:2305.11598. Cited by: §3.4.4. A. Gupta and R. K. Jha (2015) A survey of 5g network: architecture and emerging technologies. IEEE Access 3, p. 1206–1232. Cited by: §2.1. G. R. Gupta, A. Kumar, M. Rai, A. Chakraborty, A. Modi, A. Chaoub, et al. (2025) M-telco: benchmarks and multimodal large language models for telecom applications. External Links: 2511.13131 Cited by: §4.3.2, Table 5. Y. He, J. Fang, F. R. Yu, and V. C. Leung (2024) Large language models (llms) inference offloading and resource allocation in cloud-edge computing: an active inference approach. IEEE Transactions on Mobile Computing 23 (12), p. 11253–11264. External Links: Document Cited by: §4.3.2, Table 5. Md. K. Hossain and W. Aljoby (2025) NetIntent: leveraging large language models for end-to-end intent-based sdn automation. Vol. 6. External Links: Document Cited by: §4.2.1, §4.2, Table 5, Table 6. E. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations (ICLR), Cited by: §3.2, Table 2. Y. Hu et al. (2025) Reflection-driven self-optimization 6g agentic AI RAN via simulation-in-the-loop workflows. arXiv preprint arXiv:2512.20640. Cited by: §4.1.1, Table 5. L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu (2025) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst. 43 (2). External Links: ISSN 1046-8188, Link, Document Cited by: §6.1. S. Husain, A. Kunz, and J. Song (2022) 3GPP 5g core network: an overview and future directions. Array. Cited by: §2.1. IBM BeeAI (2024) Introduction to agent communication protocol (ACP). Note: https://docs.beeai.dev/acp/alpha/introductionAccessed: April 2025 Cited by: §3.5.4. B. Ifland et al. (2024) Genet: a multimodal llm-based co-pilot for network topology and configuration. External Links: 2407.08249 Cited by: §4.2.1, Table 5, Table 6. Internet Engineering Task Force (IETF) (2024a) Agentic ai standards. Note: https://w.ietf.org/blog/agentic-ai-standards/Accessed: 2026-04-17 Cited by: §5.1. Internet Engineering Task Force (IETF) (2024b) AI protocol use cases. Internet-Draft. Note: https://w.ietf.org/archive/id/draft-scrm-aiproto-usecases-02.htmlWork in progress. Accessed: 2026-04-17 Cited by: §5.1. ITU (2025a) AI and ML in 5g and beyond challenges. Note: ITU AI for Good Initiative Cited by: §5.1. ITU (2025b) AI for good standards session: multi-agent LLM frameworks for telecom. Note: ITU AI for Good Cited by: §5.1. M. Jaggavarapu (2025) The evolution of agentic AI: architecture and workflows for autonomous systems. Journal of Multidisciplinary 5, p. 418–427. Cited by: §3.4. D. Jain, A. Agarwal, S. Baliyan, and R. Kanagaraj (2025) The carbon footprint of intelligence: the environment cost of llms. In 2025 9th International Conference on Electronics, Communication and Aerospace Technology (ICECA), Vol. , p. 2069–2075. External Links: Document Cited by: §6.2. F. Jiang et al. (2025) A comprehensive survey of large AI models for future communications: foundations, applications and challenges. arXiv preprint arXiv:cs.IT. Cited by: §1.2, §1.2, Table 1. F. Jiang, C. Pan, L. Dong, K. Wang, O. A. Dobre, and M. Debbah (2025) From large AI models to agentic AI: a tutorial on future intelligent communications. arXiv preprint arXiv:cs.AI. Cited by: §1.2, §1.2, Table 1, §3.3.1, §3.3, §6. S. Kalafatidis et al. (2025) LLM-enhanced intrusion detection for containerized applications: a two-tier strategy for sdn and kubernetes environments. In Proceedings of the International Conference on Availability, Reliability and Security, p. 55–73. Cited by: §4.1.4, Table 5, Table 5. O. Kamatani et al. (2025) LLM-5GMAC: performance optimization in O-RAN split 7.2 using LLM-based MAC-layer log analysis. In Proceedings of the 2nd ACM Workshop on Open and AI RAN, Cited by: §4.1.1, §4.3.1, Table 5, Table 5. K. B. Kan, H. Mun, G. Cao, and Y. Lee (2024) Mobile-LLaMA: instruction fine-tuning open-source LLM for network analysis in 5g networks. IEEE Network 38 (5), p. 76–83. Cited by: §4.1.3, §4.3.2, Table 5, Table 5. S. Kitanov et al. (2024) Overview of research trends and challenges in 6g mobile networks and the computing continuum. Scientific and Practical Cyber Security Journal (SPCSJ) 8 (4), p. 54–65. Cited by: §2.1. T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa (2022) Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, Vol. 35, p. 22199–22213. Cited by: §3.2, Table 2. D. Kong et al. (2025) A survey of LLM-driven AI agent communication: protocols, security risks, and defense countermeasures. arXiv preprint arXiv:cs.CR. Cited by: §1.2, §1.2, Table 1, §3.3.1, §3.3, §3.5.2, §3.5. A. Ksentini et al. (2025) Lightweight resource exposure framework for efficient service and resource orchestration in the cloud-edge continuum. In Proceedings of the 2025 IEEE International Conference on Communications Workshops (ICC Workshops), p. 2081–2087. Cited by: §2.3. S. Kukliński, R. Kołakowski, and B. Mastej (2025) MLOps as a service for AI-native 6g networks. In Proceedings of the 2025 IEEE 11th International Conference on Network Softwarization (NetSoft), p. 61–66. Cited by: §2.3. D. Kwon, S. Kang, and S. Choi (2025) TrafficNetQA: question answering datasets for evaluating llm performance on traffic network files. In Proceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems, p. 1142–1145. Cited by: §4.3.1. A. Leivadeas and M. Falkner (2023) A survey on intent-based networking. IEEE Communications Surveys & Tutorials 25 (1), p. 625–655. External Links: Document Cited by: §1.1, §2.2.1. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela (2020) Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 33, p. 9459–9474. Cited by: §3.2, Table 2. E. Li, H. Du, and K. Huang (2025a) NetMCP: network-aware model context protocol platform for llm capability extension. External Links: 2510.13467 Cited by: §3.5.1, §4.1.1, §4.1.4, §4.3.2, Table 5, Table 5, Table 5. M. Li et al. (2023) API-Bank: a comprehensive benchmark for tool-augmented LLMs. arXiv preprint arXiv:2304.08244. Cited by: §3.2, Table 2. X. Li et al. (2025) The agentic-AI core: an AI-empowered, mission-oriented core network for next-generation mobile telecommunications. Engineering. Cited by: §4.1.3, §4.3.2, Table 5, Table 5. X. Li, Y. Zhang, S. Lyu, and W. Wang (2025b) Load balancing for LLM traffic via flow block. In Proceedings of the 2025 IEEE 50th Conference on Local Computer Networks (LCN), p. 1–9. Cited by: §4.1.2, Table 5. O. G. Lira, O. M. Caicedo, and N. L. da Fonseca (2024) Large language models for zero touch network configuration management. IEEE Communications Magazine. Cited by: §4.2.1, §4.2, Table 5, Table 6. B. Liu, J. Tong, and J. Zhang (2024) LLM-Slice: dedicated wireless network slicing for large language models. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, p. 853–854. Cited by: §4.1.3, Table 5. S. Lodh, I. Obaidat, F. Rustam, and A. D. Jurcut (2025) Lightweight fine-tuning of LLMs for explainable intrusion detection in SDN. In Proceedings of the 2025 21st International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob), p. 1–6. Cited by: §4.1.2, §4.3.3, Table 5, Table 5. Q. Long, Y. Chen, H. Zhang, and X. Lei (2022) Software defined 5g and 6g networks: a survey. Mobile Networks and Applications 27 (5), p. 1792–1812. Cited by: §2.1. S. Long et al. (2025) A survey on intelligent network operations and performance optimization based on large language models. IEEE Communications Surveys & Tutorials 27 (6), p. 3915–3949. Cited by: §1.2, §1.2, Table 1. M. Lupinacci, F. A. Pironti, F. Blefari, F. Romeo, L. Arena, and A. Furfaro (2025) The dark side of llms: agent-based attacks for complete computer takeover. arXiv preprint arXiv:2507.06850. Cited by: §6.4. A. Maatouk, F. Ayed, N. Piovesan, A. D. Domenico, M. Debbah, and Z.-Q. Luo (2025) TeleQnA: a benchmark dataset to assess large language models telecommunications knowledge. IEEE Network. Cited by: §4.3.2, Table 5. D. M. Manias, A. Chouman, and A. Shami (2024a) Towards intent-based network management: large language models for intent extraction in 5g core networks. In Proceedings of the 2024 20th International Conference on the Design of Reliable Communication Networks (DRCN), p. 1–6. Cited by: §4.1.3, Table 5. D. M. Manias, A. Chouman, and A. Shami (2024b) Semantic routing for enhanced performance of LLM-assisted intent-based 5g core network management and orchestration. In Proceedings of GLOBECOM 2024 - 2024 IEEE Global Communications Conference, p. 2924–2929. Cited by: §4.1.3, Table 5. T. Masterman, S. Besen, M. Sawtell, and A. Chao (2024) The landscape of emerging AI agent architectures for reasoning, planning, and tool calling: a survey. arXiv preprint arXiv:cs.AI. Cited by: §1.2, §1.2, Table 1. S. Maxenti et al. (2025) AutoRAN: automated and zero-touch open RAN systems. arXiv preprint arXiv:2504.11233. Cited by: §4.1.1, Table 5. A. Mekrache, K. Boutiba, and A. Ksentini (2023) Combining network data analytics function and machine learning for abnormal traffic detection in beyond 5g. In Proceedings of GLOBECOM 2023 - 2023 IEEE Global Communications Conference, p. 1204–1209. Cited by: §2.1, §2.3. A. Mekrache, A. Ksentini, and C. Verikoukis (2024) Intent-based management of next-generation networks: an llm-centric approach. IEEE Network 38 (5), p. 29–36. Cited by: §1.1, §2.2.1, §4.2.1, §4.2, Table 5, Table 6. A. Mekrache, A. Ksentini, and C. Verikoukis (2025a) DMO-gpt: an intent-driven framework for distributed 6g management and orchestration. IEEE Communications Magazine. Cited by: §4.2.1, Table 5, Table 5. A. Mekrache, A. Ksentini, and C. Verikoukis (2025b) DRL-enabled slo-aware task scheduling for large language models in 6g networks. In Proceedings of ICC 2025 - IEEE International Conference on Communications, p. 813–818. Cited by: §4.3.2, Table 5. A. Mekrache, A. Ksentini, and C. Verikoukis (2025c) OSS-GPT: an llm-powered intent-driven operations support system for 6g networks. In Proceedings of the 2025 IEEE 11th International Conference on Network Softwarization (NetSoft), p. 155–163. Cited by: §4.2.1, §4.2, Table 5, Table 5, Table 6. A. Mekrache and A. Ksentini (2024) LLM-enabled intent-driven service configuration for next generation networks. In Proceedings of the 2024 IEEE 10th International Conference on Network Softwarization (NetSoft), p. 253–257. Cited by: §4.2.1, Table 5. A. Mekrache et al. (2025) On combining xai and llms for trustworthy zero-touch network and service management in 6g. IEEE Communications Magazine. Cited by: §2.3, §4.2.1, §4.2.2, §4.3.3, Table 5, Table 5, Table 5. M. Mohammadi, Y. Li, J. Lo, and W. Yip (2025) Evaluation and benchmarking of LLM agents: a survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Vol. 2, p. 6129–6139. Cited by: §1.2, §1.2, Table 1. J. Moore, A. S. Abdalla, P. Khanal, and V. Marojevic (2025) Integrated LLM-based intrusion detection with secure slicing xapp for securing O-RAN-enabled wireless network deployments. arXiv preprint arXiv:2504.00341. Cited by: §4.1.1, Table 5. S. Murugesan (2025) The rise of agentic AI: implications, concerns, and the path forward. IEEE Intelligent Systems 40 (2), p. 8–14. External Links: Document Cited by: §1.1. K. Neupane et al. (2025) NetPrompt: LLM-driven programmable network policy management and optimization. In Proceedings of the 2025 34th International Conference on Computer Communications and Networks (ICCCN), p. 1–9. Cited by: §4.1.2, Table 5. L. Nguyen and Y. Xu (2025) Reasoning for translation: comparative analysis of chain-of-thought and tree-of-thought prompting for LLM translation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, p. 259–275. Cited by: §3.2, Table 2. R. Nikbakht, M. Benzaghta, and G. Geraci (2024) Tspec-llm: an open-source dataset for llm understanding of 3gpp specifications. External Links: 2406.01768 Cited by: §4.3.2, Table 5. O-RAN Alliance (2025a) Distributed GenAI agents for real-time RAN control. Technical report O-RAN Alliance, nGRG Research Contributions. Cited by: §5.1. O-RAN Alliance (2025b) R-2025-02: generative AI use cases and requirements on 6g networks. Technical report O-RAN Alliance. Cited by: §5.1. O-RAN Alliance (Technical report) O-RAN architecture description. Technical report O-RAN Alliance. Note: Release for 5G RAN Cited by: §5.1. [120] Open6G OTIC Open6G OTIC. Note: https://w.open6g.us/Accessed: 11 December 2025 Cited by: §5.2. OpenAI (2026) A practical guide to building agents. Note: https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdfAccessed: 20 January 2026 Cited by: §3.3.1, §3.3.1, §3.3.2, §3.3.2, §3.3.2, §3.3.2, §3.3.4. S. Parikh and R. Surapaneni (2025) Announcing agent payments protocol (AP2). Note: Google Cloud Blog Cited by: §3.5.5. V. B. Parthasarathy, A. Zafar, A. Khan, and A. Shahid (2024) The ultimate guide to fine-tuning LLMs from basics to breakthroughs: an exhaustive review of technologies, research, best practices, applied research challenges and opportunities. arXiv preprint arXiv:cs.LG. Cited by: §3.2. A. K. Pati (2025) Agentic AI: a comprehensive survey of technologies, applications, and societal implications. IEEE Access 13, p. 151824–151837. Cited by: §1.1, §1.2, §1.2, Table 1, §3.3.1, §3.3.1, §3.3.2, §3.3.2, §3.3.3, §3.3. J. Pellejero, L. A. H. Gómez, L. M. Tomás, and Z. F. Barroso (2025) Agentic ai for mobile network ran management and optimization. arXiv preprint arXiv:2511.02532. Cited by: §6.3. C. Prabha, A. Goel, and J. Singh (2022) A survey on SDN controller evolution: a brief review. In Proceedings of the 7th International Conference on Communication and Electronics Systems (ICCES), p. 569–575. Cited by: §2.1. A. Qayyum et al. (2025) LLM-driven multi-agent architectures for intelligent self-organizing networks. IEEE Network. Cited by: §4.2.2, §4.3.2, Table 5, Table 5, Table 6. K. Qian et al. (2024) Alibaba hpn: a data center network for large language model training. In Proceedings of the ACM SIGCOMM 2024 Conference, p. 691–706. Cited by: §4.1.4, Table 5. S. Qian, H. Wang, and C. Shi (2023) AutoGPT: autonomous GPT for complex task solving. arXiv preprint arXiv:2304.03442. Cited by: §3.4.2. L. Qin et al. (2023) Tool learning with foundation models. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36, p. 6215–6230. Cited by: §3.3.3. Z. Qu et al. (2025) LLM enabled multi-agent system for 6g networks: framework and method of dual-loop edge-terminal collaboration. External Links: 2509.04993 Cited by: §4.2.2, §4.2, §4.3.2, Table 5, Table 5, Table 6. M. Quadrini, C. Roseti, F. Zampognaro, and L. Serranti (2023) Data collection using nwdaf network function in a 5g core network with real traffic. In 2023 International Symposium on Networks, Computers and Communications (ISNCC), Vol. , p. 1–7. External Links: Document Cited by: §4.3.1, Table 5. F. Rezazadeh et al. (2024) GenOnet: generative open xg network simulation with multi-agent LLM and ns-3. In Proceedings of the 2024 3rd International Conference on 6G Networking (6GNet), p. 69–71. Cited by: §4.1.1, Table 5. M. Rigaki, C. Catania, and S. Garcia (2024) Hackphyr: a local fine-tuned llm agent for network security environments. External Links: 2409.11276 Cited by: §4.1.4, Table 5. G. Rodriguez-Navas, D. Lee, and K. Kim (2024) Leveraging LLM for evolving and declarative trace analytics towards next-generation mobile core networks. In Proceedings of the 2024 IEEE Future Networks World Forum (FNWF), p. 783–790. Cited by: §4.1.3, §4.3.1, Table 5, Table 5. W. Saad et al. (2025) Artificial general intelligence (AGI)-native wireless systems: a journey beyond 6g. Proceedings of the IEEE P (99), p. 1–39. Cited by: §1.2, §1.2, Table 1. W. Saad, M. Bennis, and M. Chen (2020) A vision of 6g wireless systems: applications, trends, technologies, and open research problems. IEEE Network 34 (3), p. 134–142. External Links: Document Cited by: §1.1. A. I. A. Said et al. (2024) 5G instruct forge: an advanced data engineering pipeline for making llms learn 5g. IEEE Transactions on Cognitive Communications and Networking. Cited by: §4.3.1, Table 5, Table 5. A. Salama et al. (2025) Poster: agentic AI meets neural architecture search: proactive traffic prediction for AI-RAN. In Proceedings of the 2nd ACM Workshop on Open and AI RAN, p. 59–61. Cited by: §4.1.1. O. Salan et al. (2025) RAG-empowered LLM-driven dynamic radio resource management in open 6g RAN. arXiv preprint arXiv:2511.22933. Cited by: §4.1.1, Table 5, Table 5. S. E. Salmi et al. (2025) AI-native O-RAN architectures for 6g: towards real-time adaptation, conflict resolution, and efficient resource management. Authorea Preprints. Cited by: §4.2.1, §4.3.2, Table 5, Table 5, Table 6. J. F. Santos et al. (2025) Managing O-RAN networks: xapp development from zero to hero. IEEE Communications Surveys & Tutorials. Cited by: §2.3. R. Sapkota, K. I. Roumeliotis, and M. Karkee (2026) AI agents vs. agentic AI: a conceptual taxonomy, applications and challenges. Information Fusion 126, p. 103599. Cited by: §1.2, §1.2, Table 1, §3.3.4, §3.3, §3.6.3. V. K. Shah and C. Shen (2025) Tele-LLM-hub: building context-aware multi-agent LLM systems for telecom networks. arXiv preprint arXiv:2511.09087. Cited by: §4.1.1, §4.2.2, Table 5. P. Sharma, S. Sun, S. Deshpande, A. Stavrou, and H. Wang (2025) Towards xapp conflict evaluation with explainable machine learning and causal inference in O-RAN. arXiv preprint arXiv:2510.13031. Cited by: §4.1.1, §4.3.3, Table 5, Table 5. M. K. Shehzad, L. Rose, M. M. Butt, I. Z. Kovacs, M. Assaad, and M. Guizani (2022) Artificial intelligence for 6g networks: technology advancement and standardization. IEEE Vehicular Technology Magazine 17 (3), p. 16–25. Cited by: §2. G. Sun et al. (2024) Large language model (LLM)-enabled graphs in dynamic networking. IEEE Network. Cited by: §4.1.2, Table 5. SUNRISE-6G Consortium (2024) SUNRISE-6G: sustainable federation of research infrastructures for scaling-up experimentation in 6g. Note: https://sunrise6g.eu/Accessed: 11 December 2025 Cited by: §5.2. M. N. Swileh and S. Zhang (2025) Proactive DDoS detection and mitigation in decentralized software-defined networking via port-level monitoring and zero-training large language models. arXiv preprint arXiv:2511.00460. Cited by: §4.1.2, Table 5. L. Tan et al. (2021) In-band network telemetry: a survey. Computer Networks 186, p. 107763. Cited by: §2.3. F. Tang et al. (2024) Large language model (llm) assisted end-to-end network health management based on multi-scale semanticization. External Links: 2406.08305 Cited by: §4.2.1, Table 5. Y. Tang et al. (2025) End-to-end edge AI service provisioning framework in 6g O-RAN. arXiv preprint arXiv:2503.11933. Cited by: §4.1.4, Table 5, Table 5. F. Tariq et al. (2020) A speculative study on 6g. IEEE Wireless Communications 27 (4), p. 118–125. External Links: Document Cited by: §1.1. TM Forum Catalyst Project (2024) Generative AI toolkit for network and service management. Technical report TM Forum. Cited by: §5.1. TM Forum (2025a) IG1251D: autonomous networks agent architecture. Technical report Technical Report v1.0.0, TM Forum. Cited by: §5.1. TM Forum (2025b) Project ONE architecture overview. Technical report TM Forum Innovation Hub. Cited by: §5.1. W. Tong et al. (2025) A-Core: a novel framework of agentic AI in the 6g core network. In Proceedings of the 2025 IEEE International Conference on Communications Workshops (ICC Workshops), p. 1104–1109. Cited by: §4.1.3, Table 5. C. Wang, Y. Wakayama, N. Yoshikane, and T. Tsuritani (2024) LLM-enabled full-stack configuration automation of SDM transport network. In Proceedings of ECOC 2024; 50th European Conference on Optical Communication, p. 1599–1602. Cited by: §4.1.2, Table 5. H. Wen et al. (2024) 6G-xSec: explainable edge security for emerging OpenRAN architectures. In Proceedings of the 23rd ACM Workshop on Hot Topics in Networks, p. 77–85. Cited by: §4.1.1, §4.3.3, Table 5, Table 5. D. Wu et al. (2024a) NetLLM: adapting large language models for networking. In Proceedings of the ACM SIGCOMM 2024 Conference, p. 661–678. Cited by: §4.3.1, §4.3.2, Table 5, Table 5, Table 5. J. Wu et al. (2021) Toward native artificial intelligence in 6g networks: system design, architectures, and paradigms. arXiv preprint arXiv:2103.02823. Cited by: §2.3. K. Wu, Q. Yu, M. Mei, R. Liu, J. Wang, K. Zhang, and Y. Bao (2025a) TN-autorca: benchmark construction and agentic framework for self-improving alarm-based root cause analysis in telecommunication networks. External Links: 2507.18190 Cited by: §4.3.2, Table 5. X. Wu, J. Farooq, Y. Wang, and J. Chen (2025b) LLM-xApp: a large language model empowered radio resource management xapp for 5g O-RAN. In Proceedings of the Symposium on Networks and Distributed Systems Security (NDSS), FutureG Workshop, Cited by: §4.1.1. Z. Wu et al. (2024b) ReFT: representation finetuning for language models. arXiv preprint arXiv:cs.CL. Cited by: §3.2, Table 2. L. Xu, H. Xie, S.-Z. J. Qin, X. Tao, and F. L. Wang (2023) Parameter-efficient fine-tuning methods for pretrained language models: a critical review and assessment. arXiv preprint arXiv:cs.CL. Cited by: §3.2, Table 2. X. Xu et al. (2024) A survey on knowledge distillation of large language models. arXiv preprint arXiv:2402.13116. Cited by: §3.2, Table 2. Z. Yang, M. Chen, K. Wong, H. V. Poor, and S. Cui (2022) Federated learning for 6g: applications, challenges, and opportunities. Engineering 8, p. 33–41. Cited by: §2.3. S. Yao, J. Zhao, D. Yu, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan (2023) ReAct: synergizing reasoning and acting in language models. In Proceedings of the International Conference on Machine Learning (ICML), Cited by: §3.4.1. X. Ye and G. Durrett (2022) The unreliability of explanations in few-shot prompting for textual reasoning. In Advances in Neural Information Processing Systems, Vol. 35, p. 30378–30392. Cited by: §3.2, Table 2. R. Younis et al. (2024) A comprehensive analysis of cloud service models: IaaS, PaaS, and SaaS in the context of emerging technologies and trend. In Proceedings of the 2024 International Conference on Electrical, Communication and Computer Engineering (ICECCE), p. 1–6. Cited by: §2.1. M. Yu, Y. Xing, X. Xia, and J. Jia (2025) AI agent based autonomous cognitive architecture for 6g core network. In Proceedings of the 2025 International Wireless Communications and Mobile Computing (IWCMC), Cited by: §4.1.3, Table 5. M. R. Zafar and N. Khan (2021) Deterministic local interpretable model-agnostic explanations for stable explainability. Machine Learning and Knowledge Extraction 3 (3), p. 525–541. Cited by: §2.3. R. Zhang et al. (2026) Toward edge general intelligence with agentic AI and agentification: concepts, technologies, and future directions. IEEE Communications Surveys & Tutorials. External Links: Document Cited by: §1.2, Table 1. Y. Zhang et al. (2023) Large language model in SD-WAN intelligent operations and maintenance. Research Briefs on Information and Communication Technology Evolution 9, p. 178–188. Cited by: §4.1.2, Table 5. Y. Zhang (2025) SafeServe: scalable tooling for release safety and push testing in multi-app monetization platforms. In Proceedings of the 2025 7th International Conference on Next-Generation Data-Driven Networks (NGDN), Shenyang, China, p. 56–59. External Links: Document Cited by: §4.1.1, §4.1.4, Table 5, Table 5. C. Zhao et al. (2025) Edge general intelligence through world models and agentic AI: fundamentals, solutions, and challenges. arXiv preprint arXiv:cs.LG. Cited by: §1.2, Table 1, §3.1. J. Zhu et al. (2025) Evolutionary perspectives on the evaluation of LLM-based AI agents: a comprehensive survey. arXiv preprint arXiv:cs.CL. Cited by: §1.2, §1.2, Table 1, §3.1.