Paper deep dive
RL-ASL: A Dynamic Listening Optimization for TSCH Networks Using Reinforcement Learning
F. Fernando Jurado-Lasso, J. F. Jurado
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 4:05:20 AM
Summary
RL-ASL is a reinforcement learning-driven adaptive listening framework for TSCH networks that dynamically optimizes radio activation to reduce idle listening. By utilizing offline-trained policies and local runtime inference, it achieves significant energy savings and latency reduction in IIoT networks without requiring global coordination.
Entities (5)
Relation Signals (3)
RL-ASL → evaluatedon → FIT IoT-LAB
confidence 95% · evaluated on the FIT IoT-LAB testbed
RL-ASL → implementedin → CONTIKI-NG
confidence 95% · RL-ASL is fully implemented in CONTIKI-NG
RL-ASL → optimizes → TSCH
confidence 95% · RL-ASL: A Dynamic Listening Optimization for TSCH Networks
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Time Slotted Channel Hopping (TSCH) is a widely adopted Media Access Control (MAC) protocol within the IEEE 802.15.4e standard, designed to provide reliable and energy-efficient communication in Industrial Internet of Things (IIoT) networks. However, state-of-the-art TSCH schedulers rely on static slot allocations, resulting in idle listening and unnecessary power consumption under dynamic traffic conditions. This paper introduces RL-ASL, a reinforcement learning-driven adaptive listening framework that dynamically decides whether to activate or skip a scheduled listening slot based on real-time network conditions. By integrating learning-based slot skipping with standard TSCH scheduling, RL-ASL reduces idle listening while preserving synchronization and delivery reliability. Experimental results on the FIT IoT-LAB testbed and Cooja network simulator show that RL-ASL achieves up to 46% lower power consumption than baseline scheduling protocols, while maintaining near-perfect reliability and reducing average latency by up to 96% compared to PRIL-M. Its link-based variant, RL-ASL-LB, further improves delay performance under high contention with similar energy efficiency. Importantly, RL-ASL performs inference on constrained motes with negligible overhead, as model training is fully performed offline. Overall, RL-ASL provides a practical, scalable, and energy-aware scheduling mechanism for next-generation low-power IIoT networks.
Tags
Links
- Source: https://arxiv.org/abs/2604.07533v1
- Canonical: https://arxiv.org/abs/2604.07533v1
Trouble viewing inline? Open PDF directly →
Full Text
73,687 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20251 RL-ASL: A Dynamic Listening Optimization for TSCH Networks Using Reinforcement Learning F. Fernando Jurado-Lasso, Member, IEEE, and J. F. Jurado Abstract—Time Slotted Channel Hopping (TSCH) is a widely adopted Media Access Control (MAC) protocol within the IEEE 802.15.4e standard, designed to provide reliable and energy- efficient communication in Industrial Internet of Things (IIoT) networks. However, state-of-the-art TSCH schedulers rely on static slot allocations, resulting in idle listening and unnecessary power consumption under dynamic traffic conditions. This paper introduces RL-ASL, a reinforcement learning–driven adaptive listening framework that dynamically decides whether to activate or skip a scheduled listening slot based on real-time network conditions. By integrating learning-based slot skipping with standard TSCH scheduling, RL-ASL reduces idle listening while preserving synchronization and delivery reliability. Experimental results on the FIT IoT-LAB testbed and Cooja network simulator show that RL-ASL achieves up to 46% lower power con- sumption than baseline scheduling protocols, while maintaining near-perfect reliability and reducing average latency by up to 96% compared to PRIL-M. Its link-based variant, RL-ASL- LB, further improves delay performance under high contention with similar energy efficiency. Importantly, RL-ASL performs inference on constrained motes with negligible overhead, as model training is fully performed offline. Overall, RL-ASL provides a practical, scalable, and energy-aware scheduling mechanism for next-generation low-power IIoT networks. Index Terms—Energy Efficiency, Internet of Things, Idle Lis- tening, Reinforcement Learning, Time Slotted Channel Hopping. I. INTRODUCTION T HE rapid evolution of Networked Embedded Sys- tems (NESs) has revolutionized the Internet of Things (IoT), enabling seamless connectivity and intelligent decision- making across diverse applications. IoT technologies are integral to critical domains such as industrial automation, healthcare monitoring, and smart grid systems [1]–[3]. These systems require highly reliable, energy-efficient communica- tion protocols to ensure sustained performance and scalability under varying conditions. Among the communication technologies that support IoT applications, the Time Slotted Channel Hopping (TSCH) protocol has emerged as a leading solution for robust and deterministic networking [4]. By combining time synchroniza- tion and channel hopping, TSCH achieves resilience against interference and multipath fading, making it particularly well- suited for Industrial Internet of Things (IIoT) deployments. Communication in TSCH is organized into synchronized time Manuscript received June 9, 2025; revised x, x. F. Fernando Jurado-Lasso is an independent researcher, based in Cali, Colombia (e-mail: fdo.jurado@gmail.com). J. F. Jurado is with the Department of Basic Science, Faculty of Engineering and Administration, Universidad Nacional de Colombia Sede Palmira, Palmira 763531, Colombia (e-mail: jfjurado@unal.edu.co). slots, with nodes adhering to predefined schedules that de- termine when to transmit, receive, or remain idle [5], [6]. This deterministic operation minimizes collisions and provides predictable performance, which is essential for industrial ap- plications requiring high reliability and low latency. Despite these advantages, TSCH networks suffer from a key inefficiency known as idle listening—a condition where nodes keep their radios active during receive (Rx) slots without any actual packet transmission. This behavior, often occurring in networks with sporadic or low traffic, leads to unnec- essary energy expenditure. In long-lived IoT deployments with battery-powered or energy-harvesting devices, mitigating idle listening is crucial to extending network lifetime and sustaining performance. The proposed solution targets industrial and environmental monitoring systems where both energy efficiency and respon- siveness are critical. Examples include process supervision and anomaly detection in production plants, or safety and condition monitoring in warehouse and logistics infrastructures. In such deployments, sensors must operate for years without mainte- nance, while still being capable of promptly reporting critical events such as temperature deviations, vibration anomalies, or air-quality alerts. Unlike traditional environmental monitoring systems, these applications require a balanced energy-latency tradeoff: energy must be conserved to extend device lifetime, but latency cannot be excessively sacrificed without degrading responsiveness and system reliability. Battery replacement is often costly or infeasible in these settings, and prolonged reporting delays may cause product degradation or safety risks. Therefore, adaptive listening mechanisms that simultaneously achieve low power consumption and reduced latency are essential for next-generation IIoT systems. Although TSCH schedules are deterministic, deciding whether a node should activate its radio in a given reception slot remains challenging due to uncertainty in effective trans- mission timing caused by retransmissions, traffic, etc. Static heuristics or explicit signaling often lead to conservative lis- tening strategies with increased latency or protocol overhead. To address this challenge, this paper proposes Reinforce- ment Learning-based Adaptive Slot Listening (RL-ASL), a RL–based robust, deployment-time adaptive listening policy designed for environments where traffic patterns may vary within a known operational range. Responsiveness is achieved through rapid, local runtime decisions based on a pre-trained policy, while energy efficiency is ensured by minimizing idle listening through learned slot-skipping strategies. RL- ASL dynamically decides whether a node should listen or skip a slot based on learned traffic and scheduling patterns, thereby reducing idle listening without compromising packet arXiv:2604.07533v1 [cs.NI] 8 Apr 2026 JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20252 reliability. Importantly, the proposed approach is agnostic to how TSCH schedules are generated. RL-ASL can operate on top of both static schedules and adaptive TSCH sched- ulers that dynamically add, remove, or relocate cells based on traffic demand or network conditions, as it relies only on locally observable slot-level information at runtime. The learning phase is performed offline using extensive simulations that cover multiple representative traffic patterns, and the resulting policies are aggregated into a single generalized Q- table using Federated Learning (FL). This generalized policy enables nodes to adapt their listening behavior at runtime based solely on locally observed state information, without requiring retraining or redeployment when traffic patterns vary within the trained regime. Here, adaptivity refers to per-slot, runtime decision-making based on observed traffic dynamics, rather than online policy retraining or network- wide reconfiguration after deployment. Scenarios involving highly unpredictable or adversarial traffic changes may require complementary mechanisms, such as centralized reconfigura- tion or online learning, which are outside the scope of this work. Unlike prior work, RL-ASL is fully implemented in CONTIKI-NG [7] and evaluated on the FIT IoT-LAB [8] testbed, ensuring both realism and reproducibility. We also provide an independent implementation of PRIL-M [9], the closest state-of-the-art adaptive listening protocol. Our exper- imental results demonstrate that policies trained on simple simulated topologies remain effective when deployed on larger and structurally different real-world networks, including relay nodes experiencing traffic loads not explicitly seen during training. All our code and experimental artifacts are publicly released to foster further research in adaptive listening for TSCH networks. A. Contributions This paper makes the following key contributions: 1) Proposing RL-ASL: We introduce a novel RL–based adaptive listening protocol for TSCH networks that sig- nificantly reduces idle listening and improves energy efficiency. 2) Formal Modeling and Constraints: We formalize the listen-receive and listen-skip decision constraints that govern adaptive listening behavior, establishing a rigorous framework for protocol design. 3) Generalized Offline Learning: We design an offline train- ing methodology based on diverse traffic patterns and FL, yielding a single Q-table that generalizes across varying network conditions without requiring online retraining. 4) Full Implementation in CONTIKI-NG: We implement RL-ASL entirely within the CONTIKI-NG operating sys- tem, ensuring compatibility with real-world TSCH de- ployments and open testbeds. 5) Experimental Evaluation on FIT IoT-LAB: We evaluate RL-ASL on the FIT IoT-LAB platform, comparing its performance with state-of-the-art protocols—particularly PRIL-M—across diverse network topologies and traffic patterns. 6) Open-Source Release: We publicly release both our RL-ASL and PRIL-M implementations, enabling repro- ducibility and future comparative research in adaptive listening mechanisms 1 . The remainder of this paper is organized as follows: Sec- tion I reviews related work on energy-efficient communication in TSCH networks. Section I presents the network model and discusses idle listening inefficiencies. Section IV details the design of the RL-ASL protocol. Section V describes implementation details and the experimental setup. Section VI outlines the performance metrics and baseline protocols used for evaluation. Section VII presents and discusses the experi- mental results. Finally, Section VIII concludes the paper and discusses future research directions. I. RELATED WORK Energy-efficient and reliable communication remains a core challenge in low-power wireless networks, especially in IIoT settings where long lifetimes and predictable performance are essential [10]. The TSCH protocol has become a leading solu- tion for such systems thanks to its deterministic time-slotting and channel-hopping mechanisms [11]. However, TSCH net- works still suffer from energy waste due to idle listening and limited adaptability to dynamic traffic conditions. A. Autonomous and Centralized Scheduling Mechanisms Early research in TSCH scheduling focused on improving slot allocation to balance reliability and energy efficiency. Orchestra [12] pioneered autonomous scheduling by letting each node independently assign its transmission and reception cells based on network topology. This approach reduced coordination overhead and improved robustness but did not explicitly address the problem of idle listening, as nodes kept their radios active during all assigned Rx slots regardless of actual traffic. Building upon this, Kim et al. proposed ALICE [13], a link-based scheduling protocol where each node allocates cells dynamically according to its neighbors’ traffic direction and updates schedules every slotframe. ALICE improved upon Orchestra in terms of throughput and energy efficiency, yet nodes still incurred energy losses when listening to unused Rx slots. Similarly, centralized and heuristic-based sched- ulers—such as those by Khan et al. [14] and Papadopoulos et al. [15]—achieved gains through global coordination or guard-time optimization but at the expense of scalability and autonomy. More recent adaptive schedulers such as TESLA [16] and OST [17] introduced elastic slotframe management, dynami- cally adjusting the slotframe length and cell allocation based on observed traffic intensity. These methods improve through- put and energy efficiency under fluctuating loads but still require the radio to remain active during allocated Rx slots, leaving idle listening largely unaddressed. 1 The code is available at https://github.com/fdojurado/contiki-ng-rl-asl JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20253 B. Recent Adaptive and Traffic-Aware Scheduling Recent advances have produced more sophisticated adaptive schedulers that monitor network traffic and topology in real time. For example, A3 [18] employs receiver-side traffic estimation to autonomously adjust the number of Tx/Rx cells in the slotframe, achieving notable adaptability and reduced congestion under dynamic traffic. DT-SF [19] follows a similar principle by adjusting slot allocations according to traffic demands and link reliability metrics. Similarly, TA-RPL [20] extends adaptability beyond the MAC layer by integrating traffic-aware routing with TSCH scheduling to optimize end- to-end network performance. It leverages the number of allo- cated transmission cells as an indicator of both link quality and traffic intensity, enabling routing decisions that balance load and improve bandwidth utilization across the network. While these methods substantially improve scheduling agility, they primarily operate at the slotframe or routing level and do not consider fine-grained energy optimization within allocated Rx slots. As a result, even adaptive schedulers may suffer from residual idle listening when nodes expect but do not receive traffic. C. Learning-Based Scheduling Approaches Reinforcement learning (RL) has recently emerged as a powerful tool for adaptive scheduling in constrained IoT environments. Pratama et al. proposed RL-SF [21] and LLQL- SF [22], which use Q-learning to optimize cell allocations and minimize latency across distributed nodes. Similarly, ELISE [23] and HRL-TSCH [24] apply hierarchical or model- based RL to adapt schedules to changing traffic patterns while maintaining deterministic communication. Although these studies demonstrate that RL can outperform heuristic approaches in adapting to dynamic conditions, they focus on slotframe configuration or global scheduling optimization, not on reducing idle listening at the link level. D. Idle Listening Reduction Mechanisms Several mechanisms have been proposed specifically to mitigate idle listening in TSCH networks. Nsabagwa et al. [25] formulated the scheduling problem as a constraint satisfaction problem (CSP) to minimize idle slots in centralized networks, achieving reduced delay and energy consumption but lim- ited scalability. PRIL-F [26] and its successor PRIL [27] introduced proactive radio-sleep coordination based on the ASN, allowing nodes to skip idle Rx slots. More recently, PRIL-M [9] improved robustness to acknowledgment loss and synchronization errors, demonstrating significant energy gains under periodic traffic. However, PRIL-based methods rely on deterministic traffic patterns and shared synchronization, making them less effective in dynamic or bursty environments. Kalita et al. proposed OASA [28], an on-the-fly adaptive scheduler that adjusts slot allocations according to instanta- neous traffic conditions, and TACTILE [29], which distributes slots across multiple channels to reduce desynchronization. While these mechanisms reduce energy consumption by adapt- ing slot allocations, their heuristic nature limits long-term adaptability in non-stationary traffic environments. Notably, PRIL and its variants constitute the existing body of work that explicitly targets adaptive unicast idle listening reduction in TSCH networks. Other recent TSCH proposals primarily address scheduling, routing, or slotframe elasticity and are therefore complementary rather than directly compa- rable to listening control mechanisms such as RL-ASL. E. Positioning of RL-ASL Despite substantial progress, no prior work directly ad- dresses unicast idle listening in distributed TSCH networks using a learning-based approach. Existing adaptive and RL- based schedulers optimize slot allocation or slotframe param- eters but leave per-slot radio activation unmanaged. RL-ASL fills this gap through a lightweight, distributed RL agent that learns per-link traffic patterns and decides whether to listen or skip each Rx slot. Unlike rule-based or centralized schemes, it autonomously adapts to stochastic traffic while maintaining near-perfect reliability. Operating at the granularity of indi- vidual receive opportunities, RL-ASL complements existing schedulers by enhancing energy efficiency without requiring global coordination or traffic predictability. In summary, previous studies focus on slotframe- or network-level adaptation, leaving per-slot energy management largely unexplored. RL-ASL bridges this gap by learning fine-grained listening behaviors that adapt to dynamic traffic, advancing link-level energy optimization in TSCH networks. I. NETWORK MODEL AND IDLE LISTENING INEFFICIENCY This section presents the network model adopted in this work and defines the concept of idle listening inefficiency in TSCH-based networks. A. Network Model We consider a TSCH network composed of N nodes operating in a synchronized slotframe of L time slots, indexed from 0 to L− 1, with each slot of duration τ seconds. The network time is represented by the Absolute Slot Number (ASN), which provides a global synchronization reference for slotframe alignment and coordinated transmissions. Each node n listens for incoming unicast transmissions during one or more designated reception slots per slotframe, denoted by the set L k n ⊆0, 1,...,L− 1 in slotframe k. Communication occurs between node n and its set of neigh- bors M n . The total number of listening activations for node n during an observation window of K consecutive slotframes is given by K X k=1 |L k n |, which represents the cumulative number of reception oppor- tunities across time. Nodes are categorized according to their functional role in the data collection process: JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20254 1) Node Classification: • Root node (G): Collects data from all other nodes. • Relay nodes (Y ): Forward data toward the root. • Leaf nodes (Z ): Generate and transmit data to relay or root nodes. 2) Packet Transmission and Reception: Packets transmitted from node n to node m are denoted as ψ n→m ∈ Ψ, while successfully received packets or those detected through control coordination (e.g., broadcast notifications) are represented as ω m→n ∈ Ω. Each transmission or reception event occurs within a specific slot l of slotframe k. 3) Idle Listening and Energy Waste: Idle listening occurs when a node keeps its radio active in a reception slot but no packet is transmitted to it. This leads to unnecessary energy expenditure without contributing to data reception. We distinguish two cases: (i) Unnecessary listening: The node is awake but no transmission is directed to it. (i) Nec- essary listening: A transmission is attempted, even if it fails; this is not considered idle. Formally, we define the idle listening indicator for node n in slotframe k, slot l, as: δ n (k,l) = ( 1 if l∈L k n and ∄ψ k,l m→n ∈ Ψ, 0 otherwise. (1) This definition ensures that only truly idle listening instances are counted as wasted energy, consistent with TSCH operation principles. 4) Traffic Generation Model: Each node generates packets at random time intervals ∆T n , modeled as ∆T n ∼N (μ n ,σ 2 n ), where μ n and σ n represent the mean and standard deviation of packet generation intervals, respectively. The distribution is truncated at zero to ensure positive intervals. This model reflects practical low-power wireless behavior where jitter is intentionally introduced to mitigate collisions and improve channel fairness, as commonly implemented in operating systems such as CONTIKI-NG [7] and RIOT [30]. 5) Network Topology and Communication Range: Nodes are placed in a two-dimensional plane, with each node n located at coordinates ρ n = (x n ,y n ), where x n ,y n ∈ [x min ,x max ] × [y min ,y max ]. The Euclidean distance between two nodes n and m is d n,m = p (x n − x m ) 2 + (y n − y m ) 2 . Successful communication occurs when d n,m ≤ d max , where d max denotes the maximum communication range. 6) Slotframe and Channel Management: We adopt the Orchestra scheduling framework to coordinate transmissions across multiple slotframes: (i) time synchronization (beacon) slotframe, (i) broadcast slotframe, and (i) unicast slotframe. Our focus is on the unicast slotframe, where receiver nodes listen in a single scheduled Rx slot per slotframe (|L k n | = 1). The network uses C channel offsets, denoted by C = 0, 1,...,C− 1, mapped to physical frequencies through the TSCH hopping sequence to mitigate interference and enable spatial reuse. IV. RL-ASL: RL FOR ADAPTIVE LISTENING The objective of RL-ASL is to minimize the overall energy cost associated with idle listening while maintaining high de- livery reliability—a balance that can be viewed as minimizing Fig. 1. Network model for a TSCH network with five nodes. a total energy cost J total under reliability constraints. RL- ASL achieves this adaptively through local Q-learning at each receiver, without global coordination. RL-ASL is a fully distributed, receiver-side policy that uses tabular Q-learning to decide, at every scheduled unicast Rx slot, whether a node should listen or skip that slot. The design couples a compact per-neighbor temporal model (Exponential Weighted Moving Average (EWMA) inter-arrival and vari- ance) with a feature-engineered aggregated state, a proba- bilistic transmission model (distance-to-nearest Gaussian), and an expected-reward computation that trades energy savings against missed receptions. Nodes learn an ε-greedy tabular policy that is updated online (training mode) or loaded as a frozen table (evaluation mode). A. Neighbor model and notation For node n and each neighbor m ∈ M n we maintain a compact statistical descriptor Θ n,m = asn last n,m , ˆμ n,m , ˆσ 2 n,m , asn exp n,m , κ n,m , where: • asn last n,m is the ASN at which n last heard m; • ˆμ n,m is an EWMA estimate of the inter-arrival (period) between receptions from m (in slots); • ˆσ 2 n,m is an EWMA estimate of the variance of those inter- arrivals; • asn exp n,m = asn last n,m + ˆμ n,m is the next expected ASN for m; • κ n,m ∈ Z ≥0 counts consecutive predicted-but-missed receptions (predicted skips). When node n observes a reception from m at ASN t i , the inter-arrival ∆ n,m (t i ) = asn n,m (t i )− asn n,m (t i−1 ) is used to update the EWMA estimates: ˆμ n,m ← (1− λ)ˆμ n,m + λ ∆ n,m (t i ), ˆσ 2 n,m ← (1− λ)ˆσ 2 n,m + λ ∆ n,m (t i )− ˆμ n,m 2 , with smoothing coefficient λ ∈ (0, 1). On a detected missed event we advance the expectation and increment κ n,m : asn exp n,m ← asn exp n,m + ˆμ n,m , κ n,m ← κ n,m + 1. These per-neighbor statistics are the primitives used by the probability model and the state encoding below. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20255 B. Probability-of-transmission model (distance-to-nearest) At current ASN a node n estimates the probability that at least one child m∈M n will transmit in that slot. Let ∆ n,m (a) = a− max asn last n,m , asn exp n,m . Denote ˆμ n,m and ˆσ n,m the EWMA mean and standard de- viation (square root of ˆσ 2 n,m ) of inter-arrivals. To ensure numerical robustness we clamp ˆσ n,m as: ˆσ n,m ← min α ˆμ n,m , max(σ min , β ˆμ n,m , ˆσ n,m ) , with configuration constants α,β ∈ (0, 1) and σ min > 0. Define the phase and distance-to-nearest: φ n,m (a) = ∆ n,m (a) mod ˆμ n,m , d n,m (a) = min φ n,m (a), ˆμ n,m − φ n,m (a) . We model a per-neighbor instantaneous transmission likeli- hood as a Gaussian kernel on the distance-to-nearest: p n,m (a) = exp − 1 2 d 2 n,m (a) ˆσ 2 n,m , p n,m (a)∈ [0, 1]. (2) Assuming conditional independence across neighbors, the probability that no child transmits is Q m∈M n (1− p n,m (a)), so the aggregated probability of at least one transmission is p n (a) = 1− Y m∈M n 1− p n,m (a) ,(3) which the implementation clamps to p n (a)∈ [ε p , 1−ε p ] with ε p = 10 −3 (approx.) to avoid degenerate expectations. This distance-to-nearest model produces a temporal “bump” slightly before and after expected instants, providing permis- sive behavior under timing uncertainty. C. State encoding (feature engineering and aggregation) To keep the Q-table compact, RL-ASL maps neighborhood statistics to a small discrete state s∈S =0,...,S− 1 via feature engineering and mixed-radix encoding. a) Per-neighbor bins.: For node n with neighborsM n = m 1 ,...,m M n , compute for each neighbor b i = bin B ∆ n,m i (a) ˆμ n,m i , i = 1,...,M n , where bin B (·) maps the normalized elapsed inter-arrival to one of B discrete bins. b) Neighborhoodaggregates.:From b i and d n,m i (a) compute: b = round 1 M n M n X i=1 b i , c short = M n X i=1 I(b i < b th ), d min (a) = min m∈M n d n,m (a), d bin min = bin D d min (a) , c near = min C max , X m∈M n I(d n,m (a)≤ ˆσ n,m ) . Here b th , C max are configuration constants, and bin D (·) maps the minimum distance to one of D discrete bins. c) Mixed-radix encoding.: We form the aggregated state index via a bijection s = f enc b, c short , d bin min , c near , where S = B · C 1 · D· C 2 (with C 1 ,C 2 the discrete ranges of c short ,c near ). All components are clipped to their declared ranges so s is guaranteed to satisfy 0≤ s < S. This compact integer index is the input to the tabular Q-learner. D. Action space At each scheduled unicast Rx slot (decision epoch) node i selects a t ∈A =a (0) ,a (1) =SKIP RX, DO NOT SKIP RX. The radio activation indicator ξ i,t ∈ 0, 1 is determined by the chosen action: ξ i,t = ( 0 if a t = SKIP RX, 1 if a t = DO NOT SKIP RX. E. Reward shaping and expected-reward computation To explicitly account for the cost of missed transmissions, RL-ASL uses an expected-reward formulation that penal- izes skipping reception slots proportionally to the estimated likelihood of an incoming transmission. Let the configured reward/penalty parameter set be R =R succ > 0, R skip > 0, C idle < 0, C miss < 0, interpreted as success reward, skip reward, idle-listen cost and miss penalty respectively. Given the aggregated transmission probability p i,t ≡ p i (a t ) computed by Eq. (3), the expected immediate reward for node i taking action a t is: E[r i,t | a t ] = p n (a t )C miss + (1− p n (a t ))R skip , p n (a t )R succ + (1− p n (a t ))C idle , (4) The implementation uses this expectation as follows: • If a t = SKIP RX there is no observation in the current slot: the Q-update uses E[r i,t | a t ] immediately. • If a t = DO NOT SKIP RX the agent listens; upon slot completion the actual outcome is observed. If a packet arrived then r i,t = R succ ; otherwise the recomputed p i,t is used and the update target is E[r i,t | a t ]. Two heuristic adjustments improve stability: 1) Missed-neighbor penalty: if any neighbor j satisfies ∆ n,j (a)≥ ˆμ n,j + ζ miss ˆσ n,j 2) Near-transmit penalty: if multiple neighbors are near their expected transmit (within one ˆσ), the expected reward for skipping is decreased proportionally to the near count (implementation uses a small logged penalty to bias learning away from skipping). JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20256 F. Tabular Q-learning and learning dynamics Each node i ∈ N maintains a local tabular action–value function Q i :S×A→ R, parameterized by learning rate α ∈ (0, 1], discount factor γ ∈ (0, 1), and exploration rate ε e . At each decision epoch t, after observing state s i,t , selecting action a i,t , and receiving or estimating the immediate reward r i,t , the Q-table is updated according to the standard Q-learning rule: Q i (s i,t ,a i,t )← Q i (s i,t ,a i,t ) + α h r i,t + γ max a ′ ∈A Q i (s i,t+1 ,a ′ )− Q i (s i,t ,a i,t ) i . (5) Action selection follows an ε-greedy policy: a i,t = ( arg max a∈A Q i (s i,t ,a), w.p. 1− ε e , Uniform(A),w.p. ε e , where ε e decays multiplicatively per episode as ε e+1 = max(ε min , η ε ε e ). Each node maintains episode statistics G e = P T e −1 t=0 r i,t and a rolling average G roll over a sliding window of W recent episodes. WheneverG roll exceeds its previous maximum, the table may be checkpointed for debugging or persistence. During training, Q-values are updated online; in evaluation mode, the Q-table is frozen and read-only. G. Runtime decision and learning loop Algorithm 1 summarizes the per-node logic of RL-ASL. At each scheduled unicast reception slot of node i at ASN a: (i) Feature extraction: the node computes per-neighbor fea- tures b n,m (a), aggregates them into ( b,c short ,d bin min ,c near ), and encodes the resulting state s i,t = f enc (·) as in Sec- tion IV-C. (i) Action selection: an action a i,t ∈ A is drawn from the ε-greedy policy on Q i (s i,t ,·). (i) Decision execution: - If a i,t = SKIP RX, the node keeps the radio off, estimates p i,t by Eq. (3), computes the expected reward E[r i,t | a i,t ] by Eq. (4), and applies the Q-update (5) imme- diately (using s i,t+1 = s i,t ). - If a i,t = DO NOT SKIP RX, the node listens; upon slot completion it observes whether a packet was received and sets r i,t = R succ if packet received; otherwise r i,t = E[r i,t | a i,t ]. before performing the same Q-update. (iv) Bookkeeping: increment the step counter, update G e and G roll , and decay ε e when the episode ends. H. Complexity and robustness considerations All computations are local and lightweight: the Q-table size is |S|·|A| entries, and per-neighbor statistics require a few floating-point scalars. The only nonlinear operations are exp(·), √ ·, and modular arithmetic for the phase calculation. Numerical stability is ensured through bounded counters, clamping of probabilities p i,t ∈ [ε p , 1−ε p ], and limiting ˆσ n,m as defined in Section IV-B. Algorithm 1 RL-ASL: Per-node online decision and learning loop 1: for each scheduled Rx slot of node i at ASN a do 2:if not associated or |M i | = 0 then 3:listen (ξ i,t ← 1); Continue 4:end if 5:extract features and encode state s i,t = f enc (·) 6:select a i,t via ε-greedy on Q i (s i,t ,·) 7:if a i,t = SKIP RX then 8:compute p i,t via Eq. (3) 9:compute expected reward ̄r i,t via Eq. (4) 10:update Q i (s i,t ,a i,t ) using ̄r i,t 11:else 12:activate radio (ξ i,t ← 1) 13:if packet received then 14:r i,t ←R succ 15:else 16:recompute p i,t ; r i,t ←E[r i,t | a i,t ] 17:end if 18:update Q i (s i,t ,a i,t ) using r i,t 19:end if 20:bookkeeping (update counters, episode return, decay ε e ) 21: end for I. Discussion on Mobility Support Mobility-induced topology changes, such as those triggered by upper-layer routing protocols (e.g., RPL), can be naturally accommodated by RL-ASL through its 2-H (see Section V-A). Routing updates in RPL-based networks may take several minutes [31]; during this period, affected nodes revert to standard TSCH behavior, listening to all scheduled Rx slots to maintain network connectivity and discover new neighbors. Once routing converges and a new parent is selected, the node reinitiates the 2H process to synchronize with the new receiver. After successful coordination, RL-ASL resumes normal op- eration, allowing the RL process to gradually adapt to the updated neighborhood statistics. This approach enables RL- ASL to preserve connectivity and progressively re-optimize performance under moderate mobility without requiring major protocol modifications. J. Routing Layer Interaction and Adaptation RL-ASL operates at the MAC layer and remains agnostic to the routing protocol above it. This modular design eases integration into existing TSCH stacks, as RL-ASL relies only on local per-neighbor statistics (e.g., reception success, inter- arrival intervals) and the node’s ASN, which are routing- independent. Our evaluation focused on static topologies to isolate MAC-layer adaptation effects. However, routing-layer dynam- ics—such as parent changes or topology updates in RPL—can transiently modify neighborhood relationships. In such cases, RL-ASL automatically reinitializes its 2-H sequence and local neighbor statistics, enabling rapid re-synchronization without global reconfiguration. Future work could strengthen routing-MAC interaction by introducing cross-layer signaling. For example, routing events (e.g., DAO updates or parent switches) could trigger a tempo- rary exploration phase where RL-ASL prioritizes listening and accelerates learning for new neighbors. This would preserve energy efficiency while enhancing robustness in mobile or time-varying networks. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20257 K. Summary RL-ASL integrates: (i) a probabilistic temporal transmis- sion model, (i) compact state aggregation via mixed-radix encoding, and (i) expected-reward-driven tabular Q-learning to balance energy efficiency and reliability on a per-node basis. The fully distributed design requires no message exchange or synchronization across nodes. Heuristics for missed-neighbor and near-transmit situations improve convergence stability and robustness under realistic TSCH dynamics. Algorithm 1 mirrors the embedded implementation and is optimized for memory- and timing-constrained IoT devices. V. IMPLEMENTATION AND EXPERIMENTAL SETUP This section describes the implementation of RL-ASL in CONTIKI-NG [7], the experimental environment on the FIT IoT-LAB [8] testbed, and the configurations used to evaluate its performance under diverse network and traffic conditions. A. Integration of RL-ASL in Contiki-NG RL-ASL was implemented natively in CONTIKI-NG (v5.0), an open-source operating system for low-power and lossy networks (LLNs) with built-in support for IEEE 802.15.4, 6LoWPAN, RPL, and the TSCH MAC protocol. This provides a realistic software stack for evaluating learning-based MAC- layer mechanisms. The RL-ASL module extends the Orchestra [12] schedul- ing framework by introducing a RL-driven adaptive listening mechanism. Each node autonomously decides whether to ac- tivate or skip its scheduled unicast receive (Rx) slot, reducing idle listening while preserving reliability. At runtime, neighboring nodes coordinate through a lightweight 2-H mechanism that synchronizes the transmitter and receiver before each transmission attempt. The handshake succeeds if 2-H n→m = ( 1, if l Tx n = l Rx m ∧ c n = c m ∧ ack n = 1, 0, otherwise, where l Tx n and l Rx m denote the scheduled slot indices of nodes n and m, c n ,c m their channel offsets, and ack n the acknowledgment flag. If a node updates its preferred next-hop, the 2-H sequence is reinitiated, allowing the previous receiver to safely skip its next Rx slot. The Q-learning component operates in two phases: 1) Offline training: performed entirely in the Cooja net- work simulator [32], using diverse traffic patterns to update Q-values and converge to an optimal slot-skipping policy. 2) On-device inference: executed in real hardware using the pre-trained Q-table, stored directly in firmware. B. Experimental Platform All experiments were conducted on the FIT IoT-LAB testbed—a large-scale open-access facility for experimentation with low-power wireless systems. The deployment utilized the IoT-LAB M3 platform, which integrates an ARM Cortex-M3 microcontroller, an Atmel AT86RF231 IEEE 802.15.4 radio, 1 2 3 4 5 Sink Relay Leaf (a) 1 2 3 4 56 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 (b) 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 X (m) 0 1 2 3 4 5 6 7 8 Y (m) 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Z (m) M3-1 M3-3 M3-5 M3-7 M3-9 M3-11 M3-13 M3-15 M3-17 M3-19 M3-21 M3-23 M3-25 M3-27 M3-29 M3-31 M3-33 M3-35 M3-37 M3-39M3-41 M3-43 M3-45 M3-47 M3-49 M3-51 M3-53 M3-55 M3-57 M3-59 M3-61 M3-63 M3-2 M3-4 M3-6 M3-8 M3-10 M3-12 M3-14 M3-16 M3-18 M3-20 M3-22 M3-24 M3-26 M3-28 M3-30 M3-32 M3-34 M3-36 M3-38 M3-40M3-42 M3-44 M3-46 M3-48 M3-50 M3-52 M3-54 M3-56 M3-58 M3-60 M3-62 M3-64 Sink Relay Leaf Unused Routing Topology (c) Fig. 2.Experimental network topologies from Strasbourg IoT-LAB site: (a) Simple 5-node topology used for training and validation, (b) Larger star topology, and (c) Deployment of the star topology in the FIT IoT-LAB. and onboard sensors for temperature, light, and acceleration. This platform provides a representative low-power embedded system with constrained computational and memory resources, while enabling detailed monitoring of energy, radio, and timing metrics at the hardware level. Each node runs CONTIKI-NG with full TSCH synchroniza- tion and fine-grained energy instrumentation, allowing precise measurement of duty cycle, latency, and packet reliability. The controlled experimental environment and large number of available nodes allow reproducible and scalable evaluation of RL-ASL under realistic operating conditions. C. Network Scenarios and Topologies Two representative network topologies were used to assess RL-ASL: • Simple topology: a compact 5-node network comprising one sink, one relay, and three leaf nodes (Fig. 2a). It supports controlled validation of per-slot decision dynam- ics and convergence behavior. All RL-ASL training was exclusively performed in this topology. • Star topology: a larger multi-hop deployment derived from the Strasbourg IoT-LAB site (Fig. 2b and 2c). Leaf nodes are positioned up to three hops from the sink, forming a hybrid star–mesh structure for scalability and latency evaluation under realistic link variability. This topology only uses the pre-trained, aggregated (FL) Q- table obtained from the simple topology. Nodes operate with standard TSCH parameters: 10 ms timeslots, 0 dBm transmission power, and network-wide slot alignment via the ASN. D. Traffic Patterns To evaluate adaptability, multiple traffic generation patterns were implemented as CONTIKI-NG processes. Each node periodically generates packets aligned with its TSCH schedule. The bursty traffic behavior evaluated in this work arises from heterogeneous periodic sources rather than from fully random or erratic traffic. In the heterogeneous traffic pattern, nodes generate packets periodically but with different inter-arrival times. When such flows converge at relay nodes, their super- position creates transient congestion and burst-like reception opportunities, especially in multi-hop topologies. This model JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20258 TABLE I TRAFFIC GENERATION PATTERNS AND PER-NODE TX INTERVALS. PatternJitteredTransmission intervals per node ID High Traffic✓All nodes: 13 s IDs 3, 11, 15, 19: 17 s IDs 4, 12, 16, 20: 30 s IDs 5, 6, 13, 17, 21: 50 s IDs 4, 18, 22: 73 s Sparse Traffic✓Alternating IDs: 60 or 73 s IDs 3, 11, 15, 19: 17 s IDs 4, 12, 16, 20: 19 s IDs 5, 13, 17, 21: 23 s IDs 4, 18, 22: 29 s Heterogeneous Traffic✓ Periodic Traffic✗ TABLE I RL-ASL Q-LEARNING CONFIGURATION PARAMETERS. ParameterValueParameterValue Episode length (T e )500Discount factor (γ)0.9 Learning rate (α)0.15Initial exploration (ε 0 )1.0 Min. exploration (ε min )0.05Decay factor (η ε )0.997 Reward success (R succ )+1.0Skip reward (R skip )+0.5 Idle cost (C idle )–0.5Miss penalty (C miss )–1.0 Terminal reward (success)+5.0Terminal penalty (failure)–5.0 Inter-arrival bin width (B)10Distance bins (D)4 Short-bin threshold (b th )2 reflects common industrial monitoring scenarios in which sub- sets of sensors temporarily report at higher rates due to local events or configuration differences. Table I summarizes the considered traffic modes and per-node transmission intervals. E. Q-learning Configuration and Hyperparameters Table I summarizes the configuration of the RL-ASL tabular Q-learning agent implemented in CONTIKI-NG. The design prioritizes simplicity and computational efficiency for real-time execution on constrained IoT motes. These parameters yield a discrete state space of N s = 10× (3+1)× 4× (3+1) = 640 states, each associated with two possible actions (listen or skip), resulting in a Q-table of size 640 × 2. The reward structure promotes energy efficiency by encouraging nodes to skip idle receive slots, while penalizing missed receptions and unnecessary listening. F. Training and On-Device Inference Offline training was conducted exclusively in both the simple topology and the Cooja network simulator. The Q- learning process is executed entirely offline, and the resulting Q-table is embedded in flash memory as a fixed decision policy, motivated by the constraints of low-power industrial IoT devices. After deployment, RL-ASL performs no online learning or exploration; runtime behavior is limited to de- terministic table lookups based on locally observable state, ensuring predictable execution and compatibility with safety- critical TSCH systems. Each run simulated 10 8 ms (approxi- mately 27.8 hours) of virtual network time, updating Q-values under an exploration rate ε decaying from 1.0 to 0.1. Learning rates α ∈ 0.15, 0.1, 0.05 were evaluated; α = 0.15 yielded 020040060080010001200 Episode 0 200 400 600 800 Reward LR=0.15 LR=0.1 LR=0.05 0.25 0.50 0.75 1.00 Epsilon Epsilon Fig. 3. Convergence of average reward during Q-learning training in the relay topology for different learning rates α. the most stable convergence (Fig. 3) and was adopted for deployment. To enhance generalization across heterogeneous traffic con- ditions, a lightweight FL aggregation was applied to the Q- tables trained independently for each traffic pattern in the simple topology. Specifically, individual modelsQ i were combined via a weighted Federated Averaging (FedAvg) step: Q global = X i w i Q i , w i = E i P j E j , where E i denotes the number of episodes used to train model i. The resulting global model was then deployed directly on real IoT-LAB nodes for inference in both topologies. The final global Q-table comprises 640× 2 = 1280 param- eters, each stored as a 32-bit floating-point value, resulting in a memory footprint of approximately 5 kB. Given the limited Flash memory available in low-power motes, this representation can be further optimized through fixed-point quantization. For instance, scaling Q-values by a factor of 10 and storing them as 16-bit integers reduces the footprint to 2.5 kB with negligible accuracy loss, since the maximum observed Q-value magnitude (|Q| max ≈ 33.4) comfortably fits within the signed 16-bit range. The global Q-table was compiled into firmware as a static lookup table. Inference reduces to a table lookup plus two integer comparisons, requiring only a few bytes of RAM for indexing. Flash overhead is platform-dependent: under 10 kB on the TI C2650 (128 kB Flash / 20 kB RAM) and about 5– 6 kB on the IoT-LAB M3 (STM32F103, 512 kB Flash / 64 kB RAM), comfortably within device constraints. 1) Energy Consumption of Training: All training of RL- ASL was conducted within the Cooja network simulator, which provides cycle-accurate emulation of CONTIKI-NG nodes with realistic radio behavior and timing, ensuring re- producibility without real energy expenditure. Training di- rectly on embedded hardware is infeasible for low-power IoT networks, as RL typically requires thousands or millions of interaction steps to converge—resulting in excessive training time and energy use from repeated transmissions and updates. In contrast, simulation in Cooja network simulator enables accelerated training under controlled yet realistic conditions, faithfully modeling network dynamics and interference. Once converged, only the compact pre-trained Q-table is deployed on the motes. During operation, inference consists solely of a table lookup, incurring negligible computational and energy cost, as demonstrated in Section VII. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 20259 TABLE I SENSITIVITY OF RL-ASL TO THE REWARD PARAMETER R skip IN THE SIMPLE TOPOLOGY. R skip Latency [ms]PDR [%]RDC [%] 0.25202.3799.911.35 0.5196.3899.971.32 0.75198.7099.971.33 G. Sensitivity Analysis Table I illustrates the sensitivity of RL-ASL to the re- ward parameter R skip in the simple topology. Across the tested range, all three performance metrics exhibit only minor variations. In particular, the PDR remains consistently above 99%, varying between 99.91% and 99.97%. The end-to-end latency ranges from 196.38 ms to 202.37 ms, corresponding to a difference of about 6 ms across all configurations, while the RDC remains low and stable between 1.32% and 1.35%. These results indicate that RL-ASL is largely insensitive to moderate changes in reward weighting and does not rely on carefully tuned values of R skip to maintain reliable and energy- efficient operation within a given deployment scenario. This insensitivity reduces the need for precise reward calibration during offline training and supports the practical deployability of RL-ASL. VI. PERFORMANCE METRICS AND BASELINE PROTOCOLS This section defines the performance metrics used to evalu- ate RL-ASL and describes the baseline protocols employed for comparison. The evaluation focuses on quantifying the protocol’s ability to improve reliability, latency, and energy efficiency under varying traffic and topology conditions. A. Performance Metrics We consider four key performance metrics to assess RL- ASL: PDR, latency, and power consumption. Together, these metrics provide a comprehensive view of the trade-offs be- tween reliability, responsiveness, and energy efficiency. 1) PDR: The PDR quantifies link reliability and is defined as the ratio of successfully received packets to the total packets transmitted. For a unicast link from node n to node m, PDR n→m = |Ω n→m | |Ψ n→m | , where Ω n→m and Ψ n→m denote the sets of successfully received and transmitted packets, respec- tively. The network-wide PDR is computed as the average over all leaf nodes: PDR = P n∈Z |Ω n→G | P n∈Z |Ψ n→G | ,(6) where Z is the set of leaf nodes and G denotes the sink. 2) Latency: Latency measures the end-to-end delay expe- rienced by packets. For a unicast link n → m, it is defined as: D n→m = 1 |Ω n→m | P p∈Ω n→m (t p rx − t p tx ), where t p tx and t p rx represent the transmission and reception timestamps of packet p. The network-wide average latency is then computed as: D = P n∈N P p∈Ω n→G (t p rx − t p tx ) P n∈N |Ω n→G | .(7) B. Power and Energy Consumption The average current consumption (κ) of each node is computed using the per-state duty cycle reported by CONTIKI- NG’s built-in Energest module, which tracks time spent in transmit, receive, idle, and low-power states. For node n at time t, κ n (t) = X s∈S D n,s (t)I s ,(8) where D n,s (t) is the fraction of time in state s and I s is the corresponding current draw. We use the current consumption values specified in the M3 platform datasheet [8]: I CPU = 14.0 mA, I LPM = 0.014 mA, I DEEP LPM = 0.002 mA, I TX = 11.6 mA, and I RX = 12.3 mA, with a supply voltage of V = 3.3 V. These parameters are used consistently to compute all energy-related metrics presented in Section VII. C. Baseline Protocols To contextualize the performance of RL-ASL, we compare it against three representative TSCH scheduling approaches spanning autonomous, link-based, and adaptive paradigms: • Orchestra (Orch.) [12]: A reference autonomous scheduling framework that assigns timeslots based on node roles. We use the receiver-based Orchestra mode with default parameters, which provides a fair baseline for static listening schedules. • Link-based Scheduling (Orch.-LB) [13], [33]: A de- terministic link-centric scheduler that allocates timeslots according to link quality metrics. This baseline highlights the gains of RL compared to structured yet non-adaptive scheduling. • PRIL-M [9]: A recent probabilistic adaptive listening protocol that dynamically adjusts reception windows based on traffic. PRIL-M represents the current state of the art in adaptive TSCH listening. We implement RL-ASL as an add-on service that operates on top of the Orchestra and link-based scheduling frame- works—denoted RL-ASL and RL-ASL-LB. As a non-invasive layer, the service dynamically skips or activates receive slots according to the learned policy while preserving the under- lying timeslot allocation and coordination logic of the base schedulers. Note that PRIL-M is evaluated only under periodic traffic conditions, as its design relies on piggybacking deterministic next-transmission timing information in packet headers. This mechanism assumes strictly periodic traffic and is therefore not applicable to heterogeneous or non-periodic workloads, for which its underlying assumptions no longer hold. VII. RESULTS AND DISCUSSION Experiments were conducted on the FIT IoT-LAB testbed using two network configurations: a simple linear topology and a complex star topology. For each topology, performance was assessed under four traffic patterns—high, heterogeneous, sparse, and periodic—as described in Section V. It is worth noting that RL-ASL is designed for static and low-mobility JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 202510 RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (a) RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (b) RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (c) PRIL-M RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (d) Fig. 4. Comparison of PDR across different traffic patterns for the simple topology: (a-d) PDR under high, heterogeneous, sparse, and periodic traffic patterns. RL-ASL-LB Orch.-LB RL-ASL Orch. 10 1 10 0 10 1 Latency [s] (a) RL-ASL-LB Orch.-LB RL-ASL Orch. 10 1 10 0 10 1 Latency [s] (b) RL-ASL-LB Orch.-LB RL-ASL Orch. 10 1 10 0 10 1 Latency [s] (c) RL-ASL-LB Orch.-LB RL-ASL Orch. PRIL-M 10 1 10 0 10 1 Latency [s] (d) Fig. 5. Latency distributions across different traffic patterns for the simple topology: (a-d) Latency under high, heterogeneous, sparse, and periodic traffic patterns. RL-ASL RL-ASL-LB Orch. Orch.-LB 0 1 Power [mW] 0.6 0.8 1.0 1.4 CPU TX RX non-UC RX UC active RX UC idle (a) RL-ASL RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 Power [mW] 0.5 0.8 1.0 1.4 (b) RL-ASL RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 Power [mW] 0.5 0.7 1.0 1.4 (c) PRIL-M RL-ASL RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 Power [mW] 0.5 0.6 0.8 1.0 1.4 (d) Fig. 6. Power consumption breakdown across different traffic patterns for the simple topology: (a-d) Power consumption under high, heterogeneous, sparse, and periodic traffic patterns. 0.751.001.25 Power [mW] 10 1 10 0 10 1 Latency [s] RL-ASL RL-ASL-LB Orch. Orch.-LB 80 85 90 95 100 PDR [%] (a) 0.500.751.001.25 Power [mW] 10 1 10 0 10 1 Latency [s] 80 85 90 95 100 PDR [%] (b) 0.500.751.001.25 Power [mW] 10 1 10 0 10 1 Latency [s] 80 85 90 95 100 PDR [%] (c) 0.500.751.001.25 Power [mW] 10 1 10 0 10 1 Latency [s] RL-ASL RL-ASL-LB Orch. Orch.-LB PRIL-M 80 85 90 95 100 PDR [%] (d) Fig. 7. Trade-off comparison across different traffic patterns for the simple topology: (a-d) Overall performance under high, heterogeneous, sparse, and periodic traffic patterns. The metrics are normalized for comparison. TSCH deployments, where network topology remains stable for sufficiently long periods to allow receiver-side listening policies to be learned and applied effectively. Such conditions are typical in many industrial, environmental monitoring, and smart infrastructure scenarios. In highly dynamic networks with continuous or fast mobility, frequent re-synchronization and parent changes may limit the time during which adaptive listening can be exploited. To ensure statistical robustness, each protocol was executed three times each lasting 1 hour, and results are reported as the average across all runs. Error bars in all figures represent 95% confidence intervals. We analyze results for each topology and traffic pattern, focusing on key performance metrics: PDR, latency, power consumption, and overall performance (via radar charts). A. Simple Topology We first analyze the results obtained from the simple topology depicted in Fig. 2a, evaluated under the four traffic patterns. Fig. 4, 5, 6, and 7 summarize the performance across PDR, latency, power consumption, and overall trade- offs, respectively. 1) Reliability: Fig. 4a-4d show that all protocols achieve near-perfect reliability, with network-wide PDRs close to 100%. PRIL-M achieves a slightly lower PDR (99.5%) in the periodic scenario. In contrast, both RL-ASL and RL-ASL- LB maintain a near-perfect PDR across all traffic patterns, demonstrating that adaptive listening does not compromise reliability. This confirms the robustness of our RL design, which effectively balances exploration and exploitation while maintaining link stability. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 202511 RL-ASL-LB Orch. Orch.-LB RL-ASL 80 90 100 PDR [%] (a) RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (b) Orch.-LB RL-ASL RL-ASL-LB Orch. 80 90 100 PDR [%] (c) PRIL-M RL-ASL RL-ASL-LB Orch. Orch.-LB 80 90 100 PDR [%] (d) Fig. 8. Comparison of PDR across different traffic patterns for the Star topology: (a-d) PDR under high, heterogeneous, sparse, and periodic traffic patterns. 2) Latency: Fig. 5a-5d illustrate latency distributions. The main latency differences stem from the underlying TSCH scheduling model. Receiver-based protocols (Orchestra and RL-ASL) exhibit slightly higher delays due to having only one Rx cell per slotframe, which increases the average waiting time. In contrast, link-based variants (Orchestra-LB and RL- ASL-LB) achieve lower latency since multiple per-link Rx cells reduce queuing delay and enable faster transmissions. Notably, RL-ASL and RL-ASL-LB achieve latency equal to or slightly lower than their Orchestra counterparts despite employing adaptive listening. This confirms that learning- based slot skipping does not disrupt scheduling continuity. RL- ASL reduces average latency by up to 96% compared to PRIL- M, while RL-ASL-LB achieves up to 98% latency reduction relative to PRIL-M. Overall, RL-ASL maintains low latency even under heterogeneous and sparse traffic, showcasing its adaptability to non-periodic workloads. 3) Energy Efficiency: Fig. 6a-6d present the power con- sumption breakdown across CPU, transmission, and reception states. RL-ASL reduces total power consumption on average by over 44% compared to Orchestra across all traffic patterns, validating the effectiveness of adaptive listening in mini- mizing idle-listening energy waste. Similarly, RL-ASL-LB achieves over 46% lower power consumption than Orchestra- LB, demonstrating that per-link scheduling combined with RL- driven listening optimization yields substantial energy savings. The most significant savings occur in the Rx unicast idle state, as adaptive listening effectively suppresses unnecessary listening intervals while preserving delivery performance. Importantly, CPU power remains nearly constant across all protocols, indicating that RL-ASL’s online inference intro- duces negligible computational overhead—an essential prop- erty for embedded devices. Transmission power also remains comparable, confirming that reduced listening does not in- cur retransmission overhead. PRIL-M achieves slightly lower power in the periodic case due to its traffic awareness; however, RL-ASL matches this efficiency while supporting arbitrary, non-periodic traffic patterns. 4) Overall Performance: Fig. 7a-7d present the overall performance trade-offs involving PDR, latency, and power consumption. Both RL-ASL and RL-ASL-LB consistently achieve the best balance across all metrics and traffic pat- terns. RL-ASL offers the most energy-efficient operation with slightly higher latency, while RL-ASL-LB provides lower latency at a modest energy cost. Compared to PRIL-M, RL- ASL delivers comparable power efficiency under periodic traffic but provides lower latency and generalizes effectively to non-periodic and heterogeneous workloads—an essential capability for real-world IoT deployments. B. Star Topology We next evaluate the star topology (Fig. 2b), which intro- duces higher contention and asymmetric traffic. The results, shown in Fig. 8-11, confirm the scalability and robustness of RL-ASL under more complex conditions. 1) Reliability and Latency: Fig. 8a-8d show that RL-ASL and RL-ASL-LB sustain a 100% PDR under all traffic patterns, while PRIL-M’s PDR drops to 83%. This demonstrates the robustness of RL-ASL in handling contention and asymmetric traffic without sacrificing reliability. Latency results (Figs. 9a-9d) further emphasize the benefits of RL. RL-ASL reduces average latency by up to 87% compared to PRIL-M, while RL-ASL-LB achieves reductions of up to 95%. Latency remains stable despite increased con- tention in the star topology, indicating that RL-ASL effectively manages slot skipping without introducing additional delay. 2) Energy Efficiency and Multi-Metric Trade-offs: Fig. 10a- 10d show that RL-ASL maintains the lowest total power consumption among all protocols, outperforming PRIL-M even in the periodic traffic scenario. This highlights that RL- ASL adapts to both temporal and spatial dynamics rather than relying on static periodicity assumptions. Reductions in idle listening power exceed 35% compared to Orchestra and reach up to 56% compared to Orchestra-LB, confirming the effectiveness of adaptive listening in complex topologies. RL- ASL-LB achieves similar savings, demonstrating that per-link scheduling combined with RL-based listening optimization remains effective under high contention. Finally, Fig. 11a–11d show that both RL-ASL and RL-ASL- LB consistently achieve the best overall performance across all evaluated metrics and traffic patterns in the star topology. RL-ASL yields the most energy-efficient operation, at the cost of a slight increase in end-to-end latency, whereas RL-ASL- LB shows lower latency with a modest increase in energy consumption. In comparison to PRIL-M, RL-ASL demon- strates substantially improved energy efficiency, lower latency, and higher reliability in the periodic traffic scenario—which represents the most favorable operating condition for PRIL- M—while also maintaining robust performance under non- periodic and heterogeneous traffic workloads. These results indicate that, despite being trained on a simple topology, the learned policy generalizes effectively and adapts to more complex network structures and traffic conditions that were not seen during training. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 202512 RL-ASL-LB Orch.-LB RL-ASL Orch. 10 0 10 1 Latency [s] (a) RL-ASL-LB Orch.-LB RL-ASL Orch. 10 0 10 1 Latency [s] (b) RL-ASL-LB Orch.-LB RL-ASL Orch. 10 0 10 1 Latency [s] (c) RL-ASL-LB Orch.-LB Orch. RL-ASL PRIL-M 10 1 10 0 10 1 Latency [s] (d) Fig. 9. Latency distributions across different traffic patterns for the Star topology: (a-d) Latency under high, heterogeneous, sparse, and periodic traffic patterns. RL-ASL RL-ASL-LB Orch. Orch.-LB 0 1 Power [mW] 0.7 0.9 1.0 1.5 CPU TX RX non-UC RX UC active RX UC idle (a) RL-ASL RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 1.5 Power [mW] 0.6 0.8 1.0 1.4 (b) RL-ASL RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 1.5 Power [mW] 0.6 0.8 1.0 1.4 (c) RL-ASL PRIL-M RL-ASL-LB Orch. Orch.-LB 0.0 0.5 1.0 1.5 Power [mW] 0.6 0.8 0.9 1.0 1.5 (d) Fig. 10. Power consumption across different traffic patterns for the Star topology: (a-d) Power consumption under high, heterogeneous, sparse, and periodic traffic patterns. 0.751.001.251.50 Power [mW] 10 0 10 1 Latency [s] RL-ASL RL-ASL-LB Orch. Orch.-LB 80 85 90 95 100 PDR [%] (a) 0.751.001.25 Power [mW] 10 0 10 1 Latency [s] 80 85 90 95 100 PDR [%] (b) 0.751.001.25 Power [mW] 10 0 10 1 Latency [s] 80 85 90 95 100 PDR [%] (c) 0.751.001.25 Power [mW] 10 0 10 1 Latency [s] RL-ASL RL-ASL-LB Orch. Orch.-LB PRIL-M 80 85 90 95 100 PDR [%] (d) Fig. 11. Trade-off comparison across different traffic patterns for the Star topology: (a-d) Overall performance under high, heterogeneous, sparse, and periodic traffic patterns. The metrics are normalized for comparison. C. Practical Energy Impact To assess the practical benefit of RL-ASL, we estimate the expected battery Lifetime (LT) from the measured average power consumption. Assuming a 3 V supply and a 220 mAh coin-cell battery (E batt ≈ 2.38 kJ), the expected LT in days is: L = E batt P avg × 86400 ,where E batt = 3× 220× 3.6 = 2376 J. Table IV summarizes the measured average power and the corresponding LT for both network topologies. In the simple topology, RL-ASL consumes only 0.56 mW, extending LT to about 174 days—a 77% gain over Orchestra and 180% over Orchestra-LB. Even the more responsive RL-ASL-LB maintains 126 days, still outperforming static scheduling. PRIL-M achieves the highest efficiency under strictly periodic traffic (0.53 mW, 184 days), while RL-ASL generalizes better to non-periodic workloads. In the star topology, RL-ASL (0.65 mW) achieves an estimated 150 days, improving LT by roughly 55% relative to Orchestra (1.01 mW, 97 days) and more than 120% over Orchestra-LB (1.45 mW, 68 days). RL-ASL-LB (0.86 mW) sustains about 115 days, balancing adaptivity and energy use. For larger nodes powered by two A batteries (≈ 21.6 kJ), RL-ASL could extend LT from roughly 2.4 years (Orchestra) TABLE IV AVERAGE POWER AND ESTIMATED LIFETIME (3 V, 220 MAH BATTERY) ProtocolSimple [mW]LT [days]Star [mW]LT [days] Orchestra0.996991.00898 Orchestra-LB1.419701.45368 PRIL-M0.5291850.751131 RL-ASL0.5611740.647152 RL-ASL-LB0.7891240.856114 TABLE V IMPACT OF MODERATE MOBILITY ON RL-ASL PERFORMANCE ProtocolPDR [%]Radio Duty Cycle [%]Latency [ms] RL-ASL93.12.8186.97 RPL Baseline96.73.4184.70 to 4.2 years, with RL-ASL-LB sustaining around 3 years—a notable advantage for long-lived IIoT deployments. D. Impact of Moderate Mobility on RL-ASL To assess the behavior of RL-ASL under mobility, we conducted a deliberately simple RPL-based simulation using the Cooja network simulator. The scenario involves a single mobile node moving at a low speed (0.2 m/s) and alternately JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 202513 attaching to one of two potential parents in a tree topology. This experiment uses the heterogeneous traffic pattern de- scribed earlier and is intended as a controlled stress test rather than a comprehensive mobility evaluation. As expected, mobility induces repeated parent changes and receiver re-synchronization phases. During these intervals, nodes temporarily revert to standard TSCH operation, lim- iting the applicability of receiver-side listening optimization. Under these conditions, the RPL baseline achieves a PDR of approximately 96%, whereas RL-ASL attains 93%. Despite this reduction in delivery ratio, RL-ASL signif- icantly reduces the radio duty cycle from 3.4% to 2.8%, corresponding to an energy saving of approximately 18%. This result indicates that RL-ASL continues to reduce idle listening whenever short periods of topology stability are present, even under moderate mobility. E. Summary of Findings Across all experiments, RL-ASL demonstrates consistent superiority in energy efficiency while maintaining perfect reliability and competitive latency. RL-ASL-LB offers slightly lower delays due to its per-link scheduling structure, while RL-ASL provides the best overall energy-delay trade-off. Compared with PRIL-M, RL-ASL achieves comparable per- formance under periodic traffic but generalizes effectively to non-periodic and heterogeneous scenarios—an essential capability for real-world IoT deployments. A targeted mobility stress test further shows that, while frequent parent changes reduce delivery ratio, RL-ASL continues to achieve substantial idle-listening reductions whenever short periods of topology stability are present. Overall, these results validate the proposed RL-ASL frame- work as a robust, adaptive, and energy-efficient solution for dynamic TSCH networks, combining the reliability of deter- ministic scheduling with the flexibility of RL. VIII. CONCLUSION This work presented RL-ASL, a RL–driven adaptive lis- tening framework for TSCH networks that enhances energy efficiency without compromising reliability or latency. By integrating learning-based slot skipping into standard TSCH scheduling, RL-ASL enables nodes to adapt their listening behavior to traffic dynamics, achieving significant reductions in idle-listening power while preserving synchronization and delivery guarantees. Experimental results from the FIT IoT- LAB testbed and Cooja network simulator show that RL- ASL consistently outperforms state-of-the-art protocols such as Orchestra and PRIL-M across multiple topologies and traffic patterns. It reduces power consumption by up to 46%, main- tains near-perfect reliability, and lowers average latency by up to 96% compared to PRIL-M. The link-based variant, RL- ASL-LB, further improves delay under contention, confirming the scalability of the proposed learning framework. Model training is performed entirely in simulation, while inference on real motes is reduced to a simple table lookup with negligible computational and energy overhead. This makes RL-ASL practical for deployment in low-power embedded networks, bridging the gap between simulation-based learning and real- world operation. A targeted mobility stress test further indi- cates that, although frequent parent changes reduce delivery ratio, RL-ASL continues to deliver substantial idle-listening reductions whenever short periods of topology stability are present. In summary, RL-ASL demonstrates that RL can be effectively integrated into TSCH scheduling to deliver a robust, adaptive, and energy-aware communication framework for next-generation IoT networks. Future extensions of RL-ASL will explore tighter integration with mobility-aware routing and scheduling mechanisms, enabling coordinated adaptation of parent selection and receiver listening behavior to further enhance performance in mobile and dynamic environments. ACKNOWLEDGMENTS The authors acknowledge the use of AI-based tools for minor language refinements, such as grammar, structure, formatting, and spelling, during manuscript preparation. All intellectual contributions, technical content, and interpretations are solely those of the authors. REFERENCES [1] D. C. Nguyen, M. Ding, P. N. Pathirana, A. Seneviratne, J. Li, D. Niyato, O. Dobre, and H. V. Poor, “6G Internet of Things: A Comprehensive Survey,” IEEE Internet Things J., vol. 9, no. 1, p. 359–383, Jan. 2022. [2] M. Soori, B. Arezoo, and R. Dastres, “Internet of things for smart factories in industry 4.0, a review,” IOTCPS, vol. 3, p. 192–204, Jan. 2023. [3] F. F. Jurado-Lasso, L. Marchegiani, J. F. Jurado, A. M. Abu-Mahfouz, and X. Fafoutis, “A Survey on Machine Learning Software-Defined Wireless Sensor Networks (ML-SDWSNs): Current Status and Major Challenges,” IEEE Access, vol. 10, p. 23 560–23 592, Feb. 2022. [4] T. Watteyne, M. R. Palattella, and L. A. Grieco, “Using IEEE 802.15.4e Time-Slotted Channel Hopping (TSCH) in the Internet of Things (IoT): Problem Statement,” IETF, Tech. Rep. RFC 7554, May 2015. [5] A. R. Urke, Ø. Kure, and K. Øvsthus, “A Survey of 802.15.4 TSCH Schedulers for a Standardized Industrial Internet of Things,” Sensors, vol. 22, no. 15, p. 1–34, Dec. 2021. [6] Y. Zhang, H. Huang, Q. Huang, and Y. Han, “6TiSCH IIoT network: A review,” Comput. Netw., vol. 254, p. 110759, Dec. 2024. [7] G. Oikonomou, S. Duquennoy, A. Elsts, J. Eriksson, Y. Tanaka, and N. Tsiftes, “The Contiki-NG open source operating system for next generation IoT devices,” SoftwareX, vol. 18, p. 101089, Jun. 2022. [8] C. Adjih, E. Baccelli, E. Fleury, G. Harter, N. Mitton, T. Noel, R. Pissard-Gibollet, F. Saint-Marcel, G. Schreiner, J. Vandaele, and T. Watteyne, “FIT IoT-LAB: A large scale open experimental IoT testbed,” in IEEE WF-IOT 2015, Milan, Italy, Dec. 2015, p. 459–464. [9] S. Scanzio, F. Quarta, G. Paolini, G. Formis, and G. Cena, “Ultralow Power and Green TSCH-Based WSNs With Proactive Reduction of Idle Listening,” IEEE Internet Things J., vol. 11, no. 17, p. 29 076–29 088, Sep. 2024. [10] R. S. Rathore, O. Kaiwartya, K. N. Qureshi, I. T. Javed, W. Nagmeldin, A. Abdelmaboud, and N. Crespi, “Towards Enabling Fault Tolerance and Reliable Green Communications in Next-Generation Wireless Systems,” Appl. Sci., vol. 12, no. 17, p. 8870, Sep. 2022. [11] A. Tabouche, B. Djamaa, and M. R. Senouci, “Traffic-Aware Reliable Scheduling in TSCH Networks for Industry 4.0: A Systematic Mapping Review,” IEEE Commun. Surveys Tuts., vol. 25, no. 4, p. 2834–2861, Aug. 2023. [12] S. Duquennoy, B. Al Nahas, O. Landsiedel, and T. Watteyne, “Orchestra: Robust Mesh Networks Through Autonomously Scheduled TSCH,” in SenSys ‘15. Seoul, South Korea: ACM, Nov. 2015, p. 337–350. [13] S. Kim, H.-S. Kim, and C. Kim, “ALICE: Autonomous link-based cell scheduling for TSCH,” in IPSN ’19. Montreal Quebec Canada: ACM, Apr. 2019, p. 121–132. [14] M. Ojo, S. Giordano, G. Portaluri, D. Adami, and M. Pagano, “An energy efficient centralized scheduling scheme in TSCH networks,” in IEEE ICC Workshops. Paris, France: IEEE, 2017, p. 570–575. JOURNAL OF L A T E X CLASS FILES, VOL. 18, NO. 9, JUNE 202514 [15] G. Z. Papadopoulos, A. Mavromatis, X. Fafoutis, R. Piechocki, T. Try- fonas, and G. Oikonomou, “Guard Time Optimisation for Energy Effi- ciency in IEEE 802.15.4-2015 TSCH Links,” in SaSeIoT 2016, InterIoT 2016. Paris, France: Springer, Feb. 2017, p. 56–63. [16] S. Jeong, J. Paek, H.-S. Kim, and S. Bahk, “TESLA: Traffic-Aware Elastic Slotframe Adjustment in TSCH Networks,” IEEE Access, vol. 7, p. 130 468–130 483, Sep. 2019. [17] S. Jeong, H.-S. Kim, J. Paek, and S. Bahk, “OST: On-Demand TSCH Scheduling with Traffic-Awareness,” in IEEE INFOCOM. Toronto, ON, Canada: IEEE, Aug. 2020, p. 69–78. [18] S. Kim, H.-S. Kim, and C.-k. Kim, “A3: Adaptive Autonomous Alloca- tion of TSCH Slots,” in IPSN ’21.New York, NY, USA: ACM, May 2021, p. 299–314. [19] O. Tavallaie, J. Taheri, and A. Y. Zomaya, “Design and Optimization of Traffic-Aware TSCH Scheduling for Mobile 6TiSCH Networks,” in IoTDI ’21. New York, NY, USA: ACM, May 2021, p. 234–246. [20] Y. Ha and S.-H. Chung, “Traffic-Aware 6TiSCH Routing Method for IIoT Wireless Networks,” IEEE Internet Things J., vol. 9, no. 22, p. 22 709–22 722, Nov. 2022. [21] Y. H. Pratama and S. Chung, “RL-SF: Reinforcement Learning based Scheduling Function for Distributed TSCH Networks,” in IEEE ICEIEC. Beijing, China: IEEE, Jul. 2022, p. 5–8. [22] Y. H. Pratama, S.-H. Chung, and D. Z. Fawwaz, “Low-Latency and Q- Learning-Based Distributed Scheduling Function for Dynamic 6TiSCH Networks,” IEEE Access, vol. 12, p. 49 694–49 707, Apr. 2024. [23] F. F. Jurado-Lasso, M. Barzegaran, J. F. Jurado, and X. Fafoutis, “ELISE: A Reinforcement Learning Framework to Optimize the Slotframe Size of the TSCH Protocol in IoT Networks,” IEEE Syst. J., vol. 18, no. 2, p. 1068–1079, Jun. 2024. [24] F. F. Jurado-Lasso, C. Orfanidis, JF. Jurado, and X. Fafoutis, “HRL- TSCH: A Hierarchical Reinforcement Learning-Based TSCH Scheduler for IIoT,” IEEE Trans. Cogn. Commun. Netw., vol. 10, no. 6, p. 2102– 2118, Dec. 2024. [25] M. Nsabagwa, J. Muhumuza, R. Kasumba, J. S. Otim, and R. Akol, “Minimal Idle-Listen Centralized Scheduling in TSCH Wireless Sensor Networks,” in TSP. Athens, Greece: IEEE, Aug. 2018, p. 1–5. [26] S. Scanzio, G. Cena, A. Valenzano, and C. Zunino, “Energy Saving in TSCH Networks by Means of Proactive Reduction of Idle Listening,” in ADHOC-NOW. Bari, Italy: Springer, Oct. 2020, p. 131–144. [27] S. Scanzio, G. Cena, and A. Valenzano, “Enhanced Energy-Saving Mechanisms in TSCH Networks for the IIoT: The PRIL Approach,” IEEE Trans. Ind. Inform., vol. 19, no. 6, p. 7445–7455, Jun. 2023. [28] A. Kalita and M. Gurusamy, “On-the-Fly Autonomous Slot Allocation in 6TiSCH-Based Industrial IoT Networks,” IEEE Trans. Ind. Inform., vol. 20, no. 7, p. 9365–9374, Jul. 2024. [29] A. Kalita and M. Khatua, “Autonomous Allocation and Scheduling of Minimal Cell in 6TiSCH Network,” IEEE Internet Things J., vol. 8, no. 15, p. 12 242–12 250, Aug. 2021. [30] E. Baccelli, O. Hahm, M. G ̈ unes, M. W ̈ ahlisch, and T. C. Schmidt, “RIOT OS: Towards an OS for the Internet of Things,” in INFOCOM WKSHPS, Turin, Italy, Apr. 2013, p. 79–80. [31] C. Vallati, S. Brienza, G. Anastasi, and S. K. Das, “Improving Network Formation in 6TiSCH Networks,” IEEE Trans. Mob. Comput., vol. 18, no. 1, p. 98–110, Jan. 2019. [32] F. Osterlind, A. Dunkels, J. Eriksson, N. Finne, and T. Voigt, “Cross- Level Sensor Network Simulation with COOJA,” in IEEE LCN 2006, Tampa, FL, USA, Nov. 2006, p. 641–648. [33] A. Elsts, S. Kim, H.-S. Kim, and C. Kim, “An Empirical Survey of Autonomous Scheduling Methods for TSCH,” IEEE Access, vol. 8, p. 67 147–67 165, Mar. 2020. F. Fernando Jurado-Lasso (GS’18–M’21) received the Ph.D. degree in engineering and the M.Eng. degree in telecommunications engineering from The University of Melbourne, Melbourne, VIC, Aus- tralia, in 2020 and 2015, respectively, and the B.Eng. degree in electronics engineering from Universidad del Valle, Cali, Colombia, in 2012. His research focuses on intelligent networked embedded systems, machine learning for low-power and time-synchronized wireless networks, cross- layer scheduling and resource optimization, and In- ternet of Things (IoT) protocols and architectures. J. F. Jurado received the doctoral degree and M.Sc. degree in physics from Universidad del Valle, Cali, Colombia, in 2000 and 1986, respectively, and the B.Sc. degree in physics from Universidad de Nari ̃ no, Pasto, Colombia, in 1984. He is currently a Professor with the Department of Basic Sciences, Faculty of Engineering and Ad- ministration, Universidad Nacional de Colombia, Palmira, Colombia. His research interests include nanomaterials, magnetic and ionic materials, nano- electronics, embedded systems, and the Internet of Things (IoT). He is a Senior Researcher recognized by Minciencias, Colombia, and has been designated as an Emeritus Researcher by Minciencias.