Paper deep dive
MASK: Multi-Agent Semantic K-Scheduling for Risk-Sensitive 6G Robotics
Ahmet Gunhan Aydin, Elif Tugce Ceran
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/9/2026, 4:17:41 AM
Summary
The paper introduces MASK (Multi-Agent Semantic K-Scheduling), a control architecture for 6G-connected robotic swarms that addresses strict bandwidth constraints in collaborative sensing. MASK employs A-SIG (Arbiter-Assisted Semantic Information Gating), a lightweight mechanism where a Central Arbiter schedules only the top-K agents based on locally computed semantic importance scores. These prioritized observations are aggregated by a self-supervised global encoder into a compact latent state, which feeds a risk-sensitive distributional policy optimized for Conditional Value-at-Risk (CVaR) to mitigate tail risks. Evaluations on the Hallway Group and Multi-Agent Car Following benchmarks demonstrate that MASK matches communication-unconstrained baselines while maintaining robustness to packet erasures and resource limitations.
Entities (9)
Relation Signals (8)
MASK → evaluatedon → Hallway Group Task
confidence 96% · We evaluate MASK across diverse benchmarks... The Hallway Group task [7, 23] is a partially observable navigation problem...
MASK → evaluatedon → Multi-Agent Car Following Task
confidence 96% · The Multi-Agent Car Following (MACF) task [16] extends a risk-sensitive driving environment... We evaluate MASK on a suite of benchmarks...
MASK → uses → A-SIG
confidence 95% · MASK, a unified architecture combining A-SIG with a self-supervised global encoder and a risk-sensitive distributional policy
MASK → optimizes → CVaR
confidence 94% · We adopt the Distortion Risk Measure (DRM) framework, and in particular optimize the Conditional Value-at-Risk [15] to ensure robustness to adverse tail outcomes.
A-SIG → employs → Central Arbiter
confidence 93% · A-SIG employs a feedback-driven architecture that decouples low-bandwidth control signaling from high-bandwidth data transmission. The gating process operates in three distinct phases... Global Arbitration at the Central Arbiter
6G Networks → imposes → bandwidth constraints
confidence 93% · In realistic collaborative sensing scenarios, spectral resources are quantized into finite physical resource blocks or orthogonal subcarriers, rendering simultaneous transmission by all agents infeasible.
Central Arbiter → enforces → top-K scheduling
confidence 92% · The Central Arbiter aggregates the scores and performs a global ranking operation to identify the subset of agents with the most critical information. A binary transmission mask... is generated based on a system-wide channel access constraint K
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Realizing the vision of 6G connected robotics requires reconciling high-performance collaborative control with the rigid spectral limitations of physical wireless channels. In realistic collaborative sensing scenarios, spectral resources are quantized into finite physical resource blocks or orthogonal subcarriers, rendering simultaneous transmission by all agents infeasible. To address this, we propose Multi-Agent Semantic K-Scheduling (MASK), a control architecture designed to sustain robust, risk-aware coordination under strict instantaneous bandwidth caps. We introduce Arbiter-Assisted Semantic Information Gating (A-SIG), a lightweight coordination mechanism that enforces hard access constraints by scheduling only the top-K agents based on locally computed semantic importance scores. By aggregating these prioritized observations into a compact latent state, a self-supervised global encoder enables a distributional policy to mitigate tail risks despite data sparsity. We evaluate MASK across diverse benchmarks, demonstrating that it matches the performance of communication-unconstrained baselines even when channel access is restricted to a small fraction of the swarm size. Furthermore, the framework exhibits inherent resilience to packet erasures, validating semantic scheduling as a critical enabler for resource-constrained 6G systems.
Tags
Links
- Source: https://arxiv.org/abs/2606.11249v1
- Canonical: https://arxiv.org/abs/2606.11249v1
Trouble viewing inline? Open PDF directly →
Full Text
36,402 characters extracted from source content.
Expand or collapse full text
MASK: Multi-Agent Semantic K-Scheduling for Risk-Sensitive 6G Robotics Ahmet Günhan Aydın, Elif Tugce Ceran Authors are with the Department of Electrical and Electronics Engineering, Middle East Technical University, Ankara, 06800, Turkey, e-mail: gunhan.aydin@metu.edu.tr, elifce@metu.edu.tr. A. G. Aydın is also with Aselsan Inc., Ankara, Turkey, e-mail: gunaydin@aselsan.com. Abstract Realizing the vision of 6G connected robotics requires reconciling high-performance collaborative control with the rigid spectral limitations of physical wireless channels. In realistic collaborative sensing scenarios, spectral resources are quantized into finite physical resource blocks or orthogonal subcarriers, rendering simultaneous transmission by all agents infeasible. To address this, we propose Multi-Agent Semantic K-Scheduling (MASK), a control architecture designed to sustain robust, risk-aware coordination under strict instantaneous bandwidth caps. We introduce Arbiter-Assisted Semantic Information Gating (A-SIG), a lightweight coordination mechanism that enforces hard access constraints by scheduling only the top-K agents based on locally computed semantic importance scores. By aggregating these prioritized observations into a compact latent state, a self-supervised global encoder enables a distributional policy to mitigate tail risks despite data sparsity. We evaluate MASK across diverse benchmarks, demonstrating that it matches the performance of communication-unconstrained baselines even when channel access is restricted to a small fraction of the swarm size. Furthermore, the framework exhibits inherent resilience to packet erasures, validating semantic scheduling as a critical enabler for resource-constrained 6G systems. I Introduction The deployment of autonomous robotic swarms in next-generation 6G networks necessitates a paradigm shift in Multi-Agent Reinforcement Learning (MARL) [9]. Whether in non-terrestrial networks (NTN) [10, 12] or industrial IoT, the fundamental challenge is coordinating decentralized agents under strict physical layer constraints. While agents must exchange local observations to resolve partial observability [13], the underlying 6G infrastructure cannot physically support simultaneous broadcasting by all nodes. In realistic edge scenarios, time-frequency resources are quantized into limited physical resource blocks (PRBs) or orthogonal subcarriers, imposing a hard instantaneous channel access constraint that limits the number of concurrent transmissions to at most K. Existing control frameworks, however, rarely account for these sharp physical limits while simultaneously ensuring safety in stochastic environments. Prior research has generally approached coordination through two disjoint lenses. The first focuses on information aggregation. Architectures like CommNet [19], TarMAC [5], and MASIA [7] aim to reconstruct global states but rely on communication-intensive protocols that assume ideal, high-capacity channels. Furthermore, these methods largely optimize for risk-neutral expected returns, ignoring the catastrophic tail risks inherent in dynamic physical environments. The second lens focuses on risk-sensitivity. To operate safely in unpredictable domains like autonomous driving, agents must account for rare, high-cost tail events. Distributional RL [1, 4] and frameworks like RiskQ [16] address this by explicitly modeling the full return distribution rather than optimizing a simple mean. These algorithms further apply risk functionals to the return distributions, such as Conditional Value-at-Risk (CVaR) [15], to emphasize adverse tail outcomes. However, these risk-aware policies are typically trained on purely local views. This creates a dangerous blind spot: an agent may behave conservatively based on its own sensor data while remaining oblivious to a critical threat visible only to a distant teammate. To bridge this disconnect, we propose a joint control-communication architecture designed explicitly for resource-constrained 6G interfaces. We argue that agents must learn not only how to act safely but which data is semantically essential to transmit when channel access is competitively arbitrated. This reflects the 6G vision of semantic communication, optimizing data exchange for task utility rather than throughput [8, 22]. We introduce Multi-Agent Semantic K-Scheduling (MASK), a framework that replaces unstructured broadcasting with Arbiter-Assisted Semantic Information Gating (A-SIG). In this scheme, a Central Arbiter strictly enforces physical channel access by granting transmission rights to the top-K agents with the highest semantic importance scores, ensuring robust global coordination without violating physical channel constraints. Our main contributions are: • We propose Arbiter-Assisted Semantic Information Gating (A-SIG), a differentiable scheduling mechanism in which a Central Arbiter enforces hard physical channel access constraints (top-K) by ranking locally computed semantic importance scores. • We introduce MASK, a unified architecture combining A-SIG with a self-supervised global encoder and a risk-sensitive distributional policy, optimizing joint safety and performance under strict channel access constraints. • We empirically demonstrate that MASK matches the performance of unconstrained baselines even when channel access is restricted to a small fraction of the swarm size. Furthermore, our results validate that the framework maintains robust performance under random packet erasures, making it suitable for realistic, unreliable 6G robotic networks. I Related Work Coordination under partial observability is a cornerstone of networked robotics. Early differentiable communication works like CommNet [19] and DIAL [6] allowed gradient propagation across channels, while attentional models (ATOC [11], TarMAC [5]) improved scalability. Recent transformer-based approaches (e.g., MACTAS [25]) further enhance state aggregation but rely on continuous, high-bandwidth exchange. For 6G edge applications, however, bandwidth is limited. Approaches like IC3Net [17] and NDQ [23] attempt to gate communication, while Tung et al. [20] address noisy channels. While MASIA [7] effectively constructs global representations, it remains communication-intensive. In contrast, our A-SIG mechanism embraces the semantic communication philosophy: we do not aim to reconstruct the full bit-stream of observations, but rather to transmit only the semantically important observations necessary for the joint risk-sensitive control task. While [18] selects specific semantic features per agent, MASK schedules the top-K agents for risk-sensitive control. Standard MARL optimizes expected returns, ignoring the variance inherent in dynamic environments [14]. Distributional RL [1] overcomes this by modeling the return distribution, enabling risk measures like CVaR [15]. In multi-agent settings, RiskQ [16] applies distributional factorization to manage collective risk. However, RiskQ lacks a communication mechanism, forcing agents to make risk assessments based solely on local views. MASK unifies these domains, using A-SIG to resolve information uncertainty, thereby providing the global context necessary for valid risk-sensitive decision-making. I Methodology: Multi-Agent Semantic K-Scheduling (MASK) We formulate the cooperative multi-agent task as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), defined by the tuple =⟨,,P,R,Ω,O,N,γ,ℳ⟩G= ,A,P,R, ,O,N,γ,M . At each timestep t, each agent i receives a local observation oit∼O(s,i)o_i^t O(s,i) from the observation space Ω along with a message mit∈ℳm_i^t from other agents. Based on these inputs, the agent selects an environmental action ait∈a_i^t and a binary communication action bit∈0,1b_i^t∈\0,1\. Here, bit=1b_i^t=1 denotes transmitting the observation, while bit=0b_i^t=0 denotes silence. The joint action AtA^t drives the state transition st+1∼P(st,At)s^t+1 P(s^t,A^t) and reward rt=R(st,At)r^t=R(s^t,A^t). Unlike prior works that optimize the expected return [∑γtrt]E[Σγ^tr^t] under idealized communication, we address a setting where channel access is explicitly constrained and agents must be robust to outcome uncertainty. Our objective is to maximize a risk-sensitive functional ρ of each agent’s return distribution ZiZ_i, defined over the discounted return ∑k=0Tγkrt+k _k=0^Tγ^kr^t+k. We adopt the Distortion Risk Measure (DRM) framework [24], and in particular optimize the Conditional Value-at-Risk [15] to ensure robustness to adverse tail outcomes. The resulting objective is ρCVaR(Zi)=1α∫0αFZi−1(ω)ω, _CVaR(Z_i)= 1α _0^αF_Z_i^-1(ω)\,dω, (1) where α∈(0,1]α∈(0,1] is the risk tolerance and FZi−1F_Z_i^-1 is the quantile function. To solve this under channel access constraint, we propose MASK, a framework following the Centralized Training with Decentralized Execution (CTDE) paradigm (see Fig. 1). The architecture unifies three key modules: (1) an A-SIG module that computes local importance scores to arbitrate channel access via a Central Arbiter; (2) a self-supervised per-agent global encoder that aggregates the sparse, gated observations into a latent global state ztz^t; (3) a Distributional Risk-Sensitive Policy that combines ztz^t with local views to estimate the return distribution ZiZ_i. We classify this execution model as decentralized because the Central Arbiter’s role is limited to low-bandwidth control signaling (scheduling), ensuring that relatively high-bandwidth semantic data transmission remains a peer-to-peer, agent-driven process. During centralized training, we minimize a joint loss Ltotal=LRL+λreprLreprL_total=L_RL+ _reprL_repr that simultaneously optimizes the distributional value factorization, the global state reconstruction, while providing communication efficiency through the A-SIG mechanism. During both training and execution, the A-SIG module enables a Central Arbiter to enforce strict channel access caps by granting transmission rights only to agents with the highest importance scores, ensuring that the latent representation ztz^t is constructed from the most semantically relevant information. Figure 1: The MASK agent and the system architecture. Agents employ an A-SIG module to selectively broadcast critical observations with the assistance of a Central Arbiter. For each agent i, the received messages mim_i are processed by a global encoder to reconstruct a latent global state z, which is combined with the encoded local observation oio_i to produce a risk-sensitive action aia_i. Solid arrows denote local computation; dashed arrows represent communication flows. Red dashed lines indicate the gating feedback (transmission decision) from the Central Arbiter. Element-wise gating filters the observations to regulate bandwidth usage during both centralized training and decentralized execution. I-A Arbiter-Assisted Semantic Information Gating (A-SIG) To address communication resource constraints while maintaining global coordination, we introduce the Arbiter-Assisted Semantic Information Gating module. Unlike fully decentralized gating mechanisms where agents make isolated transmission decisions, A-SIG employs a feedback-driven architecture that decouples low-bandwidth control signaling from high-bandwidth data transmission. The gating process operates in three distinct phases: 1. Local Importance Quantification: Each agent i is equipped with a lightweight communication scorer network, denoted by CωC_ω, parameterized as a 2-layer MLP. At each timestep t, the agent processes its local observation oito_i^t to generate a scalar importance score cit=Cω(oit)∈ℝc_i^t=C_ω(o_i^t) . This score represents the agent’s estimated utility of broadcasting its current state to the team. 2. Global Arbitration at the Central Arbiter: The importance scores t=c1t,…,cNtc^t=\c_1^t,…,c_N^t\ are transmitted to the Central Arbiter over a low-rate control channel. Compared to full observation vectors, this scalar transmission incurs negligible bandwidth overhead, despite scaling with swarm size. The Central Arbiter aggregates the scores and performs a global ranking operation to identify the subset of agents with the most critical information. A binary transmission mask t∈0,1Nb^t∈\0,1\^N is generated based on a system-wide channel access constraint K (e.g., selecting the top-K highest scores) via the indicator function bit=(cit≥RankK(t))b_i^t=I(c_i^t _K(c^t)), where bit=1b_i^t=1 indicates permission to transmit. 3. Gated Data Transmission: The binary decision bitb_i^t is sent back to agent i as immediate feedback over the control channel. Agents with bit=1b_i^t=1 broadcast their observations to the multi-agent network, while others remain silent, significantly reducing channel congestion. Each agent i receives the sparse message set mit=bjt⋅ojtj≠im_i^t=\b_j^t· o_j^t\_j≠ i. To enable end-to-end training of the scorer network CωC_ω despite the non-differentiable nature of the ranking and thresholding operation, we employ a Straight-Through Estimator (STE) [2]. During the forward pass, the hard binary mask tb^t is applied strictly to gate the information flow. However, during the backward pass, the STE allows gradients to bypass the discrete thresholding, treating the decision as a continuous function of the scores. This proxy gradient allows the scorer network to learn to assign higher scalar values to observations that yield lower global loss, effectively aligning local importance estimation with the team’s cooperative objective. I-B Self-Supervised Selective State Aggregation The global encoder EϕE_φ processes the sparsely filtered observations Ofilt=bit⋅oiti=1NO^t_filt=\b_i^t· o_i^t\_i=1^N (where non-transmitting agents contribute zero) to produce a compact latent global state zt=Eϕ(Ofilt)z^t=E_φ(O^t_filt). This process is consistent across CTDE, where agents reconstruct ztz^t from received observations OfilttO^t_filt. EϕE_φ employs a self-attention mechanism to capture agent interactions. The filtered observations are projected into query, key, and value spaces via Q,K,V=MLPQ,K,V(Ofilt)\Q,K,V\=MLP_\Q,K,V\(O^t_filt) to compute the hidden state matrix M=softmax(QK⊤/dk)VM=softmax(QK / d_k)V, where dkd_k is the key dimension. The aggregated latent vector ztz^t is derived from M via an integration network. To ensure ztz^t is globally informative and temporally predictive, we optimize a joint self-supervised objective Lrepr=LAE+λpredLpredL_repr=L_AE+ _predL_pred. The reconstruction loss LAEL_AE trains a decoder DζD_ζ to recover sts^t, while the multi-step prediction loss LpredL_pred trains a transition model LξL_ξ to predict future latents over a horizon H: LAE L_AE =[‖Dζ(zt)−st‖2], =E [\|D_ζ(z^t)-s^t\|^2 ], (2) Lpred L_pred =[∑k=1H‖Lξ(z^t+k−1,At+k−1)−zt+k‖2]. =E [ _k=1^H\|L_ξ( z^t+k-1,A^t+k-1)-z^t+k\|^2 ]. (3) This objective ensures the encoder captures essential global dynamics despite sparse inputs. I-C Risk-Sensitive Policy Head To address outcome uncertainty, the agent integrates its local embedding fo(oit)f_o(o_i^t) with the global context ztz^t (modulated by a learned weighting vector gitg_i^t) to form the policy input xit=concat(fo(oit),git⊙zt,idi)x_i^t=concat(f_o(o_i^t),\,g_i^t z^t,\,id_i). This input updates the recurrent hidden state hith_i^t. In practice, the full action–observation history τi _i is summarized by the recurrent hidden state hith_i^t. We adopt the Implicit Quantile Network (IQN) framework [3]. To approximate the return distribution, quantile fractions ω∼U(0,1)ω U(0,1) are mapped to an embedding space via ϕ(ω)φ(ω) (e.g., cosine embedding) and combined with hith_i^t to generate the per-action quantiles θi(τi,ai,ω) _i( _i,a_i,ω). During centralized training, a monotonic distributional mixer aggregates these into a joint distribution θtot _tot using weights ki≥0k_i≥ 0 to ensure decomposability: θtot(τ,a,ω)=∑i=1Nki(τ,s,ω)θi(τi,ai,ω). _tot(τ,a,ω)= _i=1^Nk_i(τ,s,ω)\, _i( _i,a_i,ω). (4) The network is trained via quantile regression minimizing the quantile Huber loss ℒκL_κ against a target yty^t constructed using a risk-sensitive Bellman operator that selects actions maximizing the risk measure ρ on the next-state distribution Z(τt+1,a′)Z(τ^t+1,a ). Here, Z implicitly represents the joint distribution of all agents (omitting indices for clarity), and the prime notation (Z′Z , β′β ) denotes the usage of the target network. The objective is defined as: LRL L_RL =[ℒκ(yt−θtot)], =E[L_κ(y^t- _tot)], (5) yt y^t =rt+γθtot(τt+1,argmaxa′ρ(Z′),ω;β′). =r^t+γ _tot(τ^t+1, _a ρ(Z ),ω;β ). (6) The final joint objective updates all components: Ltotal=LRL+λreprLreprL_total=L_RL+ _reprL_repr. The centralized training procedure is summarized in Algorithm 1. During decentralized execution, each agent independently reconstructs the latent state from received messages using its local copy of the global encoder. Algorithm 1 MASK Training Algorithm 1: Initialize agent networks Qii=1N\Q_i\_i=1^N with encoder EϕE_φ and A-SIG CωC_ω, latent model LξL_ξ, mixer QmixQ_mix 2: Initialize target networks Qi′i=1N\Q _i\_i=1^N, Qmix′Q _mix, replay buffer D, and max episodes M 3: for episode =1=1 to M do 4: Reset environment and observe global state s1s^1 and observations O1=o11,…,oN1O^1=\o_1^1,…,o_N^1\ 5: Initialize trajectory buffer τ←∅τ← 6: for timestep t=1t=1 to T do 7: Compute importance scores cit=Cω(oit),i=1…Nc_i^t=C_ω(o_i^t),i=1… N 8: Transmit scores tc^t to the Central Arbiter 9: Central Arbiter computes rank threshold: δt=RankK(t)δ^t=Rank_K(c^t) 10: Generate transmission mask bit=(cit≥δt),i=1…Nb_i^t=I(c_i^t≥δ^t),i=1… N 11: Filter observations Ofiltt=bit⋅oiti=1NO^t_filt=\b_i^t· o_i^t\_i=1^N 12: Compute latent state zt=Eϕ(Ofiltt)z^t=E_φ(O^t_filt) 13: for agent i=1i=1 to N do 14: Form input xit=[fo(oit),git⊙zt,idi]x_i^t=[f_o(o_i^t),\,g_i^t z^t,id_i] 15: Compute quantile distribution Zi(xit,⋅,a)Z_i(x_i^t,·,a) via QiQ_i 16: Compute risk-sensitive values Qi,ρ(a)=ρ(Zi(xit,⋅,a))Q_i,ρ(a)=ρ(Z_i(x_i^t,·,a)) 17: Select action ait=ϵ-greedy(Qi,ρ)a_i^t=ε-greedy(Q_i,ρ) 18: Execute At=aiti=1NA^t=\a_i^t\_i=1^N and observe rtr^t, Ot+1O^t+1, terminated 19: Append (st,Ot,At,rt,Ot+1,terminated)(s^t,O^t,A^t,r^t,O^t+1,terminated) to τ 20: Store τ in replay buffer D 21: if D contains sufficient samples then 22: Sample minibatch ℬB from D 23: Compute representation loss as Lrepr=LAE+λpredLpredL_repr=L_AE+ _predL_pred 24: Compute joint quantiles θtot _tot using Eq. (4) with online networks Qii=1N\Q_i\_i=1^N and QmixQ_mix 25: Compute target quantiles yky^k (with target networks Qi′i=1N\Q _i\_i=1^N, Qmix′Q _mix) and LRLL_RL using Eq. (6) 26: Minimize total loss Ltotal=LRL+λreprLreprL_total=L_RL+ _reprL_repr, update the network parameters 27: if episode % target_update_interval =0=0 then 28: Update Qi′i=1N\Q _i\_i=1^N and Qmix′Q _mix IV Experiments We evaluate MASK on a suite of benchmarks designed to assess performance under varying partial observability, coordination complexity, and outcome stochasticity. The Hallway Group task [7, 23] is a partially observable navigation problem emphasizing temporal coordination. Agents are divided into two groups that must reach a central goal at different, predefined times. This constraint introduces strong coupling, requiring global coordination to resolve symmetry and avoid congestion. This setup proxies 6G-enabled warehouse logistics, where fleets of autonomous robots must coordinate passage through shared bottlenecks. The requirement to stagger arrival times mirrors time-sensitive networking in 6G, where precise scheduling is critical to prevent both physical collisions and spectrum congestion in dense swarms. The Multi-Agent Car Following (MACF) task [16] extends a risk-sensitive driving environment [21] to a cooperative setting where two agents control their acceleration to platoon toward a shared goal. Under partial observability, agents receive positive rewards for maintaining formation within a limited sensing range, while simultaneously accruing negative step rewards that incentivize speed. The environment explicitly evaluates risk-sensitive control via stochastic crash dynamics: exceeding specific speed thresholds incurs probabilistic collisions, forcing agents to balance progress efficiency against catastrophic failure. This formulation effectively mimics safety-critical V2V platooning, where connected autonomous vehicles must optimize traffic flow while strictly adhering to safety margins under mechanical or sensing uncertainties. To isolate the contributions of semantic gating and risk sensitivity, we compare MASK against baselines representing distinct combinations of information handling and objective formulations. QMIX [14] serves as the foundational risk-neutral baseline, employing monotonic value factorization on local observations to optimize expected return. RiskQ [16] extends QMIX to distributional RL, enabling risk-sensitive policies via quantile regression, but remains limited to local observations. MASIA [7] addresses partial observability by aggregating a global state from all agent observations to inform local policies, yet remains risk-neutral. MASK (Ours) unifies these approaches, integrating self-supervised global state aggregation with a risk-sensitive distributional policy, while uniquely employing A-SIG to optimize communication efficiency against hard channel access constraints. IV-A Simulation Results We benchmark MASK against the baselines to evaluate its efficacy across the selected tasks. The results empirically validate that jointly addressing information and outcome uncertainty under strict communication constraints allows for robust coordination, demonstrating the core advantages of our architecture. All learning curves report the mean and standard deviation over at least three independent runs. Models are evaluated every 10k training steps, averaging results over 10 episodes for MACF and 100 episodes for Hallway Group. Figure 2: Test group win rates over training timesteps in the Hallway Group environment with N=5N=5 agents divided into two groups. For MASK, the communication parameter is set to K=4(0.8N)K=4\;(0.8N) to ensure robust and fast convergence in this high-coordination setting. We first consider the Hallway Group environment, for which the results are shown in Figure 2, a setting that poses a coordination challenge under partial observability. Agents must precisely time their actions to pass through a narrow bottleneck, making effective coordination essential. Local-observation baselines, such as QMIX and RiskQ, fail to establish stable coordination and achieve group win rates between 0.2 and 0.6. In contrast, MASK quickly converges to the optimal group win rate of 2.0, indicating that both groups consistently reach the goal. This performance matches that of the full-communication MASIA baseline, yet MASK achieves it while enforcing a hard communication constraint of K=4K=4 transmitting agents. These results demonstrate that the proposed A-SIG mechanism successfully prioritizes task-relevant information, enabling accurate global state reconstruction. Figure 3: Test returns over training timesteps on MACF with N=2N=2 agents. The communication parameter is set to K=1(0.5N)K=1\;(0.5N). Figure 3 presents the results from the MACF environment, a scenario designed to test safety and risk-aware control in the presence of stochastic collision dynamics. In this domain, risk-neutral baselines like MASIA and QMIX display instability and poor performance, as they overlook the tail risks inherent in unsafe maneuvers. Conversely, MASK demonstrates consistent stability and high returns throughout the training process. Notably, even with the channel access capped at K=1K=1, the results indicate that permitting a single agent to transmit per timestep is sufficient to effectively mitigate risk. This confirms that risk-aware decision-making can be sustained with minimal but strategic communication, making the approach highly applicable to practical, bandwidth-limited robotic networks. To evaluate robustness under realistic 6G edge conditions, we test MASK in the Hallway Group environment under a random erasure channel where each agent’s importance score reaches the Central Arbiter with probability 0.70.7, and is otherwise received as zero. Since importance scores cit∈ℝc_i^t are unbounded, a zero-filled erasure may inadvertently be ranked higher than a valid negative score. We apply Top-K scheduling (K=4K=4) based on these received values; however, agents with erased scores are physically unable to transmit observations even if selected by the Central Arbiter. Despite this stochastic mismatch between scheduling and transmission, MASK attains an average group win rate of ≈1.9≈ 1.9 (vs. 2.02.0 for perfect channels) while maintaining an effective post-erasure communication rate of ≈0.6≈ 0.6, as shown in Figure 2. This demonstrates that the global encoder reliably reconstructs the global state from stochastically incomplete inputs, supporting robust coordination under packet erasures typical of high-mobility and NTN 6G scenarios [10]. In summary, MASK with A-SIG demonstrates versatile performance across distinct challenges. It acts as a dynamic switch, matching state-of-the-art communication methods in highly partially observable tasks (Hallway Group) and setting new benchmarks for stability in safety-critical environments (MACF) where traditional risk-neutral baselines falter. IV-B Impact of Channel Access Constraint (K) Figure 4: Test group win rate as a function of the communication parameter K in the Hallway Group environment with N=5N=5 agents. For MASK and the random baseline, K denotes the number of transmitting agents. To analyze the trade-off between channel access and coordination capability, we evaluate agent performance under varying physical constraints, where the number of available channels K ranges from 0 (no communication) to N=5N=5 (full broadcast). Figure 4 illustrates the test group win rates averaged over the entire training process across different K values in the Hallway Group environment. We observe that performance improves as the number of available channels increases, with the most significant gains occurring between K=0K=0 and K=2K=2. Notably, our method reaches a performance plateau at K=2K=2, achieving a win rate comparable to the full-communication MASIA baseline. This demonstrates that MASK effectively filters semantically high-value information, achieving near-optimal coordination while utilizing only 2 channels (40% of the swarm size). To further verify that the A-SIG module effectively identifies task-critical information, we introduce MASK Rand-Baseline where K transmitting agents are selected uniformly at random. The performance gap between A-SIG and random baseline confirms that the learned importance scores successfully filter task-critical information rather than simply benefiting from increased channel access. V Conclusion This work presented MASK, a unified framework for 6G-connected robotics that jointly addresses partial observability, safety-critical decision-making, and strict channel access constraints. By synergizing Arbiter-Assisted Semantic Information Gating and a self-supervised global state encoder within a risk-sensitive distributional reinforcement learning architecture, MASK transforms communication from a passive overhead into an active, optimizable resource. Our results demonstrate that agents can learn to transmit only task-critical observations, maintaining high-performance coordination and safety even when channel access is constrained over 50% or channels are unreliable. This capability is essential for deploying robust, risk-aware robotic swarms in realistic wireless environments where spectral efficiency and mission safety are paramount. Future work will focus on distributed scheduling, scaling the global encoder for massive swarms, and adapting the framework for highly dynamic, asymmetric 6G network topologies. References [1] M. G. Bellemare, W. Dabney, and R. Munos (2017) A distributional perspective on reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, p. 449–458. Cited by: §I, §I. [2] Y. Bengio, N. Léonard, and A. Courville (2013) Estimating or propagating gradients through stochastic neurons for conditional computation. External Links: 1308.3432, Link Cited by: §I-A. [3] W. Dabney, G. Ostrovski, D. Silver, and R. Munos (2018-10–15 Jul) Implicit quantile networks for distributional reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning, J. Dy and A. Krause (Eds.), Proceedings of Machine Learning Research, Vol. 80, p. 1096–1105. Cited by: §I-C. [4] W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos (2018) Distributional reinforcement learning with quantile regression. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’18/IAAI’18/EAAI’18. External Links: ISBN 978-1-57735-800-8 Cited by: §I. [5] A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, and J. Pineau (2020) TarMAC: targeted multi-agent communication. External Links: 1810.11187, Link Cited by: §I, §I. [6] J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson (2016) Learning to communicate with deep multi-agent reinforcement learning. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, Red Hook, NY, USA, p. 2145–2153. External Links: ISBN 9781510838819 Cited by: §I. [7] C. Guan, F. Chen, L. Yuan, C. Wang, H. Yin, Z. Zhang, and Y. Yu (2022) Efficient multi-agent communication via self-supervised information aggregation. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA. External Links: ISBN 9781713871088 Cited by: §I, §I, §IV, §IV. [8] D. Gündüz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K. K. Wong, and C. Chae (2023) Beyond transmitting bits: context, semantics, and task-oriented communications. IEEE Journal on Selected Areas in Communications 41 (1), p. 5–41. External Links: Document Cited by: §I. [9] P. Hernandez-Leal, B. Kartal, and M. E. Taylor (2019-10) A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems 33 (6), p. 750–797. External Links: ISSN 1573-7454, Document Cited by: §I. [10] Z. Ji, S. Wu, and C. Jiang (2023) Cooperative multi-agent deep reinforcement learning for computation offloading in digital twin satellite edge networks. IEEE Journal on Selected Areas in Communications 41 (11), p. 3414–3429. External Links: Document Cited by: §I, §IV-A. [11] J. Jiang and Z. Lu (2018) Learning attentional communication for multi-agent cooperation. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, Red Hook, NY, USA, p. 7265–7275. Cited by: §I. [12] Y. Lyu, H. Hu, R. Fan, Z. Liu, J. An, and S. Mao (2024) Dynamic routing for integrated satellite-terrestrial networks: a constrained multi-agent reinforcement learning approach. IEEE Journal on Selected Areas in Communications 42 (5), p. 1204–1218. External Links: Document Cited by: §I. [13] S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian (2017) Deep decentralized multi-task multi-agent reinforcement learning under partial observability. External Links: 1703.06182, Link Cited by: §I. [14] T. Rashid, M. Samvelyan, C. S. de Witt, G. Farquhar, J. Foerster, and S. Whiteson (2018) QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning. External Links: 1803.11485, Link Cited by: §I, §IV. [15] R. Rockafellar and S. Uryasev (2002) Conditional value-at-risk for general loss distributions. Journal of Banking & Finance 26 (7), p. 1443–1471. External Links: ISSN 0378-4266, Document Cited by: §I, §I, §I. [16] S. Shen, C. Ma, C. Li, W. Liu, Y. Fu, S. Mei, X. Liu, and C. Wang (2023) RiskQ: risk-sensitive multi-agent reinforcement learning value factorization. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §I, §I, §IV, §IV. [17] A. Singh, T. Jain, and S. Sukhbaatar (2018) Learning when to communicate at scale in multiagent cooperative and competitive tasks. External Links: 1812.09755, Link Cited by: §I. [18] Y. Su, Y. Du, and Y. Deng (2025) Goal-oriented semantic communication in bandwidth-constrained marl. In 2025 IEEE International Conference on Communications Workshops (ICC Workshops), Vol. , p. 1274–1279. External Links: Document Cited by: §I. [19] S. Sukhbaatar, A. Szlam, and R. Fergus (2016) Learning multiagent communication with backpropagation. External Links: 1605.07736, Link Cited by: §I, §I. [20] T. Tung, S. Kobus, J. P. Roig, and D. Gündüz (2021) Effective communications: a joint learning and communication framework for multi-agent reinforcement learning over noisy channels. IEEE Journal on Selected Areas in Communications 39 (8), p. 2590–2603. External Links: Document Cited by: §I. [21] N. A. Urpí, S. Curi, and A. Krause (2021) Risk-averse offline reinforcement learning. External Links: 2102.05371, Link Cited by: §IV. [22] E. Uysal, O. Kaya, A. Ephremides, J. Gross, M. Codreanu, P. Popovski, M. Assaad, G. Liva, A. Munari, B. Soret, T. Soleymani, and K. H. Johansson (2022) Semantic communications in networked systems: a data significance perspective. IEEE Network 36 (4), p. 233–240. External Links: Document Cited by: §I. [23] T. Wang, J. Wang, C. Zheng, and C. Zhang (2020) Learning nearly decomposable value functions via communication minimization. External Links: 1910.05366, Link Cited by: §I, §IV. [24] J. Wirch and M. Hardy (2003-02) Distortion risk measures. coherence and stochastic dominance. Insurance Mathematics and Economics 32, p. 168–168. Cited by: §I. [25] M. Wojtala, B. Stefańczyk, D. Bogucki, Ł. Lepak, J. Strykowski, and P. Wawrzyński (2025) MACTAS: self-attention-based module for inter-agent communication in multi-agent reinforcement learning. External Links: 2508.13661, Link Cited by: §I.