Paper deep dive
Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach
Tianyu Pang, Hongyu Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/16/2026, 3:38:10 AM
Summary
This paper addresses energy-aware offloading and resource allocation in heterogeneous mobile edge computing (MEC) systems assisted by active beyond-diagonal reconfigurable intelligent surfaces (BD-RIS). It tackles the challenge of cross-sector energy leakage caused by reciprocal active BD-RIS devices by formulating a high-dimensional mixed-integer nonconvex optimization problem. To solve this efficiently, the authors propose DSAC-T, a refined distributional soft actor-critic algorithm that models return distributions to enhance policy stability under reward heterogeneity. Simulations demonstrate that DSAC-T outperforms baseline algorithms in energy-latency reward, achieves an 81.67% feasibility ratio, and provides rapid online decision-making at 0.0267 seconds per scenario.
Entities (5)
Relation Signals (5)
DSAC-T → achieves → 81.67% feasibility ratio
confidence 98% · Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.67%
DSAC-T → solves → Mixed-integer nonconvex problem
confidence 97% · The resulting problem is a high-dimensional mixed integer nonconvex problem... To address this challenge, we develop an end-to-end joint optimization framework based on... DSAC-T
DSAC-T → improves → Policy Stability
confidence 96% · By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity
Active BD-RIS → enables → hybrid transmitting and reflecting mode
confidence 95% · Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage
Reciprocal active BD-RIS → generates → Cross-sector energy leakage
confidence 94% · practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems. However, practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff. This paper studies energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC, where offloading decisions, CPU/GPU computation allocation, transmit powers, receive processing, and active BD-RIS are tightly coupled. The resulting problem is a high-dimensional mixed integer nonconvex problem and is difficult to solve efficiently by conventional per-instance optimization. To address this challenge, we develop an end-to-end joint optimization framework based on a refined version of the distributional soft actor--critic algorithm, named as DSAC-T. By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity. Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.67%, and a fast online decision time of 0.0267 s per scenario.
Tags
Links
- Source: https://arxiv.org/abs/2607.13160v1
- Canonical: https://arxiv.org/abs/2607.13160v1
Trouble viewing inline? Open PDF directly →
Full Text
35,630 characters extracted from source content.
Expand or collapse full text
Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach Tianyu Pang1 and Hongyu Li2 Abstract Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems. However, practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff. This paper studies energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC, where offloading decisions, CPU/GPU computation allocation, transmit powers, receive processing, and active BD-RIS are tightly coupled. The resulting problem is a high-dimensional mixed integer nonconvex problem and is difficult to solve efficiently by conventional per-instance optimization. To address this challenge, we develop an end-to-end joint optimization framework based on a refined version of the distributional soft actor–critic algorithm, named as DSAC-T. By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity. Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.6781.67%, and a fast online decision time of 0.02670.0267 s per scenario. I Introduction Mobile edge computing (MEC) enables latency-sensitive and computation-intensive services by offloading demanding workloads to nearby edge servers [10]. In practical edge intelligence scenarios, task demands are heterogeneous, since CPU- and GPU-oriented jobs coexist and compete for different edge-side resources. Therefore, task offloading is inherently a communication–computation co-design problem. Wireless communications are very often interrupted by obstacles, such as high buildings, while reconfigurable intelligent surfaces (RISs) [7] have been proven to provide a flexible way to enhance wireless connectivity and thus MEC [3]. The RIS technology has grown rapidly with the emergence of many advanced architectures, which can be broadly categorized into passive RISs and active RISs based on if or not the scattered signals can be amplified. In the passive form, beyond-diagonal (BD) RISs [5] are a universal and generalized framework that includes conventional passive RISs with diagonal phase shift matrices [7] and simultaneous transmitting and reflecting (STAR) RISs [8] as special cases, while introducing inter-element connections to open the door for supporting more advanced hybrid and multi-sector modes with more flexible wave manipulation and wider coverage. These benefits have motivated the integration of BD-RIS and MEC to improve computing performance [12]. Despite this progress, passive RISs have fundamental performance limits due to the multiplicative fading. Hence, active RISs with the ability to amplify signals using reflection-type amplifiers have been proposed to compensate for the multiplicative fading and further enhance wireless channels [9, 15]. Inspired by this point, BD-RIS has been recently upgraded to the active form, with more rigorous physics-consistent model and diverse architecture design [14], and more advanced hybrid mode [6] that includes active STAR-RIS [15] as a special case. In the family of active BD-RIS, there have been studies on active STAR-RIS-assisted MEC [1, 13], while the active STAR-RIS models used do not fully expose the coupling between transmitting sector and reflecting sector due to the inherent circuit reciprocity. This essentially means when blocked users are assisted through active RIS across sectors, the system incurs active forwarding cost and cross-sector energy leakage due to the reciprocity, which compresses the energy budget available for uplink transmission and heterogeneous edge execution. Thus, simply enhancing blocked-user connectivity may improve communication reliability but degrade overall energy efficiency or aggravate latency violation. Moreover, the aforementioned RIS-assisted MEC works use generic edge-computing abstractions without explicitly distinguishing CPU/GPU-oriented tasks or decoupled CPU/GPU edge resource pools. Deeply blocked direct links are also often omitted or over-idealized, whereas they are more appropriately modeled as severely attenuated residual paths in practice. These limitations prevent existing formulations from fully capturing the real-world heterogeneous MEC systems. Motivated by these observations, we study a physics-consistent active BD-RIS-assisted heterogeneous MEC system. The contributions are summarized as follows. First, we establish an active BD-RIS-assisted heterogeneous MEC model that captures asymmetric blockage, residual direct-link attenuation, active RIS power consumption, amplified noise, and decoupled CPU/GPU edge resources. Specifically, the active BD-RIS works on the hybrid mode (similar to active STAR-RIS) to support full-space coverage and, more importantly, is modeled rigorously based on a reciprocal circuit implementation to capture cross-sector leakage [6]. Second, we formulate and solve a mixed integer, high-dimensional, and strongly coupled nonconvex optimization problem that jointly involves offloading decisions, computation allocation, uplink transmit powers, receive processing, and the active BD-RIS. Deep reinforcement learning (DRL) offers an efficient way to solve such problems [4, 11], while the formulated system remains challenging for conventional expectation-based actor–critic methods. Blockage patterns, task heterogeneity, active RIS cost, cross-sector leakage, and latency penalties jointly induce strong reward heterogeneity and feasibility-boundary sensitivity, leading to unstable value estimation and policy updates. To address this issue, we adopt a refined version of the distributional soft actor–critic (DSAC) algorithm that is suitable for solving problems with highly heterogeneous rewards, namely DSAC-T [2], and develop an end-to-end DSAC-T-based joint optimization framework. Third, simulation results validate the effectiveness and efficiency of the proposed DSAC-T-based framework. Over 300 evaluation scenarios, DSAC-T achieves the best energy-latency reward and the highest feasibility ratio of 81.67%81.67\%, outperforming other baselines. It also requires only 0.02670.0267 s per scenario for online decision making, which is orders of magnitude faster than the mathematical optimization algorithm. I System Model and Problem Formulation In this section, we present the active BD-RIS assisted heterogeneous MEC system, including the blockage-aware channel model, the uplink transmission process, the heterogeneous local/edge computation model, and the joint offloading and resource-allocation problem. Figure 1: Diagram of an active BD-RIS assisted heterogeneous MEC system. I-A Active BD-RIS Assisted Channel Model We consider an uplink heterogeneous MEC system consisting of one edge server equipped with an N-antenna base station (BS), one M-cell hybrid mode active BD-RIS [6], and K single-antenna users. Each cell of the hybrid mode active BD-RIS is constructed by two back to back placed antenna elements connected to a 2-port active reconfigurable impedance network, which further consists of two reflection-type amplifiers connected to a 4-port passive reconfigurable impedance network. In this sense, the hybrid mode active BD-RIS has two sectors, namely transmitting sector (sector 2) and reflecting sector (sector 1), each of which contains M elements and covers half of the space, as illustrated in Fig. 1. Therefore, the BS is located within the reflecting sector of the hybrid mode active BD-RIS. In addition, the user set is partitioned into two disjoint subsets, denoted by 1K_1 with |1|=K1|K_1|=K_1 and 2K_2 with |2|=K2|K_2|=K_2, K1+K2=K_1+K_2=K, corresponding to non-blocked users located within the reflecting sector and deeply blocked users within the transmitting sector of hybrid mode active BD-RIS, respectively. Each user is associated with either a CPU-oriented task or a GPU-oriented task. The edge server maintains separate CPU and GPU resource pools to support heterogeneous task execution. I-A1 Direct Channel The direct channel from user k to BS is modeled as d,k=βk~d,k∈ℂN×1h_d,k= _k h_d,k ^N× 1, where βk=1 _k=1 if k∈1k _1 and βk≪1 _k 1 if k∈2k _2. Hence, for users in 2K_2, the direct link is approximately zero but not strictly removed. I-A2 Hybrid Mode Active BD-RIS The M-cell hybrid mode active BD-RIS is characterized by its scattering matrix [6] =[1,11,22,12,2], = [ matrix _1,1& _1,2\\ _2,1& _2,2 matrix ], (1) where i,j∈ℂM×M _i,j ^M× M. In this work, we assume the hybrid mode active BD-RIS has a reciprocal cell-wise single-connected architecture, such that = = T and i,j _i,j, ∀i,j∀ i,j are diagonal, further implying that 1,2=2,1 _1,2= _2,1. I-A3 Active BD-RIS Assisted Channel Let ∈ℂN×MH ^N× M denote the RIS-BS channel and r,k∈ℂM×1h_r,k ^M× 1 denote the user-RIS channel of user k. The effective uplink channel of user k is then given by k=d,k+1,ir,k,∀k∈i,∀i∈1,2.h_k=h_d,k+H _1,ih_r,k,∀ k _i,∀ i∈\1,2\. (2) I-B Uplink Transmission Model Let pkp_k denote the uplink transmit power of user k and kw_k denote the combiner associated with user k. The uplink signal-to-interference-plus-noise ratio (SINR) can be expressed as γk=pk|kk|2∑j≠kpj|j|2+σB2‖k‖22+σI2‖k¯‖22, _k= p_k |w_k Hh_k |^2Σ _j≠ kp_j |w_ k Hh_j |^2+ _B^2\|w_k\|_2^2+ _I^2 \|w_k HH \|_2^2, (3) where ¯=[1,11,2] =[ _1,1~ _1,2], σB2 _B^2 denote the noise power at the BS and σI2 _I^2 denote the dynamic noise power of active BD-RIS. The achievable uplink rate of user k is thus Rk=Blog2(1+γk)R_k=B _2(1+ _k), where B is the system bandwidth. I-C Heterogeneous Computation Model The task of user k is characterized by the tuple (τk,Dk,Ck,Tkmax) ( _k,D_k,C_k,T_k ), where τk∈CPU,GPU _k∈\CPU,GPU\ denotes the task type, DkD_k is the input data size, CkC_k is the required CPU cycles per bit, and TkmaxT_k is the latency deadline. Let αk∈0,1 _k∈\0,1\ denote the offloading decision, where αk=0 _k=0 indicates local execution and αk=1 _k=1 indicates edge execution. The local execution latency and offloading latency that involves transmission and edge execution components are respectively Tkloc=DkCkfkloc,Tkoff=DkRk+DkCkfkedge,T_k^loc= D_kC_kf_k^loc,~T_k^off= D_kR_k+ D_kC_kf_k^edge, (4) where fklocf_k^loc is the local computing rate and fkedgef_k^edge is the edge computing rate allocated to user k. Accordingly, the end-to-end latency is constrained by Tk=(1−αk)Tkloc+αkTkoff≤Tkmax.T_k=(1- _k)T_k^loc+ _kT_k^off≤ T_k^max. (5) To capture hardware-specific execution, the edge computation budget is decoupled into CPU and GPU resource pools. Let FcpuedgeF_cpu^edge and FgpuedgeF_gpu^edge denote the total edge CPU and GPU capacities, respectively. Then, ∑k∈cpu,offfkedge≤Fcpuedge,∑k∈gpu,offfkedge≤Fgpuedge, _k ^cpu,offf_k^edge≤ F_cpu^edge,\; _k ^gpu,offf_k^edge≤ F_gpu^edge, (6) where cpu,offK^cpu,off and gpu,offK^gpu,off denote the sets of offloaded CPU-type and GPU-type tasks, respectively. I-D Energy Consumption Model The power radiated from active BD-RIS is PRIS=∑k=1Kpk‖¯r,k‖22+σI2‖2≤PA,P_RIS= _k=1^Kp_k \| h_r,k \|_2^2+ _I^2 \| \|_ F^2≤ P_A, (7) where ¯r,k=[r,k,] h_r,k=[h_r,k T,0 T] T for k∈1k _1 and ¯r,k=[,r,k] h_r,k=[0 T,h_r,k T] T for k∈2k _2, and PAP_A denotes the power budget at active BD-RIS. Then the active BD-RIS operating power is modeled as PRIStot=Pc+1ζPRISP_RIS^tot=P_c+ 1ζP_RIS, where PcP_c denotes the static circuit power and ζ∈(0,1]ζ∈(0,1] is the amplifier efficiency. Hence, the active BD-RIS energy consumption model is ERIS=PRIStotTtx,E_RIS=P_RIS^totT_tx, (8) where Ttx=maxkαkDkRkT_tx= _k~ _k D_kR_k. The local execution energy of user k is modeled and constrained as Ekloc=(1−αk)κkDkCk(fkloc)2≤Ekmax,E_k^loc=(1- _k) _kD_kC_k(f_k^loc)^2≤ E_k^max, (9) where κk _k is the effective switched-capacitance coefficient and EkmaxE_k^max is the local energy budget for user k. The offloading energy is Ekoff=αk(pkDkRk+κedgeDkCk(fkedge)2),E_k^off= _k (p_k D_kR_k+ _edgeD_kC_k(f_k^edge)^2 ), (10) where kedgek_edge is the edge computing-energy coefficient. I-E Joint Optimization Problem Define the joint optimization variable set as Ω≜,,,, \ α,p,f,W, \, where =[α1,…,αK]T α=[ _1,…, _K]^T collects offloading decisions, =[p1,…,pK]p=[p_1,…,p_K] T collects uplink transmit powers, f collects the local/edge computation allocation variables fkedge\f_k^edge\, and =[1,…,K]W=[w_1,…,w_K] denotes the combining matrix. Then the joint offloading and resource-allocation problem to minimize the weighted system energy is formulated as minΩ _ wloc∑k=1KEkloc+woff∑k=1KEkoff+wRISERIS w_loc _k=1^KE_k^loc+w_off _k=1^KE_k^off+w_RISE_RIS (11a) s.t. Tk≤Tkmax,∀k, T_k≤ T_k , ∀ k, (11b) Ekloc≤Ekmax,∀k, E_k^loc≤ E_k^max, ∀ k, (11c) αk∈0,1,∀k, _k∈\0,1\, ∀ k, (11d) pmin≤pk≤pmax,∀k, p_ ≤ p_k≤ p_ , ∀ k, (11e) ∑k∈cpu,offfkedge≤Fcpuedge, _k ^cpu,offf_k^edge≤ F_cpu^edge, (11f) ∑k∈gpu,offfkedge≤Fgpuedge, _k ^gpu,offf_k^edge≤ F_gpu^edge, (11g) PRIS≤PA, P_RIS≤ P_A, (11h) i,j,∀i,j∈1,2 are diagonal,1,2=2,1, _i,j,∀ i,j∈\1,2\~are~diagonal, _1,2= _2,1, (11i) where wlocw_loc, woffw_off, and wRISw_RIS denote the weighting factors of local, offloading, and active BD-RIS energies, respectively, and [pmin,pmax][p_min,p_max] constraints the transmit power range. Problem (11) is a strongly coupled mixed-integer optimization involving binary offloading, heterogeneous computation allocation, uplink power control, linear reception, and active BD-RIS configuration. Its nonconvexity and strong scenario dependence motivate the development of a learning-based slot-wise decision framework as will be detailed below. I DSAC-T-Based Optimization Framework To enable efficient online decision making for problem (11), we reformulate the active BD-RIS-assisted heterogeneous MEC process as a slot-wise learning problem. The original objective and constraints are incorporated through action projection and a penalty-based reward. Figure 2: DSAC-T-based optimization framework for active BD-RIS-assisted heterogeneous MEC. I-A MDP Reformulation The slot-wise decision process is modeled as ℳ=(,,,ℛ)M=(S,A,P,R), where S, A, P, and ℛR denote the state space, action space, state transition probability, and the reward function, respectively. At slot t, the agent receives a state st∈s_t and selects an action at∈a_t , and in return gains the next state st+1∈s_t+1 and a reward rt∈ℛr_t . Such a behavior of the agent is described as a policy π(at|st)π(a_t~|~s_t) that maps from each state in S to a probability distribution over actions in A. In the considered scenario, the state st∈s_t at slot t summarizes the task-side features k,tk=1K\x_k,t\_k=1^K, channel information k,tk=1K\g_k,t\_k=1^K, and local/edge resource conditions tq_t. The action at∈a_t contains the variables generated by the actor, including the offloading component, edge-computation allocation, uplink transmit powers, and the active BD-RIS control vector ϑt _t. The raw action is processed by an environment-side mapping module: the offloading component is converted into binary decisions, the transmit powers are rescaled to the feasible range of offloaded users, and the edge-computation allocation is normalized within the corresponding CPU/GPU resource pools. The RIS-control component ϑt _t is mapped to a diagonal and reciprocal t _t, followed by power projection if the active output-power constraint is violated. Given t _t, the receive combiner tW_t is obtained by minimum mean square error (MMSE) reception. Given (st,at)(s_t,a_t), the environment computes the communication/computation outcome and returns the next state by st+1∼(⋅∣st,at).s_t+1 (· s_t,a_t). (12) The reward is designed as a penalty-based surrogate of problem (11), where the weighted energy objective is maximized in its negative form and latency violations are penalized explicitly as rt=−(wlocE~loc,t+woffE~off,t+wRISE~RIS,t)−λTVT,t.r_t=- (w_ loc E_ loc,t+w_ off E_ off,t+w_ RIS E_ RIS,t )- _TV_T,t. Here, E~loc,t=∑kEk,tlocElocref+ε E_loc,t= _kE_k,t locE_loc^ref+ , E~off,t=∑kEk,toffEoffref+ε E_off,t= _kE_k,t offE_off^ref+ , E~RIS,t=ERIS,tERISref+ε E_RIS,t= E_RIS,tE_RIS^ref+ respectively denote the normalized slot-level local, offloading, and active BD-RIS energies, with predefined normalization terms ElocrefE_loc^ref, EoffrefE_off^ref, ERISrefE_RIS^ref, and ε . VT,tV_T,t denotes the average relative latency violation, i.e., VT,t=1K∑k=1K[Tk,tTkmax−1]+,V_T,t= 1K _k=1^K [ T_k,tT_k -1 ]_+, (13) with Tk,tT_k,t being the end-to-end latency at slot t, and λT _T denotes the penalty weight. To calculate the reward rtr_t, the digital combiner matrix tW_t embedded in Eloc,tE_loc,t, Eoff,tE_off,t, and ERIS,tE_RIS,t are directly optimized using typical MMSE receivers based on the effective channels as functions of ϑt _t. Hence, the reward directly penalizes weighted energy consumption and latency violation, while preserving the optimization intention of Section I. I-B Why DSAC-T is Suitable for the Considered Problem The considered system exhibits sharp feasibility boundaries because several discrete and continuous decisions are coupled. A small change in the offloading decision may switch a user between local and edge execution, while a small change in transmit power or RIS coefficients may alter SINR, transmission latency, RIS output power, and latency feasibility. Moreover, CPU/GPU task heterogeneity and blocked/non-blocked channel conditions lead to highly non-uniform reward scales across scenarios. Therefore, estimating only the expected return may be insufficient near feasibility boundaries, where overoptimistic value estimates can lead to unstable policy updates. DSAC-T addresses this issue by modeling the state-action return of policy πϕ(at|st) _φ(a_t~|~s_t) using twin Gaussian value distributions Zθi(st,at)∼(Qθi(st,at),σθi2(st,at)),Z_ _i(s_t,a_t) \! (Q_ _i(s_t,a_t), _ _i^2(s_t,a_t) ), (14) where ϕφ and θi _i, ∀i∈1,2∀ i∈\1,2\ are parameters. This allows the critic to learn both the mean return and return uncertainty. Moreover, DSAC-T uses the smaller target mean Qθmin(st,at)=mini∈1,2Qθi(st,at),Q_ _ (s_t,a_t)= _i∈\1,2\Q_ _i(s_t,a_t), (15) to construct and update the target. The twin-distribution design and variance-aware update thus reduce overestimation bias and improve robustness under heterogeneous energy-latency rewards, making DSAC-T suitable for problem (11). I-C DSAC-T Based Learning Algorithm Based on the MDP reformulation, we adopt a DSAC-T based learning algorithm including the following two steps. Step 1: Sampling. At each slot t, the agent observes the current state sts_t, samples an action at∼πϕ(⋅∣st)a_t _φ(· s_t), receives the reward rtr_t and the next state st+1s_t+1 from the environment, and stores the transition in the replay buffer: ←∪(st,at,rt,st+1).D ∪\(s_t,a_t,r_t,s_t+1)\. (16) Step 2: Updating. Given a transition (st,at,rt,st+1)(s_t,a_t,r_t,s_t+1) sampled from the replay buffer D, the target action is sampled by at+1∼πϕ′(⋅∣st+1)a_t+1 _φ (· s_t+1), where ϕ′φ denotes a parameter for target networks. Define the target self-consistency operator as ϕ′,i=rt+γ(Zi(st+1,at+1)−αlogπϕ′(at+1∣st+1)),T_φ ,i=r_t+γ (Z_i(s_t+1,a_t+1)-α _φ (a_t+1 s_t+1) ), (17) where γ∈(0,1)γ∈(0,1) is a discount factor and α is the temperature coefficient. Then the twin value distributions θ1 _1 and θ2 _2 are updated by minimizing ℒ=ωiF(θi′(⋅|st,at),θi(⋅|st,at)),L_Z= _iE\F(Y_θ _i(·~|~s_t,a_t),Z_ _i(·~|~s_t,a_t))\, (18) where ωi=σθi2(st,at) _i=E\ _ _i^2(s_t,a_t)\ is a gradient scaling weight, F(⋅,⋅)F(·,·) denotes a distance measurement function, θi′(⋅|st,at)Y_θ _i(·~|~s_t,a_t) denotes the distribution of ϕ′,iT_φ ,i with θ1′ _1 and θ2′ _2 being parameters for target networks related to the distributions of Z1(st+1,at+1)Z_1(s_t+1,a_t+1) and Z2(st+1,at+1)Z_2(s_t+1,a_t+1). The actor ϕφ is updated by minimizing ℒπ=αlogπϕ(at∣st)−Qθmin(st,at).L_π=E\α _φ(a_t s_t)-Q_ _ (s_t,a_t)\. (19) The entropy temperature is updated by minimizing ℒα=−α(logπϕ(at∣st)+ℋtar),L_α=E\-α ( _φ(a_t s_t)+H_tar )\, (20) where ℋtarH_tar denotes the expected entropy. The update of twin value distributions, actor, and temperature is practically done by the gradient descent, i.e., θi←θi−η∇θiℒ,ϕ←ϕ−ηπ∇ϕℒπ,α←α−ηα∇αℒα, _i← _i- _Z _ _iL_Z,φ←φ- _π _φL_π,α←α- _α _αL_α, (21) where η _Z, ηπ _π, and ηα _α are learning rates and ∇θiℒ _ _iL_Z, ∇ϕℒπ _φL_π, and ∇αℒα _αL_α are first-order derivatives. Finally, the target networks are softly updated according to θi′←τθi+(1−τ)θi′,ϕ′←τϕ+(1−τ)ϕ′, _i ←τ _i+(1-τ)θ _i,~~φ ←τφ+(1-τ)φ , (22) where τ is the soft-update coefficient. The above two steps iterate with each other until convergence. To summarize the interaction between the active BD-RIS-assisted heterogeneous MEC system, the DSAC-T agent, as well as the corresponding replay-based learning pipeline, we illustrate the overall framework in Fig. 2. IV Performance Evaluation We consider a reciprocal active BD-RIS-assisted heterogeneous MEC system with K=10K=10 users (including 44 deeply blocked users), one BS with N=8N=8 antennas, and one hybrid mode active BD-RIS with M=128M=128 cells. The system bandwidth is B=16B=16 MHz. The direct link of blocked users is modeled as a severely attenuated residual path with an additional attenuation factor of 0.050.05. The wireless channels follow distance-dependent Rician fading. The direct user–BS distance is sampled from [20,120][20,120] m, the user–RIS and RIS–BS distances from [5,20][5,20] m, and the blocked-user–RIS distance from [5,15][5,15] m. The path-loss exponents of the direct, user–RIS, and RIS–BS links are 3.53.5, 2.22.2, and 2.22.2, with large-scale scaling factors 0.80.8, 1.01.0, and 1.11.1, respectively. The corresponding Rician factors are 22 dB, 66 dB, and 1010 dB. The BS and RIS noise powers are σB2=10−9 _B^2=10^-9 and σI2=1.5×10−9 _I^2=1.5× 10^-9, respectively. The uplink transmit power ranges from 0.010.01 W to 0.10.1 W for offloaded users. The active BD-RIS has static power 0.50.5 W, amplifier efficiency 0.40.4, and active output-power budget 0.010.01 W. The maximum local CPU frequency is 1.31.3 GHz, while the edge CPU and GPU pools are 2424 GHz and 4848 GHz, respectively. Each evaluation scenario contains 55 CPU-type and 55 GPU-type tasks. CPU-type tasks have data sizes in [8×104,1.8×105][8× 10^4,1.8× 10^5] bits, computation intensities in [400,650][400,650] cycles/bit, and latency deadlines in [0.35,0.8][0.35,0.8] s. GPU-type tasks have data sizes in [1.98×105,4.5×105][1.98× 10^5,4.5× 10^5] bits, computation intensities in [900,1300][900,1300] cycles/bit, and latency deadlines in [0.198,0.462][0.198,0.462] s. The reward weights are set to wloc=0.5w_loc=0.5, woff=2.0w_off=2.0, wRIS=0.5w_RIS=0.5, and λT=1.0 _T=1.0, emphasizing offloading-related energy while retaining an explicit penalty for active BD-RIS overhead induced by cross-sector support. Figure 3: Training convergence curves of different DRL methods. Fig. 3 compares the training convergence of different DRL methods. DSAC-T converges rapidly and maintains the highest reward after convergence. SAC also converges quickly but stabilizes at a lower level, while Twin Delayed Deep Deterministic Policy Gradient (TD3) improves more slowly and Deep Deterministic Policy Gradient (DDPG) achieves lower final performance. This shows that the distributional value learning and refinement mechanisms in DSAC-T improve convergence quality and policy stability. TABLE I: Performance-Complexity Comparison over 300 Evaluation Scenarios Method Reward Feasibility Ratio Runtime‡ (s/scenario) DSAC-T -2.828 81.67% 0.0267 DDPG -5.899 64.33% 0.0494 SAC -4.068 71.33% 0.0516 TD3 -5.135 70.67% 0.0533 AO–SCA -8.393 66.67% 41.865 ‡The average online decision time per scenario; training time is excluded. Table I reports the final performance and online decision complexity. Since the reward is defined as the negative energy-latency penalty, a larger value is better. DSAC-T achieves the best reward and the highest feasibility ratio, improving feasibility by 10.3410.34 percentage points over SAC and 11.0011.00 percentage points over TD3. It also requires only 0.02670.0267 s per scenario, whereas the alternating optimization (AO) method based on successive convex approximation (SCA) requires 41.86541.865 s due to iterative per-instance optimization. These results demonstrate that DSAC-T achieves the best performance-complexity tradeoff among the compared methods. Figure 4: Impact of the number of active BD-RIS cells on reward and feasibility. Fig. 4 evaluates the impact of the number of active BD-RIS cells. Increasing the number of cells improves both reward and feasibility ratio, since a larger active BD-RIS provides more wave manipulation degrees of freedom. In addition, learned RIS control consistently outperforms random RIS control under all cell number settings, confirming that explicit RIS coefficient optimization is necessary and that the gain does not simply come from increasing the aperture size. V Conclusion This paper studied energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC. A physically grounded model was established to capture asymmetric blockage, residual direct-link attenuation, active surface power consumption, amplified noise, reciprocal cross-sector leakage, and decoupled CPU/GPU edge resources. The resulting mixed integer and strongly coupled problem was reformulated as a slot-wise learning-based decision problem, where the original objective and constraints were handled through action processing and a penalty-based reward. A DSAC-T-based framework was then developed with real-valued action parameterization for joint offloading, computation allocation, power control, and active BD-RIS control. Simulation results showed that DSAC-T achieved the best energy-latency reward and the highest feasibility ratio among the compared methods, while maintaining low online decision latency. These results validate the effectiveness of distributional value learning for feasibility-sensitive and reward-heterogeneous active BD-RIS-assisted MEC optimization. Acknowledgment This work is funded by the National Natural Science Foundation of China (grant no. 62501509) and the Natural Science Foundation of Guangdong Province (grant no. 2026A1515011048). References [1] P. S. Aung, K. Kim, Y. K. Tun, E. Huh, Z. Han, and C. S. Hong (2025) Active STAR-RIS empowered edge system for enhanced energy efficiency and task management. IEEE Transactions on Mobile Computing. Cited by: §I. [2] J. Duan, W. Wang, L. Xiao, J. Gao, S. E. Li, C. Liu, Y. Zhang, B. Cheng, and K. Li (2025) Distributional soft actor-critic with three refinements. IEEE Transactions on Pattern Analysis & Machine Intelligence 47 (05), p. 3935–3946. External Links: Document Cited by: §I. [3] X. Hu, C. Masouros, and K. Wong (2021) Reconfigurable intelligent surface aided mobile edge computing: from optimization-based to location-only learning-based solutions. IEEE Transactions on Communications 69 (6), p. 3709–3725. External Links: Document Cited by: §I. [4] L. Huang et al. (2020) Deep reinforcement learning for online computation offloading in wireless powered mobile-edge computing networks. IEEE Transactions on Mobile Computing. External Links: Document Cited by: §I. [5] H. Li, M. Nerini, S. Shen, and B. Clerckx (2026) A tutorial on beyond-diagonal reconfigurable intelligent surfaces: modeling, architectures, system design and optimization, and applications. IEEE Communications Surveys & Tutorials 28 (), p. 4086–4126. External Links: Document Cited by: §I. [6] F. Liu, H. Li, and S. Shen (2026) Active beyond-diagonal reconfigurable intelligent surface with hybrid transmitting and reflecting mode. arXiv:2604.13570. Cited by: §I, §I, §I-A2, §I-A. [7] Y. Liu, X. Liu, X. Mu, T. Hou, J. Xu, M. Di Renzo, and N. Al-Dhahir (2021) Reconfigurable intelligent surfaces: principles and opportunities. IEEE Communications Surveys & Tutorials 23 (3), p. 1546–1577. Cited by: §I, §I. [8] Y. Liu, X. Mu, J. Xu, R. Schober, Y. Gao, H. V. Poor, and L. Hanzo (2021) STAR: simultaneous transmission and reflection for 360° coverage by intelligent surfaces. IEEE Wireless Communications 28 (6), p. 102–109. External Links: Document Cited by: §I. [9] R. Long, Y. Liang, Y. Pei, and E. G. Larsson (2021) Active reconfigurable intelligent surface-aided wireless communications. IEEE Transactions on Wireless Communications 20 (8), p. 4962–4975. External Links: Document Cited by: §I. [10] Y. Mao, C. You, J. Zhang, K. Huang, and K. B. Letaief (2017) A survey on mobile edge computing: the communication perspective. IEEE Communications Surveys & Tutorials 19 (4), p. 2322–2358. External Links: Document Cited by: §I. [11] H. Peng and X. Wang (2023) Energy harvesting reconfigurable intelligent surface for UAV communication based on robust deep reinforcement learning. IEEE Transactions on Wireless Communications. External Links: Document Cited by: §I. [12] X. Qin, W. Yu, Q. Ni, Z. Song, T. Hou, and X. Sun (2025) Joint resource allocation and beamforming design for BD-RIS-assisted wireless-powered cooperative mobile edge computing. IEEE Communications Letters 29 (5), p. 1042–1046. External Links: Document Cited by: §I. [13] X. Qin, W. Yu, Q. Ni, Z. Song, T. Hou, J. Wang, and X. Sun (2025) Resource allocation and beamforming design for active STAR-RIS-assisted wireless-powered MEC. IEEE Internet of Things Journal. Cited by: §I. [14] S. Shen, H. Li, M. Nerini, Q. Wu, and B. Clerckx (2026) Active beyond-diagonal reconfigurable intelligent surfaces: modeling, architecture design, and optimization. arXiv:2603.13861. Cited by: §I. [15] J. Xu, J. Zuo, J. T. Zhou, and Y. Liu (2023) Active simultaneously transmitting and reflecting (STAR)-RISs: modeling and analysis. IEEE Communications Letters 27 (9), p. 2466–2470. External Links: Document Cited by: §I.