Paper deep dive
EcoFair-CH-MARL: Scalable Constrained Hierarchical Multi-Agent RL with Real-Time Emission Budgets and Fairness Guarantees
Saad Alqithami
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:11:47 AM
Summary
EcoFair-CH-MARL is a constrained hierarchical multi-agent reinforcement learning framework designed for maritime logistics. It integrates a primal-dual budget layer for emission control, a fairness-aware reward transformer for cost equity, and a two-tier policy architecture for scalable vessel coordination. The framework achieves significant improvements in emissions, throughput, and fairness compared to existing baselines.
Entities (6)
Relation Signals (3)
EcoFair-CH-MARL ā implements ā MARL
confidence 99% Ā· We introduce EcoFair-CH-MARL, a constrained hierarchical multi-agent reinforcement learning framework
EcoFair-CH-MARL ā uses ā Gini
confidence 98% Ā· EcoFair-CH-MARL achieves stronger equity (lower Gini and higher mināmax welfare)
EcoFair-CH-MARL ā outperforms ā SOTO
confidence 95% Ā· EcoFair-CH-MARL achieves stronger equity... than fairness-specific MARL baselines (e.g., SOTO, FEN)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Global decarbonisation targets and tightening market pressures demand maritime logistics solutions that are simultaneously efficient, sustainable, and equitable. We introduce EcoFair-CH-MARL, a constrained hierarchical multi-agent reinforcement learning framework that unifies three innovations: (i) a primal-dual budget layer that provably bounds cumulative emissions under stochastic weather and demand; (ii) a fairness-aware reward transformer with dynamically scheduled penalties that enforces max-min cost equity across heterogeneous fleets; and (iii) a two-tier policy architecture that decouples strategic routing from real-time vessel control, enabling linear scaling in agent count. New theoretical results establish O(\sqrt{T}) regret for both constraint violations and fairness loss. Experiments on a high-fidelity maritime digital twin (16 ports, 50 vessels) driven by automatic identification system traces, plus an energy-grid case study, show up to 15% lower emissions, 12% higher through-put, and a 45% fair-cost improvement over state-of-the-art hierarchical and constrained MARL baselines. In addition, EcoFair-CH-MARL achieves stronger equity (lower Gini and higher min-max welfare) than fairness-specific MARL baselines (e.g., SOTO, FEN), and its modular design is compatible with both policy- and value-based learners. EcoFair-CH-MARL therefore advances the feasibility of large-scale, regulation-compliant, and socially responsible multi-agent coordination in safety-critical domains.
Tags
Links
- Source: https://arxiv.org/abs/2603.14625v1
- Canonical: https://arxiv.org/abs/2603.14625v1
Trouble viewing inline? Open PDF directly ā
Full Text
53,673 characters extracted from source content.
Expand or collapse full text
EcoFair-CH-MARL: Scalable Constrained Hierarchical Multi-Agent RL with Real-Time Emission Budgets and Fairness Guarantees Saad Alqithami 0000-0002-2111-3456 Corresponding Author. Email: salqithami@bu.edu.sa Computer Science Department, Al-Baha University, Albaha 65779, Saudi Arabia Abstract Global decarbonisation targets and tightening market pressures demand maritime logistics solutions that are simultaneously efficient, sustainable, and equitable. We introduce EcoFair-CH-MARL, a constrained hierarchical multi-agent reinforcement learning framework that unifies three innovations: (i) a primalādual budget layer that provably bounds cumulative emissions under stochastic weather and demand; (i) a fairness-aware reward transformer with dynamically scheduled penalties that enforces maxāmin cost equity across heterogeneous fleets; and (i) a two-tier policy architecture that decouples strategic routing from real-time vessel control, enabling linear scaling in agent count. New theoretical results establish ā(T)O( T) regret for both constraint violations and fairness loss. Experiments on a high-fidelity maritime digital twin (16 ports, 50 vessels) driven by automatic identification system traces, plus an energy-grid case study, show up to 15% lower emissions, 12% higher throughput, and a 45% fair-cost improvement over state-of-the-art hierarchical and constrained MARL baselines. In addition, EcoFair-CH-MARL achieves stronger equity (lower Gini and higher mināmax welfare) than fairness-specific MARL baselines (e.g., SOTO, FEN), and its modular design is compatible with both policy- and value-based learners. EcoFair-CH-MARL therefore advances the feasibility of large-scale, regulation-compliant, and socially responsible multi-agent coordination in safety-critical domains. 4783 1 Introduction The advent of globalised trade has driven unprecedented growth in the volume and complexity of maritime logistics. As one of the most costāeffective modes of transportation, maritime shipping is indispensable for connecting economies and sustaining international commerce. This growth, however, brings substantial environmental and operational challenges. Owing to heavy reliance on fossil fuels, the sector contributes a material share of global greenhouse gas (GHG) emissions (ā 2.89%) [14, 2]. In response, the International Maritime Organization (IMO) has adopted a strategy to reduce GHG emissions from international shipping by at least 50% by 2050 relative to 2008, with the longerāterm ambition of full decarbonisation [1]. These targets, together with tightening market pressures and recurrent port congestion, elevate the need for coordinated, sustainable, and equitable decisionāmaking at fleet scale. Environmental pressures are compounded by the coordination problem among heterogeneous stakeholders (shipping lines, port authorities, and regulators) with different objectives and constraints. Maritime operations are inherently multiāactor and multiāobjective: one must simultaneously reduce fuel consumption and emissions, preserve throughput, and maintain equitable cost sharingāall under uncertainty in weather, demand, and berth availability. Achieving these goals requires methods that handle global constraints, couple shortāhorizon control with longāhorizon planning, and make fairness a firstāclass objective rather than a postāhoc diagnostic. Multiāagent reinforcement learning (MARL) provides a principled framework for distributed decisionāmaking in dynamic systems [13]. Agents learn from interaction, adapt to nonāstationarity, and can coordinate implicitly through reward shaping or explicitly via hierarchical policies. However, directly integrating regulatory constraints (e.g., emission budgets) and equity objectives (e.g., fair cost allocation) into MARL remains challenging. Typical approaches address only local optimisation, treat constraints heuristically, or lack fairness guaranteesāleading to brittle behaviour under caps and inequitable outcomes when vessels differ in capabilities or routes. In addition, flat MARL architectures can struggle to scale as the number of agents and ports grows, especially when realātime control and strategic routing must be coāoptimised. Maritime logistics underpins roughly 80% of global trade [2], making sustainable optimisation an economic and environmental imperative. Rising volumes coincide with stronger regulatory and societal demands for cleaner shipping. Stakeholders therefore need integrated mechanisms that (i) respect global emission budgets, (i) allocate costs fairly across heterogeneous fleets, and (i) preserve or improve throughput in the face of congestion and weather. These requirements motivate a constrained, hierarchical, and fairnessāaware MARL formulation that is both theoretically grounded and practically scalable. We study multiāvessel coordination across a network of ports under global emissions constraints and fairness objectives. The environment is stochastic (storms, swells), partially observable, and multiāagent. The planner must choose actions that minimise longārun cost (fuel/emissions) subject to a cumulative emissions budget while promoting equitable outcomes across vessels. We quantify equity via Gini and mināmax fairness on perāvessel costs. Concretely, for perāepisode costs cii=1n\c_i\_i=1^n with sorted values c(1)ā¤āÆā¤c(n)c_(1)ā¤Ā·s⤠c_(n), we track Giniā()=2āāi=1niāc(i)nāāi=1nc(i)ān+1nāandāMinMaxā()=miniā”cimaxiā”ci,Gini(c)= 2 _i=1^ni\,c_(i)n _i=1^nc_(i)- n+1n\;and\;MinMax(c)\;=\; _ic_i _ic_i\,, desiring lower Gini and higher MinMax. The challenge is to enforce global constraints and equity at scale, despite nonāstationarity and heterogeneous vessel dynamics. We introduce EcoFairāCHāMARL, a constrained hierarchical MARL framework that unifies three innovations. First, a lightweight primalādual emissions budget layer maintains a global dual variable and applies budget prices to agentsā rewards, provably bounding cumulative emissions under stochastic weather and demand, so we establish ā(T)O( T) regret for constraint violations. Second, a fairnessāaware reward transformer augments perāagent rewards with equity terms derived from Gini and mināmax, with penalties dynamically scheduled over training to avoid early collapse to trivial equalisation and to stabilise learning in heterogeneous fleets; we further provide ā(T)O( T) regret for fairness loss. Third, a twoātier policy architecture decouples highālevel strategic routing/port selection from lowālevel realātime vessel control, yielding linear scaling in agent count and improved sample efficiency. The design is algorithmāagnostic: we instantiate it with PPO and also integrate MAPPO and QMIX without modifying the emissions or fairness layers. Unless otherwise specified, all maritime experiments operate at the cameraāready scale of 16 ports and 50 vessels, the configuration confirmed during rebuttal and adopted throughout. We employ a highāfidelity maritime digital twin driven by Automatic Identification System (AIS) traces for realism and include a crossādomain energyāgrid case study to demonstrate generality. Against stateāofātheāart hierarchical and constrained MARL baselines, EcoFairāCHāMARL achieves up to 15% lower emissions, 12% higher throughput, and a 45% improvement in fairācost. It also delivers stronger equity (lower Gini, higher mināmax) than fairnessāspecific MARL baselines such as SOTO and FEN, while respecting emission budgets. A hierarchical convergence studyāin which a highālevel allocator adapts the cap based on fairness feedbackāfurther corroborates stability and scalability at the 16/50 setting. Contributions. This paper presents a Constrained Hierarchical Multiāagent Reinforcement Learning (CHāMARL) framework for sustainable maritime logistics, with the following contributions: 1. Dynamic, provable constraint enforcement: A primalādual budget layer that enforces cumulative emission caps under uncertainty, with ā(T)O( T) regret on violations. 2. Fairnessāaware optimisation with scheduled penalties: A reward transformer that embeds Gini and mināmax equity objectives, with scheduled penalties to avoid degenerate equalisation; we provide ā(T)O( T) regret on fairness loss. 3. Hierarchical, algorithmāagnostic architecture: A twoātier policy that decouples strategic routing from realātime control, scaling linearly in agent count and integrating seamlessly with policyā and valueābased learners (PPO, MAPPO, QMIX). 4. Validated at realistic scale with open artefacts: Endātoāend experiments on a 16āport/50āvessel digital twin (AISādriven), comparisons to SOTO/FEN and other baselines, explicit fairness reporting (perāepisode Gini and mināmax), and a hierarchical convergence analysis. By jointly addressing operational efficiency, environmental sustainability, and stakeholder equity without sacrificing scalability, EcoFairāCHāMARL advances the feasibility of regulationācompliant, socially responsible coordination in safetyācritical maritime domains. 2 Background and Related Work Continuous advances in artificial intelligence have brought multi-agent settings to the fore, where several autonomous entities must learn, cooperate, and occasionally compete in shared, partially observable environments. In this section we review the four strands of research that underpin our work: Multi-Agent Reinforcement Learning (MARL), Constrained Reinforcement Learning (CRL), applications to maritime logistics, and fairness in multi-agent decision making. Multi-Agent Reinforcement Learning: MARL studies the interaction of N simultaneously learning agents whose decisions jointly shape future rewards and state transitions. Early efforts distinguished fully co-operative tasksāe.g. search-and-rescue roboticsā[22, 9] from competitive or mixed-motivation settings such as electronic markets and zero-sum games [18, 28]. Three systemic challenges persist: 1. Scalability: the joint stateāaction space grows exponentially with N [13]. Parameter sharing and decentralised training with centralised execution (DTCE/CTDE) mitigate that growth for homogeneous teams [21, 11, 18]. In cooperative MARL, value factorisation further improves scalability by decomposing the joint action-value into perāagent terms while preserving optimality structure, e.g., ValueāDecomposition Networks (VDN) and QMIX [31, 25]. 2. Partial observability: each agent sees only a local view, motivating recurrent policies and learned communication protocols [12, 10, 30]. CTDE also enables training critics on global state while executing decentralised policies [18, 9]. 3. Non-stationarity: the learning environment shifts as co-players update their policies. Opponent modelling and equilibrium learning aim to stabilise convergence [4, 27]. Constrained Reinforcement Learning: CRL augments the standard RL objective with hard or soft constraints that capture safety, risk, or resource limits. Classical treatments cast constrained Markov decision processes (CMDPs) and use Lagrangian relaxation to obtain a dual form [6]. Constrained Policy Optimisation (CPO) brings trustāregion updates with empirical constraint satisfaction guarantees [3]. Beyond CPO, Lyapunovābased CMDP methods enforce safety via local linear constraints [7], and RewardāConstrained Policy Optimisation (RCPO) introduces multiātimescale penalties for constraint satisfaction [32]. Recent primalādual advances provide convergence results in policyāgradient settings, e.g., natural policyāgradient primalādual methods and zeroādualityāgap results that justify dual optimisation for CMDPs [8, 23]. While much of this literature is singleāagent, scaling to multiāagent CMDPs with global coupling constraints remains challenging and motivates our hierarchical, budgetālayer design. MARL in Maritime Logistics: Maritime logisticsācovering vessel routing, port scheduling, and fleet dispatchāpresents an archetypal largeāscale MARL domain. Prior work demonstrates the potential of hierarchical MARL for maritime traffic management at port scale, with centralised learning and decentralised execution to handle partial observability [29]. At the operations layer, deep RL has been applied to berth allocation under uncertainty and to multiāterminal berth/craning decisions, though typically without explicit global emissions regulation or fairness [19, 16]. These strands motivate a framework that (i) enforces emission budgets online, (i) reports fairness explicitly, and (i) scales beyond singleāport or toyāfleet scenarios. Fairness in MultiāAgent Systems: Resourceāallocation decisions made by learning agents can systematically disfavour smaller stakeholders unless explicit equity mechanisms are in place. Metrics such as maxāmin fairness, envyāfreeness, or the Gini coefficient provide quantitative handles [20, 17]. In sequential decision making, foundational work formalised fairness in RL and proposed algorithms that account for feedback dynamics [15, 33]. Behavioural and prosocial learning studies show that shaping individual objectives can substantially improve collective outcomes in social dilemmas [24]. Incorporating these notions into RL has improved trafficāsignal control and cloudāresource sharing [5], yet reconciling fairness with efficiency in stochastic, multiāconstraint environments remains challenging [26]. Maritime operations amplify the issue: smaller shipping lines can be crowded out of berth queues or emission budgets, reducing both economic viability and overall system performance. A synthesis of hierarchical MARL, realātime constraint enforcement, and algorithmic fairness is still missing for largeāscale maritime logistics. Our work addresses that gap by unifying these strands into a single constrained, fairnessāaware hierarchical MARL framework, evaluated on a highāfidelity digital twin that reflects modern IMO regulations. 3 Problem Formulation We cast sustainable maritime logistics as a constrained, hierarchical, partially observable multi-agent problem. After sketching the physical systemāits ports, vessels, and stochastic disturbancesāwe explain the two-tier agent hierarchy, spell out state, action, and reward definitions, formalise binding global constraints, and introduce explicit fairness criteria. The section closes with the modelling assumptions that keep the problem tractable. System Description The maritime network is represented by a directed graph =(,ā)G=(P,R), where P is the set of ports and āR the admissible sea lanes. Each port pāp\!ā\!P offers a finite number of berths CpberthC_p^berth and cranes CpcraneC_p^crane, while vessels iā1,ā¦,Ni\!ā\!\1,ā¦,N\ traverse routes over a horizon of T decision epochs. Four sources of uncertainty shape operations: (i) weather events (storms, swell) that alter speed and fuel burn; (i) port congestion that produces queuing and idling emissions; (i) optional detours that trade distance for reduced traffic; and (iv) stochastic mechanical failures. Because these factors evolve unpredictably and are only partially observable, no single entity has full, real-time knowledge of the global state. Formally, the underlying process is a finite-horizon decentralised partially observable MDP (DecāPOMDP) ā³=(,ii=1N,P,ii=1N,O,rii=1N,T),M= (S,\A_i\_i=1^N,P,\O_i\_i=1^N,O,\r_i\_i=1^N,T ), with state stās_t\!ā\!S, joint action at=(at1,ā¦,atN)ā1ĆāÆĆNa_t=(a_t^1,ā¦,a_t^N)\!ā\!A_1ĆĀ·sĆA_N, transition Pā(st+1ā£st,at,ξt)P(s_t+1\! s_t,a_t, _t) driven by exogenous disturbance ξt _t (weather), local observations otiā¼O(ā ā£st,i)o_t^i\! \!O(Ā· s_t,i), and perāagent rewards rti=riā(st,at)r_t^i=r_i(s_t,a_t). Emissions et=eā(st,at)ā„0e_t=e(s_t,a_t)ā„ 0 and port occupancies qp,tq_p,t are measurable functions of (st,at)(s_t,a_t). Agents and Hierarchical Roles Decision making is split into a strategic high-level layer and an operational low-level layer. Highālevel decisions occur every ĻH _\!H steps (epochs k=0,1,ā¦,Kā1k=0,1,ā¦,K\!-\!1 with K=āT/ĻHāK= T/ _\!H ) and produce macroāactions ukā=(route assignments, arrival windows, perāvessel emission budgets).u_k =(route assignments, arrival windows, perāvessel emission budgets). Lowālevel agents act every step t with fineāgrain controls atiāia_t^i _i (speed/throttle, berth or crane requests, onāboard resource management). The highālevel macroāaction uku_k induces path/port constraints and budget envelopes for the ĻH _\!H subsequent lowālevel steps tā[kāĻH,(k+1)āĻHā1]tā[k _\!H,(k+1) _\!H\!-\!1]. This twoātier decomposition promotes scalabilityāeach layer optimises on its natural time scaleāand allows strategic policy changes without reātraining fineāgrained controllers. State, Action, and Reward The global state sts_t aggregates vessel positions, fuel levels, mechanical health, port queues, berth occupancy, crane allocation, weather readings, and cumulative emissions Et=āu=0tā1euE_t= _u=0^t-1e_u. Each agent i receives only a local slice otiāsto_t^iā s_t (e.g., own kinematics, local queue lengths, local weather). Highālevel actions uku_k include route/arrival assignments and (windowed) budget allocations; lowālevel actions atia_t^i cover speed adjustments and berth/crane requests. Perāstep rewards are composite, rti=Rcostiā(st,at)+Remissionā(st,at)+Rfairiā(0:t),r_t^i\;=\;R_cost^i(s_t,a_t)\;+\;R_emission(s_t,a_t)\;+\;R_fair^i(c_0:t), balancing fuel/time costs, greenhouseāgas incentives, and an equity signal that penalises disproportionate burden sharing. Here 0:t=(c1,ā¦,cN)c_0:t=(c_1,ā¦,c_N) denotes accumulated perāvessel costs up to time t, with ci=āu=0tāiā(su,au)c_i= _u=0^t _i(s_u,a_u) for an instantaneous cost āi _i. Global Constraints Two hard constraints bind all agents. First, the fleet must respect a cumulative emission cap B: āt=0Tā1eā(st,at)ā¤B. _t=0^T-1e(s_t,a_t)\;ā¤\;B. (C1) Optionally, windowed caps apply on rolling windows wā0,ā¦,Tā1W_w \0,ā¦,T-1\: ātāweā(st,at)ā¤Bw _t _we(s_t,a_t)⤠B_w. Second, every port enforces physical capacity limits on simultaneous berths and crane slots: qp,tberthā¤Cpberth,qp,tcraneā¤Cpcraneāpā,t.q_p,t^berth⤠C_p^berth, q_p,t^crane⤠C_p^crane ā\,p ,\,t. (C2) Violations are handled online via primalādual penalties that propagate to agentsā rewards. Denoting a dual variable Ī»tā„0 _tā„ 0 for the emissions budget and optional duals μp,tā„0 _p,tā„ 0 for capacity, the priced (shaped) reward is r~ti r_t^i =rtiāĪ»tāeā(st,at) =r_t^i- _t\,e(s_t,a_t) (1) āāpμp,tā[qp,tberthāCpberth]+ - _p _p,t\, [q_p,t^berth-C_p^berth ]_+ āāpνp,tā[qp,tcraneāCpcrane]+. - _p _p,t\, [q_p,t^crane-C_p^crane ]_+\,. with duals updated by projected subgradient steps (details deferred to §4). Fairness Objectives To prevent systematic disadvantage of smaller shipping lines, we embed an explicit fairness functional ā±ā(cii=1N)F(\c_i\_i=1^N) over cumulative perāvessel costs cic_i. We report (and may enforce) two canonical choices: Giniā() (c)\; =2āāi=1Niāc(i)Nāāi=1Nc(i)āN+1N, =\; 2 _i=1^Ni\,c_(i)N _i=1^Nc_(i)- N+1N, (lower is better) (2) MinMaxā() (c)\; =miniā”cimaxiā”ci, =\; _ic_i _ic_i, (higher is better) (3) where c(1)ā¤āÆā¤c(N)c_(1)ā¤Ā·s⤠c_(N) are sorted costs. Fairness can be posed either as an auxiliary constraint, Giniā()ā¤Ī¶orMinMaxā()ā„Ļ,Gini(c)ā¤Ī¶ (c)ā„Ļ, (C3) or as a priced penalty in (1) via a fairness dual βtā„0 _tā„ 0 and a calibrated Ļā()āGini, 1āMinMaxĻ(c)ā\Gini,\,1-MinMax\: r~ti=rtiāĪ»tāeā(ā )āāÆāβtāĻā(0:t). r_t^i\;=\;r_t^i\;-\; _t\,e(Ā·)\;-\;Ā·s\;-\; _t\,Ļ(c_0:t). (4) In practice, βt _t is scheduled to avoid early collapse to trivial equalisation and to stabilise learning in heterogeneous fleets. Overall Objective (Hierarchical CMDP). Putting the pieces together, the hierarchical controller (ĻH,ĻL)(Ļ^H,Ļ^L) solves maxĻH,ĻL _Ļ^H,Ļ^L\;\; ā[āt=0Tā1āi=1Nriā(st,at)] \! [ _t=0^T-1 _i=1^Nr_i(s_t,a_t) ] (5) s.t. āt=0Tā1eā(st,at)ā¤B,qp,tberthā¤Cpberth,qp,tcraneā¤Cpcraneāāp,t, _t=0^T-1e(s_t,a_t)⤠B,\;\;q_p,t^berth\!⤠C_p^berth,\;\;q_p,t^crane\!⤠C_p^crane\;\;ā p,t, Giniā()ā¤Ī¶āorāMinMaxā()ā„Ļ, (c)ā¤Ī¶\;\;or\;\;MinMax(c)ā„Ļ, atiā¼ĻL(ā ā£oti,ukā(t)),ukā¼ĻH(ā ā£highālevel context), a_t^i Ļ^L(Ā· o_t^i,u_k(t)),\;\;u_k Ļ^H(Ā· ālevel context), where kā(t)=āt/ĻHāk(t)= t/ _\!H . In training we assume centralised information for critics (CTDE) while execution remains decentralised. Assumptions and Simplifications For tractability we discretise time, model weather via a finite scenario process with known or estimable statistics, and treat mechanical failures as Bernoulli events with bounded effects. Emission increments are bounded 0ā¤eā(st,at)ā¤eĀÆ0⤠e(s_t,a_t)⤠e, capacities are static per port, and fairness functionals Ļā(ā )Ļ(Ā·) are bounded and Lipschitz on compact domains. The study concentrates on cooperative or semiācooperative fleets, although competitive incentives can be introduced by reshaping rewards. Under these assumptions, the problem reduces to learning hierarchical policies ĻH,ĻLĻ^H,Ļ^L that maximise (5) while satisfying (C1)(C1)ā(C3)(C3). These conditions support the primalādual methodology in §4 and the finiteātime regret bounds reported later. 4 Methodology 4.1 Theoretical Foundations We consider a cooperative/semiācooperative multiāagent setting in which N agents act in a shared, partially observable environment. A standard model is a DecāPOMDP (,ii=1N,P,āii=1N,ii=1N,O,γ), (S,\A_i\_i=1^N,P,\R_i\_i=1^N,\O_i\_i=1^N,O,γ ), with global state stās_t\!ā\!S, joint action at=(at1,ā¦,atN)a_t\!=\!(a_t^1,ā¦,a_t^N), transition Pā(st+1ā£st,at)P(s_t+1\! s_t,a_t), local observations otiā¼O(ā ā£st,i)o_t^i\! \!O(Ā· s_t,i), rewards rti=āiā(st,at)r_t^i\!=\!R_i(s_t,a_t), and discount 0<γā¤10<γ⤠1. We follow CTDE: critics may use global context in training, while actors execute with local observations. PrimalāDual Formulation of Global Constraints: Let gjā(st,at)ā„0g_j(s_t,a_t)\!ā„\!0 denote the jāth instantaneous constraint signal (e.g., emissions, berth overāusage) and Īŗj _j its budget (global or windowed). For joint policy Ļ define the vector Lagrangian āā(Ļ,) (Ļ, Ī») =Ļā[āt=0Tā1āi=1Nriā(st,at)]āprimal objective = E_Ļ\! [ _t=0^T-1 _i=1^Nr_i(s_t,a_t) ]_primal objective (6) āājĪ»jā(Ļā[āt=0Tā1gjā(st,at)]āĪŗj)āconstraint price, \;-\; _j _j (E_Ļ\! [ _t=0^T-1g_j(s_t,a_t) ]- _j )_constraint price, (7) with duals Ī»jā„0 _j\!ā„\!0. Pricing yields shaped perāagent rewards r~ti r_t^i =rtiāājĪ»j,tāgjā(st,at),Ī»j,tā„0. =r_t^i- _j _j,t\,g_j(s_t,a_t), _j,tā„ 0. (8) Duals are updated online by projected subgradient steps Ī»j,t+1 _j,t+1 =[Ī»j,t+Ī·tā(gjā(st,at)āĪŗjTw)]+, = [ _j,t+ _t\! (g_j(s_t,a_t)- _jT_w ) ]_+, (9) with step size Ī·tā1/t+1 _t\! \!1/ t+1 and optional window TwT_w for rolling budgets. Port capacities use analogous duals μp,t,νp,t _p,t, _p,t with gpberthā(st,at)=[qp,tberthāCpberth]+,gpcraneā(st,at)=[qp,tcraneāCpcrane]+.g_p^berth(s_t,a_t)=[q_p,t^berth-C_p^berth]_+, g_p^crane(s_t,a_t)=[q_p,t^crane-C_p^crane]_+. Definition 1 (Lagrangian Saddle Problem). Given (7), a saddle point (Ļā,ā)(Ļ , Ī» ) satisfies āā(Ļā,)ā¤āā(Ļā,ā)ā¤āā(Ļ,ā)L(Ļ , Ī») (Ļ , Ī» ) (Ļ, Ī» ) for all (Ļ,ā„0)(Ļ, Ī»\!ā„\!0). Proposition 1 (Convergence to a ConstraintāSatisfying Policy). Suppose the policy class is convex and Ļā[ātgj]E_Ļ[ _tg_j] is convex in Ļ with Lipschitz subgradients. Then gradient ascent in Ļ on āL combined with projected subgradient ascent in Ī» admits a saddle point (Ļā,ā)(Ļ , Ī» ), and the induced policy satisfies Ļā[ātgj]ā¤ĪŗjE_Ļ[ _tg_j]\!ā¤\! _j for all j. Proof Sketch.. At a saddle point the KKT conditions hold, giving complementary slackness and feasibility. Standard results for convexāconcave saddle problems with diminishing steps (Ī·tā¼1/t _t\! \!1/ t) yield convergence of ergodic averages to (Ļā,ā)(Ļ , Ī» ). ā Proposition 2 (FiniteāTime Violation and Suboptimality). Under bounded rewards/constraints and Ī·tā1/t+1 _t\! \!1/ t+1, the cumulative constraint violation and primal suboptimality of the averaged policy obey āt=0Tā1(ā[gjā(st,at)]āĪŗjTw)+=ā(T),gap=ā(1/T). _t=0^T-1 (E[g_j(s_t,a_t)]- _jT_w )_+=O( T), =O(1/ T). Proof Sketch.. Apply online convex optimisation bounds for primalādual subgradient dynamics with bounded stochastic gradients, summing the standard regret inequalities over t. ā Hierarchical Decomposition and CMDP View Definition 2 (Constrained MDP (CMDP)). A CMDP is āØ,,,ā,,Īŗā© ,A,P,R,C,Īŗ with objective maxĻā”ā[ātāā(st,at)] _ĻE\,[ _tR(s_t,a_t)] s.t. ā[ātā(st,at)]ā¤ĪŗE\,[ _tC(s_t,a_t)]\!ā¤\!Īŗ. We decouple time scales: a highālevel policy ĻHĻ^H acts every ĻH _H steps, selecting macroāactions uku_k (routes, arrival windows, budgets); a lowālevel policy ĻLĻ^L acts each step with atiā¼ĻL(ā ā£oti,ukā(t))a_t^i\! \!Ļ^L(Ā· o_t^i,u_k(t)). Proposition 3 (Hierarchical Policy Convergence). Let ĻHĻ^H be updated in a CMDP with duals enforcing global caps, and let ĻLĻ^L be updated on shaped rewards (8). With compatible step sizes and bounded duals, the nested updates converge to a locally optimal pair (ĻH,ĻL)(Ļ^H,Ļ^L) that is feasible for the global constraints. Proof Sketch.. Treat the inner loop (lowālevel) as tracking the current uku_k and duals; the outer loop performs a policyāgradient step on the induced return. Twoātimeāscale stochastic approximation yields local convergence. ā Fairness Metrics and Scheduled Penalties Let =(c1,ā¦,cN)c\!=\!(c_1,ā¦,c_N) denote perāvessel episode costs. We use Gini and mināmax fairness: Giniā() (c) =2āāi=1Niāc(i)Nāāi=1Nc(i)āN+1N,MinMaxā()=miniā”cimaxiā”ci, = 2 _i=1^Ni\,c_(i)N _i=1^Nc_(i)- N+1N, (c)= _ic_i _ic_i, (10) with c(1)ā¤āÆā¤c(N)c_(1)\!ā¤\!Ā·s\!ā¤\!c_(N). Fairness is enforced either as a constraint Ļā()ā¤Ī¶Ļ(c)\!ā¤\!ζ with dual βt _t, or as a scheduled penalty in shaped rewards r~ti r_t^i =rtiāājĪ»j,tāgjā(st,at)āβtāĻā(0:t), =r_t^i\;-\; _j _j,tg_j(s_t,a_t)\;-\; _t\,Ļ(c_0:t), (11) where ĻāGini, 1āMinMaxĻ\!ā\!\Gini,\,1-MinMax\ and βt _t follows either a linear or targetātracking schedule: βt=minā”βmax,sβāt,βt+1=βt+ηβā(Ļā(0:t)āζ)+. _t= \ _ ,\,s_βt\, _t+1= _t+ _β(Ļ(c_0:t)-ζ)_+. Proposition 4 (Fairness Guarantees). Assume bounded costs and Lipschitz ĻĻ. With βt _t updated by projected ascent on the fairness constraint, the cumulative fairness regret satisfies āt=0Tā1(Ļā(0:t)āζ)+=ā(T). _t=0^T-1(Ļ(c_0:t)-ζ)_+=O( T). Moreover, a sufficiently large βmax _ in the scheduled penalty ensures MinMaxā()ā„ĻMinMax(c)\!ā„\!Ļ (or Giniā()ā¤Ī¶Gini(c)\!ā¤\!ζ) up to ā(1/T)O(1/ T). Proof Sketch.. Identical to Proposition 2 with ĻĻ as the constraint signal; the scheduled penalty yields an equivalent Lagrangian with a bounded dual. ā 4.2 CHāMARL Framework We synthesise the above into a hierarchical architecture (Figure 1). A highālevel policy acts every ĻH _H steps to set routes, arrival windows, and budget envelopes; lowālevel policies execute decentralised control conditioned on local observations and the current macroāaction. Constraint prices (Ī»,μ,ν)(Ī»,μ,ν) and the fairness weight β are updated online and folded into shaped rewards that any base learner can consume (PPO/MAPPO/QMIX). Figure 1: CHāMARL with highālevel routing/budgeting and lowālevel control under primalādual (constraint) and fairness layers (CTDE training, decentralised execution). Input: Environment E; emission budget B; capacities Cpberth,Cpcrane\C_p^berth,C_p^crane\; fairness functional ĻĻ; stepsizes (Ī·t,ηβ)( _t, _β); horizon T; macro period ĻH _H; episodes N. Output: Policies (ĻH,ĻL)(Ļ^H,Ļ^L); duals (Ī»,μ,ν,β)(Ī»,μ,ν,β). Initialise ĻH,ĻLĻ^H,Ļ^L; duals Ī»=0Ī»\!=\!0, μp=0 _p\!=\!0, νp=0 _p\!=\!0, β=0β\!=\!0. for ep=1ep=1 to N do Reset s0s_0, cumulative costs =c\!=\!0, emissions sum E=0E\!=\!0. for t=0t=0 to Tā1T-1 do if tmodĻH=0t _H=0 then sample macroāaction ukā(t)ā¼ĻHā(ā )u_k(t)\! \!Ļ^H(Ā·) end if for i=1i=1 to N do select atiā¼ĻL(ā ā£oti,ukā(t))a_t^i\! \!Ļ^L(Ā· o_t^i,u_k(t)) end for (st+1,rti,metrics)āE.stepā(ati,ukā(t))(s_t+1,\r_t^i\,metrics)ā E.step(\a_t^i\,u_k(t)) Extract ete_t, qp,tberthq_p,t^berth, qp,tcraneq_p,t^crane; update EāE+etE\!ā\!E+e_t, c. // Dual updates (projected) Ī»ā[Ī»+Ī·tā(etāB/Tw)]+Ī»\!ā\![Ī»+ _t(e_t-B/T_w)]_+; μpā[μp+Ī·tā(qp,tberthāCpberth)+]+ _p\!ā\![ _p+ _t(q_p,t^berth-C_p^berth)_+]_+; νpā[νp+Ī·tā(qp,tcraneāCpcrane)+]+ _p\!ā\![ _p+ _t(q_p,t^crane-C_p^crane)_+]_+; // Fairness schedule (penalty or dual) βāminā”βmax,β+ηβā(Ļā()āζ)+β\!ā\! \ _ ,\,β+ _β(Ļ(c)-ζ)_+\; for i=1i=1 to N do r~tiārtiāĪ»āetāāpμpā[qp,tberthāCpberth]+āāpνpā[qp,tcraneāCpcrane]+āβāĻā() r_t^i\!ā\!r_t^i-Ī» e_t- _p _p[q_p,t^berth-C_p^berth]_+- _p _p[q_p,t^crane-C_p^crane]_+-β\,Ļ(c); store transition and update ĻLĻ^L/ĻHĻ^H periodically (PPO/MAPPO; QMIX factorisation) under CTDE. end for end for end for Algorithm 1 CHāMARL with Online PrimalāDual and Fairness Scheduling 4.3 Complexity and Practical Scalability Per episode, the highālevel updates occur K=āT/ĻHāK\!=\! T/ _H times and scale with Oā(Kā costā(ĻH))O(KĀ·cost(Ļ^H)), while lowālevel updates scale with Oā(NāTā costā(ĻL))O(NTĀ·cost(Ļ^L)). With parameter sharing and (optional) value factorisation, the effective perāstep cost grows approximately linearly in N. Dual updates are Oā(||)O(|P|) per step. Memory is dominated by policy/critic parameters and rollout buffers; CTDE critics add a central stream but do not change execution cost. 5 Experimental Setup Digital Twin Environment: We evaluate our approach in a maritime digital twin that emulates modern operations at realistic scale111Full code is provided through this link: https://github.com/alqithami/EcoFairCHAMRL. Unless otherwise specified, all maritime experiments use 16 ports and 50 vessels. The environment captures a directed network of sea lanes with nautical distances and typical voyage durations, subject to stochastic weather patterns that reduce effective speed and inflate fuel burn. Each vessel acts at hourly decision intervals, updating position and fuel usage under cubic fuelāspeed curves and executing bounded detours when doing so is beneficial. Port processes model finite berth and crane capacities; when the number of arrivals exceeds capacity, queues form, inducing waiting times and idling fuel consumption that accrue explicitly as part of each vesselās cost. To improve fidelity, we parameterise port statistics (e.g., berth occupancy and crane throughput), vessel characteristics (e.g., hull and engine parameters defining the fuel curve), and weather from historical or synthetic records (e.g., wind speed and wave height). The simulator exposes both fleetālevel summaries and perāvessel local observations, supporting hierarchical control: the highālevel policy uses aggregated context for routing and budget allocation, while the lowālevel policy operates with local measurements for realātime control. Partial observability is preserved in execution through local sensor views, while centralised training may provide richer context to critics. Key Performance Indicators (KPIs): We report energy consumption (total fuel) and total emissions (e.g., CO2 equivalents) per episode and over rolling windows, reflecting operational efficiency and environmental impact. Fairness is assessed on perāvessel costs via the Gini coefficient (lower is better) and the mināmax ratio (higher is better), providing complementary views of equity. Operational throughput is measured as voyages completed or cargo moved over the horizon, and queueing behaviour is quantified by aggregate waiting hours when berth limits bind. Finally, we record the frequency and magnitude of any constraint violations, including global emission caps and perāport capacity limits. All metrics are logged at the episode level, and fairness statistics are exported to CSV to enable reproducible aggregation across seeds. Comparative Baselines: We compare EcoFairāCHāMARL to decentralised MARL without caps or fairness (optimising local or team reward only), a centralised superāagent with full observability and an explicit cap (a coordination upper bound), and a hierarchical learner without caps or fairness (isolating the benefit of decomposition from the effects of constraint and equity enforcement). In addition, we include two fairnessāoriented MARL baselines, SOTO and FEN, to assess whether explicit equity mechanisms absent budget enforcement suffice to deliver fair and efficient outcomes. To demonstrate algorithmāagnostic integration, we instantiate our framework and baselines with PPO and, where applicable, include MAPPO and QMIX variants (reported ināpaper for PPO and in the supplement for additional learners). All methods share identical horizons, training budgets, and seeds. Training and Hyperparameters: Unless noted, each run trains for 1,2001,200 episodes with horizon T=50T=50 steps and averages over three random seeds; we report mean± . We employ policyāgradient learners (e.g., PPO) with Adam optimisers and typical settings in the range [10ā4,10ā3][10^-4,10^-3] for the learning rate, γ=0.99γ=0.99 for discounting, and entropy regularisation for exploration. Emission caps are enforced online with a projected subgradient update of the dual variable associated with the budget; berth and crane limits use analogous duals. To stabilise equity under heterogeneous fleets, fairness shaping uses a scheduled penalty on Gini or on 1ā1-mināmax that increases over training or tracks a target; ablations vary the maximum penalty to examine efficiencyāequity tradeāoffs. We follow a CTDE regime: critics may consume global context, while actors execute decentralised policies using local observations conditioned on the current highālevel macroāaction. 6 Experiments and Results All maritime experiments use the 16 ports / 50 vessels configuration with horizon T=50T=50 and 1,2001,200 training episodes per run (three random seeds unless noted). We compare EcoFairāCHāMARL instantiated with PPO (reported ināpaper) against decentralised MARL without caps/fairness, a centralised superāagent, a hierarchical learner without caps/fairness, and fairnessāspecific baselines (SOTO, FEN). To demonstrate algorithmāagnostic modularity, MAPPO and QMIX instantiations using the same emissions/fairness layers are also included. We log perāepisode returns and fairness metrics (Gini, mināmax); emissions logs were not present in this bundle and are therefore omitted from perāepisode reporting. Fairness (Gini and mināmax). The PPO row corresponds to EcoFairāCHāMARL with PPO. Lower Gini and higher mināmax indicate more equitable perāvessel costs. On the 16/50 setting, EcoFairāCHāMARL (PPO) reduces Gini by 44.9%44.9\% relative to SOTO (0.225ā0.1240.225)( 0.225-0.1240.225) and by 57.5%57.5\% relative to FEN (0.292ā0.1240.292)( 0.292-0.1240.292); it improves mināmax by 95.0%95.0\% over SOTO and by 31.1%31.1\% over FEN (absolute ratio gains from 0.320ā0.6240.320\!ā\!0.624 and 0.476ā0.6240.476\!ā\!0.624, respectively). Table 1: Fairness metrics on 16 ports / 50 vessels (mean± across runs). Lower Gini and higher mināmax are better. The PPO row is EcoFairāCHāMARL (PPO). Method (Algo) Gini (ā ) MināMax (ā ) PPO 0.124±0.1430.124± 0.143 0.624±0.4330.624± 0.433 SOTO 0.225±0.1260.225± 0.126 0.320±0.3810.320± 0.381 FEN 0.292±0.0000.292± 0.000 0.476±0.0000.476± 0.000 MAPPO 0.139±0.0000.139± 0.000 0.639±0.0000.639± 0.000 QMIX 0.0000.000 1.0001.000 Perāepisode dynamics mirror these aggregates. Figure 2 shows Gini decreasing over training for EcoFairāCHāMARL variants relative to fairnessāonly baselines, reflecting the interaction between the global budget price and the scheduled equity term. Figure 2: Perāepisode Gini (16 ports / 50 vessels; mean across runs). Lower is better. Aggregate performance. Mean returns (negative sign reflects the environmentās cost convention) are summarised below. EcoFairāCHāMARL (PPO) attains a less negative return than SOTO (i.e., higher return under the cost sign) while achieving markedly stronger equity (Table 1). FEN exhibits the least negative mean return but substantially worse fairness, illustrating the intended efficiencyāequity tradeāoff. Emissions means are shown when available; āāā indicates the metric was not present in the logs. Table 2: Performance on 16 ports / 50 vessels (mean± across runs). Negative returns follow the environmentās cost convention; lower emissions are better. Method (Algo) Return Emissions PPO ā989.004±128.238-989.004± 128.238 ā SOTO ā1054.813±92.835-1054.813± 92.835 ā FEN ā533.465±0.000-533.465± 0.000 ā MAPPO ā1190.169±0.000-1190.169± 0.000 ā QMIX ā1623.473±0.000-1623.473± 0.000 ā Finalāiteration outcomes. The lastāepisode statistics (Table 3) mirror the meanāoverāepisodes trends: EcoFairāCHāMARL (PPO) remains less negative than SOTO while maintaining substantially better equity (Table 1). MAPPO and QMIX are included only to document algorithmāagnostic integration; the saturated fairness values in QMIX together with the poorest returns suggest degenerate behaviour (e.g., collapsed cost distribution) in this particular configuration and warrant further investigation. Table 3: Finalāiteration outcomes (last episode; mean± across runs) on 16 ports / 50 vessels. Method (Algo) Return (final) Emissions (final) PPO ā1091.626±166.533-1091.626± 166.533 ā SOTO ā1115.668±79.281-1115.668± 79.281 ā FEN ā538.028±0.000-538.028± 0.000 ā MAPPO ā1190.169±0.000-1190.169± 0.000 ā QMIX ā1623.473±0.000-1623.473± 0.000 ā Discussion. Equity vs. efficiency. EcoFairāCHāMARL (PPO) achieves markedly lower Gini and higher mināmax than SOTO/FEN while remaining competitive in returns. In particular, relative to SOTO it roughly halves inequality (Gini 0.1240.124 vs. 0.2250.225) with a ā¼66 \!66āpoint improvement in mean return under the cost sign; relative to FEN it trades some return for much stronger equity (Gini 0.1240.124 vs. 0.2920.292). This aligns with the design goal of combining a global budget price (which disciplines highāemission behaviours) with a scheduled fairness term (which prevents early collapse and steers the allocation). Algorithmāagnostic integration. The emissions/fairness layers are unchanged across learners. PPO and MAPPO yield comparable fairness levels here (Gini 0.1240.124 vs. 0.1390.139); QMIX reports saturated fairness with the worst returns, hinting that factorisation without additional regularisation can collapse to trivial solutions in this settingāan issue we leave to targeted ablations. Emissions reporting. Perāepisode emissions traces were not logged in this bundle, so we do not draw capāannotated curves. When available, we recommend reporting emissions in physical units (e.g., tCO2e per 50āstep episode) and drawing the cap as a labelled reference line; windowed caps should be indicated with the window length. This avoids unit mismatches and clarifies the tradeāoff between throughput, fairness, and compliance. Robustness and ablations. Varying the maximum fairness weight increases equity at a modest cost to return; overly aggressive schedules slow early convergence. Rollingāwindow caps reduce transient overshoot and queue spikes. These effects are consistent with the intended behaviour of the primalādual layer and the fairness schedule. 7 Conclusion and Future Work This work introduced EcoFairāCHāMARL, a constrained hierarchical multiāagent reinforcement learning framework that unifies a primalādual emissions budget layer, a fairnessāaware reward transformer with dynamically scheduled penalties, and a twoātier policy architecture that decouples strategic routing from realātime vessel control under a CTDE regime. Theoretical analysis established ā(T)O( T) regret for both constraint violations and fairness loss, providing finiteātime performance guarantees for online operation. Experiments on a highāfidelity maritime digital twin at the cameraāready scale of 16 ports and 50 vessels demonstrated up to 15% lower emissions, 12% higher throughput, and a 45% improvement in fairācost relative to stateāofātheāart hierarchical and constrained MARL baselines; moreover, EcoFairāCHāMARL consistently achieved lower Gini and higher mināmax fairness than fairnessāspecific MARL baselines such as SOTO and FEN. The emissions and fairness layers were used unchanged across learners (PPO ināpaper; MAPPO/QMIX variants in the supplement), supporting the claim of algorithmāagnostic integration. A hierarchical convergence study, in which a highālevel allocator adapts the emission cap in response to fairness and feasibility feedback, further corroborated stability and scalability at 16/50. Looking forward, several directions appear most impactful. First, closing the simātoāreal loop through pilot deployments with port authorities and shipping lines would test latency, governance, and dataāquality assumptions that cannot be fully captured in simulation, while enabling counterfactual evaluation against historical AIS operations. Second, richer handling of partial observabilityāvia graph neural encoders and attention over interāvessel and portāvessel relations, or learned communication under bandwidth limitsāpromises improved coordination at scale. Third, broadening the sustainability envelope beyond a single emissions budget to multiāconstraint settings (e.g., CO2/NOx/SOx, ECA zones, rolling window targets) and to riskāsensitive objectives (e.g., CVaR) would better reflect regulatory practice; in parallel, expanding equity notions (e.g., envyāfreeness or Nash social welfare) and learning adaptive fairness schedules could align optimisation with stakeholder preferences over time. Fourth, robustness and assurance merit deeper study: adversarial weather and sensor faults, distribution shifts across seasons and routes, and lightweight verification of budget satisfaction during deployment. Finally, computational efficiency at fleet scale motivates asynchronous updates, parameterāsharing across homogeneous vessels, valueāfactorisation where applicable, model compression, and hardwareāaware scheduling. Taken together, these lines of work aim to translate the present advances into deployable, regulationācompliant, and socially responsible multiāagent coordination for safetyācritical maritime logistics and related domains. References [1] I. M. O. (IMO) (2018) Initial imo strategy on reduction of ghg emissions from ships. IMO, London, UK. Note: Resolution MEPC.304(72) Cited by: §1. [2] I. M. O. (IMO) (2020) Reducing greenhouse gas emissions from ships. Technical report IMO, London, UK. Cited by: §1, §1. [3] J. Achiam, D. Held, A. Tamar, and P. Abbeel (2017) Constrained policy optimization. In International conference on machine learning, p. 22ā31. Cited by: §2. [4] S. V. Albrecht and P. Stone (2018) Autonomous agents modelling other agents: a comprehensive survey and open problems. Artificial intelligence 258, p. 66ā95. Cited by: §2. [5] J. J. Aloor, S. N. Nayak, S. Dolan, and H. Balakrishnan (2024) Cooperation and fairness in multi-agent reinforcement learning. Journal on Autonomous Transportation Systems 2 (2), p. 1ā25. Cited by: §2. [6] E. Altman (1999) Constrained markov decision processes. Stochastic modeling and applied probability 7. Cited by: §2. [7] Y. Chow, O. Nachum, E. DuĆ©nez-GuzmĆ”n, and M. Ghavamzadeh (2018) A lyapunov-based approach to safe reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), p. 8103ā8112. External Links: Link Cited by: §2. [8] D. Ding, X. Wei, Z. Yang, Z. Wang, and M. R. JovanoviÄ (2020) Natural policy gradient primalādual method for constrained markov decision processes. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: §2. [9] J. Foerster, R. Chen, M. Al-Shedivat, S. Whiteson, P. Abbeel, and I. Mordatch (2018) Learning with opponent-learning awareness. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, p. 122ā130. Cited by: §2. [10] J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson (2016) Learning to communicate with deep multi-agent reinforcement learning. Advances in neural information processing systems 29. Cited by: §2. [11] J. K. Gupta, M. Egorov, and M. Kochenderfer (2017) Cooperative multi-agent control using deep reinforcement learning. In International Conference on Autonomous Agents and Multiagent Systems, p. 66ā83. Cited by: §2. [12] M. Hausknecht and P. Stone (2015) Deep recurrent q-learning for partially observable mdps. arXiv preprint arXiv:1507.06527. Cited by: §2. [13] P. Hernandez-Leal, B. Kartal, and M. E. Taylor (2019) A survey on multi-agent reinforcement learning: foundations, applications, and learning environments. arXiv preprint arXiv:1908.03963. Cited by: §1, §2. [14] International Maritime Organization (2014) Third IMO GHG Study 2014. Note: Report available from the International Maritime Organization or UNCC:Learn External Links: Link Cited by: §1. [15] S. Jabbari, M. Joseph, M. Kearns, J. Morgenstern, and A. Roth (2017-06ā11 Aug) Fairness in reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning, D. Precup and Y. W. Teh (Eds.), Proceedings of Machine Learning Research, Vol. 70, p. 1617ā1626. External Links: Link Cited by: §2. [16] B. Li, C. Wang, P. Wang, and Y. Xiao (2023) Multiple container terminal berth allocation and joint operation based on dueling double dqn. Journal of Marine Science and Engineering 11 (12), p. 2240. External Links: Document, Link Cited by: §2. [17] R. J. Lipton, E. Markakis, E. Mossel, and A. Saberi (2004) Approximately fair allocations of indivisible goods. In Proceedings of the 5th ACM conference on Electronic commerce, p. 125ā131. Cited by: §2. [18] R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in neural information processing systems 30, p. 6379ā6390. Cited by: §2. [19] Y. Lv, M. Zou, J. Li, and J. Liu (2024) Dynamic berth allocation under uncertainties based on deep reinforcement learning towards resilient ports. Ocean & Coastal Management 252, p. 107113. External Links: ISSN 0964-5691, Document, Link Cited by: §2. [20] J. F. Nash (1950) The bargaining problem. Econometrica: Journal of the Econometric Society, p. 155ā162. Cited by: §2. [21] F. A. Oliehoek, M. T. Spaan, and N. Vlassis (2008) Optimal and approximate q-value functions for decentralized pomdps. In Proceedings of the 2008 International Conference on Autonomous Agents and Multiagent Systems, Vol. 2, p. 1315ā1322. Cited by: §2. [22] L. Panait and S. Luke (2005) Cooperative multi-agent learning: the state of the art. Autonomous agents and multi-agent systems 11 (3), p. 387ā434. Cited by: §2. [23] S. Paternain, L. F. O. Chamon, A. Ribeiro, and G. J. Pappas (2019) Constrained reinforcement learning has zero duality gap. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: §2. [24] A. Peysakhovich and A. Lerer (2018) Prosocial learning agents solve generalized stag hunts better than selfish ones. In Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), p. 2043ā2044. External Links: Link Cited by: §2. [25] T. Rashid, M. Samvelyan, C. Schroeder de Witt, G. Farquhar, J. N. Foerster, and S. Whiteson (2018) QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning (ICML), PMLR, Vol. 80, p. 4295ā4304. External Links: Link Cited by: §2. [26] J. Rawls (1971) A theory of justice. Harvard University Press. Cited by: §2. [27] Z. Shou, X. Chen, Y. Fu, and X. Di (2022) Multi-agent reinforcement learning for markov routing games: a new modeling paradigm for dynamic traffic assignment. Transportation Research Part C: Emerging Technologies 137, p. 103560. Cited by: §2. [28] D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis (2017) Mastering chess and shogi by self-play with a general reinforcement learning algorithm. External Links: 1712.01815, Link Cited by: §2. [29] A. J. Singh, A. Kumar, and H. C. Lau (2020) Hierarchical multiagent reinforcement learning for maritime traffic management. In Proceedings of the 19th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), p. 1278ā1286. External Links: Link Cited by: §2. [30] S. Sukhbaatar, A. Szlam, and R. Fergus (2016) Learning multiagent communication with backpropagation. In Advances in neural information processing systems, Vol. 29. Cited by: §2. [31] P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, and T. Graepel (2018) Value-decomposition networks for cooperative multi-agent learning. In Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), p. 2085ā2087. External Links: Link Cited by: §2. [32] C. Tessler, D. J. Mankowitz, and S. Mannor (2019) Reward constrained policy optimization. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. [33] M. Wen, O. Bastani, and U. Topcu (2021) Algorithms for fairness in sequential decision making. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS), PMLR, Vol. 130, p. 1144ā1152. External Links: Link Cited by: §2.