Paper deep dive
Equilibria in Multiplayer Graph Games: An Algorithmic Study
Léonard Brice
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/8/2026, 5:51:59 PM
Summary
This thesis presents an algorithmic study of equilibria in multiplayer graph games, focusing on the computational complexity of the constrained existence problem. It analyzes five equilibrium concepts—Nash, subgame-perfect, strong secure, and risk-sensitive equilibria—across various game types including parity, mean-payoff, energy, discounted-sum, and stochastic games. The work establishes complexity results ranging from undecidability and co-recursive enumerability to NP-complete, PSPACE-complete, and EXPTIME-complete classifications, while also exploring related concepts like rational verification and the negotiation function.
Entities (18)
Relation Signals (15)
Léonard Brice → authored → Equilibria in Multiplayer Graph Games: An Algorithmic Study
confidence 100% · thèse présentée par Léonard Brice
Marie van den Bogaard → cosupervised → Léonard Brice
confidence 100% · et de la Professeure Marie van den Bogaard, co-promotrice
Jean-François Raskin → supervised → Léonard Brice
confidence 100% · Sous la direction du Professeur Jean-François Raskin, promoteur
Constrained Existence Problem → assesses → Nash Equilibrium
confidence 95% · we provide complexity results for the constrained existence problem, which consists of deciding whether a given game contains an equilibrium
Constrained Existence Problem → assesses → Subgame-Perfect Equilibrium
confidence 95% · We establish connections between SPEs and a function on vertex labelings... to prove that the constrained existence problem for SPEs isNP-complete
Strong Secure Equilibrium → hascomplexityin → Parity Games
confidence 95% · constrained existence problem for these equilibria isPSPACE-complete in parity games
Nash Equilibrium → hascomplexityin → Energy Games
confidence 95% · the constrained existence problem is undecidable ... in energy games.
Subgame-Perfect Equilibrium → hascomplexityin →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:To verify the robustness of a program or protocol, it is common in the computer science community to rely on the theoretical framework of game theory. In particular, if one seeks to enforce a desired property, or specification, despite an unpredictable environment, a useful abstraction is to model the situation as a two-player zero-sum game. The goal is then to find a strategy for the system that guarantees the specification against any strategy of the environment. However, to model more complex situations, such as multiple systems with different objectives or an environment composed of various agents, the richer framework of multiplayer games must be considered. In this setting, a natural question is to identify equilibria, i.e., strategy profiles that are robust in the sense that no player has an incentive to deviate. The most well-known equilibrium concept is the Nash equilibrium, but several alternatives exist. We study five such notions and, for each of them, we provide complexity results for the constrained existence problem, which consists of deciding whether a given game contains an equilibrium that ensures each player a payoff within a specified interval.
Tags
Links
- Source: https://arxiv.org/abs/2605.19954v1
- Canonical: https://arxiv.org/abs/2605.19954v1
Trouble viewing inline? Open PDF directly →
Full Text
525,543 characters extracted from source content.
Expand or collapse full text
Equilibria in Multiplayer Graph Games An Algorithmic Study thèse présentée par Léonard Brice en vue de l’obtention du grade académique de Doctorat en Sciences Année académique 2024-2025 Sous la direction du Professeur Jean-François Raskin, promoteur et de la Professeure Marie van den Bogaard, co-promotrice Jury de thèse : Jean Cardinal (Université libre de Bruxelles, Président) Jean-François Raskin (Université libre de Bruxelles) Marie van den Bogaard (Université Gustave Eiffel) Emmanuel Filiot (Université libre de Bruxelles) Patricia Bouyer-Decître (Université Paris-Saclay) Orna Kupferman (Université hébraïque de Jérusalem) Thomas Henzinger (Institute of Science and Technology Austria) arXiv:2605.19954v1 [cs.GT] 19 May 2026 2 EQUILIBRIA IN MULTIPLAYER GRAPH GAMES AN ALGORITHMIC STUDY Léonard Brice Academic year 2024-2025 2 ABSTRACT To verify the robustness of a program or protocol, it is common in the computer science community to rely on the theoretical framework of game theory. In particular, if one seeks to enforce a desired property, or specification, despite an unpredictable environment, a useful abstraction is to model the situation as a two-player zero-sum game. The goal is then to find a strategy for the system that guarantees the specification against any strategy of the environment. However, to model more complex situations, such as multiple systems with different objectives or an environment composed of various agents, the richer framework of multiplayer games must be considered. In this setting, a natural question is to identify equilibria, i.e., strategy profiles that are robust in the sense that no player has an incentive to deviate. The most well-known equilibrium concept is the Nash equilibrium, but several alternatives exist. We study five such notions and, for each of them, we provide complexity results for the constrained existence problem, which consists of deciding whether a given game contains an equilibrium that ensures each player a payoff within a specified interval. Regarding Nash equilibria in particular, we prove the following results: the constrained existence problem is undecidable (but recursively enumerable) in energy games. In discounted-sum games, that same problem is co-recursively enumerable, but at least as hard as the target discounted-sum problem, whose decidability remains a long-standing open question. The main part of our contribution focuses on subgame-perfect equilibria (SPEs), a refinement of Nash equilibria in sequential games that eliminates non-credible threats by requiring that the planned strategies of all players still form a Nash equilibrium after any possible history. We establish connections between SPEs and a function on vertex labelings with useful properties, the negotiation function, to prove that the constrained existence problem for SPEs isNP-complete in co-Büchi, parity, and mean-payoff games. Furthermore, we demonstrate that the same results as those established for Nash equilibria in energy games and discounted-sum games hold for SPEs, except for recursive enumerability in energy games, which remains an open question. We then explore the relationship between the constrained existence problem and a related question, rational verification, which involves verifying whether a given strategy enforces a specified property against all rational responses—defined in terms of Nash equilibria or SPEs. In cases where rational responses are not guaranteed to exist, such as SPEs in mean-payoff games, we propose an alternative definition of rational verification. Here, the specification must hold against every response that is as rational as possible, using the notion of휀-SPE. We show that this problem is P NP -complete in mean-payoff games. A third notion that we investigate is the strong secure equilibrium, where no coalition of players can harm another player without also harming at least one of its own members. We argue that strong secure equilibria provide a suitable model for safe protocols among untrusted agents. We further show that the constrained existence problem for these equilibria isPSPACE-complete in parity games and EXPTIME-complete in Boolean games where winning conditions are defined by parity automata. 3 4 Finally, we study the case of stochastic games, where players are allowed to randomize their strategies. In this context, the constrained existence problem for Nash or subgame-perfect equilibria is known to be undecidable, even in games with rewards on terminal vertices, if each player seeks to maximize their expected payoff. We therefore investigate risk-sensitive equilibria, where players instead maximize a risk measure that accounts for their inclination or aversion to risk. In particular, we prove that if all players exhibit extreme risk aversion or extreme risk inclination, the constrained existence problem becomes NP-complete, and even P-complete if they all have extreme risk inclination. ACKNOWLEDGEMENTS To express my feelings as sincerely as possible, I trust the reader will allow me to use my mother tongue. La rédaction d’une thèse est un grand mensonge, qui consiste à faire passer pour une matière froide et lisse le résultat chaotique et bouillonnant d’une aventure humaine. Ce document n’aurait jamais vu le jour, ou du moins n’aurait pas été le même, sans le concours d’un nombre considérable de personnes qui m’ont accompagné dans ce petit bout de chemin. En premier lieu, je veux bien sûr remercier ma promotrice et mon promoteur, Marie van den Bogaard et Jean-François Raskin. Votre présence rassurante, votre disponibilité, vos conseils avisés et votre empathie pendant ces quatre années ont constitué le socle le plus solide que je pouvais espérer pour commencer cette carrière de chercheur. Je remercie ensuite les personnes qui ont accepté de participer à mon jury de thèse, tâche dont je mesure l’ampleur. J’ai fait mon possible pour rendre la lecture de cette thèse la moins pénible possible en dépit de sa longueur, aggravée par mon incapacité maladive à faire le tri entre résultats importants et curiosités dispensables. J’espère que vous y trouverez une matière intéressante. Je remercie également tous mes autres coauteurs et coautrice, aux côtés desquels nous avons produit ce travail collectif qu’ils et elle m’ont permis de signer de mon nom. Merci à Anirban et Thomas Bruss d’avoir attiré mon attention sur le problème de Robbin, collaboration qui a abouti à un article dont les résultats ne figurent pas ici. Merci à Guillaume de m’avoir permis de mettre un petit pied dans la communauté des protocoles et de la sécurité. Merci à Mathieu pour ces longues heures de réflexion dans notre bureau commun (et pour tes efforts quotidiens pour ne pas faire de bruit à l’heure de la sieste). Merci à Thomas Henzinger de m’avoir accueilli à l’IST Austria, et d’avoir accepté de prolonger cette collaboration après le dépôt de cette thèse. Merci à Thejaswini de m’avoir fait découvrir l’Autriche—j’ai hâte que nous puissions de nouveau brainstormer dans ces fauteuils horriblement inconfortables où nous avons établi notre camp de base. Je veux aussi remercier Véronique et Emmanuel, qui ont complété cet accompagnement en participant à mon comité de thèse. Et, plus largement, toutes les personnes avec qui j’ai partagé quotidiennement les pauses-café, les déjeuners et bien plus pendant ces quatre ans : Allen, Ayrat, Damien, Debraj, Gilles, Mrudula, Sarah, Sayan et Thierry à Bruxelles, Ahad, Ali, Ana, Ehsan, Emily, Konstantin, Krishnendu, Mahyar, Marek, Maximilian, Mehrdad, Nicolas, Pawel, Raimundo, Stefanie, Valentin et Yakub à Vienne, ainsi que Aaron, Claire, Léo, Nadime, Victor ou encore Zéphyr à Marne-la-Vallée. J’ajoute un remerciement à tous les étudiants et étudiantes qui ont eu la patience de subir mes explications maladroites dans le cadre de mes tâches d’enseignement : transmettre des connaissances reste le meilleur moyen de prendre du recul sur elles, et à ce titre, j’ai appris de vous au moins autant que vous de moi. Enfin, je n’oublie pas Véronique, Maryka et Marie, nos extraordinaires secrétaires de département, sans qui rien de tout cela ne serait possible. Comme tout travail, la recherche ne se dissocie jamais parfaitement des autres aspects de nos vies. Il y a par conséquent nombre d’autres personnes, hors des murs de l’Académie, qui m’ont accompagné 5 6 pendant l’écriture de cette thèse, généralement sans comprendre un traître mot de ce que j’en racontais quand je m’essayais à la tâche. Parmi celles qui m’ont permis de me sentir chez moi en Belgique, je pense à ma marraine, Pajka, à sa fille, Alicia, ou encore à mon cousin Raoul et ma belle-cousine Ilektra, qui ont rendu ces années plus belles en amenant deux nouveaux petits membres à notre grande famille. Je pense à mes colocataires, avec lesquels j’aurais aimé pouvoir passer plus de temps—promis, dès que cette thèse est soumise, je sors la poubelle de verre. Je pense également à mes amis et amies de longue date, en particulier les deux Marie, Mathilde, Solène, Pierre ou encore Zoé qui, malgré la distance, ont invariablement été là pour moi dès que le besoin s’en faisait ressentir. Et je pense, bien sûr, à toutes ces personnes formidables avec lesquelles j’ai passé une large partie de mon temps extra-professionnel, parce que nous partagions les mêmes combats. Je n’en énumérerai pas la liste, elle serait fort longue, et je crois de toute manière qu’elles se reconnaîtront. Je me crois cependant fondé à faire une exception pour Raphaël, dont la perte brutale, quelques jours après son vingtième anniversaire, m’a profondément meurtri. Tu nous as beaucoup apporté, et nous n’avons pas eu le temps de te remercier. Je pense, enfin, à ma famille, le roc insubmersible sur lequel je n’ai fait que me hisser. Je remercie tous mes oncles, tantes, cousins et grands-parents que je ne vois pas suffisamment souvent, mais qui m’apportent à chaque fois beaucoup de bonheur. Je veux ici avoir une pensée particulière pour ma grand-mère, qui vient de nous quitter : je mesure aujourd’hui la chance incroyable que nous avons eue de t’avoir dans nos vies. Je remercie mon frère, Philémon, et ma sœur, Louve, pour toute cette complicité et cette solidarité que nous avons su construire ensemble : je n’ai pas suffisamment eu l’occasion de le dire, mais je suis immensément fier des adultes que vous êtes devenus. Et bien sûr, je remercie mes parents, qui ont toujours su, avec courage, patience et bienveillance, entendre et soutenir mes choix de vie dans tout ce qu’ils avaient d’improbable et de douteux. Je pose le point final à cette thèse dans une époque pleine d’incertitudes. Les récentes avancées de l’informatique, fruit du travail de cette communauté scientifique dans laquelle j’essaie modestement de me faire une place, sont aujourd’hui perçues comme une menace plus que comme une promesse de progrès, parce qu’elles sont mises au service d’intérêts économiques et d’idéologies réactionnaires—les mêmes qui font la sourde oreille lorsque la science sonne l’alerte sur la destruction de nos environnements. Dans mon pays d’origine, dans celui où je vis et dans celui où je m’apprête à partir, les sphères du pouvoir sont petit à petit gangrenées par une extrême droite qui répand ses idées rétrogrades et obscurantistes. Mais je ne crois pas que l’humanité soit condamnée à la barbarie. Mes derniers remerciements iront donc à toutes celles et tous ceux qui résistent et qui, à contre-courant, continuent à tisser d’autres futurs. CONTENTS INTRODUCTION11 1 Introduction13 1.1Two-player zero-sum games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .13 1.2Games played on graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .14 1.3Multiplayer games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .15 1.4Equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .16 1.5The complexity of constrained existence . . . . . . . . . . . . . . . . . . . . . . . . . . .17 1.6Related Works . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .18 1.7Structure and contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .21 1.8Previous publications and co-authorship . . . . . . . . . . . . . . . . . . . . . . . . . . .22 1.9How to read this thesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .22 I BACKGROUND27 2 Background29 2.1Writing conventions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .29 2.2Graphs games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 2.3Strategies, strategy profiles . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .34 2.4Stationary and finite-memory strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . .35 2.5Two-player zero-sum games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .36 2.6Equilibria and constrained existence problem . . . . . . . . . . . . . . . . . . . . . . . . .37 I NASH EQUILIBRIA39 3 Nash equilibria41 3.1Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 3.2Parity games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 3.2.1Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .42 3.2.2Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .42 3.3Mean-payoff games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .42 3.3.1Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 3.3.2Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 3.4Discounted-sum games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 7 8CONTENTS 3.4.1Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 3.4.2The target discounted-sum problem . . . . . . . . . . . . . . . . . . . . . . . . . .44 3.4.3Hardness result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .44 3.4.4Co-recursive enumerability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .44 3.5Energy games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .47 3.5.1Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .47 3.5.2Two-counter machines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .48 3.5.3Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .49 I SUBGAME-PERFECT EQUILIBRIA53 4 Subgame-perfect equilibria and negotiation55 4.1Subgame-perfect equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .55 4.2Negotiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .57 4.2.1Origin of the concept . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .57 4.2.2Requirements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .57 4.2.3Negotiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .58 4.3Link between negotiation and equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . .59 4.3.1Nash equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .59 4.3.2Link with (휀–)subgame-perfect equilibria . . . . . . . . . . . . . . . . . . . . . . .60 4.4Negotiation games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .64 4.4.1The abstract negotiation game . . . . . . . . . . . . . . . . . . . . . . . . . . . . .64 4.4.2The concrete negotiation game. . . . . . . . . . . . . . . . . . . . . . . . . . . . .68 5 Parity games73 5.1Reduced plays and reduced strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . .73 5.1.1Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .73 5.1.2Representatives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .74 5.2Checking that a reduced strategy is winning . . . . . . . . . . . . . . . . . . . . . . . . .76 5.3A deterministic upper bound . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .78 6 Mean-payoff games81 6.1Hardness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .81 6.2Manipulating sets of payoff vectors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .84 6.2.1Achievable payoff vectors in a mean-payoff game . . . . . . . . . . . . . . . . . .84 6.2.2Equations and inequations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .84 6.2.3Lemmas . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .85 6.3About negotiation in mean-payoff games . . . . . . . . . . . . . . . . . . . . . . . . . . .86 6.3.1Stationary optimal strategies for Challenger . . . . . . . . . . . . . . . . . . . . .87 6.3.2Steady negotiation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .88 6.3.3First application: an example with no SPE . . . . . . . . . . . . . . . . . . . . . .89 6.3.4The impossibility of an algorithm based on iterations . . . . . . . . . . . . . . . .90 6.4Size of the least 휀–fixed point . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .92 6.4.1Existence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .92 6.4.2Size . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .92 CONTENTS9 6.5Constrained existence of a 휆–consistent play . . . . . . . . . . . . . . . . . . . . . . . . .95 6.6Computing the negotiation function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .97 6.6.1A disturbing example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .97 6.6.2Reduced strategies for mean-payoff games . . . . . . . . . . . . . . . . . . . . . .97 6.7Algorithm and complexity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 7 Energy and discounted-sum games105 7.1Discounted-sum games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 7.1.1Hardness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 7.1.2Co-recursive enumerability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 7.2Energy games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 7.2.1Undecidability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 7.2.2About recursive enumerability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 8 About rational verification and its limits113 8.1Rational verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113 8.2Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 8.3Link with the constrained existence problem and complexities . . . . . . . . . . . . . . . 114 8.4A potential limit: the temptation of chaos . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 8.5Achaotic rational verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 8.5.1Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 8.5.2Coincidence with classical rational verification . . . . . . . . . . . . . . . . . . . 119 8.5.3The complexity of achaotic rational verification in mean-payoff games . . . . . . 120 8.6Rational synthesis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 IV OTHER EQUILIBRIA129 9 Strong secure equilibria and their applications131 9.1Motivating example and definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131 9.1.1With a trusted third party . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 9.1.2Beyond trust: without a trusted third party . . . . . . . . . . . . . . . . . . . . . 132 9.1.3Strong secure equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134 9.1.4Problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 9.2A tool: the deviator game . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 9.3Parity games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 9.4In 휔 -regular games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 10 Risk-sensitive equilibria149 10.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 10.1.1 About randomness in games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 10.1.2 Constrained existence of randomized NEs (and SPEs) . . . . . . . . . . . . . . . . 151 10.1.3 Randomness and risk . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151 10.2 Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 10.2.1 Probabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 10.2.2 Risk measures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 10CONTENTS 10.2.3 Simple stochastic games . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 154 10.2.4 Strategies, and strategy profiles . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155 10.2.5 Risk-sensitive equilibria . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155 10.3 Entropic risk measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156 10.4 The existence of ERSEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 10.5 Constrained existence problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 10.5.1 Undecidability in the general case . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 10.5.2 Restrictions on strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 11 Extreme risk-sensitive equilibria163 11.1 The extreme risk measure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164 11.1.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164 11.1.2 Link with the entropic risk measure . . . . . . . . . . . . . . . . . . . . . . . . . 164 11.1.3 A technical lemma . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165 11.2 The existence of XRSEs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166 11.3 Constrained existence problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171 11.3.1 Membership in NP . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172 11.3.2 Restrictions on strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182 11.3.3 Things get easier when everyone is optimistic . . . . . . . . . . . . . . . . . . . . 185 DISCUSSION197 INTRODUCTION 11 CHAPTER 1: INTRODUCTION The world is an arena, and life is a game. Every day, we interact with an environment composed of agents. They may be similar to us or vastly different, whether in terms of the range of actions they can take to influence the physical world, or the motivations that drive them—be they other human beings, animals, institutions, or machines. Given the omnipresence of interactions between such agents, each defined by their objectives and their ability to choose among multiple strategies, scientists have developed game theory—a broad conceptual framework designed to model such interactions using mathematical structures. Historically, the first advances in game theory were focused on economics, a field that remains one of its most significant, if not primary, areas of application. Naturally, its influence extended to social and political sciences, and it is frequently used to analyze international relations. However, its applications have also reached fields that might initially seem far removed from its origins. In biology, for instance, game theory is now commonly used to model and study how species adapt to their evolving environments. Finally, in recent decades, there has been a growing interest in game theory within computer science—a field to which this document belongs. 1.1TWO-PLAYER ZERO-SUM GAMES When a computer-controlled system interacts with an unpredictable environment, one of the most prominent approaches today is machine learning, often referred to—somewhat imprecisely—as artificial intelligence. For an introduction to machine learning, see [Bis06]. This approach enables the system to learn how to interact with the environment in order to maximize the satisfaction of human designers. Such a learning process builds on a substantial amount of data, which may be collected from previous similar interactions, or generated by the system itself through self-testing. For instance, while writing this document, the author occasionally used ChatGPT, a tool developed by the company OpenAI, to rephrase certain sentences in more correct English. This tool does not rely on a fixed definition of correct English, but instead interpolates it from a large set of texts generally regarded as correct. However, it is crucial to remember that every dataset represents only a partial observation of the physical world. Consequently, machine learning can at best lead to statistical satisfaction. Therefore, this approach is not suitable for all computer-controlled systems. For instance, in critical infrastructures, where errors could have severe consequences, such as in the case of a hydraulic dam, relying solely on machine learning may not be viable: such a system must absolutely guarantee certain safety properties—called its specification—regardless of the environment’s behavior. It is then convenient to model the environment as an adversarial player, whose goal is to breach the specification. Thus, this situation can be abstracted as a two-player zero-sum game: the goal is to design a strategy for the system that ensures compliance with the specification, considered as its objective, against any 13 14CHAPTER 1. INTRODUCTION Prisoner 1 confessesPrisoner 1 denies Prisoner 2 confesses(5, 5)(10, 0) Prisoner 2 denies(0, 10)(1, 1) Table 1: The prisoner’s dilemma strategy the opponent might adopt—like a chess player that, at least in theory, aims to play in a way that guarantees victory no matter how their opponent responds. This explains why the most classical applications of game theory in computer science consider two-player zero-sum games. 1.2GAMES PLAYED ON GRAPHS Game theory provides a wide range of mathematical models for games, each with its own advantages and limitations. A simple yet classical example is the model of matrix games, also known as strategic games. In this framework, each player selects a strategy from a (often finite) set of strategies, treated here as abstract choices. Each possible combination of strategies yields a given payoff to each player. A canonical example of a game that can be represented in such a form is the prisoner’s dilemma: two prisoners are arrested after committing a crime. They must choose simultaneously between confessing and denying. If one denies and the other confesses, the prisoner who denies is sentenced to 10 years in jail, while the one who confesses is immediately released as a reward. If both confess, the penalty is shared: each is sentenced to 5 years in jail. If both deny, they each spend only 1 year in jail, as a precautionary measure 1 . The matrix describing this game is given in Table 1. A pair(푥,푦)indicates that prisoner 1 is sentenced to푥years, and prisoner 2 to푦 years. However, this model fails to capture sequential dynamics, where one player’s action creates new possibilities for another, and so on. Such sequential interactions are better represented by games played on graphs, or graph games. In this model, at each time step, the game transitions to a new state, represented by a vertex in a graph, with transitions depicted as edges. A graph game is therefore similar to a snakes-and-ladders game (or a Jeu de l’oie in the French-speaking world), where all players collectively move a single token. The games considered in this document are turn-based, meaning that players never act simultaneously. Each vertex in the game is controlled by exactly one player, who decides which edge to follow when the token is located on their vertex. As an example, Figure 1 illustrates a (highly simplified) abstraction of a game modeling the decisions an automated dam must make to manage the water level. Circled vertices represent configurations where the system can act and choose among multiple (or sometimes only one) possible actions. Rectangular vertices, on the other hand, represent states where the dam is waiting for some event: a power request from the grid operator, or weather events. Thus, circled vertices are controlled by the player representing the dam, and rectangular ones are controlled by the antagonist. For example, from the vertex High Release?, the dam player chooses between moving the token to the vertex Medium or to the vertex High. If she chooses the latter, then the antagonist chooses between moving it to the vertex Flood, to the vertex High Supply?, or keeping it in the vertex High. A first natural objective for the dam player is to avoid the vertex Flood. To do so, she must also avoid the vertex High. Indeed, from such a vertex, she might be temporarily protected by a drought, or saved by 1 Although this is the usual way of presenting the prisoner’s dilemma, it is worth noting that such preventive penalties contradict the principle of the presumption of innocence, as guaranteed by the Universal Declaration of Human Rights and the International Covenant on Civil and Political Rights. 1.3. MULTIPLAYER GAMES15 Low Medium High Release? High Flood Medium Supply? High Supply? rain rain no releaserain request supply request supply release ignore ignore droughtdrought drought Figure 1: Game associated with a hydroelectric dam a request that will enable her to get back to a medium level of water by supplying electricity, but in the event of a rain, she will not be able to prevent a flood: hence the necessity of considering that this vertex is controlled by an antagonistic player, who will then move to the vertex Flood. Thus, a winning strategy for the dam player consists in, whenever the vertex High Release? is reached, releasing water, and going back to the vertex Medium. And this is indeed what we want dams to do. Note that, even though the game may be understood as over when that vertex is reached, there is still a loop from the vertex to itself. This is because the main model we will consider assumes that games have an infinite horizon. Although this might seem like a bold abstraction—since no system operates indefinitely—it is a common one in computer science, especially when specifications are properties that must be enforced over the long run. For instance, another objective for the dam player in this scenario would be to always meet every power request. From the vertex Medium Supply? and the vertex High Supply?, a winning strategy would then always supply electricity, and go back to the vertex Low or Medium. This condition applies to every request, and it is difficult to place an upper bound on the number of such requests a dam will receive over its operational lifetime—hence the infinite horizon abstraction. 1.3MULTIPLAYER GAMES Even though many situations can be captured by two-player zero-sum games, recent developments in computer science have led to an increasing need for more sophisticated models. The rise of the Internet and the proliferation of automated systems in our daily lives call for models capable of capturing interactions between an arbitrary number of agents, each driven by objectives that may not align yet are not necessarily in direct opposition. To illustrate this, let us move from the perspective of the hydraulic dam to an analysis of a complete energy system. A grid operator interacts with a large number of electricity consumers as well as multiple producers. Maintaining a constant balance between production and consumption is a challenge that must be met continuously, requiring a level of responsiveness that necessarily relies on automated processes. A 16CHAPTER 1. INTRODUCTION sudden surge in consumption or the unexpected shutdown of a power plant must immediately trigger the activation of a standby producer capable of bridging the gap, such as our dam—or, when necessary, the temporary shutdown of non-essential consumption. However, the agents involved each have their own interests in this process: consumers generally dislike power outages, and companies operating hydroelectric dams—when they are privately owned rather than public utilities—lose money if they are not given sufficient opportunities to sell their electricity. Such a situation, therefore, requires a well-defined protocol with strong guarantees ensuring that all agents adhere to it. For instance, it should not be profitable for a company operating a dam to activate it without being requested to, merely to sell electricity at the expense of a competitor. This example highlights the need for a rigorous study of equilibria, that is, agreements among all agents that remain stable against the temptation of individual agents (or sometimes coalitions of agents) to deviate for personal gain. 1.4EQUILIBRIA The literature on multiplayer games offers a wide range of equilibrium concepts. Here, we mention the ones that are studied in this work. Nash equilibria.The most fundamental notion of equilibrium (though not the only one, as we will see) is the Nash equilibrium, which is among the foundational concepts of game theory as a scientific field. It requires each player’s strategy to be such that no alternative strategy would provide a better payoff for that player, assuming all other players maintain their chosen strategies. Subgame-perfect equilibria.However, in sequential games, Nash equilibria suffer from a signifi- cant limitation known as non-credible threats: players may threaten each other with irrational behavior in the event of a deviation, even when such behavior contradicts their own objectives. Consider a simple game in which two players, Adélaïde and Barthélémy, are sequentially given the opportunity to sign a document. Both earn money if and only if they have both signed; otherwise, they earn nothing. A clear Nash equilibrium in this game is the one where both players sign. However, if both adopt the strategy of never signing, they receive no money—but improving their situation requires them to deviate together. This makes the strategy profile a Nash equilibrium as well. Yet, this outcome is highly counterintuitive, as it assumes that even if Adélaïde were to sign, Barthélémy would still refuse—which is precisely what we call a non-credible threat. To eliminate such inconsistencies, the stronger notion of subgame-perfect equilibrium (SPE) is often more relevant. A strategy profile is an SPE if, in every subgame—that is, after every possible sequence of actions—the strategies planned by the players still form a Nash equilibrium. Thus, players can still threaten each other, but only with strategies that they would have no incentive to deviate from if they actually had to follow through on them. For instance, in the game described above, only the strategy profile where both players sign is an SPE. Strong secure equilibria.In some contexts, the mere absence of profitable deviations might not be a sufficiently strong guarantee. If the primary concern is the safety of agents, then a more appropriate concept may be that of secure equilibria, where no player can deviate to harm another without also harming themselves. In the game described above, both Nash equilibria are also secure equilibria. On the other hand, considering only individual deviations may not always be sufficient, as players with overlapping interests can form coalitions. This leads to the notion of strong equilibria, which are Nash equilibria where 1.5. THE COMPLEXITY OF CONSTRAINED EXISTENCE17 no coalition of players has a collectively profitable deviation 2 . In the game described above, the strategy profile that consists in both signing is a strong equilibrium, but the one that consists in never signing is not: the coalition made of Adélaïde and Barthélémy has a profitable deviation. In this work, we do not consider secure equilibria or strong equilibria independently. Instead, we introduce strong secure equilibria, combining these two notions, and argue that this concept is useful for modeling safe protocols, particularly in the field of digital security. The strategy profile where both players sign is also a strong secure equilibrium. Risk-sensitive equilibria. Finally, when randomness is involved, Nash equilibria are typically defined using the expected payoff. This approach is justified, for instance, when the game is likely to be repeated many times, as the law of large numbers ensures that a player’s average payoff will be close to their expected payoff. However, outside of such cases, this definition has a major limitation: it does not account for the risk aversion that some players may exhibit. Assume, for instance, that a player is proposed either (1) to win€1, or (2) to play a lottery in which she loses€100 with probability 99 100 , and wins€10000 with probability 1 100 . Then, the two options yield the same expected payoff, but the author would definitely prefer option (1). This motivates the introduction of risk-sensitive equilibria, where the expectation operator is replaced by a well-chosen risk measure, which generalizes it. Many risk measures have been defined and studied in economics and finance. One particularly interesting example is the entropic risk measure, which offers useful properties, especially its flexibility due to the presence of a risk parameter휌that can take any real value: the case 휌 =0 corresponds to the classical expected payoff, the limit case휌 =+∞represents extreme risk aversion (or extreme pessimism), and the dual case휌 =−∞represents extreme risk-seeking behavior (or extreme optimism). In those two last cases, playing the lottery described above would result in a perceived payoff of€−100 for an extreme pessimist and of€10000 for an extreme optimist. We will argue in Chapter 11 that these extreme cases are particularly interesting to study, as extreme pessimists can model systems of which we know that they are designed to be completely secure (say, hydroelectric dams), while extreme optimists model agents that can restart the interaction as often as they desire (a pirate trying to hack a system, for instance). Another argument in this direction is that these extreme cases define equilibrium notions where the associated decision problems become decidable— whereas they are usually undecidable with the classical expected payoff, as well as with risk entropy, as soon as randomness is involved. 1.5THE COMPLEXITY OF CONSTRAINED EXISTENCE Throughout this document, we will often focus on computing the complexity of the following problem, referred to as the constrained existence problem: Given a gameGand two payoff vectors ̄ 푥and ̄ 푦(representing one payoff for each player), does there exist an equilibrium inGthat guarantees, for each player푖, a payoff within the interval [푥 푖 ,푦 푖 ]? To derive variations of this problem, we will explore different aspects: varying the class of games in which Gevolves, particularly by modifying the functions that define the players’ objectives; or considering 2 The definition of collectively profitable is not consistent in the literature: some definitions consider that it needs only to be profitable for one player, whereas others also require that it does not reduce any coalition member’s payoff. This ambiguity does not impact our point. 18CHAPTER 1. INTRODUCTION different equilibrium notions among those introduced in Section 1.4. By computing the complexity, we mean determining the problem’s placement within standard complex- ity classes, such as P (problems solvable in polynomial time by a deterministic algorithm),NP(problems solvable in polynomial time by a non-deterministic algorithm, i.e., an algorithm that can make guesses about the direction in which it should search),PSPACE(problems solvable using a memory of polynomial size), etc. Additionally, when possible, we aim to identify a class for which the problem is complete, meaning that the problem belongs to the class and is at least as hard as every other problem in the class—thus establishing that there is no hope that the algorithm belongs to a lower complexity class, unless the classes are actually equal. This approach comes with certain limitations. First, the constrained existence problem is only one of many decision problems that could be studied in relation to equilibria. Second, determining the complexity class for which a problem is complete may provide only a partial view of its practical difficulty. In real-world scenarios, many exponential-time algorithms are frequently used—even in cases where polynomial-time algorithms exist but are more cumbersome to implement, or turn out to be very slow on small instances. Nonetheless, we adopt this perspective based on the conviction that characterizing the complexity of the constrained existence problem (and sometimes some of its variants or subcases) for a given equilibrium notion in a given game class, necessitates the development of techniques that provide deeper insights into these concepts. Beyond direct applications of these results, we believe this contributes to a better understanding of equilibria. We hope that the techniques presented in this work will convince the reader of this claim. 1.6RELATED WORKS Non-zero-sum infinite-duration games have attracted significant attention in recent years, particularly due to their applications in reactive synthesis problems. For a comprehensive overview, we refer the reader to the survey papers [BCH + 16, Bru17, Bou19] and the references therein. From a historical perspective, the main seminal paper of this field is probably that of Nash [Nas51] in which he formalizes Nash equilibria. As for SPEs, the conceptual roots of subgame-perfection can be traced back to Zermelo’s analysis of finite, perfect-information games [Zer13], where a backward reasoning procedure is used to determine winning strategies. However, it was Kuhn [Kuh53] who formally framed this reasoning within extensive-form games, and Selten [Sel65] who ultimately defined subgame-perfect equilibrium as a refinement of Nash equilibrium applicable to more general dynamic strategic settings. Below, we discuss the contributions most relevant to our work. The complexity of Nash and subgame-perfect equilibria.Bouyer-Decître et al. [BBMU15] study (pure, i.e. non-randomized) Nash equilibria in concurrent games with휔-regular objectives, where players make choice simultaneously. In this work, we only consider turn-based game, which can be seen as a particular case of concurrent games. Brihaye et al. [BDS13] characterize Nash equilibria in quantitative games with cost-prefix-linear payoff functions based on the adversarial value. This framework encompasses parity, mean-payoff, discounted- sum objectives, and simple quantitative functions. Meunier [Meu16] develops a method based on Prover-Challenger games to decide the existence of SPEs in games with a finite number of possible payoffs. While this method results in high complexity for parity games and is inapplicable to mean-payoff and discounted-sum games (where payoffs form an infinite set), it inspired the concrete negotiation game, which we introduce in Chapter 4. 1.6. RELATED WORKS19 Le Roux and Pauly [LRP14] study conditions for the existence of Nash equilibria,휀-Nash equilibria, SPEs, and휀-SPEs. Flesch et al. [FKM + 10] prove that휀-SPEs always exist when the payoff functions are lower-semicontinuous. Flesch and Predtetchinski [FP17] propose an alternative characterization of SPEs in games with finitely many payoffs, based on a game structure that we refer to as the abstract negotiation game, without formalizing the notion of negotiation function. Brihaye et al. [BBG + 19] analyze the constrained existence problem in quantitative reachability games and show that it isPSPACE-complete. Their proof introduces a preliminary version of what we call the negotiation function. Ummels and Wojtczak [UW11b] prove that the constrained existence problem for Nash equilibria in mean-payoff games is undecidable when players can randomize strategies andNP-complete when they cannot. Furthermore, they show that the same problem remains undecidable in games with rewards on terminal vertices if stochastic vertices are allowed [UW11a]. For SPEs, Ummels and Grädel [GU08] establish their existence in games with휔-regular objectives and provides anEXPTIMEalgorithm based on tree automata to solve the constrained SPE existence problem in parity games. However, this algorithm only places the problem in the classEXPTIMEwhile proving it is NP-hard. We prove NP-completeness in Chapter 5. Brihaye et al. [BBMR15] introduce weak subgame-perfect equilibria, a relaxation of classical SPEs. This weakening is equivalent to standard SPEs when the payoff function is continuous, which holds for quantitative reachability and discounted-sum objectives but not for parity and mean-payoff objectives. The techniques introduced in [BBMR15] and extended in [BRPR17] cannot be used to characterize SPEs in non-continuous settings: this sequence of papers leaves therefore open the complexity of SPEs in games such as mean-payoff games, which will be one of our main contributions. Strategy logics, as studied for instance in [CHP10], can encode SPEs for LTL objectives, a strict subset of 휔 -regular objectives. Techniques for some game classes.Chatterjee et al. [CDE + 10] study mean-payoff automata and provide results that can be translated into an expression of all possible payoff vectors in a mean-payoff game. Brenguier and Raskin [BR15] propose an algorithm to compute the Pareto curve of multi-dimensional two-player zero-sum mean-payoff games. Techniques from these papers are used in several technical steps of our algorithm presented in Chapter 6. Energy objectives, widely studied in the context of vector addition systems with states and Petri nets, are predominantly considered in two-player zero-sum settings [BFL + 08,VCD + 15,KPV16,RSV05]. Discounted-sum objectives, introduced by Zwick and Paterson [ZP06], have also been primarily analyzed in the two-player zero-sum framework. Their connection to the target discounted-sum problem, a long- standing open problem, is discussed by Boker, Henzinger, and Otop [BHO15]. To the best of our knowledge, no algorithmic results exist for those objectives in multiplayer non-zero-sum settings prior to this work. Rational verification, rational synthesis. Rational synthesis was first introduced in a cooperative setting by Fishman, Kupferman, and Lustig [FKL10], and later extended to a non-cooperative setting— the one we refer to in Chapter 8—by Perelli, Kupferman, and Vardi [KPV16]. Non-cooperative rational synthesis can be seen as an adaptation of the classical notion of Stackelberg games [vS34] to game-theoretic problems arising in computer science. Filiot, Gentilini, and Raskin [FGR20] also study rational synthesis (under the name Stackelberg value) in two-player mean-payoff and discounted-sum games, where a leader plays optimally and the follower responds with a best response (or best response up to휀 ≥0). They consider a single follower, making their setting a special case of rational synthesis as defined in [KPV16]. 20CHAPTER 1. INTRODUCTION The notion of rational verification is explored in [GNPW23], where Gutierrez, Najib, Perelli, and Wooldridge analyze the complexity of related problems. Their setting differs from what we call rational verification in Chapter 8: they study whether all Nash equilibria (or at least one) satisfy a given specification, without any player representing the system (Leader in our setting)—such a problem is closer to what we call the constrained existence problem. Nevertheless, as we show in Corollary 2, those two problems are closely related. Secure protocols and strong secure equilibria. Several works rely on notions of equilibria to argue for the resilience of protocols, such as the RatFish peer-to-peer protocol [BCK10], although they rely on 2-player Nash equilibria. Two-player games are also considered in [KR03] to evaluate the fairness of an exchange protocol, through the logic ATL. The authors hint at using coalitions, as they solve several two-player games to take into account the fact that the infrastructure may side with one agent or the other to facilitate their cheating. In contrast, strong secure equilibria, which we study in Chapter 9, syntactically consider all possible coalitions. In [ADGH06], the authors use resilient strong Nash equilibria to show that their protocol does not require a TTP as long as the agents are all rational, but they do not provide a general framework. A more general formalism to handle rational fairness is laid out in [BHC04], where the authors point out the limit of fairness as a strong trace property. They also choose the game modeling approach and propose a variant of Nash equilibria as the appropriate solution concept to capture rationality in protocols. However, their notion of rationality of the agents is optimistic: they are not assumed to be willing to deviate from the protocol only to affect negatively the others, it must at the very least impact them strictly positively. Furthermore, since the model they choose is strictly two-player, there is no notion of robustness against malignant coalitions. Thus, they favor a sort of best Nash equilibria as a rational solution concept. Finally, while they propose a relevant general framework to model protocols in terms of multiplayer games, they provide no implementations or decision procedure. At the other end of the spectrum, a very general approach uses hyperproperties [CS10,RBS24]. These allow to express properties of sets of traces, and are therefore well suited to the verification of protocols. The drawback of such a general approach is that decision problems are often not solvable algorithmically, or non-elementary for relevant fragments; in comparison, the complexity of the most difficult problems in our formalism for safe protocols, studied in Chapter 9, is "only" EXPTIME. Although a wide array of notions of rationality in games have been defined, few works focus on the notion of secure equilibria. Two notable exceptions are [CHJ05a,CR12] which actually also motivates the use of secure equilibria with the need to model malicious agents, and [BMR14], for quantitative objectives. Both however only consider the case of two-player games. Secure equilibria are also investigated in [BMR14] by Bruyère, Meunier, and Raskin, and in [CHJ06] by Chatterjee, Henzinger, and Jurdziński. In [Bre16], Romain Brenguier studies(푘,푡)-robust equilibria, a generalization of strong equilibria. Risk measures and stochastic systems. Risk measures generalize expected payoffs by cap- turing players’ perceptions of uncertainty. Widely used in economics and finance, these include ex- pected shortfall, value at risk [Aue18], variance [Bra99], and entropic risk measures [FS02]. Their ap- plication to Markov decision processes has been extensively studied, particularly regarding variance- penalized risk measures [FK89,PSB22,MT11], expected shortfall [RRS15,KM18,Meg22], and entropic risk [HM72,BR14,BCMP24]. Among these, entropic risk stands out due to its tractability: unlike ex- pected shortfall and variance-penalized measures, which require exponential memory [HK15] and com- 1.7. STRUCTURE AND CONTRIBUTIONS21 putational resources [PSB22], entropic risk allows for optimal positional strategies in Markov decision processes [HM72], making it a prime candidate for multi-agent settings. 1.7STRUCTURE AND CONTRIBUTIONS In Chapter 2, we introduce the necessary background, particularly the game model that we use throughout this document. In Chapter 3, we study Nash equilibria. Although extensively covered in the literature, we introduce several new results for discounted-sum games and energy games. For the former, we remark that the constrained existence problem of Nash equilibria is at least as hard as the target discounted-sum problem, a decision problem whose decidability status has remained open to date. However, we prove that this problem is co-recursively enumerable. Similarly, we establish that the constrained existence problem of SPEs in energy games is undecidable, but recursively enumerable. This chapter also serves as an introduction to the classes of games studied in the next part of this work. All of Part I is dedicated to subgame-perfect equilibria, the notion of equilibrium to which most of our work has been devoted. In Chapter 4, we introduce SPEs, and present the negotiation function, a key tool used in the following two chapters. The negotiation function maps every vertex labeling휆:푉 ! R∪±∞to a new labeling nego(휆), where for each vertex푣, the valuenego(휆)(푣)represents the best payoff that the player controlling 푣can enforce from푣, assuming that the other players act rationally to minimize that player’s payoff, while still satisfying their own payoff requirements, as defined by휆. We show that the negotiation function can be computed using auxiliary two-player zero-sum game structures, and that its fixed points characterize the SPEs of a game. In Chapter 5, we apply these tools to prove that the constrained existence problem of SPEs in parity games isNP-complete, thereby resolving the complexity gap for this problem, which was previously known to be NP-hard and EXPTIME-easy. In Chapter 6, we adapt these techniques to show that the same problem isNP-complete in mean-payoff games. This result requires several intermediary results, such as a polynomial bound on the size of the least fixed point of the negotiation function, making this result the most involved of Part I, if not of this whole document. In Chapter 7, we extend our results on Nash equilibria in discounted-sum and energy games to SPEs. However, unlike the case of Nash equilibria, we are unable to establish whether the constrained existence problem in energy games is recursively enumerable. We leave this as an open question. In Chapter 8, we switch to another decision problem: rational verification, which consists in verifying that a given strategy for one player (called Leader) guarantees a payoff above some threshold against every response to that strategy that is rational—rationality being defined in terms of NEs or SPEs. We show that this problem is strongly linked to the constrained existence problem, and deduce its complexity in all the game classes we mentioned above. Then, we highlight a limitation to the classical definition of rational verification, which results in counter-intuitive outcomes when no rational response exists: we then propose an alternative definition, called achaotic rational verification, which avoids that issue. We show that this new decision problem is P NP -complete in mean-payoff games with SPEs as a rationality concept, and coincides with classical rational verification in all the other cases. Part IV explores more specialized notions of equilibria. In Chapter 9, we study strong secure equilibria in parity games, and in games where each player’s objective is specified by a parity automaton. We show 22CHAPTER 1. INTRODUCTION that the constrained existence problem isPSPACE-complete in the former andEXPTIME-complete in the latter. In Chapter 10, we consider settings involving randomness, either because the game itself contains stochastic vertices or because players are allowed to randomize their strategies. We recall several known undecidability results in such settings and introduce entropic risk-sensitive equilibria, which account for players’ risk aversion or attraction. We show that the constrained existence problem of entropic risk-sensitive equilibria is undecidable even in simple stochastic games, where payoffs are determined by the terminal vertex reached by a play. However, we establish some decidability results when players are restricted to pure, stationary, or positional strategies. Finally, in Chapter 11, we introduce extreme risk-sensitive equilibria, where players interpret random- ness with extreme pessimism or optimism. We prove that the constrained existence problem for such equilibria isNP-complete, and even P-complete when all players are optimists. This result establishes, to our knowledge, the first decidable fragment of the constrained existence problem for equilibria in stochastic games, without restrictions on the number of players or on their strategies. Figure 2 summarizes results about the complexity of the constrained existence problem. 1.8PREVIOUS PUBLICATIONS AND CO-AUTHORSHIP The results presented in Chapter 4 and part of those in Chapter 6 were previously published in [BRvdB21] and its journal version [BRvdB23b]. The results in Chapter 5 appeared in [BRvdB22b], while the main findings from Chapter 6 were published in [BRvdB22a]. The results from Chapters 3, 7 and 8 were presented in [BRvdB23a]. All the aforementioned articles are joint works with Jean-François Raskin and Marie van den Bogaard. The results in Chapter 9 stem from a collaboration with Jean-François Raskin, Mathieu Sassolas, Guillaume Scerri, and Marie van den Bogaard. They were published in [BRS + 24]. Finally, Chapters 10 and 11 compile results from a joint work with Thomas Henzinger and K.S. Thejaswini, published in [BHT25]. 1.9HOW TO READ THIS THESIS This document has been written with full awareness that few people have the time or inclination to read an entire PhD thesis. For those who do, we hope it serves as a valuable, albeit necessarily incomplete, resource for developing a strong understanding of the algorithmic aspects of equilibria in multiplayer games. More realistically, many readers will approach this document in search of specific results or proofs. While all the results presented here have previously appeared in conference papers, this version may be more convenient: some proofs have been revisited, and the material is structured in a more pedagogical manner than the rapid, often fragmented presentation imposed by conference proceeding formats. For such readers, it may be useful to know that some chapters are independent of others. Dependences between chapters. This introduction provides an overview of where this work fits within theoretical computer science and game theory, but it is not required reading before diving into a specific chapter. Chapter 2 collects definitions used throughout the document. Readers already well-versed in the field may choose to skim it or skip it entirely, but they may still find it useful as a reference whenever notation or concepts become unclear. 1.9. HOW TO READ THIS THESIS23 Constrained existence problem RSEs Simple stochastic games XRSEs all optimistsP-complete general caseNP-complete ERSEs positional strategies rational base ∃R-easy NP-hard base 푒3EXPTIME-easy stationary strategies rational base ∃R-complete base 푒3EXPTIME-easy pure strategiesundecidable [UW11a] general caseundecidable [UW11a] SSEs objectives specified by parity automata EXPTIME-complete ParityPSPACE-complete SPEs Energyundecidable Discounted-sum TDS-hard coRE Mean-payoffNP-complete ParityNP-complete NEs Energy undecidable RE Discounted-sum TDS-hard coRE Mean-payoffNP-complete [UW11a] ParityNP-complete [Umm08] Achaotic rational verification SPEsMean-payoffP NP -complete Figure 2: Our main complexity results 24CHAPTER 1. INTRODUCTION Chapter 3 Chapter 5 Chapter 4 Chapter 6Chapter 7 Chapter 8Chapter 9Chapter 10Chapter 11 Figure 3: Dependences between chapters Similarly, Chapter 3 presents results on Nash equilibria but also introduces the game classes studied in later chapters. Those already familiar with parity, mean-payoff, discounted-sum, or energy games may skip it, but for others, reading it before Part I is strongly recommended. Within Part I, Chapter 4 should be read before Chapter 5, which in turn should be read before Chapter 6. Chapter 7 is more self-contained, at least for a reader who is familiar with SPEs—otherwise, the reader should read Chapter 4 beforehand. Chapter 8 builds on results from previous chapters, especially Chapter 6, but can be read independently from Chapter 7, provided the reader accepts the stated results without proof. Chapter 9 is relatively independent, though reading Chapter 3 beforehand is advisable, especially for those unfamiliar with parity objectives. Finally, Chapters 10 and 11 are independent of the rest of the document, but Chapter 10 should be read before Chapter 11. Those dependences are depicted by Figure 3: a chapter is blue if it belongs to Part I, red if it belongs to Part I, and green if it belongs to Part IV, and an arrow between Chapter푖and Chapter푗indicates that Chapter푖should be read before reading Chapter푗—a dotted arrow indicates that reading is advised but not essential. The transitivity of the relation "should be read before" has been used to omit some arrows. Proofs.Driven by the conviction that the proof is always at least as interesting as the result itself, we have chosen to present complete proofs directly alongside each stated result. Of course, skipping a proof is always an option. Though we have done our best to avoid it, there are instances where a proof references another result or construction from within a different proof. In such cases, we explicitly indicate which proof should be read for a full understanding. The proofs in this document fall into four broad categories: 1.short proofs that merely verify a result without introducing particularly interesting arguments, such as that of Lemma 6; 2. short proofs that present key arguments, as in Theorem 3, or that conclude previous results, as in Lemma 8; 3.long (and sometimes intricate) proofs that establish an unsurprising result using unsurprising 1.9. HOW TO READ THIS THESIS25 arguments, but require significant technical development, such as Theorem 9, or, to an extreme degree, Lemma 17; 4.long proofs that are genuinely worth reading, as they introduce original ideas and non-trivial arguments. Readers may prefer to skip proofs in categories 1 and especially 3, while focusing on those in categories 2 and 4. Naturally, some proofs fall between these categories or contain sections of varying interest. Typically, hardness proofs based on reductions are most valuable for their construction steps, which we always aim to present in an intuitive manner, allowing the reader to grasp the essence of the reduction without needing to follow the full proof. The latter part of such proofs, establishing the correctness of the reduction, almost always falls into category 3. Subjectively, the author would consider the nucleus of category 4 to include the proofs of Theorems 11, 15, 18, 23, 25, 31, 32 and 34. A practical heuristic for distinguishing between proofs in categories 3 and 4 without reading them in full is as follows: the more a proof is filled with mathematical symbols and equations, the more technical it tends to be—and, in most cases, the less it introduces new conceptual insights. However, exceptions exist, and this heuristic may not align perfectly with every reader’s perspective. Now that all these introductory remarks have been made, we wish the reader a pleasant journey. 26CHAPTER 1. INTRODUCTION PART I: BACKGROUND 27 CHAPTER 2: BACKGROUND 2.1WRITING CONVENTIONS From this point, we assume that the reader is familiar with the basics of algebra, language theory, graph theory, algorithmics, and complexity theory. However, we recall some fundamental notions to establish consistency in notation. Sets of numbers. We writeN,Z,Q, andRfor the sets of, respectively, natural numbers, integers, rational numbers, and real numbers, all defined with Zermelo-Fraenkel’s theory of sets. The definition of the setNis classically contentious: although the English-speaking world, and specifically secondary education, often considers that natural numbers start with 1, the convention used in the French-speaking world, according to which 0 is also a natural number, tends to take over in the field of computer science and combinatorics. We will therefore use the latter, and write N =0, 1, 2, . . .. Tuples.For a set푋, and an indexing (usually finite, but not always) set퐼, when we have defined an element푥 푖 ∈ 푋for every푖 ∈ 퐼, we usually use the bar notation ̄ 푥 퐼 , or simply ̄ 푥 , for the tuple(푥 푖 ) 푖∈퐼 ∈ 푋 퐼 . Conversely, when we introduce a tuple ̄ 푥 퐼 , or ̄ 푥, the notation푥 푖 will refer to the element of index푖 ∈ 퐼. In some cases, we will need the double-bar ̄ ̄ 푥to denote tuples of tuples. We will implicitly use pointwise comparisons: when we write ̄ 푥 ≤ ̄ 푦, for two tuples ̄ 푥, ̄ 푦 ∈ 푋 퐼 , the reader should read푥 푖 ≤ 푦 푖 for each푖 ∈ 퐼. Tuples are often called vectors, typically when they have real values, or profiles when the set퐼is a set of players. When the set 퐼 is clear from the context, we write ̄ 0 for the vector(0) 푖∈퐼 . Words. A tuple ̄ 푥 ∈ 푋 푛 (resp. ̄ 푥 ∈ 푋 N ), for a given natural integer푛, is also called finite word (resp. infinite word) over푋. Then, we use the formalisms of language theory: the set푋is called alphabet, and the word ̄ 푥can be written푥 = 푥 0 . . .푥 푛−1 (resp.푥 = 푥 0 푥 1 . . .). The integer푛is called length of푥, and written|푥|; the length of an infinite word is+∞. To avoid confusion between length and cardinality, the cardinality of a set 푋 will always be written card푋 . When a tuple is introduced as a (finite or infinite) word푥(or a specific type of word, such as paths, histories, or plays, which will be defined later), we will still write푥 푘 for the(푘 +1)th element of푥. Moreover, we will write푥 ≤푘 , or푥 <푘+1 , for the (finite) prefix푥 0 . . .푥 푘 , and푥 ≥푘 , or푥 >푘−1 , for the (finite or infinite) suffix 푥 푘 푥 푘+1 . . . , for every integer 푘 . We write 푋 ∗ for the set of finite words, and 푋 휔 for the set of infinite words, over the alphabet 푋 . Complexity classes.We will make an extensive use of the classical complexity classes P,NP,coNP, PSPACE,EXPTIME, P NP , assuming that the reader is familiar with their definitions—the reader who is not will find a very complete overview in [Sip97]. All complexity classes are written in sans-serif capital 29 30CHAPTER 2. BACKGROUND letters, while problems are denoted using serif small capitals—for instance, the satisfiability problem is written asSat. For a given complexity classC, we callC-easy a decision problem that belongs toCand C-hard a decision problem to which everyC-easy problem can be reduced in polynomial time. For a given decision problem 푃 , we will write 푃-hard instead of푃-hard. Size. The complexity of an algorithm is defined as a function of the size∥푥∥of its input푥, itself defined as the number of bits that are necessary to describe푥by a word on the alphabet0, 1. Most of the time, we will not explain how the objects we consider are encoded by such a word, considering that there are canonical encodings with which the reader is familiar: for example, an integer푛is represented by its binary encoding, which uses⌈log 2 (푛+1)⌉bits. We give here some first definitions of sizes that may not be straightforward. The size of a rational number푟 = 푝 푞 , where푝,푞 ∈ Zare co-prime, is the quantity ∥푟∥ =1+⌈log 2 (|푝|+1)⌉+⌈log 2 (|푞|+1)⌉. The size of an irrational number is+∞. The size of infinite numbers is∥+∞∥ = ∥−∞∥ =1. The size of a tuple ̄ 푥 ∈ 푂 퐼 , where퐼is an index set and푂is a set of objects for which the notion of size has been defined, is the quantitycard퐼 + Í 푖∈퐼 ∥푥 푖 ∥. Similarly, the size of a function푓:퐼 ! 푋is the quantitycard퐼 + Í 푖∈퐼 ∥푓(푖)∥, and the size of a set푋 ⊆ 푂is the quantity card푋 + Í 푥∈푋 ∥푥∥. Players. We will give below a definition of games, in which players are abstract objects from an abstract set. However, in game theory, it is common to refer to players as persons rather than objects. We will therefore often abuse language by saying, for example, that a player wishes or intends to maximize some quantity, to say that the mentioned quantity is their payoff function. One advantage of this convention, aside from aiding intuition, is that it allows the writer to take full advantage of the flexibility of the English language, where people have genders, whereas objects do not. Thus, in examples and proofs, a gender will be implicitly assigned to all well-defined players. For instance, when an example involves players denoted as ◦ ,□, or^, players ◦ and^will be grammatically female, while player□will be grammatically male. Finally, when we refer to some undefined player of which we do not know the gender, for example when we write "let푖 ∈ ◦ ,□,^", we will use the neutral pronoun they to emphasize the fact that we are referring to some player, but not to an abstract object, for which we would have used the pronoun it. Middle dot. We sometimes use the middle dot·to denote an unspecified element. For example, if the pair ̄ 푧 is known to belong to the cartesian product 푋 ×푌 , then we write ̄ 푧 =(푥,·) to mean that there exists푦 ∈ 푌 such that ̄ 푧 =(푥,푦). Usual notations.We will endeavor to make sure that every mathematical notation is well defined or clear from the context. However, if we fail in this objective, the reader may refer to Table 2: for every line, when a mathematical object is written with the symbol presented on the left side, it is probably of the type indicated on the right side. Note that gothic capital letters are only used to denote players, while cursive capital letters are used for computing structures—games, memory structures, counter machines, automata. Operators.Operators are typically written in sans-serif typeface; for instance, the set of plays in a gameG ↾푣 0 is denoted byPlaysG ↾푣 0 . They start with an uppercase letter when their output is considered as a set, and with a lowercase letter otherwise. A notable exception is probabilistic operators, like the expectationE, for which we use the blackboard uppercase notation. A list of operators used in this thesis, along with the locations where they are introduced, can be found in Table 3. 2.1. WRITING CONVENTIONS31 푎,푏,푐,푑, . . .vertices in a graph (constants) or real numbers 푐,푑cycles 푒Euler’s number 푒 = 2.71828. . . 푓 function 푔,ℎhistories 푖, 푗players (variables) 푘,ℓ,푚,푛 integers 푝,푞,푟states of a machine, integers, or probabilities 푟reward function 푡 real number (target or threshold) or terminal vertex 푢,푣,푤vertices in a graph (variables) 푥,푦,푧real numbers, variables of a formula, or abstract objects 퐸 set of edges 퐾strongly connected component in a graph 푀risk measure 푄 set of states of a machine 푇 tree, set of terminal vertices 푈,푉,푊 sets of vertices 푋,푌,푍 abstract sets, sets of reals, or random variables 훼path 훼,훽,훾real numbers 훽 discount factor, or base of risk entropy 훿,휀(small) real numbers 훿transition function of an automaton 휅color function 휆requirement 휆 0 the vacuous requirement 휆 0 : 푣 7!−∞ 휇,휈 payoff functions 휈valuation 휋, 휒,휉plays 휌,휂runs of a machine 휌risk parameter or rationality concept 휎 푖 ,휏 푖 strategies 휑,휓formulae or functions 휔the first transfinite ordinal Δset of transitions of a machine p probability function of a stochastic game Aautomaton G,H games Kcounter machine M memory structure 픓,ℭ,픖,픚,픏,픇,프,픅,픗,픒 players (constants) ◦ ,□,^players (constants) Table 2: Notations used in this thesis 32CHAPTER 2. BACKGROUND abs 휆푖 (G) ↾푣 0 abstract negotiation gameDefinition 26 card푋cardinalSection 2.1 conc 휆푖 (G) ↾푣 c 0 concrete negotiation gameDefinition 27 휆ConsG ↾푣 (or 휆Cons(푣)) 휆-consistent playsDefinition 21 Conv푋convex hullSubsection 6.2.1 dev ̄ 푥휑 (G) ↾푣 d 0 deviator gameDefinition 41 el 푟 푖 (ℎ) (or el 푖 (ℎ))energy levelDefinition 15 ds 훽 푟 푖 (휋) (or ds 푖 (휋))discounted sumDefinition 14 first(훼)first vertexSection 2.2 HistG historiesSection 2.2 Hist 푖 Ghistories ending in푉 푖 Section 2.2 Ind ↾푣 0 (M) induced strategies (or strategy profiles)Definition 6 Inf(훼) vertices occurring infinitely oftenSection 2.2 last(ℎ)last vertexSection 2.2 mp 푟 푖 (ℎ) (or mp 푖 (ℎ))mean-payoffDefinition 13 mp 푟 푖 (휋) (or mp 푖 (휋))mean-payoffDefinition 13 nego negotiation functionDefinition 24 Occ(훼)vertices occurringSection 2.2 PlaysG playsSection 2.2 휆Rat −푖 G ↾푣 (or 휆Rat(푣)) 휆-rational strategy profilesDefinition 22 red 훽 휆푖 (G) ↾푣 r 0 reduced negotiation gameDefinition 33 (or Subsection 5.1.2) SConn(푉,퐸) strongly connected componentsSection 2.2 SCyc(푉,퐸) simple cyclesSection 2.2 Sta 푖 G ↾푣 0 stationary strategiesSection 2.4 Sta 푃 G ↾푣 0 stationary strategy profilesSection 2.4 Strat 푖 G ↾푣 0 strategiesDefinition 4 Strat 푃 G ↾푣 0 strategy profilesDefinition 4 val 푖 G ↾푣 0 (or val(푣 0 )) valueSection 2.5 ⌞ 푋 downward sealingDefinition 30 E P [푋] expectationSubsection 10.2.1 OM P [푋]optimistic risk measureDefinition 50 M P 훽휌 [푋] entropic risk measureDefinition 48 P(퐸) probabilitySubsection 10.2.1 PM P [푋]pessimistic risk measureDefinition 50 X 푖 [푋] extreme risk measureSubsection 11.1.1 Table 3: Operators 2.2. GRAPHS GAMES33 2.2GRAPHS GAMES Throughout this thesis, we will use the word game for infinite-duration turn-based quantitative games with complete information played on graphs. Graphs and paths.A graph is an ordered pair(푉,퐸), where푉is a (often finite, but not always) set of vertices and퐸 ⊆ 푉 ×푉is a set of edges. For the simplicity of writing, an edge(푣,푤) ∈ 퐸will often be written푣푤. A path in(푉,퐸)is a word훼 = 훼 0 훼 1 . . ., finite or infinite, over the alphabet푉, such that for every푘such that훼 푘 and훼 푘+1 exist, we have훼 푘 훼 푘+1 ∈ 퐸. Given a path훼, we writeOcc(훼)for the set of vertices that appear in훼, andInf(훼)for the set of vertices that appear infinitely often in훼(which is empty if 훼 is finite). We write first(훼) for the first vertex of 훼 , and last(훼) for the last vertex of 훼 (if 훼 is finite). We say that the path훼is simple if no vertex occur more than once. It is a cycle if it is finite, and if we havelast(훼)first(훼) ∈ 퐸. A lasso is a path of the form훼푐 휔 , where훼is finite and푐is a cycle. A simple lasso is a lasso 훼푐 휔 such that the path 훼푐 is simple. Given a graph(푉,퐸), we denote bySCyc(푉,퐸)the set of simple cycles in(푉,퐸). Similarly, we denote bySConn(푉,퐸)the set of strongly connected components in(푉,퐸), i.e., the set of graphs(푉 ′ ,퐸 ′ )with 푉 ′ ⊆ 푉and퐸 ′ = 퐸∩(푉 ′ ×푉 ′ )such that for every two vertices푢,푣 ∈ 푉 ′ , there is a finite path훼with |훼| ≥ 2 from푢 to 푣 . Games.A game is a graph equipped with players, each of them controlling some of the vertices, and expressing preferences using a payoff function. A play is then an infinite path in the graph, which can be seen as an infinite sequence of moves of a token on its vertices: when the token is on some vertex, the player controlling that vertex chooses to which vertex it moves, following one of the outgoing edges of that vertex, and so on, each play being associated with a payoff for each player. Definition 1 (Non-initialized game). A non-initialized game is a tupleG = ( Π,푉,(푉 푖 ) 푖∈Π ,퐸,휇 ) , with: • a finite setΠ of players; •a graph(푉,퐸), called the underlying graph ofG, in which every vertex has at least one outgoing edge; • a partition(푉 푖 ) 푖∈Π of푉 , in which푉 푖 is the set of vertices controlled by player 푖; •a mapping휇called payoff function, that maps each infinite path휋to the tuple휇(휋) =(휇 푖 (휋)) 푖∈Π ∈ R Π of the players’ payoffs. A game is called prefix-independent if for every finite pathℎand every infinite path휋satisfying last(ℎ)first(휋) ∈ 퐸, we have휇(ℎ휋) = 휇(휋). It is called Boolean if for every infinite path휋and each player푖, we have휇(휋) ∈ 0,1. In such a context, we say that player푖wins the play휋if휇 푖 (휋) =1, and loses otherwise. When a gameGis given, we will often use the notationsΠ,푉,퐸,(푉 푖 ) 푖 , and휇, without necessarily recalling them. We sometimes abuse notations and assimilate a game to its underlying graph, for example writing SCyc(G) instead of SCyc(푉,퐸). Definition 2 (Initialized game). An initialized game is a tuple(G,푣 0 ), often writtenG ↾푣 0 , whereGis a non-initialized game and 푣 0 ∈ 푉 is a vertex called initial vertex. When the context is clear, we often use the word game for both initialized and non-initialized games. 34CHAPTER 2. BACKGROUND 휇 ◦ = 휇 □ = 0 푎 휇 ◦ = 휇 □ = 0 푏 휇 ◦ = 휇 □ = 1 푐 Figure 4: A game Example 1. We give in Figure 4 a first example of game with two players, player ◦ and player□. Circled vertices belong to the former, squared ones to the latter. In this example, both players get the payoff 0 in the path 푎 휔 and in every path of the form 푎 푘 푏 휔 , and the payoff 1 in paths of the form 푎 푘 푏 ℓ 푐 휔 . Definition 3 (Play, history). A play (resp. history) in the non-initialized gameGis an infinite (resp. finite) path in the graph(푉,퐸). A play (resp. history) in the initialized gameG ↾푣 0 is a play (resp. history) path in G whose first vertex is 푣 0 . The set of plays (resp. histories) in the gameG(resp. the initialized gameG ↾푣 0 ) is denoted byPlaysG (resp.PlaysG ↾푣 0 , HistG, HistG ↾푣 0 ). We writeHist 푖 G(resp.Hist 푖 G ↾푣 0 ) for the set of histories inG(resp. G ↾푣 0 ) of the formℎ푣, where푣is a vertex controlled by player푖. We also writeHist 푃 G = Ð 푖∈푃 Hist 푖 G(and Hist 푃 G ↾푣 0 = Ð 푖∈푃 Hist 푖 G ↾푣 0 ) for every 푃 ⊆Π. In order to obtain the plays they desire, the players choose their actions following a given strategy. 2.3STRATEGIES, STRATEGY PROFILES Definition 4 (Strategy, strategy profile). A strategy for player푖in the initialized gameG ↾푣 0 is a mapping 휎 푖 :Hist 푖 G ↾푣 0 ! 푉 , such that푣휎 푖 (ℎ푣)is an edge of(푉,퐸)for every historyℎ푣. A historyℎis compatible with a strategy휎 푖 if and only ifℎ 푘+1 = 휎 푖 (ℎ 0 . . .ℎ 푘 )for all푘such thatℎ 푘 ∈ 푉 푖 . A play휋is compatible with 휎 푖 if all its finite prefixes are. A strategy profile for푃 ⊆Πis a tuple ̄ 휎 푃 = (휎 푖 ) 푖∈푃 , where each휎 푖 is a strategy for player푖inG ↾푣 0 . A play or a history is compatible with ̄ 휎 푃 if it is compatible with each휎 푖 for푖 ∈ 푃. A complete strategy profile, typically written ̄ 휎, is a strategy profile forΠ. Exactly one play is compatible with the strategy profile ̄ 휎 : we call it its outcome and write it⟨ ̄ 휎⟩. When푖is a player and when the context is clear, we will often write−푖for the setΠ\푖. When ̄ 휏 푃 and ̄ 휏 ′ 푃 ′ are two strategy profiles with푃∩푃 ′ =∅, we write( ̄ 휏 푃 , ̄ 휏 ′ 푃 ′ ) for the strategy profile ̄ 휎 푃∪푃 ′ such that 휎 푖 = 휏 푖 for푖 ∈ 푃, and휎 푖 = 휏 ′ 푖 for푖 ∈ 푃 ′ . In a strategy profile ̄ 휎 푃 , the휎 푖 ’s domains are pairwise disjoint. Therefore, we can consider ̄ 휎 푃 as one function: forℎ푣 ∈ HistG ↾푣 0 such that푣 ∈ Ð 푖∈푃 푉 푖 , we liberally write ̄ 휎 푃 (ℎ푣)for휎 푖 (ℎ푣)with푖such that푣 ∈ 푉 푖 . The set of strategies for player푖(resp. of strategy profiles for the set 푃 ) inG ↾푣 0 is written Strat 푖 G ↾푣 0 (resp. Strat 푃 G ↾푣 0 ). For a given strategy profile ̄ 휎forΠ, for a given player푖, a strategy휎 ′ 푖 will often be called deviation for player 푖 from ̄ 휎 , or sometimes simply deviation from 휎 푖 , when it is seen as an alternative to 휎 푖 . Note that our formalism does not allow the players to use randomization to decide which edge they take. Randomized strategies will only be considered in Chapters 10 and 11, in which we will give a more general definition. 2.4. STATIONARY AND FINITE-MEMORY STRATEGIES35 2.4STATIONARY AND FINITE-MEMORY STRATEGIES As we have defined them, strategies can use all the information of a history to decide which edge must be taken: they are not limited in the memory they use. Such strategies cannot, in general, be simulated perfectly by a physical computer; and they are often hard to handle in algorithms, since there is no straightforward way to describe them with a finite number of bits. In contrast, stationary strategies are much simpler objects. Definition 5. A strategy휎 푖 is stationary if for every two histories푔푢,ℎ푢 ∈ Hist 푖 G ↾푣 0 , we have휎 푖 (푔푢) = 휎 푖 (ℎ푢). When the strategy휎 푖 is stationary, we will often consider that it is defined in every gameG ↾푢 , and simply write휎 푖 (푢)for휎 푖 (ℎ푢). We will also writeG[휎 푖 ]for the non-initialized game obtained by removing all edges푢푣with푢 ∈ 푉 푖 and푣 ≠ 휎 푖 (푣 푖 ), andG ↾푣 0 [휎 푖 ]for the same game, but initialized in푣 0 , and where all vertices that are no longer accessible from푣 0 have been omitted. Finally, for every gameG ↾푣 0 and each player푖, we writeSta 푖 G ↾푣 0 , orStaG ↾푣 0 when the context is clear, for the set of stationary strategies for player 푖 inG ↾푣 0 . As an intermediary object, we sometimes consider finite-memory strategies, i.e., strategies that are induced by a memory structure. Definition 6 (Memory structure). A memory structure for player푖on a gameGis a tupleM =(푄,푞 0 ,Δ), where푄is a finite set of states, where푞 0 ∈ 푄is the initial state, and whereΔ ⊆ (푄×푉 −푖 ×푄)∪(푄×푉 푖 ×푄×푉) is a finite set of transitions, such that: • for every(푝,푢,푞,푣) ∈Δ, we have푢푣 ∈ 퐸; •and for every푝 ∈ 푄and푢 ∈ 푉, there exists a transition(푝,푢,푞)or(푝,푢,푞,푣) ∈Δ. The memory structureM is deterministic if for each 푝 and푢, there exists exactly one such transition. A strategy휎 푖 inG ↾푣 0 is induced byMif there exists a mappingℎ 7! 푞 ℎ that maps every history ℎinG ↾푣 0 to a state푞 ℎ ∈ 푄, such that for everyℎ푣 ∈ Hist −푖 G ↾푣 0 , we have(푞 ℎ ,푣,푞 ℎ푣 ) ∈Δ, and for everyℎ푣 ∈ Hist 푖 G ↾푣 0 , we have(푞 ℎ ,푣,푞 ℎ푣 ,휎 푖 (ℎ푣)) ∈Δ. The set of strategies inG ↾푣 0 induced byMis writtenInd ↾푣 0 (M) . IfMis deterministic, then there is exactly one strategy induced byM; we call it a finite-memory strategy. Remark 1. Stationary strategies are exactly the strategies that are induced by a memory structure with one state. Specialist readers may have noted that the definition of deterministic memory structures, outside of the context of game theory, corresponds to the classical notion of Mealy machine. Results about deterministic memory structures can be applied to programs, which are supposed to run deterministically; we chose to take a more general definition to capture also protocols, which may be given to an agent who would still have some room for manoeuvre in how they apply it. We define analogously memory structures that induce a set of strategy profiles for several players, including for the whole setΠ. Note that, from such a memory structure inducing a set of strategy profiles 푆, one can extract a memory structure inducing the projection휎 푖 | ̄ 휎 ∈ 푆for some given player푖, by replacing all transitions of the form(푝,푢,푞,푣)with푢 ∉ 푉 푖 by the transition(푝,푢,푞). Note also that every memory structureM can be encoded with a finite number of bits: we write∥M∥ for that number. 36CHAPTER 2. BACKGROUND 푞 0 푞 1 푎 푎 푏|푏 푏|푐 푐 푏|푏 푐 Figure 5: A non-deterministic one-player memory structure 푞 0 푞 1 푏|푏 푎|푎 푐|푐 푎|푎 푏|푐 푐|푐 Figure 6: A deterministic multiplayer memory structure Example 2. Figure 5 depicts a one-player memory structure on the game of Figure 4. Each arrow from a state푝to a state푞labeled푢|푣denotes the existence of a transition(푝,푢,푞,푣)(from the state푝, the machine reads the vertex푢, switches to the state푞and outputs the vertex푣). Each arrow from a state푝to a state푞 labeled푢denotes the existence of a transition(푝,푢,푞)(from푝, the machine reads푢, switches to푞and outputs nothing). It is a machine for player□, that is not deterministic: from the state푞 0 , reading the vertex푏, the machine stays in푞 0 but it can output either푏or푐. The strategies that are compatible with it can be described as follows: when player □ has to play, if the vertex 푎 was seen an odd number of times, then he stays in 푏; in the opposite case, he can either stay in 푏 or eventually go to 푐. Figure 6 depicts a deterministic multiplayer memory structure on the same game. The strategy profile that is compatible with it loops on the vertex푎, and after a possible deviation of player ◦ that would lead to the vertex 푏, loops once on 푏, before going to 푐. 2.5TWO-PLAYER ZERO-SUM GAMES Although this document studies multiplayer games, we will sometimes use methods that bring us back to the more classical framework of two-player zero-sum games. We will therefore need the following notions and results. Definition 7 (Zero-sum game). A gameG, withΠ =푖, 푗, is zero-sum if 휇 푗 =−휇 푖 . Most of the games we will consider, be they zero-sum or not, are Borel games. Definition 8 (Borel game). A gameGis Borel if the function휇, from the set푉 휔 equipped with the product topology to the Euclidian space R Π , is Borel, i.e., if for every Borel set 퐵 ⊆ R Π , the set 휇 −1 (퐵) is Borel. Two-players zero-sum Borel games have the following important property, called determinacy. Lemma 1 ([Mar75]). LetG ↾푣 0 be a zero-sum Borel game, withΠ = 푖, 푗. Then, we have the following equality: sup 휎 푖 inf 휎 푗 휇 푖 ⟨ ̄ 휎⟩ = inf 휎 푗 sup 휎 푖 휇 푖 ⟨ ̄ 휎⟩. That quantity is called adversarial value, or simply value, of the gameG ↾푣 0 , denoted byval 푖 G ↾푣 0 . Solving the gameG ↾푣 0 means computing its value. The concept of value can also be extended to multiplayer games: whenG ↾푣 0 is a multiplayer game and푖is a player, the quantityval 푖 G ↾푣 0 is the value of the game obtained by merging all the other players into one fictional adversarial player, i.e. a player whose payoff function 2.6. EQUILIBRIA AND CONSTRAINED EXISTENCE PROBLEM37 would be the opposite of player푖’s one. In other words, the quantityval 푖 G ↾푣 0 is the best payoff that player 푖can ensure in the gameG ↾푣 0 , whatever the other players do. WhenGis clear from the context and when 푣 0 ∈ 푉 푖 , we writeval(푣 0 )forval 푖 G ↾푣 0 , and call it simply value of the vertex푣 0 . When it exists, a strategy that ensures the payoffval 푖 G ↾푣 0 for player푖is called optimal. IfG ↾푣 0 is a Boolean game and ifval 푖 G ↾푣 0 =1, it will also be called a winning strategy. Solving two-player zero-sum games will often be much simpler in games where stationary optimal strategies exist. Let us give a sufficient condition to identify such games. Definition 9 (Shuffling). Let휋, 휒and휉be three plays in a gameG. The play휉is a shuffling of휋 and휒if there exist two infinite sequences of indices푘 0 < 푘 1 < . . .andℓ 0 < ℓ 1 < . . .such that 휒 0 = 휋 푘 0 = 휒 ℓ 0 = 휋 푘 1 = . . . , and: 휉 = 휋 0 . . .휋 푘 0 −1 휒 0 . . . 휒 ℓ 0 −1 휋 푘 0 . . .휋 푘 1 −1 휒 ℓ 0 . . . 휒 ℓ 1 −1 . . . . Definition 10 (Convexity, concavity). A payoff function휇 푖 :PlaysG ! Ris convex if for every shuffling 휉 of two plays 휋 and 휒 , we have 휇 푖 (휉) ≥ min휇 푖 (휋), 휇 푖 (휒). It is concave if the function−휇 푖 is convex. Lemma 2. In a two-player zero-sum game played on a finite graph, every player whose payoff function is concave has an optimal strategy that is stationary. Proof. According to [Kop06], this result is true for Boolean objectives. It follows that for every푥 ∈ R, if a player푖, whose payoff function is concave, has a strategy that ensures휇 푖 (휋) ≥ 푥(understood as a Boolean objective), then they have a stationary one. Hence the equality: val 푖 G ↾푣 0 =sup 휎 푖 ∈StaG ↾푣 0 inf 휎 푗 휇 푖 ⟨ ̄ 휎⟩. Since the underlying graph(푉,퐸)is assumed to be finite, there exists a finite number of stationary strategies, hence there exists a stationary strategy 휎 푖 that realizes the infimum above.□ 2.6EQUILIBRIA AND CONSTRAINED EXISTENCE PROBLEM In a game, among the strategy profiles that are available for all players, we will be interested in specific classes of strategy profiles that can be considered as more stable than others, because of guarantees they maintain when players deviate from their strategies. Such classes of strategy profiles are usually called equilibria. Our contribution focuses mostly on Nash equilibria and subgame-perfect equilibria, and later on strong secure equilibria and risk-sensitive equilibria. We will also study those notions in different classes of games. However, the question that we ask will almost always be the same: does there exist an equilibrium, in a given game, that generates a payoff in a given interval for each player? We give here a formal definition of that problem for some classCof games, and for some general classEof strategy profiles, that must be understood as some notion of equilibrium. Problem 1 (Constrained existence problem). Given a gameG ↾푣 0 ∈ Cand two vectors ̄ 푥, ̄ 푦 ∈ (Q∪±∞) Π , does there exist a strategy profile ̄ 휎 ∈ E such that the inequality ̄ 푥 ≤ 휇⟨ ̄ 휎⟩ ≤ ̄ 푦 holds? The quantities푥 푖 are called lower thresholds. They are said to be effective when they are different from −∞. Similarly, the quantities푦 푖 are called upper thresholds, and are considered effective when they are 38CHAPTER 2. BACKGROUND different from+∞. Note that thresholds are assumed to be rational, even in games where payoffs may not be, so that they can be represented using a finite number of bits. Most of this document is dedicated to studying the complexity of this problem, for various equilibrium notionsE and game classesC. PART I: NASH EQUILIBRIA 39 CHAPTER 3: NASH EQUILIBRIA The concept of Nash equilibria, though under a different name, appears to have first been introduced by the French mathematician Antoine-Augustin Cournot [Cou38]. Later, John Nash provided a more general definition [Nas51] in a setting quite different from ours, as he was not considering graph games but matrix games, in the sense of Section 1.4. In such games, he proved that what is now known as a Nash equilibrium always exists, provided that players are allowed to randomize over multiple actions—a behavior that is implicitly prohibited in our formalism. This result is now widely recognized as Nash’s theorem. Both Cournot’s and Nash’s work were motivated by economics, where their insights remain extensively used today. For instance, if a strategy profile represents the behavior of competing firms, a Nash equilibrium corresponds to a stable situation: if those companies want to coordinate but do not trust each other, following an agreement that constitutes a Nash equilibrium ensures that no company will have an incentive to deviate from that agreement. As multiplayer games became a topic of interest in computer science, researchers naturally focused on this notion. The complexity of the constrained existence problem of Nash equilibria was therefore already known for several of the most classical classes of games that we study here. We use however this chapter to present some of those results (Theorems 1 and 2), prove new ones (Theorems 3, 4 and 6), and introduce the classes of games that we will also study in Part I. 3.1DEFINITION Nash equilibria are strategy profiles from which no player can increase their payoff by deviating unilaterally. Definition 11 (Nash equilibrium). LetG ↾푣 0 be a game. A strategy profile ̄ 휎is a Nash equilibrium, or NE for short, in the gameG ↾푣 0 if and only if for each player푖and for every deviation휎 ′ 푖 for player푖, we have the inequality 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ 휇 푖 ⟨ ̄ 휎⟩. When the strategy ̄ 휎is not an NE, we call the deviations휎 ′ 푖 that do not observe the inequality above profitable deviations. 3.2PARITY GAMES Let us consider a system required to ensure a given condition, typically called a specification, over the long run. A concrete example would be a robot tasked with regularly delivering parcels arriving at a warehouse to an apartment block. As a first abstraction, we assume that the robot has infinite carrying capacity. At any given moment, it therefore has two available actions: either do nothing, or take all the parcels in the warehouse and deliver them to the apartment block. A second abstraction consists in disregarding the delivery time, as long as 41 42CHAPTER 3. NASH EQUILIBRIA every parcel that appears in the warehouse is eventually delivered. A third abstraction, which is central to all our formalisms, is to assume that the robot operates indefinitely and can thus plan its strategy over an infinite horizon. The company managing the robot and the warehouse may then choose to care about their customers only as long as they continue placing orders: if they stop ordering, delivering the remaining parcels is no longer necessary. In this framework, the robot’s objective can be expressed in a simple way: if parcels appear in the warehouse infinitely often, then the robot must infinitely often carry them to the apartment block. Now, let us extend this setting by introducing a castle, which may also receive deliveries. Since the castle owner has subscribed to a premium service, deliveries to the castle take precedence over those to the apartment block. The robot’s objective then becomes the following: if parcels for the castle appear in the warehouse infinitely often, the robot must infinitely often deliver them to the castle. Otherwise, and only in that case, it must fulfill the previously defined objective. We observe that the robot’s objective is defined using priorities among different possible situations. A natural and elegant way to formalize such an objective is through a parity condition, a concept widely used in automata theory. Indeed, it is well known that휔-regular languages are precisely those recognized by a non-deterministic automaton whose acceptance condition is a parity condition. 3.2.1 Definition Given a graph(푉,퐸), a color mapping on(푉,퐸) is a mapping 휅 : 푉 ! N. Definition 12 (Parity). The gameGis a parity game if for each player푖, there exists a color mapping휅 푖 on the underlying graph ofGsuch that for every play휋, we have휇 푖 (휋) =1 if the colormin 푣∈Inf(휋) 휅 푖 (푣) is even, and 휇 푖 (휋) = 0 if it is odd. A Büchi game is a parity game in which all colors are 0 or 1: then, each player’s objective consists simply in visiting infinitely often vertices of color 0. Dually, a co-Büchi game is a parity game in which all colors are 1 and 2: then, each player’s objective consists in, eventually, avoiding vertices of color 1. 3.2.2 Results The constrained existence problem of Nash equilibria in parity games has already been studied by Michael Ummels. Theorem 1 ([Umm08]). The constrained existence problem of Nash equilibria in parity games isNP-complete. Hardness still holds in co-Büchi games, and when there is no effective lower threshold, and only one effective upper threshold. In Büchi games, the problem is in P. Actually, in the cited article, Michael Ummels shows hardness with one effective lower (and not upper) threshold, but the result given here can be quickly showed by adding one player to his reduction—and we will need that result under this form in Chapter 8. 3.3MEAN-PAYOFF GAMES Imagine a system designed to generate a specific resource—an automated solar plant, for example, that produces green hydrogen. The plant’s objective would not be to maintain continuous hydrogen production. 3.4. DISCOUNTED-SUM GAMES43 At times, such as when there is no sunlight, production may halt, but that downtime could be used for maintenance, like cleaning the equipment. If the plant is located outside Belgium, where long periods of sunlight are possible, maintenance would still be necessary, even if it temporarily reduces production. What matters, however, is the average production over time—periods of peak output can compensate for necessary pauses. Similarly, an initially long phase of low or no production may not be an issue if it ultimately leads to higher efficiency in the long run. The key concept here is the asymptotic average production, which is precisely what mean-payoff objectives capture. 3.3.1 Definition Contrary to parity games, mean-payoff games are typically quantitative games. The players receive a reward at each turn, and they aim at maximizing their limit average reward. Given a graph(푉,퐸), a reward function on(푉,퐸) is a mapping 푟 : 퐸 ! Q. Definition 13 (Mean-payoff ). In a graph(푉,퐸), we define for each reward mapping푟the mean-payoff functionmp 푟 :ℎ 0 . . .ℎ 푛 7! 1 푛 Í 푘 푟 ( ℎ 푘 ℎ 푘+1 ) . Then, we writemp 푟 (휋) = lim inf 푛 mp 푟 (휋 ≤푛 ).The gameGis a mean-payoff game if there exists a tuple(푟 푖 ) 푖∈Π of reward mappings such that for each player푖, we have 휇 푖 = mp 푟 푖 . When the context is clear, we writemp 푖 formp 푟 푖 , andmp 푖 formp 푟 푖 . We then also writemp(휋)for the tuple(mp 푖 (휋)) 푖∈Π . 3.3.2 Results Again, the constrained existence problem of Nash equilibria in mean-payoff games has already been studied by Michael Ummels. Theorem 2 ([UW11a]). The constrained existence problem of Nash equilibria in mean-payoff games is NP-complete. Hardness still holds when all rewards are 0 and 1, and when there is no effective lower threshold and only one effective upper threshold. Again, Michael Ummel’s proved hardness with one lower threshold, but a quick modification of his proof gives us the desired result. 3.4DISCOUNTED-SUM GAMES Let us return to our example of an automated solar plant producing hydrogen, which we used to illustrate mean-payoff games. The assumption that only the asymptotic average reward matters may seem bold: in practice, as time passes, uncertainties grow, and one might prefer to secure reasonable short-term rewards rather than rely on promises of greater gains in the distant future: a bird in the hand is worth two in the bush. This idea is captured by discounted-sum objectives. 3.4.1 Definition Definition 14 (Discounted-sum). In a graph(푉,퐸), we define for each reward mapping푟and each discount factor훽 ∈ (0,1)the discounted-sum functionds 훽 푟 :ℎ 7! Í 푘 훽 푘 푟(ℎ 푘 ℎ 푘+1 ). Then, we write ds 훽 푟 (휋) = lim 푛 ds 훽 푟 (휋 ≤푛 ).The gameGis a discounted-sum game if there exists a discount factor훽 ∈ (0,1)∩Q 44CHAPTER 3. NASH EQUILIBRIA 푎 푏 □ 푏 □ 푎 □ 푏 □ 푎 Figure 7: A game constructed from an instance of the TDS problem and a tuple(푟 푖 ) 푖∈Π of reward mappings such that for each푖and every휋, we have휇 푖 (휋) = ds 훽 푟 푖 (휋) . When the context is clear, we write ds 푖 for ds 훽 푟 푖 . 3.4.2 The target discounted-sum problem Discounted-sum objectives are closely related to the following problem, which is as simple to state as it is difficult to analyze. Problem 2 (Target Discounted-Sum Problem). Given four quantities훽,푎,푏,푡 ∈ Qwith 0< 훽<1, is there a sequence(푢 푛 ) 푛∈N ∈ 푎,푏 휔 such that Í 푛∈N 푢 푛 훽 푛 = 푡 ? Although this problem naturally arises in various fields, the Target Discounted-Sum (TDS) problem is surprisingly difficult to solve, and its decidability remains an open question. For further details, the interested reader may refer to [BHO15]. We will now show that the problem we wish to study is at least as challenging. 3.4.3 Hardness result Theorem 3. The TDS problem reduces to the constrained existence problem of NEs in discounted-sum games. This hardness result still holds when there is no effective lower threshold, and only one effective upper threshold. Proof.Let푎,푏,푡 ∈ Qand let훽 ∈ (0,1)∩ Qform an instance of the TDS problem. We construct from those inputs a discounted-sum gameG ↾푎 , depicted by Figure 7, with discount factor훽. Player ◦ ’s rewards are zero on all edges. In that game, we claim that there exists an NE ̄ 휎 with푡 ≤ 휇 □ ⟨ ̄ 휎⟩ ≤ 푡, if and only if 푎,푏,푡 , and 훽 form a positive instance of the TDS problem. • Indeed, if there exists an NE ̄ 휎satisfying푡 ≤ 휇 □ ⟨ ̄ 휎⟩ ≤ 푡, i.e.휇 □ ⟨ ̄ 휎⟩ = 푡, let휋 = ⟨ ̄ 휎⟩. Then, the sequence(푢 푛 ) 푛 =(휋 푛 ) 푛 is such that Í 푛 푢 푛 훽 푛 = 휇 □ (휋) = 푡 . •Conversely, let us note that player ◦ is the only player who actually makes choices, and that her payoff is 0 in every play. Consequently, every strategy profile in that game is an NE. Thus, if there exists a sequence(푢 푛 ) 푛 ∈ 푎,푏 N with Í 푛 푢 푛 훽 푛 = 푡, then that sequence also defines a play 휋 with 휇 □ (휋) = 푡 , and every strategy profile ̄ 휎 with⟨ ̄ 휎⟩ = 휋 is an NE. To establish hardness even with only one effective upper threshold, we can extend this reasoning to the slightly more involved construction depicted by Figure 8.□ 3.4.4 Co-recursive enumerability The previous theorem suggests that designing algorithms to solve these problems lies beyond the scope of this thesis. However, we will now show that, like the TDS problem, our problem is co-recursively enumerable. The key idea is as follows: a fundamental property of discounted-sum objectives is that when 3.4. DISCOUNTED-SUM GAMES45 푣 0 푣 1 푣 2 푣 3 푎 푏 ◦ −1 □ 푏 ⋄ −푏 ◦ −1 □ 푎 ⋄ −푎 ◦ −1 □ 푏 ⋄ −푏 ◦ −1 □ 푎 ⋄ −푎 □ 푡휆(1− 휆) ⋄ −푡(1− 휆) Figure 8: Another game constructed from the same instance of TDS a play yields a payoff outside a given interval for some player, there exists a prefix of that play after which the gap becomes irrecoverable. Beyond this point, we can stop reading the play and confidently conclude that the payoff will fall outside the interval. Therefore, although strategy profiles are generally infinite objects and exist in uncountably many variations, profitable deviations can be identified by analyzing their behavior over a finite (though unbounded) number of histories. Theorem 4. The constrained existence problem of NEs in discounted-sum games is co-recursively enumerable. Proof.We present here a semi-algorithm that recognizes negative instances of our problem. Given a discounted-sum gameG ↾푣 0 with discount factor훽, a player푖and two threshold vectors ̄ 푥 and ̄ 푦, we give an algorithm that stops if and only if there exists no NE ̄ 휎inG ↾푣 0 with ̄ 푥 ≤ 휇 푖 ⟨ ̄ 휎⟩ ≤ ̄ 푦. But first, let us give a preliminary result that will justify the correctness of our result. ▶Preliminary result We show here the following lemma. Lemma 3. For every play 휋, each player 푗 and every index 푛, we have: 휇 푗 (휋) ∈ ds 푗 ( 휋 ≤푛 ) − 푀훽 푛 , ds 푗 ( 휋 ≤푛 ) + 푀훽 푛 , where: 푀 = 1 1− 훽 max 푢푣∈퐸 max 푗∈Π |푟 푗 (푢푣)| is a bound on the payoff (in absolute value) of every player. Proof.Let us proceed by induction on 푛. Base case. When 푛 = 0, for every play 휋 starting from 푣 0 and each player 푗 , we have: 휇 푗 (휋) = ∑︁ 푘 푟 푗 (휋 푘 휋 푘+1 )훽 푘 ≤ ∑︁ 푘 훽 푘 max 푢푣∈퐸 max 푗∈Π |푟 푗 (푢푣)| = 푀 and symetrically 휇 푗 (휋) ≥−푀 , which constitutes the desired interval since ds 푖 (휋 0 ) = 0. 46CHAPTER 3. NASH EQUILIBRIA Inductive case.Now, if the desired property is true for푛 ≥0, let us prove that it is true for푛+ 1. Let푣휋be a play. By induction hypothesis, we have휇 푗 (휋) ∈ ds 푗 (휋 ≤푛 )− 푀훽 푛 , ds 푗 (휋 ≤푛 )+ 푀훽 푛 . Hence: 휇 푗 (푣휋) = 푟 푗 (푣휋 0 )+ ∑︁ 푘 푟 푗 (휋 푘 휋 푘+1 )훽 푘+1 = 푟 푗 (푣휋 0 )+ 훽휇 푗 (휋) ≥ 푟 푗 (푣휋 0 )+ 훽ds 푗 (휋 ≤푛 )− 푀훽 푛+1 = ds 푗 (푣 0 휋 ≤푛 )− 푀훽 푛+1 , and analogously for the upper bound.□ ▶Algorithm Let ℎ (푛) 푛∈N be a recursive enumeration of the nonempty histories inG ↾푣 0 by increasing order of lengths (we can, for example, order histories of the same length with the lexicographic order induced by some arbitrary order on vertices). Now, let푇be the infinite tree whose nodes of depth푛+1 are all possible푛-uples ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) , and where the children of the node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) are the nodes of the form: ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ), ̄ 휎(ℎ (푛+1) ) . Thus, every node partially defines a complete strategy profile ̄ 휎, and every infinite branch entirely defines it—and conversely, every complete strategy profile is defined by an infinite branch. A node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) is called Nash-irrational if there exist two indicesℓ,푚 ≤ 푛, with |ℎ (ℓ) | =|ℎ (푚) | = 푝, an index 푘 ≤ 푝 and a player 푖 such that: •we have ℎ (ℓ) ≤푘 =ℎ (푚) ≤푘 ; •the history ℎ (ℓ) is compatible with the strategy profile ̄ 휎 that is partially defined by the node; •the history ℎ (푚) is compatible with the strategy profile ̄ 휎 −푖 ; •and finally, we have: ds 푖 ℎ (푚) − ds 푖 ℎ (ℓ) > 2푀훽 푝−1 . The same node is called off-topic if for some player푖we haveds 푖 ( ℎ ) > 푦 푖 + 푀훽 |ℎ|−1 , ords 푖 ( ℎ ) < 푥 푖 − 푀훽 |ℎ|−1 whereℎis the longest history that is compatible with ̄ 휎as partially defined by the node. Our algorithm consists in constructing the tree푇 ′ , obtained from the tree푇by cutting every branch after the first Nash-irrational or off-topic node, and declaring the instance to be negative once that construction is finished. ▶Correctness To show that our algorithm is correct, we must prove that the tree푇 ′ will be finite if and only if we have a negative instance of the constrained existence problem. By Kőnig’s lemma, that will be the case if and only if every branch is finite, i.e., if and and only if every branch of the tree푇contains either an 3.5. ENERGY GAMES47 off-topic or a Nash-irrational node. Therefore, we will be done if we prove that, given a branch of푇 and the corresponding strategy profile ̄ 휎 , the branch contains no such node if and only if the strategy profile ̄ 휎 is an NE and satisfies ̄ 푥 ≤ 휇⟨ ̄ 휎⟩ ≤ ̄ 푦. Off-topic nodes. Using Lemma 3, we know that we have ̄ 푥 ≤ 휇 푖 ⟨ ̄ 휎⟩ ≤ ̄ 푦 if and only if the corresponding branch contains no off-topic node. If the branch contains a Nash-irrational node, then the strategy profile ̄ 휎is not an NE. If the branch contains a Nash-irrational node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) . Let us use the notations푘,ℓ,푚, and푖from the definition of Nash-irrationality. Let us define휋 =⟨ ̄ 휎⟩: note that the historyℎ (ℓ) is the prefix of length푝of the play휋. Similarly, let us extend the historyℎ (푚) into some play휋 ′ , compatible with the strategy profile ̄ 휎 −푖 . By Lemma 3, the inequality: ds 푖 ℎ (푚) − ds 푖 ℎ (ℓ) > 2푀훽 푝−1 implies 휇 푗 (휋 ′ )> 휇 푗 (휋), and therefore the strategy profile ̄ 휎 is not an NE. If the strategy profile ̄ 휎is not an NE, then the branch contains a Nash-irrational node. If ̄ 휎 is not an NE, then there exists a player 푖 and a strategy 휎 ′ 푖 such that we have: 휇 푖 ⟨ ̄ 휎⟩< 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩. Then, let휋 =⟨ ̄ 휎⟩, and let휋 ′ =⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩: since휇 푖 (휋)< 휇 푖 (휋 ′ ), by Lemma 3, there exists an index푘 such that: ds 푖 ( 휋 <푘 ) + 푀훽 푘−1 < ds 푖 휋 ′ <푘 − 푀훽 푘−1 . Let nowℓand푚be the indices such thatℎ (ℓ) = 휋 <푘 , andℎ (푚) = 휋 ′ <푘 . Then, we haveds 푗 ℎ (푚) − ds 푗 ℎ (ℓ) >2푀훽 푘−1 : along the branch corresponding to ̄ 휎, the node of depthmaxℓ,푚is Nash- irrational. Which ends the proof.□ 3.5ENERGY GAMES Let us revisit the example of our green hydrogen plant, but now assume that it is connected to another system that consumes the hydrogen. For instance, the delivery robot from Section 3.2, which arrives whenever necessary to collect fuel from a tank where the station stores the hydrogen it produces. The plant’s goal is no longer to maximize hydrogen production, which would after all be a rather productivist objective. Instead, its aim is to ensure that the robot will always find the required amount of fuel available— in other words, to avoid a situation in which the robot would attempt to take more fuel than what is stored in the tank. Such scenarios are modeled with energy objectives. 3.5.1 Definition Like mean-payoff and discounted-sum objectives, energy objectives are based on reward functions. How- ever, energy objectives are Boolean. 48CHAPTER 3. NASH EQUILIBRIA Definition 15 (Energy). In a graph(푉,퐸), we associate to each reward mapping푟the energy level function el 푟 that maps a finite path ℎ =ℎ 0 . . .ℎ 푛 to the element of N∪⊥ defined as follows. • If the path has length 1, we define el 푟 (ℎ 0 ) = 0. •Otherwise, if we haveel 푟 (ℎ ≤푛 ) ≠⊥andel 푟 (ℎ ≤푛 )+푟(ℎ 푛 ℎ 푛+1 ) ≥0, we defineel 푟 (ℎ ≤푛+1 ) = el 푟 (ℎ ≤푛 )+ 푟(ℎ 푛 ℎ 푛+1 ). • In every other case, we define el 푟 (ℎ ≤푛+1 ) =⊥. The gameGis an energy game if there exists a tuple(푟 푖 ) 푖∈Π of reward mappings such that for each player푖and every play휋, we have휇 푖 (휋) =0 if we haveel 푟 푖 (휋 ≤푛 ) =⊥for some푛, and휇 푖 (휋) =1 otherwise. When the context is clear, we write el 푖 for el 푟 푖 . 3.5.2 Two-counter machines We will see in Subsection 3.5.3 that energy games can be used to encode counter machines. Let us first define that model. Definition 16 (Two-counter machine). A two-counter machine is a tupleK = 푄,푞 0 ,푞 f ,Δ 퐶 + 1 ,Δ 퐶 + 2 ,Δ 퐶 − 1 ,Δ 퐶 − 2 , with: • a finite set 푄 of states; • an initial state 푞 0 ∈ 푄 ; • a final state 푞 f ; • two setsΔ 퐶 + 1 ,Δ 퐶 + 2 ∈ 푄×푄 of incremental transitions; • and two setsΔ 퐶 − 1 ,Δ 퐶 − 2 ⊆ 푄×푄×푄 of test transitions; such that every state푞 ∈ 푄\푞 f admits exactly one outgoing transition (either incremental or test), and such that the state푞 f admits none. We call sequence of transitions a finite or infinite sequence휂 0 휂 1 휂 2 . . . such that we have휂 0 =푞 0 and that for each푘 ∈ N, there is a counter퐶 ∈ 퐶 1 ,퐶 2 such that we have either (휂 푘 ,휂 푘+1 ) ∈Δ 퐶 + ,(휂 푘 ,휂 푘+1 ,·) ∈Δ 퐶 − , or(휂 푘 ,·,휂 푘+1 ) ∈Δ 퐶 − . We define the interpretation ˆ 퐶 of each counter퐶 on sequences of transitions as follows: • on the one-state sequence of transitions, we define ˆ 퐶(푞 0 ) = 0; • if we have(푞 푛−1 ,푞 푛 ) ∈Δ 퐶 + , we define ˆ 퐶(푞 0 . . .푞 푛 ) = ˆ 퐶(푞 0 . . .푞 푛−1 )+ 1; • if we have(푞 푛−1 ,푞 푛 ,·) ∈Δ 퐶 − , we define ˆ 퐶(푞 0 . . .푞 푛 ) = ˆ 퐶(푞 0 . . .푞 푛−1 )− 1; • if we have(푞 푛−1 ,·,푞 푛 ) ∈Δ 퐶 − , we define ˆ 퐶(푞 0 . . .푞 푛 ) = ˆ 퐶(푞 0 . . .푞 푛−1 ). A run of the machineKis a sequence of transition, either infinite or ending with푞 f , such that for every index 푘 : • if we have(휂 푘 ,휂 푘+1 ,·) ∈Δ 퐶 − , then we have ˆ 퐶(휂 0 . . .휂 푘 )> 0; • and if we have(휂 푘 ,·,휂 푘+1 ) ∈Δ 퐶 − , then we have ˆ 퐶(휂 0 . . .휂 푘 ) = 0. Note that each two-counter machine admits exactly one run. A machine halts if that unique run is finite. 3.5. ENERGY GAMES49 Figure 9: A machine that halts on(0, 0) 푞 0 푞 1 푞 f 퐶 + 1 퐶 − 1 퐶 1 = 0 Figure 10: A machine that does not halt on(0, 0) 푞 0 푞 1 푞 f 퐶 + 1 퐶 − 1 퐶 1 = 0 Figures 9 and 10 depict two examples of two-counter machines. The arrow with label퐶 + 1 from푞 0 to 푞 1 indicates a transition(푞 0 ,푞 1 ) ∈Δ 퐶 + 1 . The arrows with the label퐶 − 1 from푞 1 to itself and with the label 퐶 1 = 0 from 푞 1 to 푞 f indicate a transition(푞 1 ,푞 1 ,푞 f ) ∈Δ 퐶 − 1 . We will use the following well-known result: Theorem 5 ([Min61]). The problem of deciding whether a given two-counter machine halts is undecidable, and in particular not co-recursively enumerable. 3.5.3 Results In energy games, our problem is undecidable, since a counter machine can be encoded in an energy game. However, while we showed co-recursive enumerability in discounted-sum games, we will show that in energy games, it is recursively enumerable. Theorem 6. The constrained existence problem of Nash equilibria in energy games is undecidable, even with no effective lower threshold and only one effective upper threshold, and recursively enumerable. Proof.▶Undecidability. We show undecidability by reduction from the halting problem of a two-counter machine. LetKbe a two-counter machine. We define an energy gameG ↾푞 1 0 with five players (players퐶 ⊤ 1 , 퐶 ⊥ 1 ,퐶 ⊤ 2 ,퐶 ⊥ 2 , and픚, called Witness) by assembling the gadgets presented in Figure 11. The rewards that are not presented are equal to 0, and the players controlling relevant vertices are written in black boxes. For each state ofK, we define from one to two vertices, plus the additional vertex△. Then, a play inG ↾푞 1 0 that does not reach the vertex△simulates a sequence of transitions ofK, that is a valid run if and only if each time the test gadget is visited, the player퐶 ⊤ (with퐶 ∈ 퐶 1 ,퐶 2 ) controlling 푞goes to푞 ′ if and only if her energy level is 0. Then, at each step, the counter퐶 푖 is captured by the energy level of player퐶 ⊤ 푖 , always equal to the energy level of player퐶 ⊥ 푖 . Let us now prove that the gameG ↾푞 1 0 admits an NE where Witness loses if and only if the machineK terminates. If there is an NE where Witness loses, then the machineKterminates. Let ̄ 휎be an NE such that휇 픚 ⟨ ̄ 휎⟩ =0. Let휋 =⟨ ̄ 휎⟩: that play is then lost by Witness. Since the only edge that makes Witness lose energy is the edge푞 f 푞 f , we deduce that the play휋reaches the vertex푞 f , and therefore 50CHAPTER 3. NASH EQUILIBRIA 푞 1 0 퐶 ⊤ 1 푞 2 0 퐶 ⊤ 2 (a) Initial state 푞 f 퐶 ⊥ 1 −1 퐶 ⊥ 2 −1 픚 −1 (b) Final state 푞 퐶 ⊤ 1 퐶 ⊥ 1 (c) Incrementations 푞 퐶 ⊤ (if퐶> 0) 푞 ′ 퐶 ⊥ (if퐶 = 0) △ 퐶 ⊤ −1 퐶 ⊥ −1 퐶 ⊥ −1 (d) Tests Figure 11: Gadgets simulates a terminating sequence of transitions ofK(note that one transition may be represented by several edges). We must now prove that this run is valid, i.e. that tests are simulated correctly. •Let us assume that at some point along the play휋, from a vertex푞, player퐶 ⊤ 푖 , with푖 ∈ 1,2, does not take the edge to the vertex푞 ′ while her energy level is zero. Then, her energy level drops to⊥and she loses. Therefore, she has a profitable deviation at the beginning of the play, by looping on the vertex 푞 푖 0 . •Let us now assume that she goes to the vertex푞 ′ , while her energy is positive. Then, player퐶 ⊥ 푖 ’s energy is also positive: he can go to the vertex△and win. That would be a profitable deviation, since, as the play 휋 reaches the vertex 푞 f , it is lost by player퐶 ⊥ 푖 . Therefore, the play휋does not fake any test, and simulates correctly the machineK, which terminates. If the machineKterminates, then there is an NE where Witness loses.Then, let us define a strategy profile ̄ 휎inG ↾푞 1 0 as follows: in tests of counter퐶, player퐶 ⊤ goes to푞 ′ if and only if her energy level is positive; from 푞 ′ , player퐶 ⊥ never goes to the vertex△. Let 휋 =⟨ ̄ 휎⟩: since tests are simulated correctly, the play휋simulates the run ofK, and therefore reaches the vertex푞 f . It is therefore lost by Witness, as well as by the players퐶 ⊥ 1 and퐶 ⊥ 2 . Witness has no profitable deviation, since he does not control any vertex. The only vertices controlled by each player퐶 ⊥ are the vertices of the form푞 ′ . Those vertices are reached only when player퐶 ⊤ ’s, and therefore player퐶 ⊥ ’s energy level is zero. Then, deviating and going to the vertex△ is not profitable for player퐶 ⊥ , since it makes him immediately lose. The strategy profile ̄ 휎is therefore an NE, lost by Witness. Conclusion.Every NE outcome in the gameG ↾푞 1 0 is won by Witness if and only if the machineK does not terminate. Therefore, the halting problem of two-counter machines reduces to the constrained existence problem of NEs in energy games, which is therefore undecidable. 3.5. ENERGY GAMES51 푞 0 . . . 푞 푚 . . . 푞 푛−1 푞 ′ 푚 푞 ′ 푛−1 휋 0 |휋 1 휋 푚−1 |휋 푚 휋 푚 |휋 푚+1 휋 푛−2 |휋 푛−1 휋 푛−1 |휋 푚 푣 ∈ 푉 \휋 푚 | ̄ 휏 푚 (푣) 푣 ∈ 푉 \휋 푛−1 | ̄ 휏 푛−1 (푣) 푣| ̄ 휏 푚 (푣)푣| ̄ 휏 푛 (푣) Figure 12: The memory structureM ▶Recursive enumerability Recursive enumerability will be a consequence of the following result. Lemma 4. For every NE ̄ 휎in the energy gameG ↾푣 0 , there exists a finite-memory NE ̄ 휎 ★ with휇( ̄ 휎) = 휇( ̄ 휎 ★ ). Proof.Let ̄ 휎be an NE in the gameG ↾푣 0 . Let휋 =⟨ ̄ 휎⟩. By Dickson’s lemma, there exist two indices푚 and푛, with푚< 푛, such that휋 푚 = 휋 푛 and such that for every player푖that does not lose the play휋, we haveel 푖 (휋 ≤푚 ) ≤ el 푖 (휋 ≤푛 ). Moreover, we can chose푚great enough to haveel 푗 (휋 ≤푚 ) =⊥for every player푗that loses the play휋. Thus, the players winning the play휋are exactly the players winning the play 휒 = 휋 <푚 ( 휋 푚 . . .휋 푛−1 ) 휔 . Now, let푘 ≤ 푛, and let푖be the player controlling the vertex휋 푘 . If푖is a player who loses the play휋, then since ̄ 휎is an NE, any play of the form휋 ≤푘 휒compatible with ̄ 휎 −푖 is lost by player푖; in other words, the strategy profile: ̄ 휎 −푖↾휋 ≤푘 : Hist −푖 G ↾휋 푘 ! 푉 ℎ7! ̄ 휎 −푖 (휋 <푘 ℎ) is a strategy profile against which player푖cannot win. It is known (see for example [BFL + 08, Lemma 10]) that stationary strategies are sufficient to falsify an energy objective. Therefore, let ̄ 휏 푘 −푖 be a stationary strategy profile, from the vertex휋 푘 , against which player푖cannot win. Let휏 푘 푖 be an arbitrary stationary strategy. If푖is not a player who loses the play휋, then we define ̄ 휏 푘 as an arbitrary stationary strategy profile. LetMbe the memory structure depicted by Figure 12, and defined as follows: it has 2푛states, namely푞 0 , . . .,푞 푛−1 and푞 ′ 0 , . . .,푞 ′ 푛−1 . From each state푞 푘 , the transition reading the vertex휋 푘 leads to the state푞 푘+1 (or푞 푚 if푘 = 푛−1), and outputs the vertex휋 푘+1 (or휋 푚 if푘 = 푛−1). The transition reading any other vertex푣(if푘 ≥1) leads to the state푞 ′ 푘 and outputs the vertex ̄ 휏 푘−1 (푣). From the state푞 ′ 푘 , the transition reading each vertex푣leads to푞 ′ 푘 , and outputs the vertex ̄ 휏 푘−1 (푣) (since ̄ 휏 푘−1 is stationary). Thus, the strategy profile induced by the memory structureMis the strategy profile that follows the play휒, and that punishes any player who deviates by following the stationary strategy profile ̄ 휏 푘 . That strategy profile is finite-memory, and generates the same payoff vector as ̄ 휎 , as desired.□ 52CHAPTER 3. NASH EQUILIBRIA Moreover, once we have a memory structure, there is an algorithm that says whether it induces an NE or not. Lemma 5. Given an energy gameG ↾푣 0 and a multiplayer memory structureM, deciding whether the strategy profile ̄ 휎 induced byM is an NE can be done in deterministic polynomial time. Proof. Given a player푖, that player has a profitable deviation from ̄ 휎 if and only if these two conditions are satisfied: •the play⟨ ̄ 휎⟩ is lost by player 푖; •there exists a play 휋 compatible with ̄ 휎 −푖 such that 휇 푖 (휋) = 1. The first condition can be checked in polynomial time, since the play⟨ ̄ 휎⟩is a lasso whose size is bounded by a polynomial function of∥M∥. The second condition is satisfied if and only if there exists a play giving player 푖 the payoff 1 in the one-player game: G ′ = ( 푖,푉 ×푄,(푉 ×푄),퐸 ′ , 휇 ′ ) , where휇 ′ :(휋 0 ,푞 0 )(휋 1 ,푞 1 ) . . . 7! 휇 푖 (휋), and퐸 ′ contains the transitions(푢,푝)(푣,푞)such that 훿(푝,푢) =푞and either휈(푝,푢) = 푣, or푢 ∈ 푉 푖 . Checking the existence of such a play can be done in polynomial time according to [BFL + 08, Theorem 7].□ Thus, a semi-algorithm that recognizes the positive instances of the constrained existence problem consists in enumerating the deterministic multiplayer memory structures onG ↾푣 0 , and for each of them, to check (by diagonalization): •whether the only strategy profile compatible with it is an NE: that problem is decidable (in polynomial time) by Lemma 5; • whether that strategy profile generates a payoff vector between ̄ 푥and ̄ 푦: that is recursively enumerable, by constructing step by step its outcome and computing the energy levels on the fly. By Lemma 4, we have a positive instance of the constrained problem if and only if at least one memory structure satisfies those two conditions. The constrained existence problem is therefore recursively enumerable.□ PART I: SUBGAME-PERFECT EQUILIBRIA 53 CHAPTER 4:SUBGAME-PERFECT EQUILIB- RIA AND NEGOTIATION Let us imagine the following situation: in a parking lot, cars—which can be automated, semi-automated, or driven solely by a human—must enter and exit through the same entrance, which is too narrow to allow two cars to pass simultaneously. When one car arrives and another one leaves, they (and their possibly existing drivers) must agree on who should pass first to avoid getting stuck. Since endless negotiations are undesirable, they need a protocol to make this decision. In game-theoretic terms, such a protocol would be represented by a strategy profile. For example, the protocol could be as follows: when two cars arrive at the same time, one direction (typically exiting) has priority over the other. When there are queues on both sides, cars alternate. A more challenging issue arises when a car does not follow the protocol. Ideally, such deviations should not allow the car to gain time in any way; otherwise, impatient drivers would always be tempted to deviate, making the protocol ineffective in practice. This is why it is desirable for the strategy profile representing the protocol to be a Nash equilibrium. On the other hand, let us consider the following protocol, which does satisfy the conditions of a Nash equilibrium: first, cars adhere to the priority rules defined above. Then, if a car goes through the entrance when it is not supposed to, all other cars stop driving, resulting in a complete and permanent standstill. The flaw in this approach is that the threat meant to prevent deviations is enforced by agents who have their own objectives (reaching their destinations within a reasonable time), which they will not simply disregard when a deviation occurs. This situation highlights a fundamental limitation of Nash equilibria in sequential settings: a problem known as non-credible threats. In doing so, it also advocates for a more robust equilibrium concept, one that leads to interesting algorithmic challenges: subgame-perfect equilibria. That concept was first introduced by Reinhard Selten [Sel65], and earned him the Nobel prize in economic sciences in 1994, shared with John Nash and John Harsanyi. 4.1SUBGAME-PERFECT EQUILIBRIA Formally, subgame-perfect equilibria will be defined by the fact that players play rationally not only in the outcome of the strategy profile, but also in subgames. Definition 17 (Subgame, substrategy). Letℎ푣be a history in the gameG. The subgame of the gameG after the history ℎ푣 is the initialized game Π,푉,(푉 푖 ) 푖 ,퐸,휇 ↾ℎ푣 ↾푣 , where 휇 ↾ℎ푣 maps each play to its payoff in the gameG, assuming that the historyℎ푣has already been played: formally, for every휋 ∈ PlaysG ↾ℎ푣 , we have 휇 ↾ℎ푣 (휋) = 휇(ℎ휋). 55 56CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION 휇 ◦ = 휇 □ = 0 푎 휇 ◦ = 휇 □ = 0 푏 휇 ◦ = 휇 □ = 1 푐 Figure 13: Two Nash equilibria If휎 푖 is a strategy in the gameG ↾푣 0 , its substrategy afterℎ푣is the strategy휎 푖↾ℎ푣 in the gameG ↾ℎ푣 , defined by 휎 푖↾ℎ푣 (ℎ ′ ) = 휎 푖 (ℎ ′ ) for every ℎ ′ ∈ Hist 푖 G ↾ℎ푣 . Remark 2. The initialized gameG ↾푣 0 is also the subgame ofG after the one-vertex history 푣 0 . We can now define subgame-perfect equilibria. Definition 18 (Subgame-perfect equilibrium). LetG ↾푣 0 be a game. The strategy profile ̄ 휎is a subgame- perfect equilibrium, or SPE for short, in the gameG ↾푣 0 , if and only if for every historyℎin the gameG ↾푣 0 , the strategy profile ̄ 휎 ↾ℎ is a Nash equilibrium in the subgameG ↾ℎ . Example 3. We have recalled in Figure 13 the game that was already given in Example 1. In this game, there are actually two NEs: •one, depicted in blue, where ◦ goes to the vertex푏and then player□goes to푐, and both get the payoff 1; •and one, depicted in red, where player ◦ stays in푎, and has no incentive to deviate to푏because player □ plans to stay in 푏. However, only the first one is an SPE, because in the subgame after the history푎푏, player□has a profitable deviation by going to 푐. While Nash equilibria are known to exist in all the game classes we study in this thesis, it is not the case of SPEs [SV03,BBMR15]. We will therefore, sometimes, need a quantitative relaxation of SPEs, namely휀-SPEs. That notion will be built on the notion of휀-Nash equilibrium: given a quantity휀 ≥0, a strategy profile ̄ 휎is an휀-Nash equilibrium, or휀-NE for short, if for each player푖and every deviation휎 ′ 푖 of player푖from ̄ 휎, the inequality휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ 휇 푖 ⟨ ̄ 휎⟩+휀 holds—i.e., the deviation휎 ′ 푖 is not profitable by more than 휀. We can then define 휀-SPEs. Definition 19 (휀-Subgame-perfect equilibrium). LetG ↾푣 0 be a game, and let휀 ≥0. The strategy profile ̄ 휎 is an휀-subgame-perfect equilibrium, or휀-SPE for short, in the gameG ↾푣 0 , if and only if for every historyℎ in the gameG ↾푣 0 , the strategy profile ̄ 휎 ↾ℎ is an 휀-NE in the subgameG ↾ℎ . Remark 3. The definitions of 0-NEs and 0-SPEs coincide, respectively, with those of NEs and SPEs. This part of this thesis is dedicated to finding and proving the complexities of the constrained existence problem of SPEs, and 휀-SPEs when it is relevant, in the game classes we have introduced in Part I. 4.2. NEGOTIATION57 4.2NEGOTIATION 4.2.1 Origin of the concept The idea behind the concept of negotiation seems to have emerged in two parallel ways, without a full formalization. On the one hand, János Flesch and Arkadi Predtetchinski introduced in 2017 a characterization of SPEs in Borel games where the payoff functions have a finite range (such as Boolean games), based on an iterative process corresponding to what we call here the negotiation function, defined from what we refer to as the abstract negotiation game [FP17]. When the payoff functions have an infinite range (as in mean-payoff or discounted-sum games), their work characterizes a set of plays that is known to contain only outcomes of휀-SPEs, for every휀>0, as well as all outcomes of SPEs. We will show that it actually characterizes SPE outcomes in all games with steady negotiation. Since their work was primarily conducted with the objective of proving existence results, Flesch and Predtetchinski’s approach does not provide a characterization that can be directly used for algorithmic purposes, as their iterative process may be infinite, and each iteration requires solving an infinite number of two-player zero-sum games, which can themselves be of infinite size. On the other hand, in 2019, Marie van den Bogaard, Thomas Brihaye, Véronique Bruyère, Aline Goeminne, and Jean-François Raskin proved thePSPACE-completeness of the constrained existence problem of SPEs in a class of games that we do not study here: quantitative reachability games, where players aim to reach a given set of vertices as quickly as possible [BBG + 19]. Their upper bound relied on an algorithm that labels each vertex with an integer, representing the number of steps within which the player controlling the vertex can, and therefore must, reach their target. The algorithm then updates this labeling iteratively, leveraging the fact that a player may discover new ways to reach their target more quickly when other players, while trying to prevent them from doing so, are also attempting to satisfy their own requirements. These updates, which correspond to one iteration of the negotiation function (in a case where Flesch and Predtetchinski’s work did not provide a full characterization, since quantitative reachability payoff functions have infinite range), continue until a fixed point is reached, which characterizes SPEs. In both cases, the negotiation function was not explicitly formalized as a function but was instead an implicit part of an iterative process. The reader will observe in the following chapters that reasoning about these iterations as a function, whose fixed points characterize SPEs, is useful in some cases where simply iterating the function is not the most efficient way to solve our problem, or could even fail to terminate, as is the case in mean-payoff games. 4.2.2 Requirements In the method we will develop further, we will need to analyze the players’ behaviors when they have some requirement to satisfy. Intuitively, one can see requirements as rationality constraints for the players, that is, a threshold payoff value under which a player will not accept to follow a play, because they would be better off deviating from it. Definition 20 (Requirement). A requirement on the gameG is a mapping 휆 : 푉 ! R∪±∞. For a given vertex푣, the quantity휆(푣)represents the minimal payoff that the player controlling푣will require in a play starting from 푣 . 58CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION Definition 21 (휆-consistency). Let휆be a requirement on a gameG. A play휋in the gameGis휆-consistent if and only if, for every player푖 ∈Πand integer푛 ∈ Nwith휋 푛 ∈ 푉 푖 , we have휇 푖 (휋 ≥푛 ) ≥ 휆(휋 푛 ). The set of 휆-consistent plays from a vertex 푣 is denoted by 휆ConsG ↾푣 , or 휆Cons(푣) when the context is clear. Such a requirement also induces a notion of rationality for strategy profiles aiming at punishing one player. Definition 22 (휆-rationality). Let휆be a requirement on a gameG. Let푖 ∈Π. The strategy profile ̄ 휎 −푖 is 휆-rational assuming the strategy휎 푖 if and only if for every historyℎ푣compatible with ̄ 휎 −푖 , the play⟨ ̄ 휎 ↾ℎ푣 ⟩ is휆-consistent. A strategy profile is simply휆-rational when it is휆-rational assuming some strategy. The set of휆-rational strategy profiles in the gameG ↾푣 is denoted by휆Rat −푖 G ↾푣 , or휆Rat(푣)when the context is clear. Note that휆-rationality is a property of a strategy profile for all the players but one, player푖. Intuitively, their rationality is justified by the fact that they collectively assume that player푖will, eventually, play according to the strategy 휎 푖 : if it is the case, then everyone gets their payoff satisfied. Finally, let us define a particular requirement: the vacuous requirement, that requires nothing, and with which every play is consistent. Definition 23 (Vacuous requirement). In any game, the vacuous requirement, denoted by휆 0 , is the requirement constantly equal to−∞. 4.2.3 Negotiation We will show that SPEs in prefix-independent games are characterized by the fixed points of a function on requirements. That function captures a negotiation process: when a player has a requirement to satisfy, another player can hope for a better payoff than what they can secure in general, and therefore update their own requirement. Note that we always use the convention inf∅ =+∞. Definition 24 (Negotiation function, steady negotiation). LetGbe a game. The negotiation function is the function that transforms each requirement휆onGinto a requirementnego(휆)onG, such that for each player 푖 ∈Π and every vertex 푣 ∈ 푉 푖 , we have: nego(휆)(푣) =inf ̄ 휎 −푖 ∈휆Rat(푣) sup 휎 푖 ∈Strat 푖 G ↾푣 휇 푖 ⟨ ̄ 휎⟩. If that infimum is realized for every requirement휆, player푖and vertex푣 ∈ 푉 푖 such that휆Rat(푣) ≠∅, then the gameG is called a game with steady negotiation. The quantitynego(휆)(푣)is then the worst case value that the player controlling푣can ensure, assuming that the other players play 휆-rationally. Remark 4. The negotiation function satisfies the following properties. • It is monotone: if we have 휆 ≤ 휆 ′ , then we have nego(휆) ≤ nego(휆 ′ ). • It is also non-decreasing: for every 휆, we have 휆 ≤ nego(휆). •There exists a휆-rational strategy profile from푣against the player controlling푣if and only if nego(휆)(푣) ≠+∞. 4.3. LINK BETWEEN NEGOTIATION AND EQUILIBRIA59 Example 4. Let us consider again the game depicted by Figure 13, and let us consider the vacuous requirement 휆 0 : 푣 !−∞. Let us compute 휆 1 = nego(휆 0 ). From the vertex푐, the only payoff that can be obtained is휆 1 (푐) =1. From the vertex푏, player□can go to the vertex푐and get the payoff휆 1 (푏) =1. From the vertex푎, even if player ◦ goes to푏, player□can stay there, hence she cannot get more than 휆 1 (푎) = 0. Let us now compute휆 2 = nego(휆 1 ). From푐, the only payoff possible is still휆 2 (푐) =1, and from푏, it is still possible to get휆 2 (푏) =1 and nothing more. But from the vertex푎, a휆-rational strategy profile must necessarily plan to go to the vertex 푐 if the vertex 푏 is reached, hence 휆 2 (푎) = 1. Finally, we have nego(휆 2 ) = 휆 2 . 4.3LINK BETWEEN NEGOTIATION AND EQUILIBRIA 4.3.1 Nash equilibria Although all our computational results about Nash equilibria have already been stated in Chapter 3, let us show how we can, using the negotiation function, rephrase a folklore characterization of Nash equilibria outcomes: a play is a Nash equilibrium outcome if and only if it gives to every player at least the payoff that they could enforce by deviating, if all the other players try to punish them—i.e., if for every vertex it traverses, it gives to the player controlling that vertex its adversarial value. Theorem 7. LetGbe a game with steady negotiation, and let휋be a play in the gameG. Then, the play휋is an NE outcome if and only if 휋 is nego(휆 0 )-consistent. Proof.▶If 휋 is an NE outcome, then it is nego(휆 0 )-consistent. Let ̄ 휎 be a Nash equilibrium in the gameG, and let휋 = ⟨ ̄ 휎⟩: let us prove that the play휋is nego(휆 0 )-consistent. Let푘 ∈ N, let푖 ∈Πbe such that휋 푘 ∈ 푉 푖 , and let us prove that휇 푖 ( 휋 ≥푘 ) ≥ nego(휆 0 )(휋 푘 ). For any deviation휎 ′ 푖 of player푖from the strategy profile ̄ 휎 ↾휋 ≤푘 , by definition of NEs, we have 휇 푖 ⟨ ̄ 휎 −푖↾휋 ≤푘 ,휎 ′ 푖 ⟩ ≤ 휇 푖 (휋). Therefore, we have 휇 푖 (휋) ≥ sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖↾휋 ≤푘 ,휎 ′ 푖 ⟩, hence the inequality: 휇 푖 (휋) ≥ inf ̄ 휏 −푖 sup 휏 푖 휇 푖 ⟨ ̄ 휏 −푖↾휋 ≤푘 ,휏 푖 ⟩, i.e. 휇 푖 (휋) ≥ nego(휆 0 )(휋 푘 ): the play 휋 is nego(휆 0 )-consistent. ▶If 휋 is nego(휆 0 )-consistent, then it is an NE outcome. Let휋be anego(휆 0 )-consistent play in the gameG. Let us define a strategy profile ̄ 휎generating휋 as follows. •First, we define⟨ ̄ 휎⟩ = 휋 . • Then, for each history of the form휋 ≤푘 푣with푣 ≠ 휋 푘+1 , let푖be the player controlling휋 푘 . Since the gameG is with steady negotiation, the infimum: inf ̄ 휏 −푖 ∈휆 0 Rat(휋 푘 ) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩ = nego(휆 0 )(푣) ≠+∞ 60CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION is a minimum. Let ̄ 휏 푘 −푖 be휆 0 -rational strategy profile from휋 푘 realizing that minimum, and let휏 푘 푖 be some strategy from 휋 푘 . Then, we define: ⟨ ̄ 휎 ↾휋 ≤푘 푣 ⟩ =⟨ ̄ 휏 푘 ↾휋 푘 푣 ⟩. •For every other history ℎ, the vertex ̄ 휎(ℎ) is defined arbitrarily. Let us prove that ̄ 휎 is an NE: let휎 ′ 푖 be a deviation of휎 푖 , and let휋 ′ = ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩. Since the case 휋 ′ = 휋is trivial, we assume휋 ≠ 휋 ′ and define휋 ≤푘 as the longest common prefix of휋and휋 ′ . Let 푣 = 휋 ′ 푘+1 ≠ 휋 푘+1 . Then, the play휋 ′ ≥푘+1 is compatible with the strategy profile ̄ 휎 −푖↾휋 ≤푘 푣 = ̄ 휏 −푖↾휋 푘 푣 , hence the inequality: 휇 푖 (휋 ′ ) ≤ sup 휏 푘 푖 휇 푖 ⟨ ̄ 휏 푘 ⟩ = nego(휆 0 )(휋 푘 ). On the other hand, since the play휋is휆 0 -consistent, we havenego(휆 0 )(휋 푘 ) ≤ 휇 푖 (휋), hence휇 푖 (휋 ′ ) ≤ 휇 푖 (휋): the deviation 휎 ′ 푖 is not profitable, and the strategy profile ̄ 휎 is a Nash equilibrium.□ 4.3.2 Link with (휀–)subgame-perfect equilibria The notion of negotiation will enable us to find the SPEs, but also more generally the휀-SPEs, in a game. For that purpose, we need the notion of 휀-fixed points of a function. Definition 25 (휀-fixed point). Let휀 ≥0, let퐷be a finite set and let푓:(R∪±∞) 퐷 ! (R∪±∞) 퐷 be a mapping. A tuple ̄ 푥 ∈ R 퐷 is an휀-fixed point of푓if for each푑 ∈ 퐷, for ̄ 푦 = 푓( ̄ 푥) , we have푦 푑 ∈ [푥 푑 −휀,푥 푑 +휀]. Remark 5. The definition of 0-fixed points coincide with that of fixed points. We now have all the tools required to state and prove the following theorem. We give it here in a form that applies only to prefix-independent games, even though the same result extends (at least when 휀 =0) to classes of games that do not observe that property, such as quantitative reachability games or discounted-sum games. However, since we do not use this result in such classes, we favored the following proof, less general but simpler. Theorem 8. LetGbe a game with steady negotiation, and let휋be a play in the gameG. Let휀 ≥0. Then, the play휋is an휀-SPE outcome if and only if휋is휆-consistent, for some휀-fixed point휆of the negotiation function. Proof.▶If 휋 is an 휀-SPE outcome, then it is 휆-consistent for some 휀-fixed point 휆. Let ̄ 휎 be an 휀-SPE in the gameG such that 휋 =⟨ ̄ 휎⟩. Let us define a requirement 휆 by, for each 푖 ∈Π and 푣 ∈ 푉 푖 : 휆(푣) =inf ℎ푣∈HistG ↾푣 0 휇 푖 ⟨ ̄ 휎 ↾ℎ푣 ⟩. Then, for every historyℎ푣starting in휋 0 , the play⟨ ̄ 휎 ↾ℎ푣 ⟩is휆-consistent. And in particular, the play휋 is. Let us now prove that the requirement 휆 is an 휀-fixed point of nego. Let푖 ∈Π, let푣 ∈ 푉 푖 , and let us show that we havenego(휆)(푣) ≤ 휆(푣)+ 휀. Since ̄ 휎 is an휀-SPE, for 4.3. LINK BETWEEN NEGOTIATION AND EQUILIBRIA61 • 휋 0 휋 • 푢 ∈ 푉 푖 • ̄ 휏 푢 • 푤 ∈ 푉 푗 ̄ 휏 푤 • 푣 ∈ 푉 푖 ̄ 휏 푣 (if reset) Figure 14: The construction of ̄ 휎 every history ℎ푣 ∈ HistG ↾휋 0 , we have: sup 휏 푖 휇 푖 ⟨ ̄ 휎 −푖↾ℎ푣 ,휏 푖 ⟩ ≤ 휇 푖 ⟨ ̄ 휎 ↾ℎ푣 ⟩+ 휀. By taking the infima over ℎ푣 , we deduce: inf ℎ푣∈HistG ↾휋 0 sup 휏 푖 휇 푖 ⟨ ̄ 휎 −푖↾ℎ푣 ,휏 푖 ⟩ ≤ inf ℎ푣 휇 푖 ⟨ ̄ 휎 ↾ℎ푣 ⟩+ 휀 = 휆(푣)+ 휀. Let us now notice that all the strategy profiles of the form ̄ 휎 −푖↾ℎ푣 are 휆-rational. Thus, we obtain: inf ̄ 휏 −푖 ∈휆Rat(푣) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩ ≤ 휆(푣)+ 휀, that is, we obtain nego(휆)(푣) ≤ 휆(푣)+ 휀, as desired. ▶If 휋 is 휆-consistent for some fixed point 휆, then 휋 is an 휀-SPE outcome. A particular case: if there exists푣accessible from휋 0 such that휆(푣) =+∞.In that case, for each vertex푢such that푢푣 ∈ 퐸, if the player controlling푢chooses to go to푣, no휆-consistent play can be proposed to them from there, hence there is no휆-rational strategy profile against that player from 푢, andnego(휆)(푢) =+∞. Since휀is finite and since휆is an휀-fixed point of the negotiation function, it follows that휆(푢) =+∞. Since푣is accessible from휋 0 , we can repeat this argument and show that 휆(휋 0 ) =+∞; in that case, there is no 휆-consistent play 휋 from푢, and then the proof is done. Therefore, for the rest of the proof, we assume that for all푣, we have휆(푣) ≠+∞. As a consequence, since휆is an휀-fixed point of the functionnego, for each푣accessible from휋 0 , we havenego(휆)(푣) ≠+∞; which implies that for each such푣, there exists a휆-consistent strategy profile against the player controlling 푣 , starting from 푣 . The rest of the proof constructs the strategy profile ̄ 휎and proves that it is an휀-SPE. That construc- tion is illustrated by Figure 14. Spare parts: the strategy profiles ̄ 휏 푣∗ .Let us recall that sinceGis a game with steady negotia- tion. Then, let푖 ∈Πand푣 ∈ 푉 푖 : since, by the previous point, we assume휆Rat(푣) ≠∅for each푣, we know that there exists a strategy profile ̄ 휏 푣 −푖 from푣that is휆-rational assuming a strategy휏 푣 푖 and that 62CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION satisfies the inequality: sup 휏 푖 휇 푖 ⟨ ̄ 휏 푣 −푖 ,휏 푖 ⟩ =inf ̄ 휏 −푖 ∈휆Rat(푣) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩ = nego(휆)(푣). In other words, there exists a worst휆-rational strategy profile against player푖from the vertex푣, with regards to player 푖’s payoff. Our goal in this part of the proof is to construct a strategy profile ̄ 휏 푣∗ −푖 , that is휆-rational assuming a strategy휏 푣∗ 푖 , and that will be used to punish player푖when they deviate from ̄ 휎until another player deviates. The strategy profile ̄ 휏 푣 −푖 and the strategy휏 푣 푖 are not sufficient for that purpose, because if some historyℎcompatible with ̄ 휏 푣 −푖 is such that휇 푖 ⟨ ̄ 휏 푣 ↾ℎ ⟩< 휇 푖 ⟨ ̄ 휏 푣 ⟩, then in the corresponding subgame, it may be possible for player푖to deviate and get a payoff that would be smaller than or equal to the quantity휇 푖 ⟨ ̄ 휏 푣 ⟩, but greater than휇 푖 ⟨ ̄ 휏 푣 ↾ℎ ⟩. On the contrary, the construction of the strategy profile ̄ 휏 푣∗ −푖 will ensure that each time player푖deviates, the other players punish them at least as harshly as they were planning to do before the deviation. Let us construct inductively the strategy profile ̄ 휏 푣∗ . We define it only on histories that are compatible with ̄ 휏 푣∗ −푖 , since it can be defined arbitrarily on other histories. We proceed by assembling the strategy profiles of the form ̄ 휏 푤 for various푤 ∈ 푉 푖 , and the histories after which we follow a new ̄ 휏 푤 will be called the resets of ̄ 휏 푣∗ : they will be histories of the formℎ푤 ′ , whereℎis empty or last(ℎ) =푤 . •First, we set⟨ ̄ 휏 푣∗ ⟩ =⟨ ̄ 휏 푣 ⟩: the one-vertex history 푣 is then the first reset of ̄ 휏 푣∗ −푖 . •Then, for every historyℎ푤 ′ from푣such thatℎis compatible with ̄ 휏 푣∗ −푖 , that푤 ∈ 푉 푖 , and that 푤 ′ ≠ 휏 푣∗ 푖 (ℎ푤): let us decomposeℎ푤 ′ = ℎ 1 ℎ 2 , so that the historyℎ 1 first(ℎ 2 )is the longest reset of ̄ 휏 푣∗ −푖 among the prefixes of ℎ푤 . Or, in other words, so that the strategy profile ̄ 휏 푣∗ ↾ℎ 1 first(ℎ 2 ) has been defined as equal to ̄ 휏 푢 over the prefixes ofℎ 2 until푤, where푢 = 푣ifℎ 1 is empty, or 푢 = last(ℎ 1 ) otherwise. By prefix-independence ofG and by definition of ̄ 휏 푢 and ̄ 휏 푤 , we have: inf ̄ 휏 −푖 ∈휆Rat(푤 ′ ) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩ ≤ sup 휏 푖 휇 푖 ⟨ ̄ 휏 푤 −푖 ,휏 푖 ⟩ = nego(휆)(푤). Let us now separate two cases. –Suppose first that there is equality: inf ̄ 휏 −푖 ∈휆Rat(푤 ′ ) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩ = nego(휆)(푤). Then, we choose⟨ ̄ 휏 푣∗ ↾ℎ푤 ′ ⟩ =⟨ ̄ 휏 푢 ↾푢ℎ 2 ⟩ : the coalition of players against player푖keeps following the same strategy profile. –Suppose now that the inequality is strict: inf ̄ 휏 −푖 ∈휆Rat(푤 ′ ) sup 휏 푖 휇 푖 ⟨ ̄ 휏⟩< nego(휆)(푤). Then, we choose⟨ ̄ 휏 푣 ∗ ↾ℎ푤 ′ ⟩ = ⟨ ̄ 휏 푤 ↾푤 ′ ⟩ : player푖has done something that lowers the payoff they can ensure, and therefore the other players have to update their strategy profile in order to punish them more harshly. The history ℎ푤 is a reset of ̄ 휏 푣∗ −푖 . Since there are finitely many histories of each length, this process completely defines the strategy 4.3. LINK BETWEEN NEGOTIATION AND EQUILIBRIA63 • 휋 0 휋 • 휒 푛 ∈ 푉 푖 휒 • 푤 • 휒 푝 ∈ 푉 푖 • 휒 푚 ∈ 푉 푖 휒 ′ Figure 15: The strategy profile ̄ 휎 is an SPE. profile ̄ 휏 푣∗ . Moreover, all the plays constructed are휆-consistent, hence the strategy profile ̄ 휏 푣∗ −푖 is 휆-rational assuming the strategy 휏 푣∗ 푖 , as desired. Construction of ̄ 휎.Let us now construct inductively the strategy profile ̄ 휎itself: we will prove in the next part of the proof that it is an휀-SPE. We proceed inductively, by defining all the plays⟨ ̄ 휎 ↾ℎ푣 ⟩, forℎ푣 ∈ HistG 휋 0 with푣 ≠ ̄ 휎(ℎ). We maintain the induction hypothesis that such a play is always 휆-consistent. •First, we choose⟨ ̄ 휎⟩ = 휋 , which satisfies the induction hypothesis. •Let now ℎ푢푣 be a history such that the strategy profile ̄ 휎 has been defined on all the prefixes of ℎ푢 , which we now assume to be nonempty, but not onℎ푢푣itself, and such that푣 ≠ ̄ 휎(ℎ푢). Let푖 be the player controlling the vertex푢. Then, we define⟨ ̄ 휎 ↾ℎ푢푣 ⟩ = ⟨ ̄ 휏 푢∗ ↾푢푣 ⟩, and inductively, for every historyℎ ′ 푤starting from푣and compatible with ̄ 휎 −푖↾ℎ푢푣 , we define⟨ ̄ 휎 ↾ℎ푢ℎ ′ 푤 ⟩ =⟨ ̄ 휏 푢∗ ↾푢ℎ ′ 푤 ⟩. The strategy profile ̄ 휎 ↾ℎ푢푣 is then equal to ̄ 휏 푣∗ ↾푢푣 on any history compatible with ̄ 휏 푣∗ −푖 . Since there are finitely many histories of each length, this process completely defines ̄ 휎 . The strategy profile ̄ 휎is an휀-SPE. Consider a historyℎ 0 푤 ∈ HistG ↾휋 0 , a player푖 ∈Π, and a deviation휎 ′ 푖 of휎 푖 . Let휒 =ℎ 0 ⟨ ̄ 휎 ↾ℎ 0 푤 ⟩, and let휒 ′ =ℎ 0 ⟨ ̄ 휎 −푖↾ℎ 0 푤 ,휎 ′ 푖↾ℎ 0 푤 ⟩. We wish to prove the inequality 휇 푖 (휒 ′ ) ≤ 휇 푖 (휒)+ 휀. The different notations in this proof are illustrated by Figure 15. First, if the play휒 ′ is compatible with휎 푖 , then we have휒 ′ = 휒and the proof is immediate. Now, if it is not, we let푛denote the least index such that휒 ′ 푛 ∈ 푉 푖 and휒 ′ 푛+1 ≠ 휎 푖 (휒 ′ ≤푛 ) , and such that the play휒 ′ ≥푛 is compatible with the strategy profile ̄ 휎 −푖↾휒 ′ ≤푛 . Thus, the edge휒 ′ 푛 휒 ′ 푛+1 marks the time when player푖begins to deviate unilaterally from휎 푖 . However, note that휒 ′ ≤푛 can be both longer or shorter than ℎ 0 푤 : player 푖 may have already deviated in ℎ 0 푤 , or may wait afterwards to effectively deviate. Be that as it may, the history휒 ′ ≤푛 is a common prefix of the plays휒and휒 ′ , and the substrategy profile ̄ 휎 ↾휒 ′ ≤푛+1 has been defined during the construction of ̄ 휎as equal to ̄ 휏 푣∗ ↾휒 ′ 푛 휒 ′ 푛+1 , where푣 = 휒 ′ 푛 , on any history compatible with ̄ 휎 −푖↾휒 ′ ≤푛+1 . By construction of ̄ 휏 푣∗ , the sequence: nego(휆)(휒 ′ 푘 ) 푘≥푛,휒 ′ 푘 ∈푉 푖 64CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION is non-increasing. It is therefore ultimately constant (or is finite), because it can take only a finite number of values. Consequently, there is a finite number of resets along the play휒 ′ ≥푛 . Let then 휒 ′ 푛 . . . 휒 ′ 푚+1 be the last (longest) one. Afterwards, the play휒 ′ ≥푚+1 is compatible with the strategy profile ̄ 휏 휒 ′ 푚 −푖 . By definition of that strategy profile, we have the inequality휇 푖 (휒 ′ ) ≤ nego(휆)(휒 ′ 푚 ). We need now to prove the inequality nego(휆)(휒 ′ 푚 ) ≤ 휇 푖 (휒)+ 휀. Let now휒 ≤푝 = 휒 ′ ≤푝 denote the longest common prefix of휒and휒 ′ . Note that, then, we have푛 ≤ 푝 and휒 푝 ∈ 푉 푖 . Moreover, we have휒 ≥푝 =⟨ ̄ 휎 ↾휒 ≤푝 ⟩, which is휆-consistent. As a consequence, we have the inequality 휇 푖 (휒) ≥ 휆(휒 푝 ). Finally, since the sequence of the quantitiesnego(휆)(휒 ′ 푘 ) with휒 ′ 푘 ∈ 푉 푖 is non-increasing for푘 ≥ 푛, we also have nego(휆)(휒 ′ 푚−1 ) ≤ nego(휆)(휒 푝 ). Consequently, we have: 휇 푖 (휒 ′ ) ≤ nego(휆)(휒 ′ 푚 ) ≤ nego(휆)(휒 푝 ) ≤ 휆(휒 푝 )+ 휀 ≤ 휇 푖 (휒)+ 휀. The strategy profile ̄ 휎 is an 휀-SPE.□ Example 5. Let us consider again the game of Figure 13. The two NE outcomes that we identified in Example 3 are indeed the two plays from푎that are휆 1 -consistent: the play푎 휔 and the play푎푏푐 휔 . However, only the latter is 휆 2 -consistent, and 휆 2 is a fixed point of the negotiation function. This theorem will enable us to design efficient algorithms for the (휀-)SPE constrained existence problems. Indeed, given a gameGand two thresholds ̄ 푥and ̄ 푦, one can decide whether there exists an (휀-)SPE generating a payoff vector between ̄ 푥 and ̄ 푦 by: • guessing a requirement 휆; •checking that there exists a휆-consistent play in the gameGthat generates a payoff vector between ̄ 푥 and ̄ 푦; • checking that 휆 is an (휀-)fixed point of the negotiation function. The last point will usually be the one that induces more substantial work, since it requires a general method to compute the negotiation function. Such a method will be provided by tools called negotiation games. 4.4NEGOTIATION GAMES 4.4.1 The abstract negotiation game Given a requirement휆and a vertex푣 0 ∈ 푉, the quantitynego(휆)(푣 0 )can be characterized as the value of a negotiation game, a two-player zero-sum game opposing the player Prover, who simulates a휆-rational strategy profile and wants to minimize player푖’s payoff (where푖is the player controlling푣 0 ), and the player Challenger, who simulates player푖’s reaction by accepting or refusing Prover’s proposals. First defined by János Flesch and Arkadi Predtetchinski [FP17] (without being linked to the concept of negotiation function), the abstract negotiation game, written abs 휆푖 (G) ↾푣 0 unfolds as follows. • From the vertex푣 0 , Prover chooses a휆-consistent play휋and proposes it to Challenger. If Prover has no play to propose, the game is over and Challenger gets the payoff+∞. 4.4. NEGOTIATION GAMES65 •Once a play휋has been proposed, Challenger can accept it. Or he can deviate, and choose a prefix 휋 ≤푘 with 휋 푘 ∈ 푉 푖 and a new edge 휋 푘 푣 ∈ 퐸. • In the former case, the game is over. In the latter, it starts again from the vertex 푣 . In the play that Prover and Challenger construct together, Challenger’s objective consists in maximizing player푖’s payoff, and Prover’s objective in minimizing it. More formally, the abstract negotiation game abs 휆푖 (G) ↾푣 0 is a game with an uncountable vertex space, defined as follows. Definition 26 (Abstract negotiation game). LetG ↾푣 0 be a game, let푖 ∈Π, and let휆be a requirement on G. The abstract negotiation game ofG ↾푣 0 for player푖with requirement휆is the two-player zero-sum game: abs 휆푖 (G) ↾푣 0 = 픓,ℭ,푉 a ,(푉 a 픓 ,푉 a ℭ ),퐸 a , 휇 a ↾푣 0 , where: • 픓 denotes the player Prover and ℭ the player Challenger. • Prover’s vertices are the vertices ofG, and two special vertices⊤ and⊥. • Challenger’s vertices are the vertices of the form[ℎ푣], whereℎ푣can be any history in the gameG, and the vertices of the form [휋], where 휋 is a 휆-consistent play in the gameG. • From the vertex⊥, the only available edge is the self-loop to⊥, and similarly from⊤. From a vertex of the form 푣 , Prover can move to: – any vertex [휋] where 휋 starts from 푣 (she proposes the play 휋 ); – the vertex⊥ (she gives up, last option if she cannot propose any 휆-consistent play). • From a vertex of the form[휋], Challenger can move to the vertex⊤(he accepts the play휋), or to any vertex of the form[휋 ≤푘 푣], where푘 ∈ N,휋 푘 ∈ 푉 푖 and푣 ≠ 휋 푘+1 (he deviates to푣). From a vertex of the form [ℎ푣], he can only go to the vertex 푣 . • In a play휒in this game, if Prover gives up and takes an edge to the vertex⊤, Challenger’s payoff is defined by휇 a ℭ (휒) = +∞ . If Challenger finally accepts a proposal휋, and takes the edge[휋]⊤, then his payoff is defined by휇 a ℭ (휒) = 휇 푖 (휋) . If he deviates infinitely often, then Prover’s proposals and his deviations construct a play¤휒 = 휋 (0) ≤푘 0 휋 (1) ≤푘 1 휋 (2) ≤푘 2 . . .. Then, Challenger’s payoff is defined by 휇 a ℭ (휒) = 휇 푖 (¤휒). In all those cases, Prover’s payoff is the opposite of Challenger’s one. The abstract negotiation game constitutes an alternative definition of the negotiation function. Note that the following theorem does not use the notationval ℭ , because the abstract negotiation game is not guaranteed to be Borel. Theorem 9. LetG ↾푣 0 be a game, let휆be a requirement onGand let푖 ∈Πbe such that푣 0 ∈ 푉 푖 . Then, we have: inf 휏 픓 ∈Strat 픓 abs 휆푖 (G) ↾푣 0 sup 휏 ℭ ∈Strat ℭ abs 휆푖 (G) ↾푣 0 휇 a ℭ ⟨ ̄ 휏⟩ = nego(휆)(푣 0 ). Proof.Let 훼 ∈ R, and let us prove that the following statements are equivalent. 1.There exists a strategy 휏 픓 such that for every strategy 휏 ℭ , we have 휇 a ℭ ⟨ ̄ 휏⟩< 훼 . 2. There exists a휆-rational strategy profile ̄ 휎 −푖 in the gameG ↾푣 0 such that for every strategy휎 푖 , we have 휇 푖 ⟨ ̄ 휎⟩< 훼 . 66CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION ▶Assertion 1 implies Assertion 2. Let 휏 픓 be such that for every strategy 휏 ℭ , we have 휇 a ℭ ⟨ ̄ 휏⟩< 훼 . In what follows, any historyℎcompatible with an already defined strategy profile ̄ 휎 −푖 in the game G ↾푣 0 will be decomposed in: ℎ = 푣 0 ℎ (0) 푣 1 ℎ (1) . . .ℎ (푛−1) 푣 푛 ℎ (푛) , so that there exist plays 휋 (0) , . . .,휋 (푛−1) , 휒 and a history: [푣 0 ] 휋 (0) 푣 1 ℎ (1) 푣 2 . . . 푣 푛−1 ℎ (푛−1) 푣 푛 푣 푛 ℎ (푛) 휒 in the gameabs 휆푖 (G) 푣 0 compatible with휏 픓 : the existence and the unicity of that decomposition can be proved by induction. Intuitively, the historyℎis cut in histories which are prefixes of plays that can be proposed by Prover. Then, let us define inductively the strategy profile ̄ 휎 −푖 by ̄ 휎 −푖 (ℎ) = 휒 0 , with휒defined fromℎas above, for everyℎsuch that ̄ 휎 −푖 has been defined on the prefixes ofℎ, and such that the last vertex of ℎ is not controlled by player 푖. Let us prove that ̄ 휎 −푖 is the desired strategy profile. The strategy profile ̄ 휎 −푖 is휆-rational.Let us define휎 푖 so that for every historyℎ푣compatible with ̄ 휎 −푖 , the play⟨ ̄ 휎 ↾ℎ푣 ⟩ is 휆-consistent. For each history: ℎ = 푣 0 ℎ (0) 푣 1 ℎ (1) . . .ℎ (푛−1) 푣 푛 ℎ (푛) compatible with ̄ 휎 −푖 and ending in푉 푖 , let휎 푖 (ℎ) = 휒 0 with휒corresponding to the decomposition ofℎ, so that by induction: ⟨ ̄ 휎 ↾푣 0 ℎ (0) 푣 1 ℎ (1) ...ℎ (푛−1) 푣 푛 ⟩ = 푣 푛 ℎ (푛) 휒. Let nowℎ푣be a history in the gameG ↾푣 0 , and let us show that the play⟨ ̄ 휎 ↾ℎ푣 ⟩is휆-consistent. If we decompose: ℎ푣 = 푣 0 ℎ (0) 푣 1 ℎ (1) . . .ℎ (푛−1) 푣 푛 ℎ (푛) with the same definition of휒(note that the vertex푣is now included in the decomposition), then we have⟨ ̄ 휎 ↾ℎ푣 ⟩ = 푣휒, and by definition of the abstract negotiation game, the play푣 푛 ℎ (푛) 휒is휆-consistent, and therefore so is the play 푣휒 . The strategy profile ̄ 휎 −푖 keeps player푖’s payoff under the value훼.Let휎 푖 be some strategy for player 푖, and let 휋 =⟨ ̄ 휎⟩. We want to prove that 휇 푖 (휋)< 훼 . Let us define two finite or infinite sequences 휋 (푘) 푘∈퐾 and ℎ (푘) 푣 푘 푘∈퐾 , where퐾 =1, . . .,푛or 퐾 = N\0, by for every 푘 ∈ 퐾 : 휋 (푘) =휏 픓 [푣 0 ] 휋 (0) . . . 휋 (푘−1) ℎ (푘) 푣 푘 and so that for every푘, the historyℎ (푘) 푣 푘 is the shortest prefix of휋that is not a prefix ofℎ (1) . . .ℎ (푘−1) 휋 (푘−1) (or equivalently, the history ℎ (푘) is the longest common prefix of 휋 and ℎ (1) . . .ℎ (푘−1) 휋 (푘−1) ). Then, the length of the longest common prefix ofℎ (1) . . .ℎ (푘−1) 휋 (푘) and휋increases with푘, and the set 퐾 is finite if and only if there exists 푛 such that ℎ (1) . . .ℎ (푛−1) 휋 (푛) = 휋 . 4.4. NEGOTIATION GAMES67 In the infinite case, let: 휒 =[푣 0 ] 휋 (0) ℎ (1) 푣 1 . . . 휋 (푘) ℎ (푘) 푣 푘 . . . . The play 휒 is compatible with 휏 픓 , hence 휇 a ℭ (휒)< 훼 , that is to say: 휇 푖 ℎ (1) ℎ (2) . . . < 훼, i.e. 휇 푖 (휋)< 훼 . In the finite case, let: 휒 =[푣 0 ] 휋 (0) ℎ (1) 푣 1 . . . 휋 (푛) ⊤ 휔 . For the same reason, we have 휇 a ℭ (휒)< 훼 , i.e., we have 휇 푖 ℎ (1) . . .ℎ (푛) 휋 (푛) = 휇 푖 (휋)< 훼 . ▶Assertion 2 implies Assertion 1. Let ̄ 휎 −푖 be a strategy profile that is휆-rational assuming a strategy휎 푖 , and that maintains player푖’s payoff below the quantity 훼 . Let us define a strategy 휏 픓 for Prover in the abstract negotiation game. Let푔 = [푣 0 ] 휋 (0) ℎ (1) 푣 1 휋 (1) . . . ℎ (푛) 푣 푛 be a history in the abstract game, ending in푉 a 픓 . Then, we define: 휏 픓 (푔) = ⟨ ̄ 휎 ↾ℎ (1) ...ℎ (푛) 푣 푛 ⟩ . If푔is a history ending in⊤, then we define휏 픓 (푔) =⊤, and similarly, if푔ends in⊥, we define 휏 픓 (푔) =⊥. Let us show that휏 픓 is the strategy we were looking for. Let휒be a play compatible with휏 픓 . Let us first note that the vertex⊥cannot appear in휒. Then, the play휒can only have two forms: a play in which Challenger eventually accepts Prover’s proposal, or an infinite sequence of proposals and deviations. If Challenger eventually accepts Prover’s proposal.If휒 =[푣 0 ] 휋 (0) ℎ (1) 푣 1 . . . 휋 (푛) ⊤ 휔 , then we have: 휋 (푛) =⟨ ̄ 휎 ↾ℎ (1) ...ℎ (푛) 푣 푛 ⟩, and the history ℎ (1) . . .ℎ (푛) 푣 푛 in the gameG ↾푣 0 is compatible with ̄ 휎 −푖 . By hypothesis, we have: 휇 푖 ℎ (1) . . .ℎ (푛) 휋 (푛) < 훼, hence 휇 a ℭ (휒)< 훼 . If Challenger deviates infinitely often.If휒 =[푣 0 ] 휋 (0) . . . ℎ (푛) 푣 푛 휋 (푛) . . . , then the play 휋 =ℎ (1) ℎ (2) . . . is compatible with ̄ 휎 −푖 , and by hypothesis we have 휇 푖 (휋)< 훼 , hence 휇 a ℭ (휒)< 훼 . □ Of course, the abstract negotiation game does not provide a straightforward way to compute the negotiation function, since its vertex space is infinite, and even uncountable, in the general case. However, it gives a first intuition. 68CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION Example 6. Let us consider again the game depicted by Figure 13, and consider the requirement휆 1 , defined by휆 1 (푎) =0, and휆 1 (푏) = 휆 1 (푐) =1. From each of the three vertices, let us present a play of the abstract negotiation game where both Prover and Challenger play optimally. From the vertex푐, Prover proposes the play푐 휔 , and Challenger cannot deviate from it: he may deviate from any vertex controlled by player ◦ , which is the case of the only vertex traversed by the play, but he has no alternative edge to use. Therefore, he accepts the play, and gets the payoff 1, hencenego(휆 2 )(푐) =1. From the vertex푏, Prover is not allowed to propose the play푏 휔 , which is not휆 1 -consistent (player□ should get the payoff 1, and gets only the payoff 0). Therefore, she proposes the play푏푐 휔 , and Challenger accepts, hence nego(휆 1 )(푏) = 1. From the vertex푎, Prover may propose the play푎 휔 . But then, Challenger can deviate to the vertex푏, and from there, again, Prover has to propose a play that eventually reaches the vertex푐, providing player ◦ , and therefore Challenger, the payoff 1, hence nego(휆 1 )(푎) = 1. Thus, checking that a given requirement휆is an (휀-)fixed point of the negotiation function can be done by guessing a finite representation of a strategy for Prover in the abstract negotiation game, from each vertex푢, that forces the player controlling the vertex푢to get, at most, the payoff휆(푢)(or휆(푢)+ 휀). Sometimes (when휀is clear from the context), we will abuse language and call such a strategy winning, assimilating the abstract negotiation game with a Boolean version in which Prover’s goal is to keep Challenger’s payoff under휆(푢)(or휆(푢)+ 휀)—proving, then, that휆is an (휀-)fixed point of the function nego. To obtain effective algorithms from that idea, we need therefore to prove that if such a strategy exists, there exists one that is simple, i.e. that has a finite, and small, representation. 4.4.2 The concrete negotiation game. In the general case, the use that can be made of the abstract negotiation game is limited by its infinite (and uncountable) vertex space. However, under the hypothesis that the gameGis prefix-independent, it can be turned into a game on a finite graph if Prover does not propose plays as a whole, but edge by edge. In the concrete negotiation gameconc 휆푖 (G) ↾푣 c 0 , the vertices controlled by Prover have the form(푣,푀), where 푀 ⊆ 푉memorizes the vertices seen since the last time Challenger did deviate, in order to control that the play Prover is constructing from that point is휆-consistent: for each푢 ∈ 푀, Prover has to give to the player controlling푢at least the payoff휆(푢). Similarly, the vertices controlled by Challenger are of the form(푣 ′ ,푀), where푣 ′ ∈ 퐸is an edge proposed by Prover. The game unfolds as follows, initially from the vertex 푣 c 0 =(푣 0 ,푣 0 ). • From the vertex(푣,푀), Prover chooses an edge푣 ′ and proposes it to Challenger. She therefore moves to Challenger’s vertex(푣 ′ ,푀). • Once an edge푣 ′ has been proposed, Challenger can accept it, moving to Prover’s vertex(푣 ′ ,푀∪푣 ′ ). Or, if 푣 ∈ 푉 푖 , he can deviate, and choose a new edge 푣푤 , moving then to Prover’s vertex(푤,푤). • Then, the game starts again from that new vertex. Again, those proposals and deviations draw a play in the gameG ↾푣 0 , in which Challenger intends to maximize player 푖’s payoff, while Prover intends to minimize it. We give a formal definition below. Definition 27 (Concrete negotiation game). LetGbe a prefix-independent game played on a finite graph, let푖 ∈Πand푣 0 ∈ 푉 푖 , and let휆be a requirement onG. The concrete negotiation game ofG ↾푣 0 is the two-player zero-sum game conc 휆푖 (G) ↾푣 c 0 = 픓,ℭ,푉 c ,(푉 c 픓 ,푉 c ℭ ),퐸 c , 휇 c ↾푣 c 0 , defined as follows: 4.4. NEGOTIATION GAMES69 • player 픓 is called Prover, and player ℭ is called Challenger. •The set of vertices controlled by Prover is푉 c 픓 = 푉 × 2 푉 , where the vertex푣 c = (푣,푀)contains the information of the current vertex푣on which Prover has to define the strategy profile, and the memory푀of the vertices that have been traversed so far since the last deviation, defining the requirements Prover has to satisfy. The initial vertex is 푣 c 0 =(푣 0 ,푣 0 ). •The set of vertices controlled by Challenger is푉 c ℭ = 퐸×2 푉 , where in the vertex푣 c =(푢푣,푀), the edge푢푣 is the edge proposed by Prover. • The set 퐸 c contains three types of edges: proposals, acceptations and deviations. – Proposals are edges in which Prover proposes an edge of the gameG: Prop = (푣,푀)(푣 ′ ,푀) 푣 ′ ∈ 퐸,푀 ∈ 2 푉 . –Acceptations are edges in which Challenger accepts to follow the edge proposed by Prover (it is in particular his only possibility when that edge begins on a vertex that is not controlled by player 푖): Acc = (푣 ′ ,푀) ( 푣 ′ ,푀∪푣 ′ ) 푣 ′ ∈ 퐸,푀 ∈ 2 푉 . Note that the memory is updated. – Deviations are edges in which Challenger refuses to follow the edge proposed by Prover, as he can if that edge begins in a vertex controlled by player푖. The memory is then erased, and only the new vertex the deviating edge leads to is memorized: Dev = (푣 ′ ,푀)(푤,푤) 푣 ∈ 푉 푖 ,푤 ≠ 푣 ′ ,푣 ′ ,푣푤 ∈ 퐸,푀 ∈ 2 푉 . • Let푔 =(ℎ 0 ,푀 0 )(ℎ 0 ℎ ′ 0 ,푀 0 ) . . .(ℎ 푛 ℎ ′ 푛 ,푀 푛 )be a history inconc 휆푖 (G): the projection of the history푔is the history ¤푔 =ℎ 0 . . .ℎ 푛 in the gameG. That definition is naturally extended to plays. •The payoff function휇 c ℭ =−휇 c 픓 measures player푖’s payoff, with a winning condition if the constructed strategy profile is not휆-rational, that is to say if after finitely many player푖’s deviations, it generates a play which is not 휆-consistent: –we define휇 c ℭ (휋) =+∞if after some index푛 ∈ N, the play휋 ≥2푛 contains no deviation, and if the projection ¤휋 ≥푛 is not 휆-consistent 1 ; – and 휇 c ℭ (휋) = 휇 푖 (¤휋) otherwise. Like in the abstract negotiation game, the goal of Challenger is to find a휆-rational strategy profile that forces the worst possible payoff for player푖, and the goal of Prover is to find a possibly deviating strategy for player 푖 that gives them the highest possible payoff. Remark 6. The concrete negotiation game has the following properties. • If the gameG is Borel, then the game conc 휆푖 (G) is Borel. 1 When we combine the notations¤휋and휋 ≥푛 , the notation¤휋is applied first; that is, the play¤휋 ≥푛 is the projection of the play 휋 ≥2푛 , not 휋 ≥푛 . 70CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION •When휋 ≥2푛 contains no deviation, the memory of its vertices is increasing, and therefore eventually equal to the memory푀 = Occ(¤휋 ≥푛 ). If it is the longest such suffix of휋, it means that the projection ¤휋 ≥푛 is 휆-consistent if and only if for each player 푗 and each vertex 푣 ∈ 푉 푗 , we have 휇 푗 (¤휋) ≥ 휆(푣). We can now prove the equivalent of Theorem 9. For convenience, we prove it only when the gameG is Borel. Theorem 10. LetGbe a Borel prefix-independent game played on a finite graph. Let휆be a requirement, let 푖be a player and let푣 0 ∈ 푉 푖 . Then, we have the equalityval ℭ conc 휆푖 (G) ↾푣 c 0 = nego(휆)(푣 0 ). Moreover, if for each player푖and every state푣 0 ∈ 푉 푖 , Prover has an optimal strategy inconc 휆푖 (G) ↾푣 c 0 , thenGis a game with steady negotiation. Proof.▶First direction: the inequality nego(휆)(푣 0 ) ≤ val ℭ conc 휆푖 (G) ↾푣 c 0 . Let휏 픓 be a strategy such thatsup 휏 ℭ 휇 c ℭ ⟨ ̄ 휏⟩ ≠+∞ , and let us define inductively the strategy profile ̄ 휎as follows: for every historyℎ ∈ Hist −푖 G ↾푣 0 compatible with ̄ 휎 −푖 , there exists (by induction) exactly one history푔compatible with휏 픓 such that¤푔 =ℎ. Let then ̄ 휎(¤푔) =푤for every history푔compatible with휏 픓 with휏 픓 (푔) = (푣푤,푀)for some푀. The strategy profile ̄ 휎is arbitrarily defined on other histories. We prove that the strategy profile ̄ 휎 −푖 is휆-rational assuming the strategy휎 푖 , and that sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ sup 휏 ℭ 휇 c ℭ ⟨ ̄ 휏⟩. The strategy profile ̄ 휎 −푖 is휆-rational, assuming the strategy휎 푖 . Indeed, let us assume it is not. Then, there exists a historyℎ =ℎ 0 . . .ℎ 푛 in the gameG ↾푣 0 compatible with ̄ 휎 −푖 such that the play ⟨ ̄ 휎 ↾ℎ ⟩ is not 휆-consistent. Then, let: 푔푣 c = ( ℎ 0 ,푀 0 )( ℎ 0 ̄ 휎(ℎ 0 ),푀 0 ) . . . ( ℎ 푛 ,푀 푛 ) be the only history inconc 휆푖 (G) ↾푣 c 0 compatible with휏 픓 such that¤푔 =ℎ. Let휏 ℭ be a strategy constructing the history ℎ, defined by: 휏 ℭ ( 푔 0 . . .푔 2푘−1 ) =푔 2푘 for every 푘 , and: 휏 ℭ ( 푔 ′ (푣푤,푀) ) =(푤,푀∪푤) for any other history푔 ′ (푣푤,푀). Then, the play휋 =⟨ ̄ 휏⟩contains finitely many deviations (Challenger stops the deviations after having constructed the historyℎ), and the projection¤휋 ≥푛 is not휆-consistent. Therefore, we have 휇 c ℭ (휋) =+∞, which is false by hypothesis. The inequalitysup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ sup 휏 ℭ 휇 c ℭ ⟨ ̄ 휏⟩ holds. Let휎 ′ 푖 be a strategy for player푖, and let 휋 =⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩. Let 휏 ℭ be a strategy such that for every 푘 : 휏 ℭ ( (휋 0 ,·)(휋 0 ·,·) . . .(휋 푘 ·,·) ) =(휋 푘+1 ,·), i.e. a strategy forcing the play휋against휏 픓 . Then, since휇 c ℭ ⟨ ̄ 휏⟩ ≠+∞by hypothesis on휏 픓 , we have 휇 푖 (휋) = 휇 c ℭ ⟨ ̄ 휏⟩, hence 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ sup 휏 ℭ 휇 c ℭ ⟨ ̄ 휏⟩, hence the desired inequality. 4.4. NEGOTIATION GAMES71 푎,푎푏,푎,푎 푏,푎,푏,푏,푎,푏,푏 푏푐,푎,푏푐,푏 푐,푎,푏,푐,푏,푐,푎,푏,푐,푏,푐 Figure 16: A concrete negotiation game Steady negotiation.Moreover, if휏 픓 is optimal, then the휆-rational strategy profile ̄ 휎 −푖 realizes the infimum: inf ̄ 휎 −푖 ∈휆Rat(푣 0 ) sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩, hence if there exists such an optimal strategy for every vertex푣 0 , then the gameGis with steady negotiation. ▶Second direction: the inequality val ℭ conc 휆푖 (G) ↾푣 c 0 ≤ nego(휆)(푣 0 ). Let ̄ 휎 −푖 be a휆-rational strategy profile from푣 0 , assuming the strategy휎 푖 ; let us define a strategy 휏 픓 , by휏 픓 (푔(푣,·)) = ( 푣 ̄ 휎(¤푔푣),· ) for every history푔and for every푣 ∈ 푉. Let us prove the inequality sup 휏 ℭ 휇 c ℭ ⟨ ̄ 휏⟩ ≤ sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩. Let휏 ℭ be a strategy for Challenger, and let휋 =⟨ ̄ 휏⟩. If휇 c ℭ (휋) =+∞ , then there exists푛such that the play휋 ≥2푛 contains no deviation, i.e.¤휋 ≥푛 = ⟨ ̄ 휎 ↾¤휋 ≤푛 ⟩, and that play is not휆-consistent, which is impossible. Therefore, we have휇 c ℭ (휋) ≠+∞, and as a consequence휇 c ℭ (휋) = 휇 푖 (¤휋) = 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩for some strategy 휎 ′ 푖 , hence 휇 c ℭ (휋) ≤ sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩, hence the desired inequality.□ Example 7. Consider again the game of Figure 13. The arena of the concrete negotiation game from the vertex푎is depicted by Figure 16. The blue vertices belong to Prover, the orange ones to Challenger. The dashed arrows depict the deviations, and optimal strategies for both Prover and Challenger, for the requirement 휆 1 defined by 휆 1 (푎) = 0 and 휆 1 (푏) = 휆 1 (푐) = 1, are defined by the thick arrows. Contrary to the abstract negotiation game, the concrete negotiation game has a finite vertex space. However, the size of that vertex space is exponential in the size of the original game: constructing that game and solving it is therefore costly. The concrete negotiation game will nevertheless be used to prove the fixed-parameter tractability of the SPE constrained existence problem in parity games (Theorem 12), 72CHAPTER 4. SUBGAME-PERFECT EQUILIBRIA AND NEGOTIATION and to establish an intermediary result about the fixed points of the negotiation function in mean-payoff games (Lemma 16). CHAPTER 5: PARITY GAMES In this chapter, we use the tools given in Chapter 4 to proveNP-completeness for the constrained existence problem of SPEs in parity games. Note that, here, we are only interested in SPEs and not휀-SPEs, since parity games are Boolean games, in which the notion of 휀-SPE has little relevance. In order to design an efficient algorithm for that problem, we define an equivalence relation between histories and between plays; and, then, we show that in the abstract negotiation game, Prover can propose only plays that are simple representatives of their equivalence class, and propose always the same play from each vertex. 5.1REDUCED PLAYS AND REDUCED STRATEGIES 5.1.1 Definitions The equivalence relation that we use is based on the order in which vertices appear. Definition 28 (Occurrence-equivalence). Two historiesℎandℎ ′ are occurrence-equivalent, writtenℎ ≈ ℎ ′ , if and only if first(ℎ) = first(ℎ ′ ), last(ℎ) = last(ℎ ′ ) and Occ(ℎ) = Occ(ℎ ′ ). Two plays휋and휋 ′ are occurrence-equivalent, written휋 ≈ 휋 ′ , if and only if the three following conditions are satisfied: • we have Inf(휋) = Inf(휋 ′ ); • for each history prefix of 휋 , there exists a occurrence-equivalent history prefix of 휋 ′ ; • for each history prefix of 휋 ′ , there exists a occurrence-equivalent history prefix of 휋 . Example 8. Let us consider the game of Figure 17. In that game, the play푎푏(푐푑푐푒) 휔 is occurrence- equivalent to the play푎푏푐푑(푐푑푐푒) 휔 , but not to the play푎푏(푐푒푐푑) 휔 . Indeed, the latter has the history푎푏푐푒 푎 푏 푐 푑 푒 ◦ 0 □ 1 1 ◦ 0 □ 0 1 ◦ 1 □ 1 1 ◦ 1 □ 0 1 ◦ 1 □ 1 0 (휆 ∗ ) 0 11 0 0 Figure 17: A Büchi game. 73 74CHAPTER 5. PARITY GAMES as a prefix, which is not occurrence-equivalent to any prefix of푎푏(푐푑푐푒) 휔 , in which the vertex푒occurs only when the vertex 푑 has already occurred. Remark 7. The operatorsOccandInf, and any parity payoff function, are stable by occurrence- equivalence. 5.1.2 Representatives The interest of that equivalence relation lies in the finite number of its equivalence classes, and by the existence of simple representatives for each of them. Lemma 6. Let휋be a play ofG. There exists a lassoℎ푐 휔 ≈ 휋where the historyℎhas length|ℎ| ≤ 푛 3 +푛 2 and the cycle 푐 has length|푐| ≤ 푛 2 , where 푛 = card푉 . Proof. Let us write푊 0 ⊂ · ⊂ 푊 푡 for all the sets of the formOcc(휋 ≤푘 )with푘 ∈ N, without repetition. Note that for each index푠 ∈ 0, . . .,푡 −1, the set푊 푠+1 contains the set푊 푠 plus one additional vertex. Let us construct the historyℎand the cycle푐as follows, maintaining the hypothesis that for all푝, the set Occ(ℎ ≤푝 ) is equal to푊 푠 for some 푠. Base case. First, we define ℎ 0 = 휋 0 , andℎ 0 =푊 0 . Inductive case.Then, when the prefixℎ ≤푝 is constructed: let푠be such thatOcc(ℎ ≤푝 ) =푊 푠 , and let푘be the minimal integer such thatOcc(휋 ≤푘 ) =푊 푠 . Let푈be the set of all the vertices푢such that there existsℓwithOcc(휋 ≤ℓ ) =푊 푠 and휋 ℓ =푢: any suchℓis greater than or equal to푘, and푈 ⊆ 푊 푠 . Then, there exists at least one path from휋 푘 =ℎ 푝 that traverses all the vertices of푈and only them. Let ℎ 푝 . . .ℎ 푞 be such a path with minimal length: it has at most length푛 2 . If푠< 푡, let nowℓbe the minimal index greater than푘such thatOcc(휋 ≤ℓ ) = 푊 푠+1 . Then, there exists a path fromℎ 푞 ∈ 푊 푠 to휋 ℓ that uses only vertices of푈: letℎ 푞 . . .ℎ 푟 be such a path with minimal length. Then, it traverses all vertices at most once, and has therefore length at most푛. If푠 = 푡, letℎ 푞 . . .ℎ 푟 be a path of minimal length from ℎ 푞 to a vertexℎ 푟 ∈ Inf(휋): for the same reasons as above, such a path exists and has length at most 푛. Then, we can stop here the construction ofℎ, and observe that the vertexℎ 푟 belongs to the graph (Inf(휋),퐸∩ Inf(휋) 2 ), which is strongly connected. We can therefore choose a cycle푐that traverses all its vertices and only them, and that has length at most 푛 2 . Conclusion. By construction, the lassoℎ푐 휔 is occurrence-equivalent to휋, and satisfies the desired size conditions.□ We call such lassos reduced plays. For each play휋, we write ̃ 휋for an arbitrary occurrence-equivalent reduced play. Then, givenGand a reduced play ̃ 휋, operations such as computing the vector휇( ̃ 휋), the sets Occ( ̃ 휋) and Inf( ̃ 휋), or checking whether the play ̃ 휋 is 휆-consistent, can be done in time 푂(푛 3 ). Example 9. In the game of Figure 17, the play푎푏(푐푑푐푒) 휔 is reduced, but the occurrence-equivalent play 푎푏 150 ( 푐푑푐푒 ) 휔 is not. Definition 29 (Reduced strategy). A strategy휏 픓 for Prover inabs 휆푖 (G)is reduced if and only if it is stationary, and for each vertex 푣 , the play 휋 with [휋] =휏 픓 ([푣]) is a reduced play. 5.1. REDUCED PLAYS AND REDUCED STRATEGIES75 If we have휋 ≈ 휋 ′ , and if Challenger can deviate from the play휋after the historyℎ푣, then he can also deviate in휋 ′ after some historyℎ ′ 푣that traverses the same vertices. Thus, Prover can play optimally while proposing only reduced plays, and by proposing always the same play from each vertex; that is, by following a reduced strategy. Note that since we will use the abstract negotiation game to check whether a given requirement is a fixed point of the negotiation function, we abuse language here by considering that Prover wins if she maintains player푖’s payoff below휆(푣 0 ), even though the abstract negotiation game is not exactly Boolean. Lemma 7. Prover has a winning strategy in the abstract negotiation game if and only if she has a reduced one. Proof.Let us consider the reduced negotiation game, i.e. the abstract negotiation game in which one would have removed all the vertices but: •Prover’s vertices, i.e., the set푉 a 픓 =푉 ∪⊤,⊥; •those of the form [ ̃ 휋], where ̃ 휋 is a reduced play; •those of the form[ℎ푣], whereℎis a prefix of a휆-consistent reduced play ̃ 휋 , and has minimal length among the occurrence-equivalent prefixes of ̃ 휋 . If the gameGhas푛vertices, then this game has at most푛 푛 3 +2푛 2 (푛 3 +3푛 2 )+1 vertices. Let us now notice that it has another interesting property. Sublemma 1. In the reduced negotiation game, either Prover or Challenger has a stationary winning strategy. Proof.By Lemma 2 (and since the reduced game is Borel), that is the case if both Prover’s and Challenger’s objectives are convex. Let therefore휋and휒be two plays in the reduced negotiation game, and휉be a shuffling of휋and휒. Then, we observe that the play ¤ 휉is also a shuffling of the plays¤휋and¤휒. The convexity of Prover’s and Challenger’s objectives is then a consequence of the convexity of parity objectives: the minimal color seen infinitely often by player푖in ¤ 휉is the minimum of the minimal colors seen infinitely often in ¤휋 and in ¤휒 .□ We can now prove our lemma by using the equivalence between the abstract and the reduced negotiation game. If Prover has a winning strategy in the abstract negotiation game, she has one in the reduced negotiation game. We proceed by contraposition: if Prover has no winning strategy in the reduced game, then, by Sublemma 1, Challenger has a stationary one: let us write it 휏 ℭ . Now, let us extend 휏 ℭ into a stationary winning strategy 휏 a ℭ in the abstract negotiation game. Let[휋] ∈ 푉 a ℭ , and let[ℎ푣푤] = 휏 ℭ ([ ̃ 휋]). By occurrence-equivalence, there exists푘 ∈ Nsuch that휋 푘 = 푣, andOcc(휋 ≤푘 ) = Occ(ℎ푣). We then set휏 a ℭ ([휋]) = [휋 ≤푘 푤] . Note that the vertices 휏 a ℭ ([휋]) =[휋 ≤푘 푤] and휏 ℭ ([ ̃ 휋]) =[ℎ푣푤]can be different, but they both have as unique successor the vertex [푤]. Let us prove, now, that휏 a ℭ is winning: let휒 a be a play compatible with휏 a ℭ . When Prover proposes a play휋, in the abstract game, against the strategy휏 a ℭ , and when she proposes the play ̃ 휋in the reduced game against the strategy휏 ℭ , the same thing happens in both cases: either Challenger accepts in both 76CHAPTER 5. PARITY GAMES games, or he deviates, and Prover has to propose a new play from the same vertex푤. Therefore, we can define from휒 a a play휒compatible with휏 ℭ in the reduced game, in which each Challenger’s vertex [휋] is replaced by the vertex [ ̃ 휋], and Prover’s vertices are replaced accordingly. Since the play휒is compatible with휏 ℭ , it is winning for Challenger: let us then prove that so is 휒 a . If휒 a has the form푔[휋]⊤ 휔 , then휒has the form푔 ′ [ ̃ 휋]⊤ 휔 , hence ̃ 휋 is winning for player푖, and therefore휋is winning for player푖and휒 a is winning for Challenger. If휒never reaches the vertex⊤, then we have¤휒 =ℎ 0 ℎ 1 . . .and¤휒 a =ℎ 0a ℎ 1a . . .where, for every푘, the historiesℎ 푘 andℎ 푘 a are possibly different, but contain exactly the same vertices. Then, the set of player푖’s colors appearing infinitely often in¤휒and¤휒 a are the same, and the play¤휒 a is winning for player푖, i.e. the play휒 a is winning for Challenger: the strategy 휏 a ℭ is winning. If Prover has a winning strategy in the reduced negotiation game, she has a reduced one in the abstract negotiation game.Indeed, let휏 픓 be a winning strategy for Prover in the reduced negotiation game. As said above, we can assume that휏 픓 is stationary. Since in the abstract game, the only vertices controlled by Prover where she has several possible choices are the ones of the form[푣], for푣 ∈ 푉, we can see휏 픓 as a reduced strategy in the abstract game: let us write it휏 a 픓 in that case. We now have to prove that 휏 a 픓 is also a winning strategy. Let휒 a be a play in the abstract negotiation game compatible with휏 a 픓 . For any sequence of vertices [푣][ ̃ 휋][ℎ푤][푤]that appears in휒 a , the historyℎis occurrence-equivalent to some prefix ̃ ℎof ̃ 휋such that[ ̃ ℎ푤]is a vertex of the reduced game. Therefore, we can transform the play휒 a into a play휒of the reduced game, compatible with휏 픓 , where each vertex of the form[ℎ푤]have been replaced by[ ̃ ℎ푤]. Since휏 픓 is a winning strategy in the reduced negotiation game, the play휒is winning for Prover. Let us prove that so is 휒 a . If휒 a has the form푔[ ̃ 휋]⊤ 휔 , then휒has the form푔 ′ [ ̃ 휋]⊤ 휔 , and since휒is winning for Prover, the play ̃ 휋 is losing for player 푖, and therefore the play 휒 a is winning for Prover. If 휒 a has the form: 휒 a =[푣 0 ][ ̃ 휋 0 ][ℎ 0 푣 1 ][푣 1 ][ ̃ 휋 1 ][ℎ 1 푣 2 ] . . . then we have: 휒 =[푣 0 ][ ̃ 휋 0 ][ ̃ ℎ 0 푣 1 ][푣 1 ][ ̃ 휋 1 ][ ̃ ℎ 1 푣 2 ] . . . and since for each푘, we haveOcc(ℎ 푘 ) = Occ( ̃ ℎ 푘 ), we findInf(ℎ 0 ℎ 1 . . .) = Inf( ̃ ℎ 0 ̃ ℎ 1 . . .)and therefore, if the play 휒 is winning for Prover, so is the play 휒 a . The strategy 휏 a 픓 is a reduced winning strategy. Therefore, to conclude, if Prover has a winning strategy in the abstract negotiation game, she has a reduced one.□ 5.2CHECKING THAT A REDUCED STRATEGY IS WINNING We have established that Prover is winning the abstract negotiation game if and only if she has a reduced winning strategy. Such a strategy has polynomial size and can thus be guessed in nondeterministic polynomial time. We must now show that we can verify in deterministic polynomial time that a guessed strategy is winning. Lemma 8. Given a parity gameG, a requirement휆, a player푖, a vertex푣 0 ∈ 푉 푖 , and a reduced strategy 휏 픓 in the abstract negotiation gameabs 휆푖 (G) ↾푣 0 , deciding whether휏 픓 is a winning strategy can be done in 5.2. CHECKING THAT A REDUCED STRATEGY IS WINNING77 polynomial time. Proof. Let us consider the induced gameabs 휆푖 (G) ↾푣 0 [휏 픓 ], as defined in Section 2.4. In that game, there remain only: •Prover’s vertices, i.e., the set푉 ∪⊤,⊥; •Challenger’s vertices of the form[ ̃ 휋], where there exists푣 ∈ 푉from which the strategy휏 픓 proposes the reduced play ̃ 휋 ; • Challenger’s vertices of the form[ℎ푣]accessible from one vertex ̃ 휋as defined in the previous item, that have minimal length among occurrence-equivalent prefixes of ̃ 휋 ; and the edges binding those vertices. It therefore contains at most푛+푛+푛(푛 3 +2푛 2 )vertices (at most 푛of the form푣, at most푛of the form[ ̃ 휋], and at most푛(푛 3 +2푛 2 )of the form[ℎ푣]: assuming each play ̃ 휋 has the formℎ푐 휔 with|ℎ푐| ≤ 푛 3 +2푛 2 , a minimal prefix is necessarily a prefix ofℎ푐). Deciding whether Challenger can win against the strategy휏 픓 amounts, then, to looking for a play in that game that is winning for him, i.e. either (assuming휆(푣 0 ) =0—otherwise, Challenger cannot win except if Prover gives up, which is excluded by the definition of a reduced strategy): •that reaches a vertex [ ̃ 휋] where ̃ 휋 is won by player 푖 and then goes to the vertex⊤; •that has the form [푣 0 ][ ̃ 휋 0 ][ℎ 0 푣 1 ][푣 1 ][ ̃ 휋 1 ] . . . where the play ℎ 0 ℎ 1 . . . is won by player 푖. In the first case, the existence of such a play can clearly be decided in polynomial time. In the second case, looking for such a play amounts to looking for a play satisfying the parity condition defined by휅([푣]) = 푚for each vertex푣 ∈ 푉, by휅([ ̃ 휋]) = 푚for each play ̃ 휋 , and by휅([ℎ푣]) = min 푘 휅 푖 (ℎ 푘 ), where푚 is the maximal color in the gameG. That can also be done in polynomial time.□ We have then all the necessary ingredients for an algorithm solving the SPE constrained existence problem. Indeed, given a gameG ↾푣 0 , two thresholds ̄ 푥, ̄ 푦 ∈ (Q∪±∞) Π and a requirement휆, if there exists a휆-consistent play that generates a payoff vector between ̄ 푥and ̄ 푦, then it is the case of every occurrence-equivalent play, and therefore, there exists a reduced one that satisfies those conditions. Therefore, an algorithm solving the SPE constrained existence problem consists in: • guessing a requirement 휆 onG with values 0 or 1; • guessing a reduced strategy휏 푢 픓 in each abstract negotiation gameabs 휆푖 (G) ↾푢 with푖 ∈Πand푢 ∈ 푉 푖 ; • guessing a reduced play ̃ 휋 inG ↾푣 0 ; • checking that ̃ 휋 is 휆-consistent and that ̄ 푥 ≤ 휇( ̃ 휋) ≤ ̄ 푦; • checking that each 휏 푣 픓 is winning (and thus that nego(휆) = 휆). By Theorems 8 and Lemmas 7 and 8, that algorithm is correct and runs in non-deterministic polynomial time. Finally, it is already known, since the work of Erich Grädel and Michael Ummels [GU08], that the constrained existence problem of SPEs in parity (and co-Büchi) games isNP-hard, even with only one effective lower threshold. Once again, a quick modification of their proof, by adding one player, entails hardness with only one effective upper threshold. Hence the following. Theorem 11. In parity games, the constrained existence problem of SPEs isNP-complete. Hardness still holds in co-Büchi games, and when there is no effective lower threshold, and only one effective upper threshold. 78CHAPTER 5. PARITY GAMES While the NE constrained existence problem in Büchi games has been proved to be decidable in polynomial time, the problem is, to the best of our knowledge, still open for the SPE constrained existence problem. Open Problem 1. Is the SPE constrained existence problem in Büchi games decidable in polynomial time? 5.3A DETERMINISTIC UPPER BOUND We end this chapter by mentioning an additional complexity result on the constrained existence problem of SPEs in parity games: it is fixed-parameter tractable. Theorem 12. The SPE constrained existence problem on parity games is fixed-parameter tractable when the number of players and the number of colors are fixed. More precisely, there exists a deterministic algorithm that solves that problem in time푂(2 2 푝푚 푛 12 ), where푛is the number of vertices, where푝is the number of players and where푚 is the number of colors. Proof. LetGbe a parity game, let푖 ∈Π, and let푢 ∈ 푉 푖 . Let us assume, without loss of generality, that푚 is even, and that the vertices ofGor labeled by the colors 0, . . .,푚−1. Let휆be a requirement. The value ofnego(휆)(푢)can be computed as the value of the corresponding concrete negotiation game. Let us recall that in that game, Challenger wins a play 휒 if and only if either: •player 푖 wins the projection ¤휒 ; •or Challenger stops deviating, and the play proposed by Prover after the last deviation is not 휆-consistent. Moreover, since parity games are Boolean, a vertex푣that is visited after the last deviation induces either no constraint (if휆(푣) ≤0), or the constraint that the player controlling푣must win (if휆(푣)>0). Thus, the game need only memorize the set of players푖such that a vertex푣 ∈ 푉 푖 with휆(푣)>0 has been visited. We therefore slightly modify the definition of the concrete game, to use a version in which the vertices are of the form(푣,푀)or(푢푣,푀)with푀 ⊆Π, instead of푀 ⊆ 푉, which gives it size at most(푛 2 +푛)2 푝 instead of(푛 2 +푛)2 푛 . Thus, that game can be seen as a multi-parity game, where a color function is defined on edges (and not on vertices as it usually is) for each dimension 푑 ∈Π∪★: •if 푑 =★, then for each edge(푢,푀)(푢푣,푀) or(푢푣,푀)(푤,푀 ′ ), we define: ˆ 휅 ★ ((푢,푀)(푢푣,푀)) = ˆ 휅 ★ ((푢푣,푀)(푤,푀 ′ )) =휅 푖 (푢); •if 푑 = 푗 ∈Π, then for each edge of the form(푢,푀)(푢푣,푀) we define: ˆ 휅 푗 ((푢,푀)(푢푣,푀)) =푚− 1, and for each edge of the form(푢푣,푀)(푤,푀 ′ ) we define ˆ 휅 푗 ((푢푣,푀)(푤,푀 ′ )) = 1 if푤 ≠ 푣(i.e., if the edge is a deviation) and ˆ 휅 푗 ((푢푣,푀)(푤,푀 ′ )) =휅 푗 (푢푣)+1 if푤 = 푣(i.e., if it is not a deviation); and where Challenger’s goal consists in satisfying at least one of the corresponding parity conditions. 5.3. A DETERMINISTIC UPPER BOUND79 It is a multi-parity game with colors on edges, but one can easily transform it into a game with colors on vertices, up to adding one vertex in the middle of each edge, i.e. decomposing each edge 푢 c 푣 c into two edges푢 c 훿 푢 c 푣 c and 훿 푢 c 푣 c 푣 c . Then, that game can also be interpreted as a Boolean Büchi game in the sense of [BHR18a], i.e. a two-player zero-sum game in which the objective of the first player (Prover) is to validate a Boolean formula whose atoms are Büchi conditions. Indeed, Prover’s objective can be written: Û 푑∈Π∪★ 푚 2 Ü 푘=0 ( B훿 푢 c 푣 c | 휅 푑 (푢 c 푣 c ) = 2푘∧¬B훿 푢 c 푣 c | 휅 푑 (푢 c 푣 c )< 2푘 ) , where B(푊)is the Büchi objective associated to the set푊, i.e. the objective of visiting infinitely often at least one vertex of푊. That formula has at most푑푚atoms, and has size푑푚, in a game of size 푂(푛 2 2 푝 ), hence by Proposition 5 from [BHR18a], there exists a deterministic algorithm that decides which player has a winning strategy in time: 푂 2 2 (푝+1)푚 (푝+ 1)푚+ 2 (푝+1)푚2 (푝+1)푚 푂(푛 2 2 푝 ) 5 = 2 2 푂(푝푚) 푛 10 . Therefore, given a requirement휆, by constructing the concrete negotiation game and applying that algorithm on each vertex, it is possible to compute the requirementnego(휆)in time 2 2 푂(푝푚) 푛 11 . Thus, it is possible to compute the iterations of the negotiation function from the vacuous requirement 휆 0 : since the functionnegois non-decreasing, its least fixed point휆 ∗ (whose existence is guaranteed by Tarski’s fixed point theorem, since the set of requirements is a complete lattice) will be reached in at most 푛 steps, and will therefore be found in time 2 2 푂(푝푚) 푛 12 . Once휆 ∗ has been computed, given two thresholds ̄ 푥and ̄ 푦, the SPE constrained existence problem can be solved by searching, for each tuple ̄ 푧 ∈ 0, . . .,푚−1 Π where푧 푖 is even whenever푥 푖 =1 and odd whenever푦 푖 = 0, a play 휋 inG that avoids the set: 푊 ̄ 푧 =푣 ∈ 푉 푖 | 푧 푖 ∈ 2Z and 휆 ∗ (푣) = 1, and such that for each푖, we havemin휅 푖 (Inf(휋)) = 푧 푖 . For a given tuple ̄ 푧 , the existence of a play can be decided by removing all the vertices of푊 ̄ 푧 and the vertices푣such that휅 푖 (푣)< 푧 푖 for some푖, then looking for a strongly connected component that contains at least one vertex푣with휅 푖 (푣) = 푧 푖 for each푖, and finally check whether the vertices of that strongly connected component are accessible from the initial vertex inGwithout visiting the vertices of푊 ̄ 푧 . All those computations can be done in time푂(푛). Thus, once휆 ∗ has be computed, the SPE constrained existence problem can be solved in time 푂(푚 푝 푛). Given a parity gameG ↾푣 0 and two thresholds ̄ 푥and ̄ 푦, solving the SPE constrained existence problem can be done in time 2 2 푂(푝푚) 푛 12 , and is therefore fixed-parameter tractable with parameters푚 and 푝.□ 80CHAPTER 5. PARITY GAMES CHAPTER 6: MEAN-PAYOFF GAMES Let us now move to the study of mean-payoff games. Here, a similar algorithm will be used but with additional steps, due to the fact that optimal strategies in the abstract negotiation game may not be stationary. Moreover, we will not only consider the constrained existence of SPEs, but also of휀-SPEs, where휀 ≥0 is given with the instance. We will see that this relaxation does not require specific additional tools, nor entail additionnal complexity. 6.1HARDNESS We first show a lower bound for our problem. Lemma 9. The constrained existence problem of휀-SPE in mean-payoff games isNP-hard, even when휀is fixed equal to 0, and when there is no effective lower threshold and only one effective upper threshold. Proof.The structure of this proof is strongly inspired from several proofs from Michael Ummels, in articles that have already been cited above. We proceed by reduction from theNP-complete problem Sat. Let휑 = Ó 푛 푖=1 Ô 푚 푗=1 퐿 푖푗 be a formula from propositional logic, written in conjunctive normal form, over the finite variable set푋. We construct a mean-payoff gameG 휑 ↾푣 0 that admits an SPE where the player 픚 gets the payoff 1, if and only if 휑 is satisfiable. ▶Construction of the gameG 휑 First, we define the set of playersΠ =픖∪ 푋: every variable of휑is a player and there are two additional players 픖 and 픚, called Solver and Witness, respectively. Then, let us define the vertex space: for each clause퐶 푖 , with푖 ∈ Z/푛Z, of휑, we define a vertex퐶 푖 that is controlled by Solver, and for each literal퐿 푖푗 of퐶 푖 we define a vertex(퐶 푖 ,퐿 푖푗 ), that is controlled by the player푥such that퐿 푖푗 = 푥or¬푥. Witness controls no vertex. We add an edge from퐶 푖 to(퐶 푖 ,퐿 푖푗 ), and another one from(퐶 푖 ,퐿 푖푗 )to퐶 푖+1 . Moreover, we add a sink vertex⊥, with an edge from it to itself, and edges from all the vertices of the form(퐶,¬푥) to⊥. We define the reward function 푟 on this game as follows: • 푟 픖 (⊥) = 0, and 푟 픖 (푢푣) = 1 for any other edge 푣푤 ; • 푟 픚 (⊥) = 1, and 푟 픚 (푢푣) = 0 for any other edge 푣푤 ; •for each player푥, we have푟 푥 (푢푣) =0 for every edge leading to a vertex of the form푣 =(퐶,푥), and 푟 푥 (푢푣) = 1 for any other edge. 81 82CHAPTER 6. MEAN-PAYOFF GAMES 퐶 1 푥 1 ¬푥 1 퐶 2 푥 2 ¬푥 2 퐶 3 푥 3 ¬푥 3 퐶 4 푥 4 ¬푥 4 퐶 5 푥 5 ¬푥 5 퐶 6 푥 6 ¬푥 6 ⊥ 푥 1 0 푥 2 0 푥 3 0 푥 4 0 푥 5 0 푥 6 0 픖 0 픚 1 Figure 18: The gameG 휑 Note that Solver and Witness can only get the payoffs 0 or 1, and that in every play, exactly one of them gets the payoff 1. Another player푥gets the payoff 1 in a play that never visits (or finitely often, or infinitely often but with negligible frequence) a vertex of the form(퐶,푥). Otherwise, he may get any payoff between 0.5 and 1, depending on the frequence with which such a vertex is visited. Finally, we initialize that game in 푣 0 =퐶 1 . Example 10. The gameG 휑 , when휑is the tautology(푥 1 ∨¬푥 1 )∧·∧(푥 6 ∨¬푥 6 ), is represented by Figure 18. The rewards that are not written are equal to 1, or to 0 in the case of Witness. Now, let us prove that there is an SPE inG 휑 ↾퐶 1 in which Solver gets the payoff 1, if and only if the formula 휑 is satisfiable, that is, if there exists a valuation 휈 : 푋 !0, 1 that satisfies it. ▶If there is an SPE inG 휑 ↾퐶 1 in which Witness gets the payoff 0, then 휑 is satisfiable. Let us write ̄ 휎for such an SPE, and let휋 = ⟨ ̄ 휎⟩. Since휇 픚 (휋) =0, the sink vertex⊥is never visited. Let us define a valuation휈on푋as follows: for each variable푥, we have휈(푥) =1 if and only if 휇 푥 (휋)< 1. Now, let퐶be a clause of휑: since퐶, as a vertex, is necessarily visited infinitely often and with a 6.1. HARDNESS83 fixed frequence in the play휋(because no player ever go to the sink vertex⊥), one of its successors, say (퐶,퐿), is visited with a non-negligible frequence (more formally, the time between two occurrences of (퐶,퐿)is bounded). If퐿is a positive literal, say푥, then by definition of휈, we have휈(푥) =1 and the clause퐶 is satisfied. If퐿has the form¬푥, then each time the vertex(퐶,¬푥)is traversed, player푥has the possibility to deviate and to go to the sink vertex⊥, where he is sure to get the payoff 1. Since ̄ 휎 is an SPE, it means that he already gets the payoff 1 in the play휋. By definition of휈, we then have휈(푥) =0, hence the literal¬푥 is satisfied, hence so is the clause퐶. The valuation 휈 satisfies all the clauses of 휑 , and therefore satisfies the formula 휑 itself. ▶If 휑 is satisfiable, then there is an SPE in which Witness gets payoff 0. Let 휈 be a valuation satisfying 휑 , and let us define a strategy profile ̄ 휎 by: • 휎 픖 (ℎ퐶) = (퐶,퐿)for each historyℎ퐶where퐶is a clause of휑, where퐿is a literal of퐶that is satisfied in the valuation 휈 ; •and휎 푥 (ℎ(퐶,¬푥)) =⊥if and only if휈(푥) =1 for each historyℎ(퐶,¬푥)where퐶is a clause of휑 and 푥 is a variable. Any other vertex has only one successor, hence we now have completely defined a strategy profile. Now, let us prove it is an SPE, in which Witness gets the payoff 0. Letℎ퐶be a history, where퐶is a clause of휑. We want to prove that ̄ 휎 ↾ℎ퐶 is a Nash equilibrium, in which Witness gets the payoff 0. Let휋 =⟨ ̄ 휎 ↾ℎ퐶 ⟩. If휇 픚 (휋) =1, i.e. if휋is of the formℎ퐷(퐷,¬푥)⊥ 휔 , then by definition of ̄ 휎 we have휈(푥) =0. But then, we cannot have휎 픖 (퐷) =(퐷,¬푥): contradiction. The play휋never reaches the vertex⊥, and Witness gets the payoff 0. Since Witness controls no vertex, he has no profitable deviation. On her side, Solver gets the payoff 1, and as a consequence she does not have any profitable deviation. Now, if another player푥has a profitable deviation, it means that he does not get the payoff 1 in휋, and therefore that some vertex of the form(퐷,푥)is visited infinitely often. But then, if Solver chooses to go to the vertex(퐷,푥), it means that the literal푥is satisfied in휈, i.e. that휈(푥) =1. In that case, if some clause퐷 ′ contains the literal¬푥, it is not a literal satisfied by휈, and therefore the strategy휎 픖 , as we defined it, never chooses the edge to the vertex(퐷 ′ ,¬푥), where player푥could have the possibility to deviate from his strategy. Contradiction. Finally, after a history of the form ℎ(퐶,퐿), either: •we have퐿 =¬푥with휈(푥) =1, and in that case, we have⟨ ̄ 휎 ↾ℎ(퐶,퐿) ⟩ =(퐶,퐿)⊥ 휔 , player푥gets the payoff 1, and no player has a profitable deviation; • or퐿is a positive literal, and then there exists only one edge from the vertex(퐶,퐿)to another clause 퐷 , and we go back to the previous case; •or we have퐿 =¬푥with휈(푥) =0, and in that case, we have휎 푥 (퐶,퐿) = 퐷where퐷is the following clause, and by the first case the strategy profile ̄ 휎 ↾ℎ(퐶,¬푥)퐷 is a Nash equilibrium. Moreover, since the literal¬푥is not satisfied in휈, the play⟨ ̄ 휎 ↾ℎ(퐶,¬푥)퐷 ⟩does never traverse again any vertex of the form(퐷 ′ ,¬푥), hence player푥wins, and therefore has no profitable deviation: the strategy profile ̄ 휎 ↾ℎ(퐶,¬푥) is a Nash equilibrium. 84CHAPTER 6. MEAN-PAYOFF GAMES The constrained existence problem of SPEs, and therefore of휀-SPEs, isNP-hard in mean-payoff games.□ The rest of this chapter is dedicated to prove that the same problem isNP-easy, and thereforeNP- complete. 6.2MANIPULATING SETS OF PAYOFF VECTORS This section provides us with a toolbox to compute and manipulate the sets of payoff vectors that can be achieved in a mean-payoff game. 6.2.1 Achievable payoff vectors in a mean-payoff game A first important result that we need is the characterization of the set of possible payoff vectors in a mean-payoff game, which has been introduced in [CDE + 10]. Let us recall that given a graph(푉,퐸), we denote bySCyc(푉,퐸)the set of simple cycles it contains. Given a finite set퐷of dimensions and a set 푋 ⊆ R 퐷 , we writeConv푋for the convex hull of푋, i.e., the set of vectors ̄ 푦 such that there exists a family of positive real numbers(훼 ̄ 푥 ) ̄ 푥∈푋 that satisfies the equalities Í ̄ 푥∈푋 훼 ̄ 푥 =1 and Í ̄ 푥∈푋 훼 ̄ 푥 ̄ 푥 = ̄ 푦. We will often use the subscript notation Conv 푥∈푋 푓(푥) for the set Conv푓(푋). Definition 30 (Downward sealing). Given a set 푌 ⊆ R 퐷 , the downward sealing of 푌 is the set: ⌞ 푌 = min ̄ 푧∈푍 푧 푑 푑∈퐷 푍 is a finite subset of 푌 . Lemma 10 ([CDE + 10]). LetGbe a mean-payoff game, whose underlying graph is strongly connected. Then, we have the equality: 휇 ( PlaysG ) = ⌞ Conv 푐∈SCyc(푉,퐸) mp(푐) . Example 11. If the set푌is depicted by the blue area in Figure 19b, then the set ⌞ 푌 is obtained by adding the gray area. As a consequence, ifGis the game depicted by Figure 19a, then the blue and the gray area in Figure 19b form the set of achievable payoffs inG. Indeed, following exclusively one of the three simple cycles푎,푎푏and푏of the game graph during a play yields the payoffs 01,10 and 22, respectively. By combining those cycles with well chosen frequencies, one can obtain any payoff in the convex hull of those three points. It is also possible to obtain the point 00 by using the properties of the limit inferior: it is for instance the payoff vector of the play 푎 2 푏 4 푎 16 푏 256 . . .푎 2 2 푛 푏 2 2 푛+1 . . . . Then, by combining that play with the simple cycles, one can construct a play that yields any payoff in the convex hull of the four points(0, 0),(0, 1),(1, 0), and(2, 2). 6.2.2 Equations and inequations As the experienced reader might have guessed after reading the previous subsection, our algorithms for mean-payoff games will require cautious manipulation of sets of real numbers. The proof of Theorem 17 below, in particular, requires manipulation of polytopes, e.g. downward sealings of convex hulls (from Lemma 10), expressed as solution sets of systems of linear inequations. 6.2. MANIPULATING SETS OF PAYOFF VECTORS85 푎 푏 ◦ 2 □ 2 ◦ 2 □ 2 ◦ 0 □ 1 ◦ 1 □ 0 (a) The gameG ◦ □ 1 2 012 (b) The payoffs of plays and SPE outcomes inG Figure 19: An example of mean-payoff game Definition 31 (Linear equations, inequations, systems). Let퐷be a finite set. A linear equation inR 퐷 is a pair( ̄ 푎,푏) ∈ R 퐷 \ ̄ 0 ×R . The solution set of the equation( ̄ 푎,푏)is the setSol = ( ̄ 푎,푏) = ̄ 푥 ∈ R 퐷 | ̄ 푎· ̄ 푥 =푏, where·denotes the canonical scalar product on the euclidian spaceR 퐷 . A set푋 ⊆ R 퐷 is a hyperplane of R 퐷 if it is the solution set of some linear equation. A system of linear equations is a finite setΣof linear equations. The solution set of the systemΣis the setSol = Σ = Ñ ( ̄ 푎,푏)∈Σ Sol = ( ̄ 푎,푏). A set푋 ⊆ R 퐷 is a linear subspace of R 퐷 if it is the solution set of some system of linear equations. A linear inequation inR 퐷 is also a pair( ̄ 푎,푏) ∈ R 퐷 \ ̄ 0 × R. The solution set of the inequation ( ̄ 푎,푏)is the setSol ≥ ( ̄ 푎,푏) = ̄ 푥 ∈ R 퐷 | ̄ 푎· ̄ 푥 ≥ 푏. A set푋 ⊆ R 퐷 is a half-space ofR 퐷 if it is the solution set of some linear inequation. A system of linear inequations is a finite setΣof linear inequations. The solution set of the systemΣis the setSol ≥ Σ = Ñ ( ̄ 푎,푏)∈Σ Sol ≥ ( ̄ 푎,푏). A set푋 ⊆ R 퐷 is a polyhedron ofR 퐷 if it is the solution set of some system of linear inequationsΣ. A vertex of푋is a point ̄ 푥 ∈ R 퐷 such that ̄ 푥 = Sol = (Σ ′ ) for some subsetΣ ′ ⊆Σ. A polytope is a bounded polyhedron. Remark 8.• Polyhedra are closed sets. • The polytopes of R 퐷 are exactly the sets of the form Conv(푆), where 푆 is a finite subset of R 퐷 . • A same pair( ̄ 푎,푏) may alternatively be seen as a linear equation or a linear inequation. 6.2.3 Lemmas A first important result is the following: when the set푋is a polyhedron, so is ⌞ 푋 . Given a system of inequations defining the polyhedron푋, it is therefore possible to find a system of inequations defining ⌞ 푋 . An exponential blowup might appear in the cardinality of that system, but not in the size of individual inequations. Let us recall that what we mean with the size of a number, or of a set, has been defined precisely in Section 2.1. Lemma 11 ([CDE + 10]). LetΣbe a system of inequations, and let푋 = Sol ≥ (Σ). The set ⌞ 푋is itself a polyhedron, and there exists a system of inequationsΣ ′ such that ⌞ 푋 = Sol ≥ (Σ ′ ) and that for every ( ̄ 푎 ′ ,푏 ′ ) ∈Σ ′ , there exists( ̄ 푎,푏) ∈Σ with∥( ̄ 푎 ′ ,푏 ′ )∥ ≤ ∥( ̄ 푎,푏)∥. The previous lemma will be completed by the following, which bounds polynomially the size of the vertices of a polyhedron given the system of inequations that defines it. Lemma 12 ([BR15], Theorem 1). There exists a polynomial푃 1 such that, for every system of equationsΣ, there exists a point ̄ 푥 ∈ Sol = Σ, such that∥ ̄ 푥∥ ≤ 푃 1 max ( ̄ 푎,푏)∈Σ ∥( ̄ 푎,푏)∥ . 86CHAPTER 6. MEAN-PAYOFF GAMES Corollary 1. For every system of inequationsΣ, each vertex ̄ 푥 of the polyhedron Sol ≥ (Σ) has size: ∥푥∥ ≤ 푃 1 max ( ̄ 푎,푏)∈Σ ∥( ̄ 푎,푏)∥ . Note that in Lemma 11, in Lemma 12 and in Corollary 1, the number of equations or inequations has no influence. In the reverse direction, the size of the inequations defining a polyhedron can be polynomially bounded given the size of its vertices. Lemma 13. There exists a polynomial푃 2 such that, for each finite set퐷and every finite subset푋 ⊆ R 퐷 , there exists a system of linear inequationsΣ, such thatSol ≥ (Σ) = Conv(푋)and∥( ̄ 푎,푏)∥ ≤ 푃 2 (∥푋∥)for every( ̄ 푎,푏) ∈Σ. Proof.First, let us recall the notion of facet: a facet of a polytope푃is a subset of푃of dimensiondim푃−1 and of the form푃∩퐻, where퐻is a hyperplane defined by an equation( ̄ 푎,푏), such that푃 ⊆ Sol ≥ ( ̄ 푎,푏). Each facet of the polytopeConv(푋)is of the formConv ̄ 푥 1 , . . ., ̄ 푥 푛 , where ̄ 푥 1 , . . ., ̄ 푥 푛 are vertices of 푋 and 푛 ≥ 푑 = card퐷 . LetΦ be the set of those facets. We can then write (see for example [Brø83] for a proof ): Conv(푋) = Sol ≥ ( ̄ 푎 퐹 ,푏 퐹 ) | 퐹 ∈Φ , where for each퐹, the equation( ̄ 푎 퐹 ,푏 퐹 )defines the hyperplane to which the facet퐹belongs. Let us now study the complexity of each of those equations (or inequations). For a given facet퐹 = Conv ̄ 푥 1 , . . ., ̄ 푥 푛 , let us choose푑points ̄ 푦 1 , . . ., ̄ 푦 푑 ∈ ̄ 푥 1 , . . ., ̄ 푥 푛 that are linearly independent. The equation( ̄ 푎 퐹 ,푏 퐹 )has the points ̄ 푦 1 , . . ., ̄ 푦 푑 among its solutions, i.e. it satisfies: ∀푖 ∈ 1, . . .,푑, 푑 ∑︁ 푗=1 푎 퐹 푗 푦 푖푗 −푏 퐹 = 0. Those푑equalities can themselve be understood as equations that are satisfied by the pair( ̄ 푎 퐹 ,푏 퐹 ). Let us add a(푑+1)-th equation: for some dimension푖 0 , we have푎 퐹푖 0 ≠ 0, and by multiplying if needed by a nonzero factor, we can assume푎 퐹푖 0 = 1. By Lemma 12, there exists a pair( ̄ 푎,푏)satisfying that system of푑+1 equations and the inequality∥( ̄ 푎,푏)∥ ≤ 푃 1 2푑+ 5+ max 푖 Í 푗 ∥푥 푖푗 ∥ ≤ 푃 1 (∥푋∥). That pair is an equation of the hyperplane containing 퐹 . As a consequence, the polynomial 푃 2 = 푃 1 satisfies the desired inequalities.□ 6.3ABOUT NEGOTIATION IN MEAN-PAYOFF GAMES We now give some observations about the behavior of the negotiation function in mean-payoff games. We first show a useful result: in the concrete negotiation game, stationary strategies are sufficient for Challenger (Lemma 14). We infer from that result that mean-payoff games are games with steady negotiation (Lemma 15), a result that is not as trivial as it was for parity games, and that is still necessary to apply Theorem 8, and use the negotiation function as a tool to characterize SPEs and휀-SPEs. We end the section with two important examples, which should help the reader forge intuition on how the negotiation function can be used on mean-payoff games, before moving toward an algorithm. 6.3. ABOUT NEGOTIATION IN MEAN-PAYOFF GAMES87 6.3.1 Stationary optimal strategies for Challenger We first give a useful result for the sequel: in the concrete negotiation game, Challenger has a stationary optimal strategy. Lemma 14. LetG ↾푣 0 be a mean-payoff game, let푖be a player, and let휆be a requirement. In the corresponding concrete negotiation game, there exists a stationary strategy 휏 ℭ that is optimal for Challenger. Proof.The structure of that proof is inspired from the proof of Lemma 14 in [VCD + 15]. Let 휈 ℭ be the payoff function defined by: • 휈 ℭ (휋) =+∞ if there exists푛such that휋 ≥2푛 contains no deviation, and such that the play¤휋 ≥푛 is not 휆-consistent. • 휈 ℭ (휋) = lim sup 푛 mp 푖 (¤휋 ≤푛 ) otherwise. The payoff function휈 ℭ is then defined as휇 c ℭ , but with a limit superior instead of inferior. We will first prove that Challenger has a stationary optimal strategy to maximize휈 ℭ , and then show that such a strategy is also optimal to maximize 휇 c ℭ . ▶The payoff function 휈 ℭ is concave. Indeed, let휋and휒be two plays inconc 휆푖 (G) ↾푣 0 , and let휉be a shuffling of those plays. Let us check that 휈 ℭ (휉) ≤ max휈 ℭ (휋),휈 ℭ (휒). If either휈 ℭ (휋) = +∞or휈 ℭ (휒) = +∞, it is immediate. Otherwise, we also have휈 ℭ (휉) ≠ +∞: if either휋or휒contains infinitely many deviations, then so does휉. If both contain finitely many deviations, then so does휉: the vertices of휉have therefore eventually the same memory푀, which is also the memory of, eventually, the vertices of both휋and휒. Now, since휈 ℭ (휋),휈 ℭ (휒) ≠+∞, we have 휇 푗 (¤휋),휇 푗 (¤휒) ≥ 휆(푣) for each player 푗 and every 푣 ∈ 푀∩푉 푗 . Since mean-payoff functions are convex, it is also the case for the play ¤ 휉 , which is a shuffling of ¤휋 and ¤휒 . Hence 휈 ℭ (휉) ≠+∞. Therefore, we have휈 ℭ (휉) = lim sup 푛 mp 푖 ( ¤ 휉 ≤푛 ), as well as휈 ℭ (휋) = lim sup 푛 mp 푖 (¤휋 ≤푛 )and휈 ℭ (¤휒) = lim sup 푛 mp 푖 (¤휒 ≤푛 ) . Since, as shown in [VCD + 15], mean-payoff functions defined with a limit superior are concave, it implies 휈 ℭ (휉) ≤ max휈 ℭ (휋),휈 ℭ (휒): the payoff function 휈 ℭ is concave. Therefore, by Lemma 2 Challenger has a stationary strategy that is optimal with regards to the payoff function 휈 ℭ : let us write it 휏 ℭ . ▶The stationary strategy 휏 ℭ is also optimal with regards to 휇 c ℭ . Note that for every play휋, we have휇 c ℭ (휋) ≤ 휈 ℭ (휋) , and thereforeval ℭ conc 휆푖 (G) ↾푣 c 0 ≤ 훼, where 훼is the value of the gameconc 휆푖 (G) ↾푣 c 0 with the payoff function휈 ℭ instead of휇 c ℭ . Therefore, we have proven that 휏 ℭ is optimal with regards to 휇 c ℭ if we prove that inf 휏 P 휇 c ℭ ⟨ ̄ 휏⟩ ≥ 훼 . Let휋be a play compatible with휏 ℭ , i.e., a play in the gameconc 휆푖 (G) ↾푣 c 0 [휏 ℭ ]. If휇 c ℭ (휋) =+∞ , then clearly 휇 c ℭ (휋) ≥ 훼 . Otherwise, we have 휇 c ℭ (휋) = 휇 푖 (¤휋), and by Lemma 10, we have: 휇 푖 (¤휋) ≥ min mp 푖 (¤푐) | 푐 ∈ SCyc(conc 휆푖 (G) ↾푣 c 0 [휏 ℭ ]) . 88CHAPTER 6. MEAN-PAYOFF GAMES Now, for each cycle푐 ∈ SCyc(conc 휆푖 (G) ↾푣 c 0 [휏 ℭ ]) , there exists a history푔such that the play푔푐 휔 is compatible with the strategy휏 ℭ , and therefore satisfies휈 ℭ (푔푐 휔 ) ≥ 훼, and consequentlymp 푖 (¤푐) ≥ 훼. Therefore, we have휇 푖 (¤휋) ≥ 훼, and the strategy휏 ℭ is optimal with regards to the payoff function휇 c ℭ , which concludes the proof.□ 6.3.2 Steady negotiation Using this lemma, computingnego(휆)for any given휆amounts to looking for an optimal play for Prover in the gameconc 휆푖 (G) ↾푣 c 0 [휏 ℭ ]. We show that there exists an optimal such play; and that this proves that mean-payoff games are games with steady negotiation, which will enable us to apply Theorem 8. Lemma 15. Mean-payoff games are games with steady negotiation. Proof. According to Theorem 10, to prove that mean-payoff games are games with steady negotiation, it suffices to prove that Prover always has an optimal strategy in every concrete negotiation game constructed from a mean-payoff game. We have already noted that mean-payoff games are Borel, and that when a game is Borel, every concrete negotiation game constructed from it is Borel. By Lemma 1, it is then equivalent to prove that, in every concrete negotiation gameconc 휆푖 (G) ↾푣 c 0 built from a mean-payoff gameG, against every strategy휏 ℭ of Challenger, Prover has a strategy to obtain at least the payoffval 픓 (conc 휆푖 (G) ↾푣 c 0 ), i.e. a strategy so that Challenger obtains at most the payoff val ℭ (conc 휆푖 (G) ↾푣 c 0 ). Finally, by Lemma 14, it is then sufficient to prove that result only for stationary strategies of Challenger. Let therefore휏 ℭ be a stationary strategy of Prover. By definition of the adversarial value, we have: inf n 휇 c ℭ (휋) | 휋 ∈ Plays conc 휆푖 (G) ↾푣 c 0 [휏 ℭ ] o ≤ val ℭ (conc 휆푖 (G) ↾푣 c 0 ). We must now prove that this infimum is a minimum, i.e. that Prover can choose a worst play for Challenger in the game conc 휆푖 (G) ↾푣 c 0 [휏 ℭ ]. For any play휋in that graph, there exists a strongly connected component퐾ofconc 휆푖 (G) ↾푣 c 0 [휏 ℭ ] such that after a finite number of steps, the play휋remains in퐾. Since there are finitely many strongly connected components in a graph, there is a worst such strongly connected component, i.e. a strongly connected component 퐾 such that we have: inf n 휇 c ℭ (휋) | 휋 ∈ Plays conc 휆푖 (G) ↾푣 c 0 [휏 ℭ ] o = inf 휇 c ℭ (휋) | 휋 ∈ Plays(퐾) . We must now prove that there is a worst play in 퐾 . Let us distinguish two cases. If there is at least one deviation in퐾.Then, for every play휋in퐾, it is possible to transform the play휋into a play휋 ′ with휇(¤휋 ′ ) = 휇(¤휋), which contains infinitely many deviations: it suffices to add round trips to a deviation, endlessly, but less and less often. Therefore, the outcomes휇 c ℭ (휋) of plays in퐾are exactly the mean-payoffs휇 푖 (¤휋)of plays in퐾(plus possibly+∞); and in particular, the lowest payoff that can be given to Challenger in 퐾 is the quantity (using Lemma 10): min 푐∈SCyc(퐾) mp 푖 (¤푐), which is achieved by the play 푐 휔 , where 푐 is a simple cycle realizing that minimum. 6.3. ABOUT NEGOTIATION IN MEAN-PAYOFF GAMES89 푎푐 푏푑 ◦ 0 □ 3 ◦ 0 □ 3 ◦ 2 □ 2 ◦ 1 □ 1 (휆 0 ) −∞−∞−∞−∞ (휆 1 ) 1212 (휆 2 ) 2212 (휆 3 ) 2 3 12 (휆 4 ) +∞+∞ 12 Figure 20: A game without SPE If there is no deviation in퐾. Then, let us notice that deviations are the only edges that go from a vertex of the form(·,푀)to a vertex of the form(·,푀 ′ )with푀 ′ ⊂ 푀. Therefore, necessarily, there exists a set 푀 ⊆ 푉 such that all vertices in 퐾 have the form(·,푀). By Lemma 10, the set of possible values of 휇(¤휋) for all plays 휋 in 퐾 is exactly the set: 푋 = ⌞ Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ) . Since all plays in퐾contain finitely many deviations (actually none), for every play휋in퐾, we have휇 c ℭ (휋) =+∞ if and only if there exists푗 ∈Πand푢 ∈ 푉 푗 ∩ 푀such that휇 푗 (¤휋)< 휆(푢). Then, the lowest payoff that can be given to Challenger in 퐾 is the quantity: inf 푥 푖 | ̄ 푥 ∈ 푋,∀푗 ∈Π,∀푢 ∈ 푉 푗 ∩ 푀,푥 푗 ≥ 휆(푢) , and since that set is closed, this infimum is a minimum, which concludes the proof.□ Note that the proof is constructive, suggesting an algorithm to compute the negotiation function. We will however not use the concrete negotiation game for that purpose, because of its prohibitive size: it will only be used for intermediary results. 6.3.3 First application: an example with no SPE It is a well-known fact that there are mean-payoff games that do not contain SPEs, while they are guaranteed to exist in Borel Boolean games such as parity or energy games [GU08], and while NEs are guaranteed to exist in mean-payoff games [BDS13]. We present here a classical example of such a game, along with a proof of the non-existence of SPEs based on the negotiation function. Theorem 13 ([SV03]). There exists a mean-payoff game with no SPE. Proof.Let us consider the game depicted by Figure 20. On the two first lines below the vertices, we present the requirements휆 0 and휆 1 = nego(휆 0 ). The latter is easy to compute since for each푣, we have the equality 휆 1 (푣) = val(푣). Let us now compute the following iterations. 90CHAPTER 6. MEAN-PAYOFF GAMES ▶Computation of 휆 2 = nego(휆 1 ) From the vertex푐, there exists exactly one휆 1 -rational strategy profile ̄ 휎 −◦ = 휎 □ , which is the empty strategy since player□never has to choose anything. Against that strategy, the best and the only payoff player ◦ can get is 1, hence 휆 2 (푐) = 1. For the same reasons, we have 휆 2 (푑) = 2. From the vertex푏, player ◦ can force□to get the payoff 2 or less, with the strategy (profile) 휎 ◦ :ℎ 7! 푐. Such a strategy (profile) is휆 1 -rational, assuming the strategy휎 □ :ℎ 7! 푑. Therefore, we have 휆 2 (푏) = 2. Finally, from the vertex푎, player□can force ◦ to get the payoff 2 or less, with the strategy profile 휎 □ :ℎ 7! 푑. Such a strategy is휆 1 -rational, assuming the strategy휎 ◦ :ℎ 7! 푐. However, he cannot force her to get less than the payoff 2, because she can force the access to the vertex푏, and the only 휆 1 -consistent plays from 푏 are the plays with the form(푏푎) 푘 푏푑 휔 . Therefore, we have 휆 2 (푎) = 2. ▶Computation of 휆 3 = nego(휆 2 ) The computation of 휆 3 (푐) = 1 and 휆 3 (푑) = 2 is immediate. From the vertex푎, player ◦ cannot obtain more than the payoff 2 against the strategy휎 □ :ℎ 7! 푑. And the strategy profile composed of that strategy is still휆 2 -rational, assuming the strategy휎 ◦ :ℎ 7! 푐. We therefore have 휆 3 (푎) = 2. From the vertex푏, on the other hand, the only휆 2 -rational strategy profile for player ◦ is composed of the strategy휎 ◦ :ℎ푎 7! 푏, which is휆 2 -rational assuming the strategy휎 □ :ℎ 7! 푑: taking the edge from푎to푐is now unacceptable for player ◦ . Consequently, using the strategy휎 ′ □ :ℎ푏 7! 푎, player□ can force the payoff 3, hence 휆 3 (푏) = 3. ▶Computation of 휆 4 = nego(휆 3 ) and conclusion Finally, let us observe that there is no휆 3 -consistent play, and therefore no휆 3 -rational strategy profile, from the vertices푎and푏, hence휆 4 (푎) = 휆 4 (푏) =+∞. The requirement휆 4 is then a fixed point of the negotiation function. By a quick induction, we can show that every fixed point of the negotiation function is always greater than or equal to all the iteratednego 푘 (휆 0 ), and therefore every SPE outcome from the vertices 푎or푏is necessarily휆 4 -consistent. Since there is no such play, there is no SPE that starts from one of those vertices.□ 6.3.4 The impossibility of an algorithm based on iterations The two previous subsections suggest an algorithm based on the computation of the negotiation sequence, i.e. the iterations of the negotiation function on the vacuous requirement휆 0 , until a fixed point is reached. Such an algorithm would be analogous to the one described in the proof of Theorem 12. However, such an approach cannot be used on mean-payoff games, because there are mean-payoff games in which the negotiation sequence does never reach any fixed point. Theorem 14. There exists a mean-payoff game on which the negotiation sequence is not ultimately constant. 6.3. ABOUT NEGOTIATION IN MEAN-PAYOFF GAMES91 푐 푑 푎 푏 푒 푓 ◦ 2 □ 2 0 ◦ 2 □ 2 0 ◦ 0 □ 1 0 ◦ 0 □ 1 0 ◦ 2 □ 2 0 ◦ 1 □ 0 0 ◦ 1 □ 0 0 ◦ 2 □ 2 0 Figure 21: A game where the negotiation sequence is not ultimately constant Proof.LetGbe the game of Figure 21, and let us compute the negotiation sequence(휆 푛 ) 푛∈N . Since all player^’s rewards are equal to 0, for all푛>0, we have휆 푛 (푐) = 휆 푛 (푑) = 휆 푛 (푒) = 휆 푛 (푓) =0. Moreover, by symmetry of the game, we always have휆 푛 (푎) = 휆 푛 (푏). Therefore, to compute the negotiation sequence, it suffices to compute휆 푛+1 (푎)as a function of휆 푛 (푏), knowing that휆 1 (푎) = 휆 1 (푏) =1, and therefore that for all푛>0, we have휆 푛 (푎) = 휆 푛 (푏) ≥1. For convenience, we compute here the negotiation function using the abstract negotiation game. From푎, the worst play that Prover can propose would be a combination of the cycles푐푑and푑 giving her exactly 1. But then, Challenger would deviate to go to푏, from which, if Prover always proposes plays in the strongly connected component containing푐and푑, Challenger will always deviate and generate the play(푎푏) 휔 , and then get the payoff 2. Then, in order to give player ◦ a payoff lower than 2, Prover has to go to the vertex푒. Since player ◦ does not control any vertex in that strongly connected component, the play that Prover will propose will be accepted: she will, then, propose the worst possible combination of the cycles푒푓and 푓for player ◦ , such that player□gets at least his requirement휆 푛 (푏). The payoff휆 푛+1 (푎)is then the minimal solution of the system: 휆 푛+1 (푎) = 푥 + 2(1− 푥) 2(1− 푥) ≥ 휆 푛 (푏) 0≤ 푥 ≤ 1 that is to say 휆 푛+1 (푎) = 1+ 휆 푛 (푏) 2 = 1+ 휆 푛 (푎) 2 . By induction, we obtain for all 푛> 0: 휆 푛 (푎) = 휆 푛 (푏) = 2− 1 2 푛−1 which converges to 2 but never reaches it.□ Along with the prohibitive size of the concrete negotiation game, this theorem suggests that, if the constrained existence problem of SPEs in mean-payoff games is fixed-parameter tractable, a proof of that result must use a completely different approach. We leave that problem open. Open Problem 2. Is the constrained existence problem of (휀-)SPEs in mean-payoff games fixed-parameter tractable, when the number of players is fixed? 92CHAPTER 6. MEAN-PAYOFF GAMES 6.4SIZE OF THE LEAST 휀–FIXED POINT Now that the necessary conceptual tools are given, we present a technical but important result about mean-payoff games: the size of the least휀-fixed point of the negotiation function in a mean-payoff game Gis bounded by a polynomial of the size ofGand휀. The first piece of the witnesses identifying positive instances of the휀-SPE constrained existence problem will then be an휀-fixed point of the negotiation function of polynomial size. 6.4.1 Existence Before proving anything about its size, we must first prove that the negotiation function always has a least 휀-fixed point. Lemma 16. Let 휀 ≥ 0. On every game, the function nego has a least 휀-fixed point. Proof.The following proof is a generalization of a classical proof of Tarski’s fixed point theorem. Let Λ 휀 be the set of휀-fixed points of the negotiation function. The setΛ 휀 is not empty, since it contains at least the requirement 푣 7!+∞. Let 휆 휀 be the requirement defined by: 휆 휀 : 푣 7! inf 휆∈Λ 휀 휆(푣). For every휀-fixed point휆of the negotiation function, and for each vertex푣, we have then 휆 휀 (푣) ≤ 휆(푣), and thereforenego(휆 휀 )(푣) ≤ nego(휆)(푣)sincenegois monotone; and therefore, we have nego(휆 휀 )(푣) ≤ 휆(푣)+ 휀. As a consequence, we have: nego(휆 휀 )(푣) ≤ inf 휆∈Λ 휀 휆(푣)+ 휀 = 휆 휀 (푣)+ 휀. The requirement휆 휀 is an휀-fixed point of the negotiation function, and is therefore the least of them.□ 6.4.2 Size Lemma 17. There exists a polynomial푃 3 such that for every mean-payoff gameG, the least휀-fixed point휆 휀 of the negotiation function has size∥휆 휀 ∥ ≤ 푃 3 (∥G∥+∥휀∥). Proof. This proof will use a characterization of휆 휀 as a vertex (or, more exactly, a well-chosen projection of a vertex) in a union of polyhedra, to which we will apply Corollary 1. Those polyhedra will be defined using the concrete negotiation games, and the fact that, by Lemma 14, stationary strategies are optimal for Challenger in those games. ▶Reminders about the concrete negotiation game Without loss of generality, we assume that all the rewards푟 푖 (푢푣)for푖 ∈Πand푢푣 ∈ 퐸are integers— otherwise, we can multiply all the rewards by the least common denominator, compute the least휀-fixed point in the resulting game, and finally divide each of its values by the least common denominator. Let now푖 ∈Πand푣 ∈ 푉 푖 , let휆be a requirement, and let us consider the corresponding concrete negotiation gameconc 휆푖 (G) ↾(푣,푣) . By Lemma 14, we know that Challenger has a stationary optimal 6.4. SIZE OF THE LEAST 휀–FIXED POINT93 strategy in that game. Let then휏 ℭ be a stationary strategy for Challenger inconc 휆푖 (G) ↾(푣,푣) and let us consider the gameconc 휆푖 (G) (푣,푣) [휏 ℭ ]. Facing the strategy휏 ℭ , Prover must choose an optimal play in that game, i.e., since the concrete negotiation game is prefix-independent, a strongly connected component퐾of its underlying graph, and a good combination of cycles in퐾. There are, then, two possibilities: either Challenger does not deviate anywhere in퐾, and then, Prover must find a play that minimizes player푖’s payoff while remaining휆-consistent. Or, there is such a deviation, and it is then possible for Prover to visit that deviation infinitely often (if necessary, with negligible frequency), and thus to ignore the constraints defined by 휆. ▶Size of mean-payoffs of cycles Let therefore퐾be a strongly connected component ofconc 휆푖 (G) (푣,푣) [휏 ℭ ]. Let푐 =푐 1 . . .푐 푛 be a simple cycle of퐾: necessarily, we have푛 ≤ card푉 c = card푉2 card푉 . Therefore, we have, for each player 푗 : ∥mp 푗 (¤푐)∥ = 1 푛 ∑︁ 푘∈Z/푛Z 푟 푗 (¤푐 푘 ¤푐 푘+1 ) = 1+ log 2 © « ∑︁ 푘∈Z/푛Z 푟 푗 (¤푐 푘 ¤푐 푘+1 ) + 1 ª ® ¬ +⌈log 2 (푛)⌉ ≤ 1+ & log 2 2 card푉 ∑︁ 푒∈퐸 푟 푗 (푒) + 1 !' +⌈card푉 log 2 (card푉)⌉ ≤ 1+ card푉 + ∑︁ 푒∈퐸 log 2 |푟 푗 (푒)|+ 1 +(card푉) 2 ≤ (card퐸) 3 + ∑︁ 푒∈퐸 log 2 |푟 푗 (푒)|+ 1 ≤ ∥푟 푗 ∥ 3 ≤ ∥G∥ 3 . ▶Feasible payoff vectors in a strongly connected component Let us now define the polytope: 퐹 퐾 = ⌞ Conv 푐∈SCyc(퐾) mp(¤푐) . By Lemma 10, the polytope퐹 퐾 is exactly the set of the payoff vectors of the projections of plays in the strongly connected component 퐾 . By Lemma 13, the polytope: Conv 푐∈SCyc(퐾) mp(¤푐) is the solution set of a system of inequations whose sizes are bounded by a polynomial function of ∥mp(¤푐) | 푐 ∈ SCyc(퐾)∥, i.e., according to the previous point, of∥G∥. By Lemma 11, this is also the case for the polytope 퐹 퐾 . 94CHAPTER 6. MEAN-PAYOFF GAMES ▶Consistent payoff vectors in a strongly connected component If the strongly connected component퐾contains no deviation, then all the vertices of퐾share the same memory: there is a set푀 퐾 ⊆ 푉such that all the vertices of퐾have the form(·,푀 퐾 ). If퐾contains a deviation, it might not be the case: then, we define푀 퐾 =∅. Thus, in both cases, the set푀 퐾 defines the set of constraints a play must satisfy so that Challenger does not get the payoff+∞: the vertices memorized so far if there is no deviation, and nothing if there is a deviation. Given a requirement 휆, we can now define the polytope of consistent payoff vectors: 퐶 퐾휆 = ̄ 푥 ∈ 퐹 퐾 |∀푗,∀푢 ∈ 푉 푖 ∩ 푀 퐾 ,푥 푖 ≥ 휆(푢) . That polytope is the set of tuples ̄ 푥 such that there exists a play휋in퐾realizing휇(¤휋) = ̄ 푥, and 휇 c ℭ (휋) ≠+∞. We can then write: nego(휆)(푣) = sup 휏 ℭ inf 퐾 inf ̄ 푥∈퐶 퐾휆 푥 푖 . The inequations defining퐶 퐾휆 are the same as those defining퐹 퐾 , plus the inequations of the form 푥 푖 ≥ 휆(푢). ▶Union, intersection and product Let us now work in the space R 푉×Π . Given a requirement 휆, we define the set: 푋 휆 = Ö 푖∈Π 푣∈푉 푖 Ù 휏 ℭ ∈Sta ( conc 휆푖 (G) ↾(푣,푣) ) Ø 퐾 퐶 " 퐾휆 , where푌 7! 푌 " denotes the upward closure operation, i.e.푌 " =푧 |∃푦 ∈ 푌,푧 ≥ 푦. Then, for every tuple of tuples ̄ ̄ 푥 ∈ R 푉×Π , we have ̄ ̄ 푥 ∈ 푋 휆 if and only if for each푖and푣 ∈ 푉 푖 , for every stationary strategy of Challenger, there exists a play휋compatible with휏 ℭ satisfying휇 c ℭ (휋) ≠+∞ and휇(¤휋) ≤ ̄ 푥 푣 . Therefore, for each 푖 and 푣 ∈ 푉 푖 , we have nego(휆)(푣) = inf 푥 푣푖 | ̄ ̄ 푥 ∈ 푋 휆 . In terms of inequations, the set푋 휆 is a union of polyhedra which are all defined by inequations that are inequations defining some퐶 퐾휆 , padded with 0 to fit with the dimension change. ▶ 휀-fixed points To each tuple of tuples ̄ ̄ 푥, we associate the requirement휆 ̄ ̄ 푥 defined, for each푖 ∈Πand푣 ∈ 푉 푖 , by 휆 ̄ ̄ 푥 (푣) = 푥 푣푖 − 휀. Now, let us consider the diagonal set: 푋 = ̄ ̄ 푥 ∈ R 푉×Π ̄ ̄ 푥 ∈ 푋 휆 ̄ ̄ 푥 . In terms of inequations, the set푋is a union of polyhedra defined by the same inequations as푋 휆 , but where those of the form푥 푢푖 ≥ 휆(푣)are replaced by equations of the form푥 푢푖 ≥ 푥 푣푖 , each of size 2+ 4card푉 cardΠ. We can now link the set 푋 to 휀-fixed points of the negotiation function. 6.5. CONSTRAINED EXISTENCE OF A 휆–CONSISTENT PLAY95 Sublemma 2. Let휆be a requirement. Then, it is an휀-fixed point of the negotiation function if and only if we have 휆 = 휆 ̄ ̄ 푥 for some ̄ ̄ 푥 ∈ 푋. Proof.• If there exists ̄ ̄ 푥 ∈ 푋 such that휆 = 휆 ̄ ̄ 푥 , then we have ̄ ̄ 푥 ∈ 푋 휆 ̄ ̄ 푥 , and by the previous point, for each푖and푣 ∈ 푉 푖 , we havenego(휆 ̄ ̄ 푥 )(푣) ≤ 푥 푣푖 = 휆 ̄ ̄ 푥 (푣)+ 휀 . Therefore,휆is an휀-fixed point of the negotiation function. •Conversely, if for each푖and푣 ∈ 푉 푖 , we havenego(휆)(푣) ≤ 휆(푣)+ 휀, then according to the previous point and since the set푋 휆 is closed, there exists a tuple of tuples ̄ ̄ 푥 (푣) ∈ 푋 휆 such that 푥 (푣) 푣푖 = 휆(푣)+ 휀 . Then, since푋 휆 is defined as a cartesian product over푣, the tuple of tuples ̄ ̄ 푥 = ̄ 푥 (푣) 푣 푣 does also belong to 푋 휆 , and satisfies 휆 = 휆 ̄ ̄ 푥 . □ Then, in particular, the least휀-fixed point휆 휀 is the (unique) minimal element in the set 휆 ̄ ̄ 푥 | ̄ ̄ 푥 ∈ 푋 . The set푋is itself a union of polyhedra: consequently, the linear mapping ̄ ̄ 푥 7! Í 푣 휆 ̄ ̄ 푥 (푣) finds its minimum over푋on some vertex ̄ ̄ 푥of one of those polyhedra, which is therefore such that휆 휀 = 휆 ̄ ̄ 푥 . By Corollary 1, such a vertex has size bounded by a polynomial function of the maximal size of the inequations defining 푋 , and therefore a polynomial function of∥G∥+∥휀∥.□ With this proof, we are now done with our utilization of the concrete negotiation game. Like in the case of parity games, the negotiation function will rather be computed using simple strategies in the abstract negotiation game. 6.5CONSTRAINED EXISTENCE OF A 휆–CONSISTENT PLAY We claim that a non-deterministic algorithm can recognize the positive instances of the휀-SPE constrained existence problem by guessing an휀-fixed point휆of the negotiation function, as first piece of a three-piece witness. Once휆has been guessed, following Theorem 8, two assertions must be checked: on the one hand, that there exists a휆-consistent play between the two desired thresholds, and on the other hand, that휆is actually an휀-fixed point of the negotiation function. The latter will be handled later through a new notion of reduced strategy. Now, we tackle the former, and provide the second piece of our notion of witness: to prove the existence of a휆-consistent play휋with ̄ 푥 ≤ 휇(휋) ≤ ̄ 푦, we need to guess the sets푊 = Inf(휋)and 푊 ′ = Occ(휋), and a tuple of tuples ̄ ̄ 훼 ∈ [0,1] Π×SCyc(푊) indicating how휋combines the cycles of푊(we liberally assimilate the set푊to the induced subgraph(푊,퐸∩(푊 ×푊))to simplify notations), i.e. such that we have: 휇(휋) = © « min 푗∈Π ∑︁ 푐∈SCyc(푊) 훼 푗푐 mp 푖 (푐) ª ® ¬ 푖 . Lemma 18. There exists a polynomial푃 4 such that for every mean-payoff gameG ↾푣 0 , for every ̄ 푥, ̄ 푦 ∈ (R ∪ ±∞) Π , and for every requirement휆onG, there exists a휆-consistent play휋inG ↾푣 0 satisfying ̄ 푥 ≤ 휇(휋) ≤ ̄ 푦if and only if there exist two sets푊 ⊆ 푊 ′ ⊆ 푉and a tuple of tuples ̄ ̄ 훼 ∈ [0,1] Π×SCyc(푊) such that: • the set푊 is strongly connected in(푉,퐸), and accessible from the vertex 푣 0 using only vertices of푊 ′ ; • for each player 푖, we have Í 푐 훼 푖푐 = 1, and: 푥 푖 ≤ min 푗∈Π ∑︁ 푐∈SCyc(푊) 훼 푗푐 mp 푖 (푐) ≤ 푦 푖 ; 96CHAPTER 6. MEAN-PAYOFF GAMES • for each player 푖 and 푣 ∈ 푊 ∩푉 푖 , we have: min 푗∈Π ∑︁ 푐∈SCyc(푊) 훼 푗푐 mp 푖 (푐) ≥ 휆(푣); • and finally, we have∥ ̄ ̄ 훼∥ ≤ 푃 4 (∥G, ̄ 푥, ̄ 푦,휆∥). Proof. Let us first notice that given a set푋 ⊆ R Π , the elements of the set ⌞ (Conv푋) are exactly the tuples of the form: min 푗∈Π ∑︁ 푥∈푋 훼 푗푥 푥 푖 ! 푖∈Π for some tuple ̄ ̄ 훼 ∈ R Π×푋 satisfying Í 푥 훼 푖푥 = 1 for each 푖. Now, let us assume that the sets푊,푊 ′ and the tuple ̄ ̄ 훼exist. Then, there exists a play휒 with Occ(휒) = Inf(휒) =푊 with payoff vector: 휇(휒) = © « min 푗∈Π ∑︁ 푐∈SCyc(푊) 훼 푗푥 mp 푖 (푐) ª ® ¬ 푖∈Π . Moreover, since푊is accessible from푣 0 using only vertices of푊 ′ , there exists a historyℎ휒 0 from푣 0 to 휒 0 with Occ(ℎ) =푊 ′ . Then, the play 휋 =ℎ휒 is 휆-consistent and satisfies ̄ 푥 ≤ 휇(휋) ≤ ̄ 푦. Conversely, let us assume that the play휋exists. Let푊 = Inf(휋)and푊 ′ = Occ(휋). The polytope: 푍 = 휇(휒) 휒 ∈ 휆ConsG ↾푣 0 , Inf(휒) =푊, Occ(휒) =푊 ′ , and ̄ 푥 ≤ 휇(휒) ≤ ̄ 푦 = ( ̄ 푧 ∈ ⌞ Conv 푐∈SCyc(푊) mp(푐) ̄ 푥 ≤ ̄ 푧 ≤ ̄ 푦, and ∀푖,∀푣 ∈ 푊 ′ ∩푉 푖 ,푧 푖 ≥ 휆(푣) ) (the equality holds by Lemma 10) is nonempty (it contains at least the vector휇(휋)). By Lemma 13, the set Conv 푐∈SCyc(푊) mp(푐) is defined by a system of inequations which all have size: ∥( ̄ 푎,푏)∥ ≤ 푃 2 max 푐 ∥mp(푐)∥ . Since by Lemma 11, the inequations defining the polytope ⌞ Conv 푐∈SCyc(푊) mp(푐) are not larger, there exists a polynomial푃 6 , independent ofG, ̄ 푥, ̄ 푦and휆, such that the polytope푍is defined by a system of inequationsΣsuch that for every( ̄ 푎,푏) ∈Σ, we have∥( ̄ 푎,푏)∥ ≤ 푃 6 (∥(G, ̄ 푥, ̄ 푦,휆)∥). Therefore, by Corollary 1, the polytope 푍 admits a vertex ̄ 푧 of size∥ ̄ 푧∥ ≤ 푃 1 (푃 6 (∥(G, ̄ 푥, ̄ 푦,휆)∥)). Then, since we have ̄ 푧 ∈ ⌞ Conv 푐∈SCyc(푊) mp(푐) , that vertex is, according to Definition 30, of the form: ̄ 푧 = min 푗 ∑︁ 푐 훼 푗푐 mp 푖 (푐) ! 푖 6.6. COMPUTING THE NEGOTIATION FUNCTION97 for some tuple of tuples ̄ ̄ 훼 ∈ [0, 1] Π×SCyc(푊) with Í 푐 훼 푖푐 = 1 and having, by Corollary 1 again, size: ∥ ̄ ̄ 훼∥ ≤ 푃 1 © « max 푖∈Π ∑︁ 푐∈SCyc(푊) ∥mp 푖 (푐)∥+∥푧 푖 ∥ ª ® ¬ , i.e. ∥ ̄ ̄ 훼∥ ≤ 푃 4 (∥(G, ̄ 푥, ̄ 푦,휆)∥) for some polynomial 푃 4 independent ofG, ̄ 푥, ̄ 푦 and 휆.□ Now, we need the third and last piece of our witness, which will be the evidence of the fact that the requirement 휆 is an 휀-fixed point of the negotiation function. 6.6COMPUTING THE NEGOTIATION FUNCTION 6.6.1 A disturbing example Let us study more deeply the gameGdepicted by Figure 19a: consider the requirement휆, which maps both푎and푏to the value 1. We would like to know whether휆is a fixed point of the negotiation function; or in other words, using the symmetry of the game, whether there exists a strategy for Prover, in the abstract negotiation game from vertex푎, that forces player ◦ to get at most the payoff 1. Prover can propose, for example, the play(푎푏) 휔 , that is휆-consistent, and in which player ◦ gets exactly the payoff 1. But then, Challenger can make player ◦ deviate and go directly to푏. That strategy will give player ◦ a payoff greater than 1 against every finite-memory strategy of Prover. Now, consider the following strategy of Prover: propose the play(푎푏) 휔 . Then, whenever Chal- lenger deviates, and reaches prematurely the vertex푏, propose the play푏 |ℎ| 2 (푎푏) 휔 , whereℎis the history inG ↾푎 that has already been drawn by Prover and Challenger’s previous interactions. Such a play is휆-consistent, and even if Challenger deviates to go to the vertex푏whenever he has the opportunity to do so, the play that will be generated is푎푏푎푏 9 푎푏 169 푎푏 33489 . . ., in which player ◦ ’s payoff is still 1. The requirement휆is therefore a fixed point of the negotiation function (hence the set of SPE payoff vectors is the red-circled area in Figure 19b), but Prover requires infinite memory to maintain it. The algorithm that we used for parity games in Chapter 5 can therefore not be used as such. However, we can see in this example that the plays proposed by Prover are still very similar: only the number of repetitions of the loop푏does increase. More generally, one observes that Prover can play optimally while always proposing a play of the formℎ푐 푛 휋, where the historyℎ, the cycle푐and the play 휋are constant, and only the number푛increases, quadratically with time—so that Challenger’s payoff is dominated by the mean-payoffmp 푖 (푐)if he deviates infinitely often. This will define a new notion of reduced strategy, inspired with what we did in Chapter 5 but more sophisticated. 6.6.2 Reduced strategies for mean-payoff games Let us first define the notion of punishment family, the brick from which our reduced strategies will be built. Definition 32 (Punishment family). A punishment family is a set of plays of the form: ℎ푐 푛 휋 | 푛> 0,휇(휋) = ̄ 푥, Occ(휋) =푊 whereℎis a simple history, where푐is a (nonempty) simple cycle, and where푊 ⊆ 푉and ̄ 푥 ∈ R Π . The cycle푐is called its punishing cycle. For every훽 ∈ N, a훽-punishment family is a punishment family with 98CHAPTER 6. MEAN-PAYOFF GAMES ∥ ̄ 푥∥ ≤ 훽. A훽-punishment family is represented by the dataℎ,푐, ̄ 푥and푊, and that representation has a size smaller than or equal to the quantity 3card푉⌈log 2 (card푉 + 1)⌉+ 훽. We writeℎ푐 ∞ 휋for the punishment familyℎ푐 푛 휋 ′ | 푛>0, 휇(휋 ′ ) = 휇(휋), Occ(휋 ′ ) = Occ(휋). Beware that the play휋matters only for its payoff vector and the vertices it traverses: if we haveOcc(휋) = Occ(휋 ′ ) and휇(휋) = 휇(휋 ′ ), then we haveℎ푐 ∞ 휋 = ℎ푐 ∞ 휋 ′ . We write휇(ℎ푐 ∞ 휋)for the common payoff vector of all elements of the setℎ푐 ∞ 휋, and we will say thatℎ푐 ∞ 휋is휆-consistent if all its elements are (which is the case as soon as one of its elements is). Furthermore, let us clarify that a punishment family is not an equivalence class: for example, in the game of Figure 19a, the play푎푏 휔 belongs to both sets푎 ∞ 푏 휔 and 푎푏 ∞ 푏 휔 , which are distinct. We can now define the reduced negotiation game, where Prover proposes훽-punishment families instead of plays. Definition 33 (Reduced negotiation game). LetGbe a mean-payoff game, let휆be a requirement, let푖be a player, let푣 0 ∈ 푉 푖 and let훽be a natural integer. The corresponding reduced negotiation game is the game red 훽 휆푖 (G) ↾푣 0 =(픓,ℭ,푉 r ,(푉 r 픓 ,푉 r ℭ ),퐸 r , 휇 r ) ↾푣 0 , where: • the vertices controlled by Prover are the vertices ofG, i.e. the set푉 r 픓 =푉 ; • the vertices controlled by Challenger are the vertices: – of the form [ℎ푐 ∞ 휋], where ℎ푐 ∞ 휋 is a 휆-consistent 훽-punishment family; – of the form(푐,푢), whereℎ푐 ∞ 휋is a휆-consistent훽-punishment family, and where there exists a vertex 휋 푘 ∈ 푉 푖 along the play 휋 such that 휋 푘 푢 ∈ 퐸; –of the form[ℎ ′ 푣], whereℎ푐 ∞ 휋is a휆-consistent훽-punishment family, and the historyℎ ′ 푣is such that ℎ ′ is a prefix of the history ℎ푐, and last(ℎ ′ ) ∈ 푉 푖 ; – ⊤ and⊥; • with the same notations, the set 퐸 r contains the edges of the form: – 푣[ℎ푐 ∞ 휋] (Prover proposes a punishment family); – 푣⊥ (Prover gives up); – [ℎ푐 ∞ 휋]⊤ (Challenger accepts Prover’s proposal); – [ℎ푐 ∞ 휋][ℎ ′ 푣] (Challenger deviates before the punishing cycle—pre-cycle deviation); – [ℎ푐 ∞ 휋](푐,푢) (Challenger deviates after the punishing cycle—post-cycle deviation); – [ℎ ′ 푣]푣 and(푐,푢)푢 (Prover has now to propose a new play); – ⊤ and⊥ (the play is over); • given a history푔 = 푔 0 . . .푔 푛 ∈ Hist red 훽 휆푖 (G)that does not reach the vertex⊥, we denote by ¤푔 =ℎ (1) . . .ℎ (푛) the history or play inG defined by, for each 푘 : – if 푔 푘−1 푔 푘 = 푣[ℎ푐 ∞ 휋], then ℎ (푘) is empty; –if푔 푘−1 푔 푘 =[ℎ푐 ∞ 휋]⊤, thenℎ (푘) . . .ℎ (푛) =ℎ푐 | ℎ (1) ...ℎ (푘−1) ℎ | 2 휋 (the number of times the cycle푐is repeated depends quadratically on time); – if 푔 푘−1 푔 푘 =[ℎ푐 ∞ 휋][ℎ ′ 푣], then ℎ (푘) =ℎ ′ ; –if푔 푘−1 푔 푘 =[ℎ푐 ∞ 휋](푐,푣), thenℎ (푘) =ℎ푐 | ℎ (1) ...ℎ (푘−1) ℎ | 2 ℎ ′ , whereℎ ′ is one of the shortest histories such that Occ(ℎ ′ ) ⊆ Occ(휋), last(ℎ ′ ) ∈ 푉 푖 and last(ℎ ′ )푣 ∈ 퐸; 6.6. COMPUTING THE NEGOTIATION FUNCTION99 푎 푎푏 ∞ (푎 3 푏 3 ) 휔 푎 (푏,푏) 푏 . . . . . . . . .. . . ⊤⊥ Figure 22: The reduced negotiation game – if 푔 푘−1 푔 푘 =[ℎ ′ 푣]푣 or(푐,푣)푣 , then ℎ (푘) is empty; and that definition is naturally extended to plays: for example, ifGis the game of Figure 19a and if 휋 = 푎[푎푏 ∞ 푎 휔 ][푎]푎[푎푏 ∞ 푎 휔 ](푏,푏)푏[푏 ∞ 푎 휔 ]⊤ 휔 , then we write ¤휋 = 푎· 푎푏 2 2 · 푎푏 7 2 푎 휔 = 푎 2 푏 4 푎푏 49 푎 휔 ; • the payoff function휇 r is defined, for each play휋, by휇 r ℭ (휋) =−휇 r 픓 (휋) =+∞if휋reaches the vertex ⊥, and by 휇 r ℭ (휋) =−휇 r 픓 (휋) = 휇 푖 (¤휋) otherwise. Example 12. Figure 22 illustrates a (small) part of the gamered 2 휆◦ (G) ↾푎 , whereGis the game of Figure 19a, and휆(푎) = 휆(푏) =1. Blue vertices are owned by Prover, orange ones by Challenger. When Prover proposes the punishment family푎푏 ∞ (푎 3 푏 3 ) 휔 , the function휇 r ℭ interprets it as the play푎푏 |ℎ| 2 (푎 3 푏 3 ) 휔 , whereℎis the history that has already been constructed so far. Remark 9. Reduced negotiation games are Borel, and are played on a finite graph. We will now prove that the reduced negotiation game captures the negotiation function, as do the abstract and concrete ones. For that purpose, we first need the following key result. Lemma 19. In a reduced negotiation game, Prover has a stationary optimal strategy. Proof. This lemma is a consequence of Lemma 2: the payoff function휇 r ℭ is concave. Indeed, let휉 be a shuffling of two plays휋and휒. If either휋or휒reaches the vertex⊥(in which case both do), then we immediately have휇 r ℭ ( ¤ 휉) ≤ max휇 r ℭ (¤휋), 휇 r ℭ (¤휒) =+∞. Otherwise, the play ¤ 휉 is a shuffling of¤휋and¤휒, and since mean-payoff objectives defined with a limit inferior are convex, we have 휇 r ℭ ( ¤ 휉) ≤ max휇 r ℭ (¤휋),휇 r ℭ (¤휒).□ This lemma enables us to prove that the reduced negotiation game is equivalent to the other negotiation games. Lemma 20. There exists a polynomial푃 5 such that for every mean-payoff gameG, every requirement휆 with rational values, each player푖and each푣 0 ∈ 푉 푖 , for every훽 ≥ 푃 5 (∥G∥+∥휆∥), we havenego(휆)(푣 0 ) = val ℭ red 훽 휆푖 (G) ↾푣 0 . Proof.For every mean-payoff gameG and every requirement 휆, we assume: 훽 ≥ 푃 1 ◦ 푃 2 ( ∥ mp(푐) | 푐 ∈ SCyc(G) ∥ ) 100CHAPTER 6. MEAN-PAYOFF GAMES and for each 푣 ∈ 푉 : 훽 ≥ ∥휆(푣)∥+ 3, which are indeed quantities that are bounded by a polynomial of∥G∥+∥휆∥. Contrary to what we did with parity games in the proof of Lemma 7, here, the equivalence between the reduced and the abstract game cannot be established directly (the proof would be too technical): we need to get back to the definition of the negotiation function. ▶First direction: nego(휆)(푣 0 ) ≥ val ℭ red 훽 휆푖 (G) ↾푣 0 . Let ̄ 휎 −푖 be a strategy profile inGthat is휆-rational assuming a strategy휎 푖 , and let푥 = sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ . We wish to prove that there exists a strategy휏 픓 in the reduced negotiation game such thatsup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩ ≤ 푥 . Thus, we will have proved that the quantityval ℭ red 훽 휆푖 (G) ↾푣 0 is smaller than or equal to every such 푥 , and therefore smaller than or equal to nego(휆)(푣 0 ). Let us define simultaneously the strategy휏 픓 and a mapping휑:Hist 픓 red 훽 휆푖 (G) ↾푣 0 ! HistG ↾푣 0 , such that for each history푔, the punishment family휏 픓 (푔)will be defined from the play⟨ ̄ 휎 ↾휑(푔) ⟩. We guarantee inductively that if푔 ∈ Hist 픓 red 훽 휆푖 (G) ↾푣 0 is compatible with휏 픓 , then휑(푔) ∈ HistG ↾푣 0 is compatible with ̄ 휎 −푖 . First, let us define 휑(푣 0 ) = 푣 0 . Let푔 ∈ Hist 픓 red 훽 휆푖 (G) ↾푣 0 be a history compatible with휏 픓 as it has been defined so far, and such that휑(푔)has already been defined. Let휒 0 = ⟨ ̄ 휎 ↾휑(푔) ⟩. By induction hypothesis, the history휑(푔)is compatible with ̄ 휎 −푖 , hence the play 휒 0 is 휆-consistent, and satisfies 휇 푖 (휒 0 ) ≤ 푥 . Let휒 0 ≤ℓ be the shortest prefix of휒 0 that is not simple, i.e. such that there exists푘< ℓwith휒 0 푘 = 휒 0 ℓ . If we have mp 푖 (휒 0 푘+1 . . . 휒 0 ℓ ) ≤ 푥 , then we define: 휏 픓 (푔) = 휒 0 ≤푘 휒 0 푘+1 . . . 휒 0 ℓ ∞ 휋 , where휋is a play such thatOcc(휋) = Occ 휒 0 >ℓ , that휇 푖 (휋) ≤ 푥, and that∥휇(휋)∥ ≤ 훽. Such a play exists, because the polytope: 푍 = ( 휇(휋) ∀푗,∀푣 ∈ 푉 푗 ∩ Occ(휒 0 ), 휇 푗 (휋) ≥ 휆(푣), and Occ(휋) = Occ 휒 0 >ℓ ) is nonempty (it contains the vector휇(휒 0 >ℓ ) ), and has at least one vertex ̄ 푧with푧 푖 ≤ 푥(because we have 휇 푖 (휒 0 >ℓ ) ≤ 푥 ), which by Lemma 13 and Corollary 1 has size∥ ̄ 푧∥ ≤ 훽. Otherwise, if we havemp 푖 (휒 0 푘+1 . . . 휒 0 ℓ )> 푥 , we define휒 1 = 휒 0 ≤푘 휒 0 >ℓ , and we iterate the process, which does necessarily terminate since we have휇 푖 (휒 0 ) ≤ 푥. As a consequence, it effectively defines the proposal: 휏 픓 (푔) = 휒 푛 ≤푘 휒 푛 푘+1 . . . 휒 푛 ℓ ∞ 휋 , for some 푛. Then, for each prefix ℎ푣 , we define: 휑 푔 휒 푛 ≤푘 휒 푛 푘+1 . . . 휒 푛 ℓ ∞ 휋 [ℎ푣]푣 = 휑(푔)휒 0 ≤푚 , where휒 0 ≤푚 is the prefix of휒 0 from which푛simple cycles have been pulled out to obtain the prefixℎ 6.6. COMPUTING THE NEGOTIATION FUNCTION101 of 휒 푛 ; and similarly, for each pair(푐,푣), we define: 휑 푔 휒 푛 ≤푘 휒 푛 푘+1 . . . 휒 푛 ℓ ∞ 휋 (푐,푣)푣 = 휑(푔)휒 0 ≤푝 , where휒 0 ≤푝 is the prefix of휒 0 from which푛simple cycles have been pulled out to obtain the shortest prefix 휒 푛 ≤푝 of 휒 푛 such that 휒 푛 푝 ∈ 푉 푖 and 휒 푛 푝 푣 ∈ 퐸. Thus, the mapping휑is defined on every history compatible with휏 픓 , and the image of such a history is always a history compatible with ̄ 휎 −푖 . We define it arbitrarily on other histories. Note that for each history푔, the history¤푔can be obtained from휑(푔)by pulling out cycles푐satisfyingmp 푖 (푐)> 푥, and adding cycles푑withmp 푖 (푑) ≤ 푥. As a consequence, if we havemp 푖 (휑(푔)) ≤ 푥, then we have mp 푖 (¤푔) ≤ 푥 —and the same result is true when we naturally extend the mapping 휑 to plays. Let us now prove that we havesup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩ ≤ sup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ . Let휉be a play compatible with휏 픓 : •the vertex⊥ does not appear in 휉 , because Prover’s strategy does never use a edge to it. •If the play 휉 has the form 휉 =푔[ℎ푐 ∞ 휋]⊤ 휔 : then, we have 휇 r ℭ (휉) = 휇 푖 (휋) ≤ 푥 . •If the play휉is made of infinitely many deviations: the play휑(휉)is compatible with ̄ 휎 −푖 , hence 휇 푖 (휑(휉)) ≤ 푥 ; which implies 휇 푖 (¤휒) ≤ 푥 , i.e. 휇 r ℭ (휉) ≤ 푥 . ▶Second direction: nego(휆)(푣 0 ) ≤ val ℭ red 훽 휆푖 (G) ↾푣 0 . Let휏 픓 be a stationary strategy for Prover in the reduced negotiation game, and let푦 = sup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩. We want to show thatnego(휆)(푣 0 ) ≤ 푦: by Lemma 19, it will be enough to conclude. If we have푦 =+∞, it is clear. Let us assume푦 ≠+∞. Then, we will define a strategy profile ̄ 휎, where ̄ 휎 −푖 is휆-rational assuming휎 푖 , such thatsup 휎 ′ 푖 휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ 푦. We proceed inductively by defining the play⟨ ̄ 휎 ↾ℎ푣 ⟩for each historyℎ푣compatible with ̄ 휎 −푖 such thatℎis empty, orlast(ℎ) ∈ 푉 푖 and푣 ≠ 휎 푖 (ℎ). Such a history is called a bud history. After other histories, the strategy profile can be defined arbitrarily. To that end, we construct a mapping휓which maps each bud history to a history휓(ℎ푣) ∈ Hist 픓 red 훽 휆푖 (G) ↾푣 0 that is compatible with휏 픓 . This mapping will induce a definition of ̄ 휎 : since푦 ≠+∞, we have휏 픓 (휓(ℎ푣)) ≠⊥: let then[ℎ ′ 푐 ∞ 휋] =휏 픓 (휓(ℎ푣)). We then define⟨ ̄ 휎 ↾ℎ푣 ⟩ =ℎ ′ 푐 |ℎ ′ | 2 휋 , which is a휆-consistent play since ℎ ′ 푐 ∞ 휋 is a 휆-consistent punishment family, by definition of the reduced negotiation game. Let nowℎ 0 푣be a bud history: we assume that ̄ 휎has been defined on every prefix ofℎ 0 , but not onℎ 0 푣itself. Ifℎ 0 is empty, that is ifℎ 0 푣 = 푣 0 , then we define휓(ℎ 0 푣) = 휓(푣 0 ) = 푣 0 . Otherwise, let us writeℎ 0 = ℎ 1 푤ℎ 2 , whereℎ 1 푤is the longest prefix ofℎ 0 that is a bud history—that is, its longest prefix such that휓(ℎ 1 푤)has been defined, or its shortest prefix such that푤ℎ 2 is compatible with ̄ 휎 ↾ℎ 1 푤 . Let푔 = 휓(ℎ 1 푤), and let[ℎ푐 ∞ 휋] = 휏 픓 (푔). We have defined⟨ ̄ 휎 ↾ℎ 1 푤 ⟩ = ℎ푐 |ℎ 1 ℎ| 2 휋, and consequently, the history푤ℎ 2 is a prefix of that play. If it is a prefix of the historyℎ푐, then we define 휓(ℎ 0 푣) =푔[ℎ푐 ∞ 휋][푤ℎ 2 푣]푣 . Otherwise, we define휓(ℎ 0 푣) =푔[ℎ푐 ∞ 휋](푐,푣)푣 . Now, the strategy profile ̄ 휎has been defined, and since all the punishment families proposed by Prover are휆-consistent, the strategy profile ̄ 휎 −푖 is휆-rational assuming휎 푖 . Let휒be a play compatible with ̄ 휎 −푖 , and let us prove that휇 푖 (휒) ≤ 푦. If휒has finitely many prefixes that are bud histories, then let휒 ≤푛 be the longest one: we have휒 ≥푛 = ℎ푐 푛+|ℎ| 휋, where[ℎ푐 ∞ 휋] = 휏 픓 (휓(휒 ≤푛 )). Then, we have 휇 푖 (휒) = 휇 푖 (휋) ≤ 푦. Now, if휒has infinitely many such prefixes, then there exists a unique play휋in the reduced negotiation game such that for any prefix휒 ≤푛 of휒that is a bud history, the history휓(휒 ≤푛 )is a prefix 102CHAPTER 6. MEAN-PAYOFF GAMES of휋. Then, if휋contains finitely many post-cycle deviations, then there exist two indices푚and푛such that 휒 ≥푚 = ¤휋 ≥푛 , hence 휇 푖 (휒) = 휇 푖 (¤휋) ≤ 푦. Finally, if휋contains infinitely many post-cycle deviations, i.e. infinitely many occurrences of a vertex of the form(푐,푣), then let us choose such vertex that minimizes the quantitymp 푖 (푐). The play 휒 has the form: 휒 =ℎ 0 푐 푘 2 0 ℎ 1 푐 푘 2 1 ℎ 2 . . ., where for each 푛, we have 푘 푛 = ℎ 0 푐 푘 2 0 . . .푐 푘 2 푛−1 ℎ 푛 . Then, if we write 푀 = max푟 푖 , we have: mp 푖 ℎ 0 푐 푘 2 0 . . .ℎ 푛 푐 푘 2 푛 ≤ 1 푘 푛 +푘 2 푛 |푐|− 1 푘 푛 푀+ 푘 2 푛 |푐|− 1 mp 푖 (푐) , which converges tomp 푖 (푐)when푛tends to+∞, hence휇 푖 (휒) ≤ mp 푖 (푐). Now, since휏 픓 is stationary, there exists a play of the form푔푑 휔 that is compatible with it, and such that(푐,푣) ∈ Occ(푑) ⊆ Inf(휋); and by definition of푦, we havemp 푖 ( ¤ 푑) = 휇 r ℭ (푔푑 휔 ) ≤ 푦. By minimality ofmp 푖 (푐), we havemp 푖 ( ¤ 푑) = mp 푖 (푐), hence mp 푖 (푐) ≤ 푦, and therefore 휇 푖 (휒) ≤ 푦.□ Thus, a given requirement휆is an휀-fixed point of the negotiation function if and only if for each푖 and푣 ∈ 푉 푖 , there exists a stationary strategy휏 픓 in the gamered 훽 휆푖 , with훽 = 푃 5 (∥G∥+∥휆∥), such that sup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩ ≤ 휆(푣)+ 휀 . The reduced negotiation game has an exponential size, but it contains onlycard푉 vertices that are controlled by Prover: stationary strategies for Prover are therefore objects of polynomial size. Such strategies will be called reduced strategies, and can alternatively be seen as simple strategies in the asbtract negotiation game, as suggested in the end of Subsection 6.6.1. They constitute the third and last piece of our notion of witness. 6.7ALGORITHM AND COMPLEXITY We are now in a position to define formally our notion of witness. Definition 34 (Witness). Let퐼 =(G ↾푣 0 , ̄ 푥, ̄ 푦,휀) be an instance of the휀-SPE constrained existence problem. A witness for 퐼 is a tuple 푊,푊 ′ , ̄ ̄ 훼,휆,(휏 푣 픓 ) 푣 , where: • 푊 ⊆ 푊 ′ ⊆ 푉 ; • ̄ ̄ 훼 ∈ [0, 1] Π×SCyc(푊) ; • 휆 is a requirement; • and each 휏 푣 픓 is a stationary strategy in the game red 훽 휆푖 (G) ↾푣 , where 훽 = 푃 5 (∥G∥+∥휆∥). A witness is valid if: • each strategy 휏 푣 픓 satisfies the inequality sup 휏 ℭ 휇 r ℭ ⟨휏 푣 픓 ,휏 ℭ ⟩ ≤ 휆(푣)+ 휀; • the sets푊 and푊 ′ and the tuple of tuples ̄ ̄ 훼 satisfy the hypotheses of Theorem 18. The휀-SPE constrained existence problem will beNP-easy if we show, first, that there exists a valid witness of polynomial size if and only if the instance is positive, and second, that the validity of a witness can be decided in polynomial time. The former is a consequence of Theorem 8 and Lemmas 17 to 20. 6.7. ALGORITHM AND COMPLEXITY103 Lemma 21. There exists a polynomial푃 6 such that an instance퐼of the휀-SPE constrained existence problem admits a valid witness of size at most 푃 6 (∥퐼∥) if and only if it is a positive instance. Let us now tackle the latter. Lemma 22. Given an instance of the휀-SPE constrained existence problem and a witness for it, deciding whether that witness is valid is P-easy. Proof.The validity of a witness has been defined by two conditions. As regards the second one, all the hypotheses of Lemma 18 can be checked in polynomial time with classical algorithms. Let us now show how the first condition can also be checked in polynomial time. Let us recall that in our algorithm for parity games, at the same point, we constructed the game induced by each strategy guessed for Prover, and searched for a strategy for Challenger, i.e. for a play in that game, that would contradict the claim that Prover’s strategy is winning. Such a play could be seen as a winning path for either a reachability or a parity objective. Here, the idea will be similar, but we need to consider Prover’s strategy in the reduced negotiation game, so that it indeed induces a finite graph; and then, the search of a path for Challenger will be slightly more complicated. Let푛 = card푉. Given a stationary strategy휏 푣 픓 of Prover in a reduced negotiation game, one can construct in a time polynomial in∥휏 푣 픓 ∥the gamered 훽 휆푖 (G) ↾푣 [휏 푣 픓 ]. The underlying graph of that game has indeed a polynomial size, because it is composed only of: •at most 푛 vertices of the form 푤 ∈ 푉 ; •at most 푛 vertices of the form 휏 픓 (푤) (either equal to⊥ or of the form [ℎ푐 ∞ 휋]); •at most 푛 2 vertices of the form(푐,푤); •at most 2푛 2 vertices of the form[ℎ ′ 푤 ′ ], whereℎ ′ is a prefix of the historyℎ푐for some punishment family [ℎ푐 ∞ 휋] =휏 픓 (푤); •possibly the vertices⊤ and⊥. We call this connected graph the deviation graph. Note that if among those vertices, there is the vertex⊥, then since the vertices that are not accessible have been removed, we havesup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩ =+∞ and the problem can be solved immediately. In what follows, we assume that it is not the case, i.e. that for each푤, the vertex휏 픓 (푤)has the form[ℎ푐 ∞ 휋]. Deciding whethersup 휏 ℭ 휇 r ℭ ⟨ ̄ 휏⟩ ≤ 훼 is then equivalent to deciding whether there exists a path휋, in that graph, such that we have휇 r ℭ (휋)> 훼. Such a play can have three forms. •It can end in the vertex⊤, i.e. with Challenger accepting Prover’s proposal. The existence of such a play can be decided immediately, by checking whether in the deviation graph, there exists a vertex of the form [ℎ푐 ∞ 휋] with 휇 푖 (휋)> 훼 . • It can avoid the vertex⊤, and comprise finitely many post-cycle deviations. This is the case if and only if there exists a cycle푑in the deviation graph, without post-cycle deviations, such that we havemp 푖 ( ¤ 푑)> 훼. The existence of such a cycle can be decided in polynomial time with Karp’s algorithm [Kar78]. • It can avoid the vertex⊤, and comprise infinitely many post-cycle deviations. In that case, we have휇 r ℭ (휋) ≤ mp 푖 (푐)for each vertex of the form(푐,푤)appearing infinitely often along휋; then, there exists a cycle푑in the deviation graph, such that every vertex of the form(푐,푤)along푑 104CHAPTER 6. MEAN-PAYOFF GAMES satisfiesmp 푖 (푐)> 훼. Conversely, if such a cycle exists, then휋exists. The existence of such a cycle can be decided in polynomial time with Karp’s algorithm. Therefore, the existence of such a play is decidable in polynomial time.□ Thus, given an instance of the휀-SPE constrained existence problem, a valid witness can be guessed and checked in polynomial time. Applying Lemma 9, we finally obtain the following theorem. Theorem 15. The휀-SPE constrained existence problem in mean-payoff games isNP-complete. Hardness still holds when there is only one effective upper threshold. CHAPTER 7:ENERGY AND DISCOUNTED- SUM GAMES We now tackle subgame-perfect equilibria in energy and in discounted-sum games. Here, the negotiation function will not be used, and our results and proofs will mostly be extensions of those presented in Part I. 7.1DISCOUNTED-SUM GAMES 7.1.1 Hardness We showed in Part I that the constrained existence problem of NEs in discounted-sum games was at least as hard as the target discounted-sum problem, whose decidability remains an open problem. That result also holds when we replace NEs with SPEs. Theorem 16. The TDS problem reduces to the constrained existence problem of SPEs in discounted-sum games. This hardness result still holds when there is no effective lower threshold, and only one effective upper threshold. Proof.The proof is analogous to that of Theorem 3.□ 7.1.2 Co-recursive enumerability Since, by Theorem 16, we know that finding an algorithm for the SPE constrained existence problem would imply solving a long-standing open problem, we only present a semi-algorithm that recognizes its negative instances, leaving decidability as an open question. Theorem 17. In discounted-sum games, the constrained existence problem of SPEs is co-recursively enumer- able. Proof.We present here a modified version of the semi-algorithm presented in the proof of Theorem 4, so that it recognizes the negative instances of the SPE constrained existence problem. ▶Algorithm Let ℎ (푛) 푛∈N be a recursive enumeration of the nonempty histories inG ↾푣 0 by increasing order of lengths (we can, for example, order histories of the same length with the lexicographic order induced by some arbitrary order on vertices). 105 106CHAPTER 7. ENERGY AND DISCOUNTED-SUM GAMES Now, let푇be the infinite tree whose nodes of depth푛+1 are all possible푛-uples ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) , and where the children of the node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) are the nodes of the form: ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ), ̄ 휎(ℎ (푛+1) ) . Thus, every node partially defines a complete strategy profile ̄ 휎 , and every infinite branch entirely defines it—and every complete strategy profile is defined by an infinite branch. A node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) is called subgame-irrational if there exist two indicesℓ,푚 ≤ 푛, with |ℎ (ℓ) | =|ℎ (푚) | = 푝, an index 푘 ≤ 푝 and a player 푖 such that: •we have ℎ (ℓ) ≤푘 =ℎ (푚) ≤푘 ; •the history ℎ (ℓ) ≥푘 is compatible with the strategy profile ̄ 휎 ↾ℎ (ℓ) ≤푘 , as partially defined by the node; •the history ℎ (푚) ≥푘 is compatible with the strategy profile ̄ 휎 −푖↾ℎ (푚) ≤푘 ; •and finally, we have: ds 푖 ℎ (푚) − ds 푖 ℎ (ℓ) > 2푀훽 푝−1 . The same node is called off-topic if for some player푖we haveds 푖 ( ℎ ) > 푦 푖 − 푀훽 |ℎ|−1 , ords 푖 (ℎ)< 푥 푖 + 푀훽 |ℎ|−1 , where ℎ is the longest history that is compatible with ̄ 휎 as defined so far. Our algorithm consists in constructing the tree푇 ′ , obtained from the tree푇by cutting every branch after the first subgame-irrational or off-topic node, and terminating once that construction is finished. ▶Correctness Our correctness proof will use Lemma 3, that was proved inside the proof of Theorem 4. To show that our algorithm is correct, we must prove that the tree푇 ′ will be finite if and only if we have a negative instance of the constrained existence problem. By Kőnig’s lemma, that will be the case if and only if every branch is finite, i.e., if and and only if every branch of the tree푇contains either an off-topic or a Nash-irrational node. Therefore, we will be done if we prove that, given a branch of푇 and the corresponding strategy profile ̄ 휎, the branch contains no such node if and only if the strategy profile ̄ 휎 is an SPE and satisfies ̄ 푥 ≤ 휇⟨ ̄ 휎⟩ ≤ ̄ 푦. Off-topic nodes. Using Lemma 3, we know that we have ̄ 푥 ≤ 휇 푖 ⟨ ̄ 휎⟩ ≤ ̄ 푦if and only if the corresponding branch contains no off-topic node. If the branch contains a subgame-irrational node, then the strategy profile ̄ 휎is not an SPE. Let us assume that some branch contains a subgame-irrational node ̄ 휎(ℎ (0) ), . . ., ̄ 휎(ℎ (푛) ) . Let us use the notations푘,ℓ,푚,푝and푖from the definition of subgame-irrationality. Let us define휋 =⟨ ̄ 휎 ↾ℎ (ℓ) ≤푘 ⟩ : note that the historyℎ (ℓ) ≥푘 is a prefix of length푝−푘of the play휋. Similarly, let us extend the history ℎ (푚) ≥푘 into some play 휋 ′ , compatible with the strategy profile ̄ 휎 −푖↾ℎ (푚) ≤푘 . By Lemma 3, the inequality: ds 푖 ℎ (푚) − ds 푖 ℎ (ℓ) > 2푀훽 푝−1 implies 휇 푖 (ℎ (ℓ) <푘 휋 ′ )> 휇 푖 (ℎ (ℓ) <푘 휋), and therefore the strategy profile ̄ 휎 is not an SPE. 7.2. ENERGY GAMES107 Conversely, if ̄ 휎is not an SPE, then there exists a historyℎ푣, a player푖and a strategy휎 ′ 푖 such that we have: 휇 푖 (ℎ⟨ ̄ 휎 ↾ℎ푣 ⟩)< 휇 푖 (ℎ⟨ ̄ 휎 −푖↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ⟩). Then, let휋 = ⟨ ̄ 휎 ↾ℎ푣 ⟩, and let휋 ′ = ⟨ ̄ 휎 −푖↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ⟩: since we have휇 푖 (ℎ휋)< 휇 푖 (ℎ휋 ′ ), by Lemma 3, there exists an index 푞 such that: ds 푖 ℎ휋 ≤푞 + 푀훽 |ℎ|+푞−1 < ds 푖 휋 ′ ≤푞 − 푀훽 |ℎ|+푞−1 . Let nowℓand푚be the indices such thatℎ (ℓ) = ℎ휋 ≤푞 , andℎ (푚) = ℎ휋 ′ ≤푞 . Let푝 =|ℎ (ℓ) | =|ℎ (푚) |, and let푘 =|ℎ|. Then, we haveds 푖 ℎ (푚) ≥푘 − ds 푖 ℎ (ℓ) ≥푘 > 2푀훽 푝−1 : along the branch corresponding to ̄ 휎 , the node of depth maxℓ,푚 is subgame-irrational. Which ends the proof.□ 7.2ENERGY GAMES 7.2.1 Undecidability We proved in Part I that the constrained existence problem of Nash equilibria in energy games is undecidable, since counter machines can be encoded as energy games. It should therefore not come as a surprise that the same problem for SPEs is also undecidable. We prove however a slightly stronger result: undecidability holds even when the game contains only two players. Theorem 18. In energy games, the SPE constrained existence problem is undecidable, even on games with only two players, and even with no effective lower threshold and only one effective upper threshold. Proof.We proceed by reduction from the halting problem of a two-counter machine. LetKbe a two-counter machine. We define an energy gameG ↾푞 0 with two players, player퐶 1 and player퐶 2 , by assembling the gadgets presented in Figures 23a to 23e. The rewards that are not presented are equal to 0, the round vertices are those controlled by player퐶 1 , and the square ones are those controlled by player퐶 2 . For each state of the machineK, we define from one to three vertices, plus six additional vertices, written 푎, 푏, 푐,△, ▽, and ▽ 1 . Then, a play inG ↾푞 0 that does not reach one of the sink vertices△,▽, or▽ 1 simulates a sequence of transitions of the machineK, that can be a run or not: at each step, the counter퐶is captured by the energy level of player퐶. Let us now prove that the gameG ↾푞 0 admits an SPE where player퐶 2 loses if and only if the machineK terminates. ▶If such an SPE exists, then the machineK terminates. Let us write ̄ 휎for such an SPE, and let휋 =⟨ ̄ 휎⟩. Let us show that the play휋simulates a run ofK. Since 휋 is lost by player퐶 2 , there are a priori three possibilities. The play휋reaches the vertex▽.Then, player퐶 1 loses, and has therefore a profitable deviation by looping on 휋 0 =푞 0 , which is impossible. 108CHAPTER 7. ENERGY AND DISCOUNTED-SUM GAMES 푞 0 (a) Gadget for the initial state 푞 f 퐶 2 −1 (b) Gadget for the final state 푞 퐶 1 (c) Gadget for incrementations of counter퐶 ∈ 퐶 1 ,퐶 2 푞 (if퐶 1 > 0) 푞 ′ (if퐶 1 = 0) 푎 △ ▽ 퐶 1 −1 퐶 1 −1 퐶 1 −1 퐶 2 −1 (d) Gadget for tests of counter퐶 1 푞 푞 ′ (if퐶 2 = 0) △ 푞 ′ (if퐶 2 > 0) 푏 푐 △ ▽ 1 ▽ 퐶 2 −1 퐶 2 −1 퐶 2 −1 퐶 1 −1 퐶 2 −1 퐶 1 −1 (e) Gadget for tests of counter퐶 2 Figure 23: Gadgets 7.2. ENERGY GAMES109 Player퐶 1 makes player퐶 2 lose by going to a vertex of the form푞 ′ , when his energy level is zero. Thus, the play휋simulates a spurious run ofK, since such an action amounts to faking a test of퐶 2 above zero. But then, player퐶 2 has a profitable deviation by going to푏: there, player퐶 1 cannot go to the vertex▽, because it would make her lose, while she can win by going to the vertex푐; indeed, from there, player퐶 2 cannot go to the vertex▽ 1 because it would make him lose, since he has no more energy, while he can go to the vertex△and win—and let player퐶 1 win. Therefore, this case is also impossible. The play휋reaches the vertex푞 f . Then, it simulates a correct run of the machineK, that reaches the state푞 f . Indeed, we have already shown that휋cannot fake a test of퐶 2 above 0. It cannot fake a test of퐶 1 above 0 either, because then, player퐶 1 would lose, while she can win by looping on 푞 0 . It cannot fake a test of퐶 2 to 0, because then, from the vertex푞 ′ , if player퐶 2 ’s energy is greater than 0, he has a profitable deviation by going to the vertex△. Finally, it cannot fake a test of퐶 1 to 0, because then, from the vertex푞 ′ , if player퐶 1 ’s energy is greater than 0, player퐶 2 has a profitable deviation by going to푎, from where player퐶 1 cannot go to the vertex▽, because it would make her lose while she can win by going to△. Therefore, the machineK terminates. ▶If the machineK terminates, then there is an SPE where player퐶 2 loses. If the machineK terminates, let us define a strategy profile ̄ 휎 inG ↾푞 0 as follows. • In tests of counter퐶 1 , player퐶 1 goes to푞 ′ if and only if her energy level is zero or⊥; from푞 ′ , player퐶 2 goes to the vertex 푎 if and only if player퐶 1 ’s energy level is positive. •In tests of counter퐶 2 , player퐶 1 goes to푞 ′ if and only if the energy level of player퐶 2 is zero or ⊥; from푞 ′ , player퐶 2 goes to the vertex△if and only if his energy level is positive, and from푞 ′ , he goes to 푏 if and only if it is zero or⊥. •From the vertex 푎, player퐶 1 goes to the vertex ▽ if and only if her energy level is zero or⊥. •From the vertex푏, player퐶 1 goes to the vertex▽if and only if the energy level of player퐶 2 is positive. •From the vertex 푐, player퐶 2 goes to the vertex ▽ 1 if and only if his energy level is positive. Let us show that the strategy profile ̄ 휎is an SPE. Letℎ푣be a history from the vertex푞 0 : let us prove that ̄ 휎 ↾ℎ푣 is an NE. If푣is a vertex of the form푞.Then, the play⟨ ̄ 휎 ↾ℎ푣 ⟩simulates the correct run of the machineK from푞when the counters are initialized toel 퐶 1 (ℎ푣)andel 퐶 2 (ℎ푣)(if one of those energy levels is⊥, then it simulates the correct run ofKwhere that counter is locked to 0). Therefore, player퐶 1 wins (or has already lost), hence she cannot have a profitable deviation. As for player퐶 2 , he cannot have a profitable deviation from a vertex of the form푞 ′ : when such a vertex is reached, player퐶 1 ’s energy is zero or⊥, hence if player퐶 2 chooses to go to푎, player퐶 1 will go to▽and he will lose. He cannot have a profitable deviation from a vertex of the form푞 ′ either: when such a vertex is reached, he has a zero or⊥energy level, hence if he chooses to go to the vertex△, he loses. Finally, he cannot have a profitable deviation from a vertex of the form푞 ′ : when such a vertex is reached, he has a positive energy level, hence if he goes to the vertex 푏, player퐶 1 will go to the vertex ▽ and he will lose. 110CHAPTER 7. ENERGY AND DISCOUNTED-SUM GAMES If푣is a vertex of the form푞 ′ . Then, if player퐶 1 ’s energy level is positive, player퐶 2 goes to푎, then player퐶 1 goes to△, and both player퐶 1 and퐶 2 win—and have therefore no profitable deviation. Otherwise, the substrategy profile ̄ 휎 ↾ℎ푣 simulates a correct run ofK, and we can use the same arguments as in the previous point. If푣is a vertex of the form푞 ′ . Then, if player퐶 2 ’s energy level is positive, he goes to△and wins—and no player has a profitable deviation. Otherwise, the substrategy profile ̄ 휎 ↾ℎ푣 simulates a correct run ofK, and we can use the same arguments than in the first point. If푣is a vertex of the form푞 ′ .Then, if player퐶 2 ’s energy level is zero or⊥, player퐶 2 goes to 푏, then player퐶 1 goes to푐, and finally player퐶 2 goes to△, and both player퐶 1 and퐶 2 win—they have therefore no profitable deviation. Otherwise, the substrategy profile ̄ 휎 ↾ℎ푣 simulates a correct run ofK, and we can use the same arguments than in the first point. If 푣 = 푎. Then, either player퐶 1 has a positive energy level, and then she goes to△ and wins, or she has a zero or⊥ energy level, and then she cannot win. If푣 = 푏. Then, either player퐶 2 has a zero or⊥energy level, and then the play ends in△and both players win (or have already lost), or he has a positive energy level, and then player퐶 1 cannot win, since player퐶 2 plans to go to ▽ 1 . If푣 =푐.Then, either퐶 2 has a positive energy level, and he goes to▽ 1 and wins, or he has a zero energy level, and he goes to△ and wins, or he has already lost. If 푣 = ▽, ▽ 1 , 푞 f or△. Then, the proof is immediate. Therefore, the strategy profile ̄ 휎is an SPE, that simulates the run of the machineK. That run terminates, hence the play⟨ ̄ 휎⟩reaches the vertex푞 f , and is lost by player퐶 2 . Which concludes the proof.□ 7.2.2 About recursive enumerability Like for NEs, this proof shows that, in particular, the SPE constrained existence problem is not co- recursively enumerable in energy games. It might still be the case that it is recursively enumerable. That would in particular be the case if finite memory was sufficient for an SPE to achieve a given payoff vector, when that is possible, as in the case of NEs. Unfortunately, one cannot follow this approach, because that statement is false: in order to be able to punish some player, without making another player lose, an SPE may have to memorize their energy levels, and therefore require infinite memory. Theorem 19. There exists an energy game in which there exists an SPE such that no finite-memory SPE generates the same payoff vector. Proof. In the energy game presented in Figure 24, let us show that there exists an SPE that makes player □ lose, but that no finite-memory SPE can achieve that result. 7.2. ENERGY GAMES111 푎 푏 푐 푑 푒 ◦ 1 □ 1 1 ◦ 1 □ 1 1 ◦ 1 ◦ 1 ◦ −1 □ −1 −1 Figure 24: A game where infinite memory is necessary to make player □ lose ▶There is an SPE in which player □ loses. Consider the strategy profile ̄ 휎 defined by ̄ 휎(푎) = 푐 , by ̄ 휎(ℎ푐푎) = 푏 for every historyℎ푐푎, by ̄ 휎(ℎ푏) =푑for every historyℎ푏, by ̄ 휎(ℎ푏푎) =푐for every historyℎ푏푎, by ̄ 휎(ℎ푐) =푑for every historyℎ푐, and finally by⟨ ̄ 휎 ↾ℎ푑 ⟩ =푑 el ◦ (ℎ푑)+1 푒 휔 for every history ℎ푑 . Intuitively: in every subgame, player ◦ makes player□and player^lose. But to do so, she needs one of those players to cooperate with her: for example, in the main subgameG ↾푎 , she wants to make player□lose, and to do so, she traverses the vertex푐to go to the vertex푑. But then, player^may deviate, and go back to the vertex푎. Then, player^has to be punished: and to do so, player ◦ must go to the vertex푑through the vertex푏, with player□’s cooperation. . . and so on. Once the vertex푑 is reached (which will eventually be the case in every subgame), player ◦ ’s energy level is equal to player□’s and player^’s one, plus 1. Thus, to make those players lose without losing herself, player ◦ loops on the vertex 푑 exactly the right number of times, before going to the vertex 푒. Thus, the strategy profile ̄ 휎 is an SPE in which player □ loses—but it is not a finite-memory SPE. ▶There is no finite-memory SPE in which player □ loses. Let ̄ 휏be a finite-memory SPE inG ↾푎 , compatible with a memory structureM, and let us assume toward contradiction that 휇 □ ⟨ ̄ 휏⟩ = 0. Since we have 휇 □ ⟨ ̄ 휏⟩ = 0 (and therefore 휇 ⋄ ⟨ ̄ 휏⟩ = 0, since players □ and^always receive the same rewards), we also have, by induction and because ̄ 휏is an SPE, the equality 휇 □ ⟨ ̄ 휏 ↾ℎ푎 ⟩ = 휇 ⋄ ⟨ ̄ 휏 ↾ℎ푎 ⟩ = 0 for every history ℎ that is compatible with 휎 ◦ . Let now푛be the number of states of the memory structureM, and let us consider a historyℎ푎 with|ℎ| =2푛(and thereforeel ◦ (ℎ푎) = el □ (ℎ푎) = el ⋄ (ℎ푎) = 푛). By the previous proposition, we know that휇 □ ⟨ ̄ 휏 ↾ℎ푎 ⟩ = 휇 ⋄ ⟨ ̄ 휏 ↾ℎ푎 ⟩ =0, i.e. that the play⟨ ̄ 휏 ↾ℎ푎 ⟩reaches the vertex푑, and takes the edge푑more than푛times. Since the memory structureMhas only푛states, it means that that play actually loops on the vertex푑infinitely often, and is therefore also lost by player ◦ . Then, player ◦ has a profitable deviation in that subgame, by going to the vertex 푒: contradiction. There exists an SPE that makes player □ lose in that game, but no finite-memory one.□ We leave therefore the question open. Open Problem 3. Is the constrained existence problem of SPEs in energy games recursively enumerable? 112CHAPTER 7. ENERGY AND DISCOUNTED-SUM GAMES CHAPTER 8:ABOUT RATIONAL VERIFICA- TION AND ITS LIMITS We have now established all our results concerning the constrained existence problem of NEs and SPEs. Before turning to more specialized notions of equilibria, as we will do in Part IV, we use this chapter to discuss a possible application of our results: rational verification. Our observations will lead us to define a new decision problem, for which we also provide tight complexity bounds. 8.1RATIONAL VERIFICATION Verification, without further precision, usually refers to the study of computational systems (typically abstracted as collections of automata or machine models) with the goal of verifying whether they satisfy some property, called a specification. A particular case of verification is reactive verification, already described in Section 1.1, where the computational model under study interacts with an environment that may behave unpredictably. In such a case, one seeks to check whether the system satisfies the specification against every possible behavior of the environment. Illustrations of this concept often involve coffee machines; for the sake of detoxification, let us instead consider an elevator. Its specification could be that, for every푖, if at some point in time a user presses button푖, then the elevator must eventually reach floor푖(or reach it within a specified duration). If the elevator is located in a place where its use is essential, and where a malfunction could have serious consequences—such as in a hospital—it is reasonable to require that the elevator always guarantees this specification, even if a mischievous child persistently presses button 1. Such an example motivates the study of two-player zero-sum games, which serve as a relevant metaphor for reactive verification: if the system (the elevator) and the environment (the child) are two players, we can define the system’s objective as the specification, and we then want to verify whether the system’s strategy is winning. Now, let us imagine that our system is a self-driving car. It would have an obvious Boolean specification: it must not cause the death of any human, whether the passenger or any other person nearby. This specification can be refined into a quantitative objective by defining priorities: for example, we may require that it does not kill any human, and if that condition is satisfied, we may further wish for it to reach a given destination from location퐴to location퐵, if possible within a reasonable duration. In such a framework, it might seem excessively cautious to demand that the system satisfies its specification under every possible behavior of the environment—or equivalently, to treat the environment as a purely adversarial entity whose sole aim is to cause the specification to fail. If we assume, for instance, that all drivers in the world suddenly conspire to assassinate the passenger, then the most rational strategy, given the defined priorities, would be to keep the car safely locked in a garage. This example illustrates that, in some cases, it might be useful to employ the game metaphor without modeling the environment as a 113 114CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS purely adversarial player, but rather as a player or coalition of players with their own objectives, acting rationally according to some notion of rationality. This is the motivation of rational verification. 8.2DEFINITIONS Let us consider a gameG ↾푣 0 with one specific player, called Leader, and denoted by픏 ∈Π. Our decision problem will take as input a strategy휎 픏 for Leader (given as a memory structure), and ask whether that strategy ensures a given specification (encoded by Leader’s payoff function and a threshold) against all rational responses of the other players. We must therefore define an adapted notion of rationality, that does not qualify a strategy profile for all players, but for all players except Leader. Definition 35 (픏-fixed NE,픏-fixed SPE). A strategy profile ̄ 휎is a픏-fixed Nash equilibrium, or픏-fixed NE, if for each player푖 ≠ 픏and every strategy휎 ′ 푖 , we have휇 푖 ⟨ ̄ 휎 −푖 ,휎 ′ 푖 ⟩ ≤ 휇 푖 ⟨ ̄ 휎⟩ . It is a픏-fixed subgame-perfect equilibrium, or 픏-fixed SPE, if for every history ℎ, the strategy profile ̄ 휎 ↾ℎ is a 픏-fixed NE. We can then define the rational verification problem, for a given class of gameCand a given rationality concept휌 ∈ Nash, SubgamePerfect. A (픏-fixed)휌-equilibrium is a Nash equilibrium if휌 = Nash, and an SPE if휌 = SubgamePerfect. Let us recall that memory structures have been defined in Definition 6, and that a non-deterministic memory structureMfor a player푖, from a given vertex푣 0 , induces a set Ind ↾푣 0 (M) of strategies for player 푖. Problem 3 ((Deterministic)휌-rational verification problem in the classC). Given a gameG ↾푣 0 ∈ C , a threshold푡 ∈ Qand a (deterministic) memory structureMfor Leader onG, does every픏-fixed 휌 -equilibrium ̄ 휎 with 휎 픏 ∈ Ind ↾푣 0 (M) satisfy the inequality 휇 픏 ⟨ ̄ 휎⟩> 푡 ? 8.3 LINK WITH THE CONSTRAINED EXISTENCE PROBLEM AND COMPLEXITIES Rational verification problems are phrased in a way that makes them particularly annoying to treat directly, since their instances contain two graph structures: the gameG, and the memory structureM. However, the experienced reader may have noticed that this annoyance can easily be removed, by considering the product of those two structures, or equivalently, by incorporating the memory structureMin the game, so that Leader’s moves are constrained by the game structure itself. To capture non-determinism, a fictional player must be added, who will choose which strategy휎 픏 ∈ Ind ↾푣 0 (M)Leader will follow, without having any interest in that choice (his payoff function will be constantly zero). Since that player may choose the very strategy that does not guarantee the specification, if there is one, we call him Demon. We shall see that such a construction links the rational verification problem to the constrained existence problem, as studied in the previous chapters. Definition 36 (Product game). LetG ↾푣 0 be a game, and letMbe a memory structure for Leader inG. Their product game is the gameG ↾푣 0 ⊗M =(Π∪픇,푉 × ,(푉 × 푖 ) 푖 ,퐸 × , 휇 × ) ↾(푣 0 ,푞 0 ) where the player픇, called Demon, chooses how the memory structureM will run. Formally: • the vertex set is푉 × =(푉 ×푄)∪(푉 ×푄×푄); 8.3. LINK WITH THE CONSTRAINED EXISTENCE PROBLEM AND COMPLEXITIES115 푏,푞 1 푎,푞 0 ,푞 1 푎,푞 0 푏,푞 1 ,푞 1 푎,푞 1 푎,푞 1 ,푞 0 푏,푞 0 ,푞 0 푏,푞 0 푐,푞 0 푐,푞 0 ,푞 0 ◦ 1 □ 1 ◦ 1 □ 1 ◦ 1 □ 1 ◦ 1 □ 1 Figure 25: A product game •Leader controls the set푉 × 픏 =∅, each player푖 ≠ 픏,픇controls the set푉 × 푖 =푉 푖 ×푄×푄 , and Demon controls the set푉 × 픇 =(푉 ×푄)∪(푉 픏 ×푄×푄); • the set 퐸 × contains: –the edge(푢,푝)(푢,푝,푞)for each transition(푝,푢,푞) ∈Δ(if푢 ∉ 푉 픏 ), or(푝,푢,푞,푣) ∈Δ(if푢 ∈ 푉 픏 ); – the edge(푢,푝,푞)(푣,푞) for each transition(푝,푢,푞,푣) ∈Δ (if푢 ∈ 푉 픏 ); – the edge(푢,푝,푞)(푣,푞) for each transition(푝,푢,푞) ∈Δ, and each edge푢푣 ∈ 퐸 (if푢 ∉ 푉 픏 ); •each payoff function휇 × 푖 maps every play(휋 0 ,푞 0 )(휋 0 ,푞 0 ,푞 1 )(휋 1 ,푞 1 ) . . .to the payoff휇 푖 (휋 0 휋 1 . . .)if 푖 ≠ 픇, and to the payoff 0 if 푖 = 픇. Example 13. Figure 25 depicts the gameG ↾푣 0 ⊗M, whenG ↾푣 0 is the game of Figure 4 (Section 2.2) and M the memory structure of Figure 5 (Section 2.4). Leader is then assimilated to player□, and Demon’s vertices are represented by dark red boxes. The unreachable vertices have been omitted, and we have given only the non-zero rewards. Since, from the vertex(푎,푞 0 ,푞 1 ), player ◦ has always the possibility to go to the vertex(푏,푞 1 )and to get the payoff 1, it can be shown that every NE and every SPE in that game gives player□the payoff 1. As we will see now, that means that the strategies induced by the memory structureMguarantee the payoff 1 to player□against Nash-rational or subgame-perfect rational responses, i.e., that the gameG ↾푣 0 , the memory structureM, and the quantity 1− 휀for every휀>0, form a positive instance of the Nash and subgame-perfect rational verification problems. Theorem 20. Let휌 ∈ Nash, SubgamePerfect. LetG ↾푣 0 be a game, letMbe a memory structure for Leader inG, and let푡 ∈ Q. Then, every픏-fixed휌-equilibrium ̄ 휎with휎 픏 ∈ Ind ↾푣 0 (M) satisfies휇 픏 ⟨ ̄ 휎⟩> 푡if and only if every 휌-equilibrium ̄ 휏 in the gameG ↾푣 0 ⊗M satisfies 휇 × 픏 ⟨ ̄ 휏⟩> 푡. Proof.We present the proof when 휌 = SubgamePerfect; the proof for 휌 = Nash is analogous. ▶ If every픏-fixed SPE ̄ 휎with휎 픏 ∈ Ind ↾푣 0 (M)satisfies휇 픏 ⟨ ̄ 휎⟩> 푡, then every SPE ̄ 휏in G ↾푣 0 ⊗M satisfies 휇 × 픏 ⟨ ̄ 휏⟩> 푡 . Let ̄ 휏be an SPE in the gameG ↾푣 0 ⊗M. Let us define a strategy profile ̄ 휎inG ↾푣 0 as follows: for every historyℎ =ℎ 0 . . .ℎ 푛 inG ↾푣 0 , let푔 =(ℎ 0 ,푞 0 )(ℎ 0 ,푞 0 ,푞 1 ) . . .(ℎ 푛 ,푞 푛−1 ,푞 푛 )be the unique history of that form inG ↾푣 0 ⊗Msuch that we have(ℎ 푘 ,푞 푘 ,푞 푘+1 ) = 휏 픇 ((ℎ 0 ,푞 0 ) . . .(ℎ 푘 ,푞 푘 ))for each푘, and let (푣,푞 푛 ) = ̄ 휏(푔). Then, we define ̄ 휎(ℎ) = 푣 . Since the only edges available for Demon in the gameG ↾푣 0 ⊗Mare those that are induced by M, we have휎 픏 ∈ Ind ↾푣 0 (M) . Moreover, the strategy profile ̄ 휎 is a픏-fixed SPE: letℎ =ℎ 0 . . .ℎ 푛 be a 116CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS history from푣 0 that is compatible with the strategy휎 픏 , let푖 ∈Π\픏and let휎 ′ 푖 be a deviation of휎 푖 . We want to prove the inequality휇 푖 ⟨ ̄ 휎 −푖↾ℎ ,휎 ′ 푖↾ℎ ⟩ ≤ 휇 푖 ⟨ ̄ 휎 ↾ℎ ⟩ . Then, let us define the history푔as above, and let휏 ′ 푖 be the strategy that simulates휎 ′ 푖 in the gameG ↾푣 0 ⊗M, i.e., that maps each history of the form푔(푣 1 ,푞 푛 )(푣 1 ,푞 푛 ,푞 푛+1 ) . . .(푣 푘 ,푞 푛+푘−1 ,푞 푛+푘 )to the vertex(휎 ′ 푖 (ℎ푣 1 . . .푣 푘 ),푞 푛+푘 ) . Since the strategy profile ̄ 휏 is an SPE, we have: 휇 × 푖 (푔 <2푛+1 ⟨ ̄ 휏 −푖↾푔 ,휏 ′ 푖↾푔 ⟩) ≤ 휇 × 푖 (푔 <2푛+1 ⟨ ̄ 휏 ↾푔 ⟩), and therefore: 휇 푖 (ℎ <푛 ⟨ ̄ 휎 −푖↾ℎ ,휎 ′ 푖↾ℎ ⟩) ≤ 휇 푖 (ℎ <푛 ⟨ ̄ 휎 ↾ℎ ⟩). Thus, the strategy ̄ 휎 is a픏-fixed SPE that satisfies휎 픏 ∈ Ind ↾푣 0 (M). By hypothesis, it comes that we have 휇 픏 ⟨ ̄ 휎⟩> 푡 , and therefore 휇 × 픏 ⟨ ̄ 휏⟩> 푡 . ▶ If every SPE ̄ 휏 inG ↾푣 0 ⊗ Msatisfies휇 × 픏 ⟨ ̄ 휏⟩> 푡, then every픏-fixed SPE with휎 픏 ∈ Ind ↾푣 0 (M) satisfies 휇 픏 ⟨ ̄ 휎⟩> 푡 . Indeed, let ̄ 휎be a픏-fixed SPE inG ↾푣 0 with휎 픏 ∈ Ind ↾푣 0 (M) . We writeℎ 7! 푞 ℎ for the mapping establishing the fact that the strategy휎 픏 is induced by the memory structureM, as defined in Definition 6. Let us define a strategy profile ̄ 휏inG ↾푣 0 ⊗M as follows: for every history of the form 푔 =(ℎ 0 ,푞 0 )(ℎ 0 ,푞 0 ,푞 1 ) . . .(ℎ 푛 ,푞 푛−1 ), we define ̄ 휏(푔) =(ℎ 푛 ,푞 푛−1 ,푞 ℎ 0 ...ℎ 푛 ) . For every history of the form 푔 =(ℎ 0 ,푞 0 ) . . .(ℎ 푛 ,푞 푛−1 ,푞 푛 ), we define ̄ 휏(푔) =( ̄ 휎(ℎ 0 . . .ℎ 푛 ),푞 푛 ). Then, the strategy profile ̄ 휏is an SPE: let푔 =푔 0 . . .푔 푚 be a history inG ↾푣 0 ⊗M , let푖be a player and let 휏 ′ 푖 be a deviation of 휏 푖 . If 푖 = 픏, then we have: 휇 × 푖 (푔 <푚 ⟨ ̄ 휏 −푖↾푔 ,휏 ′ 푖↾푔 ⟩) ≤ 휇 × 푖 (푔 <푚 ⟨ ̄ 휏 ↾푔 ⟩), because Leader does not control any vertex inG ↾푣 0 ⊗M, hence actually휏 ′ 푖 = 휏 푖 . Likewise if푖 = 픇, because Demon gets the payoff 0 in every play. Now, if푖 ≠ 픇,픏, let us consider without loss of generality that푔has the form푔 = (ℎ 0 ,푞 0 )(ℎ 0 ,푞 0 ,푞 1 ) . . .(ℎ 푛 ,푞 푛 )(if the last vertex is controlled by Demon, it can be removed). Let: (휋 0 ,푞 푛 )(휋 0 ,푞 푛 ,푞 푛+1 )(휋 1 ,푞 푛+1 )· =⟨ ̄ 휏 −푖↾푔 ,휏 ′ 푖↾푔 ⟩. The play휋 = 휋 0 휋 1 . . .is compatible with the strategy profile ̄ 휎 −푖↾ℎ . Therefore, since the strategy profile ̄ 휎 is a 픏-fixed SPE, we have the inequality 휇 푖 (ℎ <푛 휋) ≤ 휇 푖 (ℎ <푛 ⟨ ̄ 휎 ↾ℎ ⟩), i.e., the inequality: 휇 × 푖 (푔 <2푛+1 ⟨ ̄ 휏 −푖↾푔 ,휏 ′ 푖↾푔 ⟩) ≤ 휇 × 푖 (푔 <2푛+1 ⟨ ̄ 휏 ↾푔 ⟩). Thus, the strategy profile ̄ 휏is an SPE. Then, by hypothesis, we have휇 × 픏 ⟨ ̄ 휏⟩> 푡 , and therefore 휇 픏 ⟨ ̄ 휎⟩> 푡 .□ As we will show, this result entails that the휌-rational verification problem in the classCis computa- tionally equivalent to the following problem. Problem 4 (휌-universal threshold problem in the classC). Given a gameG ↾푣 0 ∈ C, a player푖 ∈Π, and a threshold 푡 ∈ Q, is every 휌 -equilibrium ̄ 휎 inG ↾푣 0 such that 휇 푖 ⟨ ̄ 휎⟩> 푡 ? 8.3. LINK WITH THE CONSTRAINED EXISTENCE PROBLEM AND COMPLEXITIES117 Corollary 2. LetC be a game class among parity games, mean-payoff games, discounted-sum games, and energy games. Then, in the classC, for a given rationality concept휌 ∈ Nash, SubgamePerfect, the휌- universal threshold problem, the휌-rational verification problem, and the deterministic휌-rational verification problem are reducible to each other in polynomial time. Proof.▶The deterministic휌-rational verification problem reduces to the휌-rational verification problem. This result is true because a deterministic memory structure is a memory structure. ▶The휌-universal threshold problem reduces to the deterministic휌-rational verifica- tion problem. Let the gameG ↾푣 0 , the player푖and the threshold푡form an instance of the휌-universal threshold problem. We define the gameG ′ ↾푣 0 as equal to the gameG ↾푣 0 , where Leader has been added to the player set, but controls no vertex. We define휇 픏 = 휇 푖 . IfGbelongs to the classC, so doesG ′ . LetMbe the one-state deterministic memory structure onG ′ that never outputs anything. Then, a strategy profile ̄ 휎inG ′ ↾푣 0 is a픏-fixed휌-equilibrium, if and only if it is a픏-fixed휌-equilibrium with휎 픏 ∈ Ind ↾푣 0 (M) , if and only if the strategy profile ̄ 휎 −픏 is a휌-equilibrium in the gameG ↾푣 0 . As a consequence the game G ↾푣 0 , the player푖, and the threshold푡form a positive instance of the휌-universal threshold problem, if and only if the gameG ′ ↾푣 0 , the memory structureM, and the threshold푡form a positive instance of the deterministic휌-rational verification problem. Moreover, the latter can be constructed from the former in polynomial time. ▶The 휌 -rational verification problem reduces to the 휌 -universal threshold problem. This result is a consequence of Theorem 20, since the product gameG ↾푣 0 ⊗M can be constructed fromG ↾푣 0 andMin polynomial time, and since for each of the four game classes considered, if the gameG belongs to the classC, so does the gameG ↾푣 0 ⊗M.□ It is now time to note that this problem is the complement of a subproblem of the constrained existence problem of휌-equilibria in the classC, where the threshold vectors ̄ 푥and ̄ 푦are such that푥 푗 =−∞for all푗, and푦 푗 =+∞for all푗except one. This strong connection between the two problems explains that, in the literature, rational verification is sometimes the name given to what we call here the constrained existence problem (with possibly different types of constraints)—for a recent example, see [GNPW23]. Therefore, the complexity of Nash and subgame-perfect rational verification in parity, mean-payoff, discounted-sum or energy games can immediately be infered from Theorems 1 to 4, 6, 11, 12 and 15 to 18. Theorem 21. The Nash rational verification problem is: • coNP-complete in parity games; • coNP-complete in mean-payoff games; • coTDS-hard and recursively enumerable in discounted-sum games; • undecidable and co-recursively enumerable in energy games. 118CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS 푎푐 푏푑 픏 0 ◦ 0 □ 3 픏 0 ◦ 0 □ 3 픏 0 ◦ 2 □ 2 픏 0 ◦ 1 □ 1 Figure 26: The temptation of chaos: an illustration The subgame-perfect rational verification problem is: • coNP-complete and fixed-parameter tractable (with the number of players and the number of colors as parameters) in parity games; • coNP-complete in mean-payoff games; • coTDS-hard and recursively enumerable in discounted-sum games; • undecidable, and in particular not recursively enumerable, in energy games. 8.4A POTENTIAL LIMIT: THE TEMPTATION OF CHAOS It is now worth noting that the definition we gave of rational verification entails, in the case of mean-payoff games, results that may be considered as counter-intuitive. Example 14. Consider the mean-payoff game depicted by Figure 26, where Leader owns no vertex, and consider the only (vacuous) strategy available for Leader. Does that strategy guarantee a payoff greater than 1? Intuitively, it does not, since Leader always receives the payoff 0. But still, that strategy, that game, and that threshold form a positive instance of the subgame-perfect rational verification problem, because no 픏-fixed SPE exists in that game (see the proof of Theorem 13). More generally, the definition we give of rational verification considers that a good strategy for Leader is a strategy such that for every response of the environment that is rational, the generated outcome observes some specification. But a strategy is then good, in that sense, if no rational response of the environment exists: that is the phenomenon that we can call temptation of chaos. While that case does never occur in all the other settings we have been studying, because rational responses are always guaranteed to exist (as we will see below), it must be considered specifically for subgame-perfect rational verification in mean-payoff games. 8.5ACHAOTIC RATIONAL VERIFICATION 8.5.1 Definitions To avoid such phenomena, we generalize the approach used in a two-player context in [FGR20], where best responses, when they are not guaranteed to exist, are replaced by휀-best responses—i.e. best responses up to 휀, where 휀 ≥ 0 is as small as possible. We therefore introduce an alternative definition of rational verification, achaotic rational verification: a good strategy for Leader is a strategy that guarantees the given threshold against every response that is as rational as possible. A quantitative relaxation of our rationality concepts is therefore necessary: one 8.5. ACHAOTIC RATIONAL VERIFICATION119 is offered by the notions of휀-NEs and휀-SPEs, already defined in Chapter 4. For convenience, we write 휀SPE 픏 G ↾푣 0 for the set of픏-fixed휀-SPEs in the gameG ↾푣 0 . Later, we will also write휀SPEG ↾푣 0 for the set of 휀-SPEs in the same game. Finally, given a strategy휎 픏 in the gameG ↾푣 0 , we denote by휀SPR(휎 픏 )the set of 휀-best responses to휎 픏 , i.e., the set of strategy profiles ̄ 휎 −픏 such that the strategy profile ̄ 휎is a픏-fixed휀-SPE. Problem 5 (Achaotic (deterministic) subgame-perfect rational verification in the classC). Given a game G ↾푣 0 ∈ C, a threshold푡 ∈ Q, and a memory structure (resp. a deterministic memory structure)MonG, do we have: inf 휎 픏 ∈Ind ↾푣 0 (M) sup 휀≥0 SPR(휎 픏 )≠∅ inf ̄ 휎 −픏 ∈휀SPR(휎 픏 ) 휇 픏 ⟨ ̄ 휎⟩> 푡 ? In contrast with classical rational verification, the supremum and the infimum above are taken on sets that are never empty: if one considers휀large enough (for example휀 = max 푖∈Π max 푢푣,푢 ′ 푣 ′ ∈퐸 |푟 푖 (푢푣)− 푟 푖 (푢 ′ 푣 ′ )|in a mean-payoff game), there always exists an휀-SPE in every game with a bounded set of payoff vectors, which is the case of all the game classes we considered so far. The definition of achaotic (deterministic) Nash rational verification is analogous. Example 15. Let us consider again the example depicted by Figure 26. In that game, there exist (픏-fixed) 휀-SPEs for every휀 ≥1 (for example, the stationary strategy profile ̄ 휎defined by휎(푎) =푐and휎(푏) =푑), but not for휀<1 (for the same reasons as there is no SPEs). Achaotic subgame-perfect rational verification amounts then to wondering whether all픏-fixed 1-SPEs ̄ 휎satisfy the inequality휇 픏 ⟨ ̄ 휎⟩> 푡, which is the case if and only if we have 푡< 0. In the general case, a proof analogous to that of Theorem 20 and Corollary 2 can show that for each 휌 ∈ Nash, SubgamePerfectand every game classCamong parity, mean-payoff, discounted-sum and energy games, the achaotic (deterministic or not)휌-rational verification problem reduces in polynomial time to the following problem, and that a reverse reduction in polynomial time also exists. Problem 6 (Achaotic휌-universal threshold problem in the classC). Given a gameG ↾푣 0 ∈ C , a player 푖 ∈Π, and a threshold 푡 ∈ Q, do we have: sup 휀≥0 휀SPEG ↾푣 0 ≠∅ inf ̄ 휎∈휀SPEG ↾푣 0 휇 픏 ⟨ ̄ 휎⟩> 푡 ? 8.5.2 Coincidence with classical rational verification Among the problems we study here, this new definition is relevant in only one case: subgame-perfect rational verification in mean-payoff games. In all other cases, the rational verification problems are equivalent to their achaotic versions, because Nash and subgame-perfect responses are guaranteed to exist. Theorem 22. LetCbe a class of games, among the classes of parity, energy games and discounted-sum games. Let휌 ∈ Nash, SubgamePerfect. Then, the positive instances of the achaotic휌-universal threshold problem inC are exactly the positive instances of the 휌-universal threshold problem inC. 120CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS Proof.▶Parity, energy, and discounted-sum games The result is a consequence of the fact that in those three classes, SPEs are always guaranteed to exist—and therefore so are NEs. Indeed, it has been shown in [GU08] that every game with Borel Boolean winning condition contains an SPE. This covers parity, but also energy games. Thus, ifCis the class of parity or of energy games, then the gameG ↾푣 0 ⊗M ∈ C contains an SPE. IfCis the class of discounted-sum games, then the gameG ↾푣 0 , and therefore the gameG ↾푣 0 ⊗M, is a game with payoff functions that are continuous for the canonical distance on infinite words. By [BBMR15], every such game contains an SPE. ▶Nash-universal threshold problem in mean-payoff games In this second case, SPEs may not exist in mean-payoff games, but NEs still always do—see for example [BDS13].□ 8.5.3 The complexity of achaotic rational verification in mean-payoff games We now focus on the complexity of the achaotic subgame-perfect-universal threshold problem in mean- payoff games, knowing that it will also give us the complexity of the achaotic subgame-perfect rational verification problem. An optimal algorithm for this problem requires the following lemma: in each game, there exists a least 휀such that휀-SPEs exist. Moreover, that quantity can be written with a polynomially bounded number of bits. Lemma 23. There exists a polynomial푃 1 such that in every mean-payoff gameG ↾푣 0 , there exists휀 min with ∥휀 min ∥ ≤ 푃 1 (∥G∥) such that 휀 min -SPEs exist inG ↾푣 0 , and 휀-SPEs, for every 휀< 휀 min , do not. Proof. This proof will rely on some new technical considerations on the negotiation function in mean- payoff games. ▶A preliminary result We will need to prove that the negotiation function is continuous from below. Let us first recall a definition of that notion. Definition 37 (Continuity from below). A mapping푓:R 퐷 ! R 퐷 is continuous from below if for every non-decreasing sequence( ̄ 푥 푛 ), we have 푓(sup 푛 ̄ 푥 푛 ) = sup 푛 푓( ̄ 푥 푛 ). Sublemma 3. In mean-payoff games, the negotiation function is continuous from below Proof. Let(휆 푛 ) 푛 be a non-decreasing sequence of requirements on a mean-payoff gameG, and let휆 = sup 푛 휆 푛 . We want to prove the equalitynego(휆) = sup 푛 nego(휆 푛 ). Since the negotiation function is monotone, we already havenego(휆) ≥ sup 푛 nego(휆 푛 ). Let us prove thatnego(휆) ≤ sup 푛 nego(휆 푛 ). Let훿>0: we want to find푛such thatnego(휆 푛 )(푣) ≥ nego(휆)(푣) − 훿for each 푣 ∈ 푉 . 8.5. ACHAOTIC RATIONAL VERIFICATION121 We will use for that purpose, for a given player푖and a given vertex푣 ∈ 푉 푖 , the concrete negotiation games: conc 휆푖 (G) ↾푣 c = 픓,ℭ,푉 c ,(푉 c 픓 ,푉 c ℭ ),퐸 c , 휇 c ↾푣 c and: conc 휆 푛 푖 (G) ↾푣 c = 픓,ℭ,푉 c ,(푉 c 픓 ,푉 c ℭ ),퐸 c , 휇 푛 ↾푣 c for some requirement휆 푛 . Let us note that both games have the same underlying graph, and that the only difference is between the payoff functions 휇 c and 휇 푛 . Let us define: 훾 푛 = sup 푤∈푉 (휆(푤)− 휆 푛 (푤)). Then, the sequence(훾 푛 ) 푛 is non-increasing and converges to 0. Let휏 ℭ be a stationary strategy for Challenger in the gameconc 휆푖 (G) ↾푠 0 ; it can also be considered as a stationary strategy in the gameconc 휆 푛 푖 (G) ↾푠 0 . Now, let퐾be a strongly connected component of the underlying graph of the gameconc 휆푖 (G)[휏 ℭ ], accessible from the vertex푣 c 0 . Against the strategy휏 ℭ , Prover may go to the strongly connected component퐾, and give to Challenger the payoff that he gets in any play in that strongly connected component. If the strongly connected component퐾contains some deviation.Then, the least payoff that Challenger can obtain in퐾is the same when the payoff function is휇 c and when it is휇 푛 : it corresponds to the least payoff of the form 휇 푖 (¤푐 휔 ) where 푐 is a simple cycle of 퐾 . If the strongly connected component퐾contains no deviation.Then, the vertices of퐾 share a common memory푀. Let푅 = ̄ 푥 ∈ R Π |∀푖 ∈Π,∀푣 ∈ 푀∩푉 푖 ,푥 푖 ≥ 휆(푣), and similarly, let 푅 푛 = ̄ 푥 ∈ R Π |∀푖 ∈Π,∀푣 ∈ 푀∩푉 푖 ,푥 푖 ≥ 휆 푛 (푣). If we have(Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ))∩ 푅 =∅, since the setsConv 푐∈SCyc(퐾) 휇(¤푐 휔 )and푅are closed, then by taking the quantity훾 푛 small enough, we also have(Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ))∩ 푅 푛 =∅. In this case, whatever Prover does in the strongly connected component퐾, Challenger will always get the payoff+∞. If that is not the case, the polytope(Conv 푐∈SCyc(퐾) 휇(¤푐 휔 )) ∩ 푅is included in the polytope (Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ))∩ 푅 푛 , and the difference between the infima of those polytopes’ projections on the dimension푖can be bounded by taking the biggest slope of those polytopes’ edges, which are in finite number. Formally, we have: inf 푥 푖 ̄ 푥 ∈ Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ) ∩ 푅 푛 ≥ inf 푥 푖 ̄ 푥 ∈ Conv 푐∈SCyc(퐾) 휇(¤푐 휔 ) ∩ 푅 −훾 푛 max 푐,푑∈SCyc(퐾) ∑︁ 푗∈Π, 휇 푗 (¤푐 휔 )>휇 푗 ( ¤ 푑 휔 ) 휇 푖 (¤푐 휔 )− 휇 푖 ( ¤ 푑 휔 ) 휇 푗 (¤푐 휔 )− 휇 푗 ( ¤ 푑 휔 ) and if훾 푛 is small enough, the right-hand subtrahend can be made smaller than 훿 . Conclusion.In both cases, we find that there exists훾 푛 small enough, i.e.푛large enough, to ensure that the least payoff that Prover can give to Challenger in퐾varies by less than the quantity 훿 when the requirement varies between 휆 and 휆 푛 . 122CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS We can find such푛for each strongly connected component퐾, and there exists a finite number of such components: we can therefore find some푛that satisfies the same property for all퐾. Similarly, we can find some푛that also ensures it for all퐾and for all stationary strategies휏 ℭ . Since by Lemma 14 stationary strategies are optimal in the concrete negotiation game, we infer that there exists 푛 ∈ N such that: nego(휆 푛 )(푣) ≥ nego(휆)(푣)−훿. Finally, since there are finitely many states푣 ∈ 푉, we can conclude to the existence of푛 ∈ N such that for each 푣 ∈ 푉 , we have: nego(휆 푛 )(푣) ≥ nego(휆)(푣)−훿, which concludes the proof.□ ▶Existence of 휀 min Let us recall that, by Lemma 16, for every휀 ≥0, the negotiation function has a least휀-fixed point, which we will write휆 휀 . Let us also recall that by Theorem 8, a play in the gameG ↾푣 0 is an휀-SPE outcome if and only if it is 휆 휀 -consistent. Let us define: 휀 = inf훿 ≥ 0| 훿 SPEG ↾푣 0 ≠∅. We only need to show that휀is also such that휀-SPEs exist, or in other words, that휆 휀 (푣 0 ) ≠+∞. Let us note that for every훿,훿 ′ with훿 ≤ 훿 ′ , the requirement휆 훿 is itself a훿 ′ -fixed point of the negotiation function, and therefore satisfies 휆 훿 ≥ 휆 훿 ′ . Let us now define the requirement 휆 : 푣 7! sup 훿>휀 휆 훿 (푣). For each 푣 and every 훿> 휀, we have: nego(휆 훿 )(푣) ≤ 휆 훿 (푣)−훿. Using Sublemma 3, we obtain: nego(휆)(푣) ≤ 휆(푣)− 휀, i.e., the requirement휆is an휀-fixed point of the negotiation function, and therefore휆 = 휆 휀 . Since 휆(푣 0 ) = sup 훿 휆 훿 (푣 0 ) ≠+∞, we obtain that 휀-SPEs exist inG ↾푣 0 , and therefore that 휀 min = 휀 exists. ▶Size Now, let us recall that we showed in the proof of Lemma 16 that one can define, for every휆, a finite union of polyhedra푋 휆 ⊆ R 푉×Π , where each of those polyhedra is defined by a set of inequations either of the form푥 푣푖 ≥ 휆(푤)or of size bounded by a polynomial function of∥G∥, and such that for each 푖 ∈Π and 푣 ∈ 푉 푖 , we have nego(휆)(푣) = min 푥 푣푖 | ̄ ̄ 푥 ∈ 푋 휆 . Let us now define the set: 푌 = ( 휆, ̄ ̄ 푥 ) | ̄ ̄ 푥 ∈ 푋 휆 ⊆ R 푉 × R 푉×Π . 8.5. ACHAOTIC RATIONAL VERIFICATION123 This set is itself a finite union of polyhedra defined by the same inequations than푋 휆 , or by inequations of bounded size (those of the form푥 푣푖 ≥ 휆(푤), where휆(푤)is no longer a constant but a coordinate of ( 휆, ̄ ̄ 푥 ) ). Let us also define the mapping: 푓 : 푌 !R ( 휆, ̄ ̄ 푥 ) 7!max 푖∈Π,푣∈푉 푖 ( 푥 푣푖 − 휆(푣) ) . Then, we have: 휀 min = min ( 휆, ̄ ̄ 푥 ) 푓 ( 휆, ̄ ̄ 푥 ) , and since푓is a piecewise linear mapping on a finite union of polyhedra that has a minimum, that minimum is reached on a vertex ( 휆, ̄ ̄ 푥 ) of one of those polyhedra. By Corollary 1, such a vertex has coordinates bounded by a polynomial of the maximal size of the inequations defining the polyhedron to which it belongs, i.e., by a polynomial of∥G∥. Therefore, that is also the case of푓(휆, ̄ ̄ 푥) = 휀 min .□ We are now equipped to prove the following theorem. Theorem 23. In the class of mean-payoff games, the achaotic subgame-perfect-universal threshold problem, deterministic or not, is P NP -complete. Proof.▶Easiness By Theorem 15 (taking ̄ 푥 = (−∞) 푖∈Π and ̄ 푦 = (+∞) 푖∈Π ), there is anNPalgorithm that decides, given휀andG ↾푣 0 , whether there exists an휀-SPE inG ↾푣 0 , i.e. whether휀 ≥ 휀 min . Using Lemma 23, a dichotomous search can therefore compute휀 min by a polynomial number of calls to that algorithm. Then, one last call to that same algorithm can decide whether there exists an휀 min -SPE ̄ 휎such that 휇 푖 ⟨ ̄ 휎⟩ ≤ 푡 . ▶Hardness We proceed by reducing, to the complement of our problem, the following P NP -complete problem. Problem 7 (Lexicographic minimum problem). Given a Boolean formula휑in conjunctive normal form over the ordered variables푥 1 , . . .,푥 푛 , is the lexicographically first valuation휈 min satisfying휑 such that 휈 min (푥 푛 ) = 1? (and in particular, does such a valuation exist?) Construction.Let us write휑 = Ó 푝 푗=1 퐶 푗 . We construct a gameG ↾푎 , with a player called Witness and written픚, in which there exists an휀 min -SPE ̄ 휎such that휇 픚 ⟨ ̄ 휎⟩ ≤0 if and only if휑is satisfiable and is such that휈 min (푥 푛 ) =1. That game, depicted in Figure 27, has 2푛+푝+4 players: the literal players 푥 1 ,¬푥 1 , . . .,푥 푛 ,¬푥 푛 ; the clause players퐶 1 , . . .,퐶 푝 ; the player Solver, written픖; the player Witness, written픚; the player Adélaïde, written프; and the player Barthélémy, written픅. It contains 3푛+푝+4 vertices: •the initial vertex 푣 0 = 푎, controlled by Adélaïde; •two vertices 푏 and 푐, controlled by Barthélémy; •for each variable 푥 푖 , a vertex ?푥 푖 ∈ 푉 픖 , a vertex 푥 푖 ∈ 푉 푥 푖 , and a vertex¬푥 푖 ∈ 푉 ¬푥 푖 ; 124CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS ?푥 1 픖 . . . . . . ?푥 푖 픖 푥 푖 푥 푖 ¬푥 푖 ¬푥 푖 . . . 푥 푛 푥 푛 ¬푥 푛 ¬푥 푛 퐶 1 퐶 1 . . . 퐶 푝 퐶 푝 푎 프 푏 픅 푐 ▽ ▽ ▽ 프 0 픅 3 픚 1 프 0 픅 3 픚 1 프 2 픅 2 픚 1 푥 푖 2푚 퐶 푚 프 2− 푚 2 푖+1 ¬푥 푖 2푚 퐶 푚 프 2 푥 푖 2푚 퐶 푚 프 2− 푚 2 푖+1 ¬푥 푖 2푚 퐶 푚 프 2 푥 푛 2푚 퐶 푚 프 2− 푚 2 푛+1 ¬푥 푛 2푚 퐶 푚 프 2 픚 1 프 2 프 2 프 2 프 1 푥 1 ,¬푥 1 ,...,푥 푛 ,¬푥 푛 4 퐶 1 ,...,퐶 푝 2 픚 1 Figure 27: The gameG ↾푎 •for each clause퐶 푗 , a vertex퐶 푗 ∈ 푉 퐶 푗 ; •a sink vertex ▽ (drawn three times in Figure 27 for convenience). Those vertices are connected by the following edges (unmentioned rewards are equal to 0, and we write푚 = 2푛+ 푝): • from the vertex푎to the vertex푏and from the vertex푏to the vertex푎, two edges that give Adélaïde the reward 0, Barthélémy the reward 3, and Witness the reward 1; •from the vertex 푎 to the vertex ?푥 1 and from 푏 to 푐, an edge; •from the vertex푐to itself, an edge giving both Adélaïde and Barthélémy the reward 2, and giving Witness the reward 1; •from each vertex ?푥 푖 to the vertex¬푥 푖 and from each vertex¬푥 푖 to the vertex ?푥 푖+1 (or to the vertex퐶 1 if 푖 = 푛), an edge giving: –the reward 2푚 to player¬푥 푖 , –the reward푚 to every player퐶 푗 such that the clause퐶 푗 contains the literal¬푥 푖 , –the reward 2 to Adélaïde; –and if 푖 = 푛, the reward 1 to Witness; • from each vertex ?푥 푖 to the vertex푥 푖 and from each vertex푥 푖 to the vertex ?푥 푖+1 (or to the vertex 퐶 1 if 푖 = 푛), an edge giving: –the reward 2푚 to player 푥 푖 , –the reward푚 to every player퐶 푗 such that the clause퐶 푗 contains the literal 푥 푖 , –and the reward 2− 푚 2 푖+1 to Adélaïde; 8.5. ACHAOTIC RATIONAL VERIFICATION125 •from each vertex퐶 푗 to the vertex퐶 푗+1 (or the vertex ?푥 1 if푗 = 푝), an edge giving the reward 2 to Adélaïde; •from the sink vertex▽to itself, an edge giving the reward 1 to Adélaïde, the reward 2 to each clause player, the reward 4 to each literal player, and the reward 1 to Witness. We will show that in this game, when휈 min exists, it defines a binary encoding for the quantity휀 min . To do so, let us present a more general correspondence between 휀-SPEs and valuations satisfying 휑 . The strategy profile ̄ 휎 휈 . Let휈be a valuation satisfying휑. We define the stationary strategy profile ̄ 휎 휈 as follows: from the vertex푎, Adélaïde always goes to?푥 1 ; from the vertex푏, Barthélémy always goes to푐; from each vertex?푥 푖 , Solver always goes to푥 푖 if휈(푥 푖 ) =1, and to¬푥 푖 otherwise; from their vertices, the clause players and the literal players that are satisfied by휈do never go to the vertex ▽; and the literal players that are not satisfied by 휈 do whenever they have the opportunity. Sublemma 4. The strategy profile ̄ 휎 휈 is an휀-SPE, where휀 = Í 푛 푖=1 휈(푥 푖 ) 2 푖 , and is not a훿-SPE for any훿< 휀. Proof.Solver’s payoff is constant, and Witness does not control any vertex, hence they have no profitable deviation. In every subgame where they have actions to choose, all the clause players get at least the payoff 2, since at least one of their literals is satisfied (and therefore the corresponding vertex is visited at each turn). Similarly, from the vertex he controls, each literal player gets the payoff 4, either because he is satisfied by휈, or by going to the vertex▽; they have therefore no profitable deviation. As for Barthélémy, in every subgame where he has an action to choose, he gets the payoff 2, and cannot get a better one, since Adélaïde always plans to go to ?푥 1 . Let us finally focus on Adélaïde: from the vertex푎, the only one that she controls, she could get the payoff 2 by going to 푏. By going to the vertex ?푥 1 , she gets the payoff: 1 2푛+ 푝 푛 ∑︁ 푖=1 2 2−(2푛+ 푝) 휈(푥 푖 ) 2 푖+1 + 2푝 ! = 2− 푛 ∑︁ 푖=1 휈(푥 푖 ) 2 푖 = 2− 휀, hence ̄ 휎 휈 is an 휀-SPE, and is not a 훿 -SPE for any 훿< 휀.□ The valuation휈 ̄ 휎 .Let ̄ 휎 be a 1-SPE such that the play휋 =⟨ ̄ 휎⟩traverses the vertex ?푥 1 , but does never reach the vertex▽. We define the valuation휈 ̄ 휎 by휈 ̄ 휎 (푥 푖 ) =1 if and only if휋traverses the vertex 푥 푖 . Let us note that for every vertex푥 푖 that is traversed by휋, the player푥 푖 can deviate and go to the vertex▽, where he can get the payoff 4. Since ̄ 휎is a 1-SPE, we have therefore휇 푥 푖 (휋) ≥3, and therefore 휇 ¬푥 푖 (휋) ≤1. As a consequence, the vertices푥 푖 and¬푥 푖 cannot both be traversed; and if the vertex¬푥 푖 is traversed, then we have 휈 ̄ 휎 (푥 푖 ) = 0. Sublemma 5. The valuation 휈 ̄ 휎 satisfies 휑, and we have: 휇 프 (휋) = 2− 푛 ∑︁ 푖=1 휈 ̄ 휎 (푥 푖 ) 2 푖 . 126CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS Proof.Let us consider a clause퐶 푗 of휑. Since the vertex퐶 푗 is visited (infinitely often), there is a deviation available for player퐶 푗 that grants her the payoff 2. Since ̄ 휎 is a 1-SPE, we have휇 퐶 푗 (휋) ≥1. Therefore, there is at least one literalℓof퐶 푗 such that the vertexℓis visited by휋, i.e. such that휈 ̄ 휎 satisfies ℓ . Consequently, the valuation 휈 ̄ 휎 satisfies퐶 푗 , and therefore 휑 . Moreover, since the play휋does never traverse both푥 푖 and¬푥 푖 for any푖, it grants Adélaïde the payoff: 1 2푛+ 푝 푛 ∑︁ 푖=1 2 2−(2푛+ 푝) 휈 ̄ 휎 (푥 푖 ) 2 푖+1 + 2푝 ! = 2− 푛 ∑︁ 푖=1 휈 ̄ 휎 (푥 푖 ) 2 푖 . □ Consequence. We can then prove the following proposition. Proposition 1. if휈 min is the least valuation satisfying휑(and in particular if such a valuation exists), then we have 휀 min = Í 푛 푖=1 휈 min (푥 푖 ) 2 푖 . Proof.Let 훿 = Í 푛 푖=1 휈 min (푥 푖 ) 2 푖 < 1. By Sublemma 4, the strategy profile ̄ 휎 휈 min is a훿-SPE. Let us prove that there is no휀-SPE with 휀< 훿 . Let 휀< 1 be such that there exists an 휀-SPE ̄ 휎 inG ↾푎 , and let us prove that 휀 ≥ 훿 . If ̄ 휎 is an휀-SPE with휀<1, then necessarily, there exist infinitely many integers푘such that 휎 프 ((푎푏) 푘 푎) = ?푥 1 , because otherwise each play⟨ ̄ 휎 ↾(푎푏) 푘 푎 ⟩would either end in the vertex푐(and then Barthélémy would have a deviation profitable by more than 1 by refusing to go to푐), or would be equal to(푎푏) 휔 (and then Adélaïde would have a deviation profitable by more than 1 by going to ?푥 1 ). Then, for each such푘 ≠0, we have휎 픅 ((푎푏) 푘 ) = 푐(if Barthélémy does not go to푐, then he gets only the payoff 1, while by going to푐, he gets the payoff 2). Therefore, if we choose some푘 such that휎 프 ((푎푏) 푘 푎) = ?푥 1 , we also have휇 프 ⟨ ̄ 휎 ↾(푎푏) 푘 푎푏 ⟩ =2. Since ̄ 휎is an휀-SPE with휀<1, we have therefore휇 프 ⟨ ̄ 휎 ↾(푎푏) 푘 푎 ⟩>1, which means that the play⟨ ̄ 휎 ↾(푎푏) 푘 푎 ⟩does never reach the vertex ▽. Then, the strategy profile ̄ 휎 ↾(푎푏) 푘 푎 is a 1-SPE such that the play⟨ ̄ 휎 ↾(푎푏) 푘 푎 ⟩traverses the vertex ?푥 1 , but does never reach the vertex▽: the valuation휈 ̄ 휎 ↾(푎푏) 푘 푎 is defined, and by Sublemma 5, it satisfies the formula 휑 , and is such that: 휇 프 ⟨ ̄ 휎 ↾(푎푏) 푘 푎 ⟩ = 2− 푛 ∑︁ 푖=1 휈 ̄ 휎 ↾(푎푏) 푘 푎 (푥 푖 ) 2 푖 . By going to the vertex푏instead of ?푥 1 , Adélaïde has therefore a deviation that is profitable by Í 푛 푖=1 휈 ̄ 휎 ↾(푎푏) 푘 푎 (푥 푖 ) 2 푖 . Moreover, since the valuation휈 ̄ 휎 ↾(푎푏) 푘 푎 satisfies휑, it is lexicographically greater than or equal to 휈 min , hence the inequality: 푛 ∑︁ 푖=1 휈 ̄ 휎 ↾(푎푏) 푘 푎 (푥 푖 ) 2 푖 ≥ 푛 ∑︁ 푖=1 휈 min (푥 푖 ) 2 푖 = 훿, and therefore 휀 ≥ 훿 , as desired.□ 8.6. RATIONAL SYNTHESIS127 푎 푏 푐 ◦ 1 ◦ 1 ◦ 2 ◦ 0 ◦ 0 ◦ 2 Figure 28: The temptation of chaos in a discounted-sum game Conclusion. We can now conclude that there exists an휀 min -SPE ̄ 휎in this game such that 휇 픚 ⟨ ̄ 휎⟩ ≤ 0 if and only if 휑 is satisfiable and is such that 휈 min (푥 푛 ) = 1. • If there exists such an휀 min -SPE, then necessarily the play⟨ ̄ 휎⟩visits infinitely often the vertex푥 푛 (every other play gives Witness a positive payoff). Therefore, the valuation휈 min = 휈 ̄ 휎 is such that 휈 min (푥 푛 ) = 1. •Conversely, if휈 min (푥 푛 ) =1, then the휀 min -SPE ̄ 휎 휈 min traverses the vertex ?푥 1 and does never traverse the vertex¬푥 푛 , nor reach the vertex ▽; it grants therefore Witness the payoff 0. □ 8.6RATIONAL SYNTHESIS Let us close this chapter with some considerations on the natural continuation of rational verification, where the strategy휎 픏 is not given: rational synthesis, a concept due to Giuseppe Perelli, Orna Kupferman, and Moshe Vardi [KPV16]. Such a problem can be defined as follows, for a given game classCand rationality concept 휌 . Problem 8 (휌-rational synthesis problem in the classC). Given a gameG ↾푣 0 ∈ C and a threshold푡 ∈ Q, does there exist a strategy휎 픏 such that every픏-fixed휌-equilibrium ̄ 휏with휏 픏 = 휎 픏 satisfies the inequality 휇 픏 ⟨ ̄ 휏⟩> 푡 ? Another common presentation of this problem in the literature (used, for example, in [FGR20]) is in terms of the computation of the Stackelberg value. Given a gameG |푣 0 , the Stackelberg value can be defined as the supremum of all thresholds푡for which the pair(G |푣 0 ,푡)is a positive instance of the decision problem described above. Variants of this problem can be obtained by restricting Leader’s strategies to finite-memory ones (which makes sense if Leader’s strategy is an abstraction for a program), or by considering its achaotic version. Both variants might be relevant given the fact that if휎 픏 can use infinite memory, then the temptation of chaos can now be observed in all classes of quantitative games where payoff functions can have infinite range. The game depicted by Figure 28, where violet vertices belong to Leader, illustrates that fact. If Leader is allowed to use infinite memory, then she can design her strategy so that player ◦ has an incentive to eventually go to the vertex푏, but as late as possible. That will be the case for example with the strategy 휎 픏 defined by: ⟨ ̄ 휎 ↾푎 푘 푏 ⟩ = 푏 푘 푐 휔 , for every푘 ∈ Nand independently of휎 ◦ . Then, every strategy of player ◦ admits a profitable deviation, by postponing the moment where she takes the edge푎푏. By contrast, let us recall that we showed in the proof of Theorem 22 that in discounted-sum games, there is always a subgame-perfect, and a fortiori a Nash response to every finite-memory strategy. 128CHAPTER 8. ABOUT RATIONAL VERIFICATION AND ITS LIMITS This document does not bring any significant result about rational verification. Let us simply mention that the case of parity games, Nash rational synthesis has been proved to beEXPTIME-easy andPSPACE- hard, andPSPACE-easy,NP-hard, andcoNP-hard when the number of players is fixed [CFGR16]. As for subgame-perfect rational synthesis, it has recently been proved to beNP-hard andcoNP-hard, 2EXPTIME- easy in general, and EXPTIME-easy when the number of players is fixed [BRRvdB24]. It is also worth noting here that we have the following general result: Lemma 24. For each휌 ∈ SubgamePerfect, Nashand every game classCamong parity games, mean- payoff games, discounted-sum games and energy games, the휌-universal threshold problem in the classC reduces in polynomial time to the 휌-rational synthesis problem inC. The reduction consists simply in adding a player Leader who does not control any vertex. Consequently, using Theorem 21, Nash and subgame-perfect rational synthesis are at least as hard as the target discounted- sum problem in discounted-sum games, and undecidable in energy games. Finally, to our knowledge, little is known about rational synthesis in mean-payoff games with arbitrary many players in the environment (the case where the environment is made of one player has been studied in [FGR20]). When Leader is restricted to finite-memory strategies, Theorems 2 and 15 entail recursive enumerability. Similarly, one could obtain an analogous result for achaotic rational synthesis by proving that finite-memory strategies are optimal for Leader, which seems to be a reasonable conjecture (while it is clearly not true for classical rational verification, as examplified by the game depicted by Figure 28). But all those problems might still be undecidable. Open Problem 4. Are (achaotic) Nash and subgame-perfect rational synthesis decidable in mean-payoff games? Are they decidable when Leader is restricted to finite-memory strategies? PART IV: OTHER EQUILIBRIA 129 CHAPTER 9:STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS In the previous sections, our examples have primarily dealt with systems interacting with an unpredictable or partially predictable environment. As we have argued, the most classical game-theoretic framework for modeling such situations is that of two-player zero-sum games. However, when the assumption of full adversity is too restrictive, multiplayer settings can be more appropriate, incorporating relevant notions of collective rationality, typically captured by equilibria concepts such as Nash equilibria or subgame-perfect equilibria. In this context, the environment—potentially consisting of multiple agents—may be better represented as acting rationally according to its own objectives. Here, we focus on another class of computer-science-related problems that can be modeled using multiplayer games: protocol design, where a set of agents interact according to a predefined protocol that must remain robust against deviations. We argue that in such cases, a particularly relevant notion is that of strong secure equilibria. That notion combines the well-established concepts of strong equilibria (resilient to deviations by coalitions) and secure equilibria (where no player can deviate in a way that harms another player without also harming themself ). However, to the best of our knowledge, it has not been studied as a standalone concept before this work. 9.1MOTIVATING EXAMPLE AND DEFINITION The aim of designing a protocol is to establish a set of behavior rules that yield a specific outcome if thoroughly followed, and that are self-enforceable: every participating agent must adhere to these rules, lest they do not get anything out of the interactions. Before moving on the more formal definitions and results of the paper, let us run through the obstacles and possible solutions to the design of a good fair exchange protocol. Towards this, let us consider the simple setting of only two agents, Adélaïde and Barthélémy, that wish to exchange items online. (For instance, Adélaïde could want to purchase digital art from Barthélémy.) A first naive approach to designing such a protocol could be “let both send their part to the other one”: see Figure 29 for a depiction of the possible sequences of actions following this basic guideline. If both of them do this, even asynchronously, then the exchange is performed. Of course, there could be a malicious third party involved here, that could derail the execution of the protocol for any reason—an external attacker. Another source of trouble could be the failure of the communication channels. Here, we assume that the channels are resilient: they deliver every message perfectly, albeit maybe with an arbitrary but finite delay (every message is eventually delivered). However, another critical aspect to take into account is what assumptions we make about the motiva- tions of the agents. Indeed, one can argue that there is no reason for an agent to trust the other one. In 131 132CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS 프 : item 1 픅 : item 2 프 : item 1 , item 2 픅 : ∅ 프 : ∅ 픅 : item 1 , item 2 프 : item 2 픅 : item 1 프 sends item 1 픅 sends item 2 프 sends item 1 픅 sends item 2 Figure 29: Exchange between Adélaïde and Barthélémy only. fact, both may even prefer an outcome where they receive something from the other as intended, but do not send their share. On the other hand, the agents wish in priority to avoid being wronged, i.e. sending their share without receiving anything. Depending on such assumptions, the solutions for designing a fair protocol may vary: clearly, the naive approach sketched above fails as soon as one agent is not fully trustworthy. Let us briefly describe two approaches to ensure robustness against these possibly untrustworthy agents. 9.1.1 With a trusted third party A first possible solution to ensure fairness while making no assumption on the trustworthiness of the participants is to involve a trusted third party (TTP) in the protocol. This particular agent is assumed to be trustworthy, and to have no other interest than the completion of the exchange. Furthermore, it is supposed to have sufficient leverage on the agents to enforce the completion of the exchange, if one of them tries to deflect from the intended course of actions. The TTP is thus considered to be an agent with absolute authority. One could consider performing the exchange only via the TTP, sending both agents’ parts to the TTP that would, upon full reception, redistribute them appropriately. This, however foolproof, is extremely costly, as the TTP is used in every instance of the protocol. Usually, this is avoided by amending the naive protocol to add the rule “if an agent does not receive anything while they have sent their share, they contact and alert the TTP, who then makes sure the other agent complies”. This is what is called an optimistic version of a protocol involving a TTP. It is assumed that the participating agents desire the intended outcome of the protocol, and thus that in most cases, the TTP will not need to intervene at all. A depiction of a such a TTP moderated protocol exchange can be found in Figure 30. This is also the approach of the Zhou-Gollmann protocol [ZG97] that uses a TTP only in case of a (suspected or real) cheating attempt (and is encompassed by our framework). 9.1.2 Beyond trust: without a trusted third party The TTP approach is well-known and widely accepted as a necessity (as well-justified by the impossibility result of [PG99]). What can be done when this approach is not applicable? Indeed, there are many cases where the trustworthiness of a third party is not clear enough to reasonably rely upon it. For those cases, can we design protocols which rules are self-enforceable and that eliminate the need of a TTP? In other terms, what if the incentives to complete the protocol properly were stronger than the ones to deviate 9.1. MOTIVATING EXAMPLE AND DEFINITION133 프 : item 1 픅 : item 2 프 : item 1 , item 2 픅 : ∅ 프 : ∅ 픅 : item 1 , item 2 프 : item 2 픅 : item 1 프 : item 1 , item 2 픅 : ∅ 픗 alerted 프 : ∅ 픅 : item 1 , item 2 픗 alerted 프 sends item 1 픅 sends item 2 프 sends item 1 픅 sends item 2 프 alerts 픗 픅 alerts 픗 픗forces프to send item 1 픗forces픅to send item 2 Figure 30: Exchange between Adélaïde and Barthélémy with a trusted third party. from the intended behaviors? An alternative, and the approach we choose in this work, is to rely, not on the trustworthiness of agents, but instead on their rationality. While not amenable to every possible context where fair exchange is needed, it can provide a viable solution in relevant situations. We will therefore consider that a protocol is safe if no alliance of agents can wrong another agent without wronging a member of the alliance itself. The security of the protocol stems from the assumption that all agents, even though they might be malicious, would nonetheless behave rationally. For instance, in peer-to-peer settings, no one can be absolutely trusted in theory, but all agents have an incentive to be honest, since keeping a good reputation enable them to perform further exchanges. As another example, a ride-share application can be used as a trusted third party between a driver and a traveler without being trusted per se, but because both the driver and the traveler know that if the application had dishonest behaviors in some situations, then no one would trust it anymore, and the company that holds it would stop their profits. In both these examples, important common features are their multi-agent nature and the repeatability of the exchange, albeit with different agents. To illustrate further and start going towards a more abstract model, consider the following setting, depicted by Figure 31. Adélaïde, Barthélémy and Capucine are three agents who wish to exchange an infinity of items: each agent has infinitely many items to send to both other agents. Again, as shown in [PG99], there is no protocol that could satisfy the usual definitions of fairness. Indeed, there is an agent (say, for instance, Adélaïde) that can scam another agent (say Barthélémy), i.e. that can send only finitely many items to Barthélémy and receive at least one more item from him. That is the case if the agents exchange their items directly, but also if they use the third agent (Capucine) as an intermediary, because Capucine is not a TTP and can therefore choose to help Adélaïde, forming a (malicious) coalition, to scam Barthélémy (or the reverse). However, since the process is repeating for further exchanges, if Capucine does so, Barthélémy may react by stopping exchanges not only with Adélaïde, but also with Capucine. Adélaïde would then be satisfied (she actually scammed Barthélémy), but Capucine would be worse off (she did not scam anyone herself, and can no longer exchange with Barthélémy). So if Capucine behaves rationally, she has no incentive to help another agent scam the third. In practice, therefore, such a protocol can be considered as reasonably safe: Adélaïde, Barthélémy, and Capucine all take, alternatively, the role of the TTP and enable an exchange of items between the two other agents; and if an agent deviates from the protocol, the other 134CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS 프픅 ℭ ( 프 ! 픅 푛 ) 푛 ∈ N ( 픅 ! 프 푛 ) 푛 ∈ N ( 프 ! ℭ 푛 ) 푛 ∈ N ( ℭ ! 프 푛 ) 푛 ∈ N ( 픅 ! ℭ 푛 ) 푛 ∈ N ( ℭ ! 픅 푛 ) 푛 ∈ N Figure 31: Expected exchanged items between Adélaïde, Barthélémy, and Capucine. ones stop all exchanges with them. 9.1.3 Strong secure equilibria At this point, we hope the reader can be convinced that a relevant equilibrium notion for capturing protocols that ensure a reasonable level of safety is the following. Definition 38 (Strong secure equilibrium). LetG ↾푣 0 be a game. Let ̄ 휎be a strategy profile inG ↾푣 0 : a harmful deviation of the coalition퐶 ⊆Πfrom the strategy profile ̄ 휎is a strategy profile ̄ 휎 ′ 퐶 ∈ Strat 퐶 G ↾푣 0 such that for every player푖 ∈ 퐶, we have휇 푖 ⟨ ̄ 휎 −퐶 , ̄ 휎 ′ 퐶 ⟩ ≥ 휇 푖 ⟨ ̄ 휎⟩ , and for some player푖 ∉ 퐶, we have 휇 푖 ⟨ ̄ 휎 −퐶 , ̄ 휎 ′ 퐶 ⟩< 휇 푖 ⟨ ̄ 휎⟩. A strong secure equilibrium inG ↾푣 0 , or SSE for short, is a strategy profile from which there is no harmful deviation. Before moving to computational aspects, let us make two remarks to give an insight on how expressive a model based on SSEs is. Indeed, one may believe that SSEs provide a very specific type of potential attackers—those that may participate in an attack as long as they are not wronged in the process. But such a model can actually capture more classical types of agents, such as fully hostile ones, or on the contrary trusted ones. •A fully untrusted agent, on which no rationality assumptions can be made, can be modeled as an agent that is impervious to harm—i.e., one that always wins. •Conversely, in a protocol that includes an agent who is fully trusted and will never participate in an attack, such an agent can be modeled as a player who is wronged as soon as any other player is wronged. As a result, if the protocol contains an attack, captured by the harmful deviation of a coalition, the trusted agent cannot be part of that coalition. We provide a more detailed discussion in [BRS + 24] regarding the protocol and attack models that can be captured by this notion, as well as its relevance to classical protocols such as the Zhou-Gollmann optimistic protocol [ZG97]. Since these security aspects diverge significantly from the game-theoretic focus of this document, we choose to concentrate here on algorithmic aspects. 9.1. MOTIVATING EXAMPLE AND DEFINITION135 9.1.4 Problem Modeling secure protocols with strong secure equilibria is particularly relevant in Boolean games, where each player wins if and only if they are not wronged by other players. In this setting, determining whether a given strategy profile (where every player wins) is an SSE reduces to checking whether there exists a coalition that can deviate and wrong some player without harming any of its own members. Since such wronging conditions are typically represented by휔-regular conditions (e.g., "at some point, Adélaïde sends the itemitem 1 , but Barthélémy never sends the itemitem 2 "), we restrict our study to parity games, as defined in Chapter 3, and, more generally, to 휔-regular games, a broader class defined as follows. Definition 39 (Parity automaton). A parity automaton over alphabet푋is a tupleA = (푋,푄,푞 0 ,훿,휅) where푄is a set of states, where훿:푄×푋 ! 푄is a deterministic transition function, and where휅:푄 ! N is a color mapping. Given a word푤 =푤 0 푤 1 · ∈ 푋 휔 , the run ofAon푤is휌 푤 =푞 0 푞 1 ·where for푖 ∈ N, 푞 푖+1 = 훿(푞 푖 ,푤 푖 ). The set of infinite words accepted byA is L(A) = 푤 ∈ 푋 휔 | min휅(Inf(휌 푤 )) is even . Definition 40 (휔-regular game). An휔-regular game is a multi-player gameG ↾푣 0 such that there exists, for each player푖, a parity automatonA 푖 over alphabet푉, such that for each play휋, we have휇 푖 (휋) =1 if and only if 휋 ∈ L(A 푖 ). Moreover, since the primary objective in this framework is to design a protocol in which no player is wronged, we do not focus on the constrained existence problem of SSEs. Instead, we study a variant of this problem, the fixed-payoff SSE problem, where we seek a strategy profile that not only produces a given payoff vector but also satisfies a specified correctness condition. This condition is of the same type as the players’ objectives: a color function on the vertices of the game in the case of a parity game, or a parity automaton in the case of an휔-regular game. In both cases, we denote by휋 ⊨ 휑the fact that the play휋 satisfies the condition휑. It is worth noting that, when considering Nash equilibrium or subgame-perfect equilibrium, such a correctness condition can be encoded as the objective of a fictional player. However, this approach is not applicable here, as it may introduce harmful deviations against that player. Problem 9 (Fixed-payoff SSE problem). Given a gameG ↾푣 0 , a payoff vector ̄ 푥 ∈ Q Π , and a condition휑, does there exist an SSE ̄ 휎 inG ↾푣 0 such that 휇⟨ ̄ 휎⟩ = ̄ 푥 and⟨ ̄ 휎⟩ ⊨ 휑 ? Algorithms for this problem can also be used to solve the constrained existence problem, at least in Boolean games (or more generally games with payoff functions ranging in a finite set), by checking each possible payoff vector ̄ 푥between the two threshold vectors (with휑 =⊤). Of course, in the general case, the number of such calls grows exponentially with the number of players. However, since the lowest complexity we will obtain in this chapter isPSPACE, this additional cost will not be prohibitive, and our complexity results for the fixed-payoff SSE problem will immediately entail results for the constrained existence problem of SSEs. The remainder of this chapter is therefore dedicated to establishing tight complexity bounds for the fixed-payoff SSE problem in parity games and휔-regular games. In particular, we will pay close attention to the complexities that arise when the number of players and the number of colors are small, as this is often the case in practical applications. 136CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS 9.2A TOOL: THE DEVIATOR GAME The main tool we use to solve our problem takes the form of a new game structure, in which one player, Prover, tries to prove that an SSE generating the desired payoff exists, while another one, Challenger, tries to prove that the strategy profile she constructs is actually not an SSE. That game is the deviator game, very similar to a construction with the same name proposed in [Bre16], in a slightly different context. Definition 41 (Deviator game). LetG ↾푣 0 be a game, let ̄ 푥be a payoff vector, and let휑a correctness condition. The deviator game is the initialized game: dev ̄ 푥휑 (G) ↾푣 d 0 = 픓,ℭ,푉 d ,퐸 d , 푉 d 픓 ,푉 d ℭ , 휇 d ↾푣 d 0 , where: • the player 픓 is called Prover, and the player ℭ Challenger. • Prover controls the set푉 d 픓 =푉 × 2 Π , and Challenger the set푉 d ℭ = 퐸× 2 Π . • The initial vertex is 푣 d 0 =(푣 0 ,∅). • The edge set퐸 d is defined as follows: from the vertex(푢,퐷), Prover can go to every vertex of the form(푢푣,퐷) ∈ 퐸(she proposes the edge푢푣). From the vertex(푢푣,퐷), Challenger can go to the vertex (푣,퐷)(he accepts the edge푢푣), or to every vertex(푤,퐷∪푖), with푤 ≠ 푣and푢푤 ∈ 퐸, and where푖 is the player controlling푢(he deviates from Prover’s proposal, and player푖is added to the set of deviators). • Given a play 휒 is this game, we write ¤휒 for the play inG constructed by the actions of Prover and Challenger: if휒 = (푢 0 ,퐷 0 )(푢 0 푣 0 ,퐷 0 )(푢 1 ,퐷 1 )(푢 1 푣 1 ,퐷 1 ) . . ., then we define¤휒 = 푢 0 푢 1 . . .. We also define D(휒) = Ð 푘 퐷 푘 , the set of players who deviated along the play 휒 . •Then, the Boolean payoff function휇 d is defined as follows. A play휒is won by Challenger if and only if either: – we have D(휒) =∅ and 휇(¤휒) ≠ ̄ 푥 ; – we have D(휒) =∅ and ¤휒 ⊭ 휑 ; – or for every player푖 ∈D(휒), we have휇 푖 (¤휒) ≥ 푥 푖 , and there exists a player푗 ∈Πsuch that 휇 푗 (¤휒)< 푥 푗 . Theorem 24. Prover has a winning strategy in the gamedev ̄ 푥휑 (G) ↾푣 d 0 if and only if there exists an SSE ̄ 휎in G ↾푣 0 with 휇⟨ ̄ 휎⟩ = ̄ 푥 and⟨ ̄ 휎⟩ ⊨ 휑. A proof of this theorem can easily be obtained by adapting the one presented in [Bre16]. 9.3PARITY GAMES Let us now consider the case of parity games: then, similarly to what was done in [Bre16], a cautious way to solve the deviator game leads to an algorithm that uses only polynomial space. For convenience, we now writeParity(휅)for the set of plays satisfying the parity condition associated to each color function휅, andParity(휅)for its complement. Similarly, we write B(푣)for the set of plays that visit infinitely often the vertex 푣 , andB(푣) for its complement. The condition 휑 is defined by the color mapping 휅 휑 . 9.3. PARITY GAMES137 Theorem 25. In parity games, the fixed-payoff SSE problem isPSPACE-complete. Hardness still holds in co-Büchi games and with correctness condition휑 =⊤. With two players, the problem is at least as hard as solving a two-player zero-sum parity game. If both the number of colors and the number of players are fixed, then the problem is fixed-parameter tractable. Proof.Given a parity gameG ↾푣 0 , we write푛for the number of players and푘for the highest color that appears in the automataA 푖 andA 휑 . ▶PSPACE-easiness By Theorem 24, deciding the fixed-payoff SSE existence problem amounts to deciding which player, among Prover and Challenger, has a winning strategy in the corresponding deviator game. That game has an exponential size, but has a specific shape that makes it possible to solve it re- gion by region, without using exponential space. Indeed, let us observe that along a play휒 = (푢 0 ,퐷 0 )(푢 0 푣 0 ,퐷 0 )(푢 1 ,퐷 1 )(푢 1 푣 1 ,퐷 1 ) . . . , the sequence of sets(퐷 ℓ ) ℓ∈N is non-decreasing. Let us therefore recursively define an algorithm that decides, given the gameG ↾푣 0 , the vector ̄ 푥 and a vertex(푢 0 ,퐷)of the gamedev ̄ 푥휑 G , whether Prover has a winning strategy from that vertex. Let 푊 =푖 ∈Π | 푥 푖 = 1 be the set of players that win according to the payoff vector ̄ 푥 . Definition of the gameH 퐷 ↾(푢 0 ,퐷) . First, construct the gameH 퐷 ↾(푢 0 ,퐷) as follows: construct the region of the gamedev ̄ 푥휑 Gthat is made of vertices of the form(푢,퐷)or(푢푣,퐷). Define in that region the color mappings so that player푖’s payoff, in that region, is always equal to their payoff in the corresponding play inG: i.e., for all푖 ∈Π, define휅 ′ 푖 (푢푣,퐷) =휅 ′ 푖 (푢,퐷) =휅 푖 (푢) (and similarly for the correctness condition, define 휅 ′ 휑 (푢푣,퐷) =휅 ′ 휑 (푢,퐷) =휅 푖 (푢)). For each edge(푢푣,퐷)(푤,퐷∪푖)that leaves that region, add to the constructed game the edge (푢푣,퐷)(푤,퐷∪푖)and the loop(푤,퐷∪푖)(푤,퐷∪푖), and recursively decide whether Prover has a winning strategy from the vertex(푤,퐷∪푖). If she does, call the vertex(푤,퐷∪푖)a Prover’s leaf, and define the color mappings so that only players in푊would win on that loop: for each푖 ′ ∈ 푊, define휅 ′ 푖 ′ (푤,퐷∪푖) = 0; for푖 ′ ∉푊, define휅 ′ 푖 ′ (푤,퐷∪푖) = 1; and finally, define휅 ′ 휑 (푤,퐷∪푖) =0. If she does not, call the vertex(푤,퐷∪푖)a Challenger’s leaf, and define the color mappings so that only players in퐷∪푖would win on that loop: for each푖 ′ ∈ 퐷∪푖, define휅 ′ 푖 ′ (푤,퐷∪푖) = 0; for 푖 ∉ 퐷∪푖, define 휅 ′ 푖 ′ (푤,퐷∪푖) = 1; and finally, define 휅 ′ 휑 (푤,퐷∪푖) = 1. Now, define the payoff functions as follows. If퐷 ≠∅, then a play휋is won by Challenger if and only if we have: 휋 ∈ Ø 푖∈푊 Parity(휅 ′ 푖 ) ! ∩ Ù 푖∈퐷∩푊 Parity(휅 ′ 푖 ) ! (i.e., if the play휋gives to at least one player푖a payoff worse than푥 푖 , and to every player who deviated at least the same payoff as in ̄ 푥 ). If 퐷 =∅, then a play 휋 is won by Challenger if and only if we have: 휋 ∈ Ø 푖∈푊 Parity(휅 ′ 푖 )∪ Ø 푖∉푊 Parity(휅 ′ 푖 )∪Parity(휅 ′ 휑 ) (i.e., if the payoff vector generated by Prover, without any deviation, is not equal to ̄ 푥, or if the correctness constraint is not met). Note that in both cases, Challenger wins if he reaches a Challenger’s leaf, and loses if he reaches a Prover’s leaf. 138CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS Equivalence with the gamedev ̄ 푥휑 G. Prover has a winning strategy in the gameH 퐷 ↾(푢 0 ,퐷) if and only if she has a winning strategy from the vertex(푢 0 ,퐷)in the gamedev ̄ 푥휑 G. Indeed, if Prover has a winning strategy inH 퐷 ↾(푢 0 ,퐷) , then she can follow that strategy indev ̄ 푥휑 G. Then, she either stays in the region where the deviating coalition is퐷, or she reaches a vertex of the form(푤,퐷∪푖)from which she has a winning strategy, that she can then follow. Conversely, if she has a winning strategy from(푢 0 ,퐷)indev ̄ 푥휑 G, she can follow it inH 퐷 ↾(푢 0 ,퐷) : she will then either stay in the region where the deviating coalition is퐷, or reach a Prover’s leaf (since the strategy she is following afterwards in dev ̄ 푥휑 G is still winning), which makes her win. Resolution.This game can be seen as an Emerson-Lei game [AL87]: Prover’s and Challenger’s winning conditions are Boolean combinations of Büchi conditions. Indeed, we have the equality: Parity(휅 ′ 푖 ) = Ø 2푚<푘 © « Ø 휅 ′ 푖 (푣 d )=2푚 B(푣 d )∩¬ Ø ℓ<2푚 Ø 휅 ′ 푖 (푤 d )=ℓ B(푤 d ) ª ® ¬ . It was shown in [HD05] that such games can be solved using a space polynomial in the size of the game and of the objectives. Here, both are themselves polynomial in the size ofG. Thus, the last step of our recursive algorithm consists in solvingH 퐷 ↾(푢 0 ,퐷) . Then, applying this recursive algorithm from the vertex(푣 0 ,∅)decides our problem: there is an SSE ̄ 휎inG ↾푣 0 such that휇⟨ ̄ 휎⟩ = ̄ 푥and⟨ ̄ 휎⟩ ⊨ 휑if and only if Prover has a winning strategy from that vertex. Moreover, that algorithm uses polynomial space: each recursive call does, and the recursion stack has size at mostcardΠ, i.e. polynomial. Therefore, the fixed-payoff SSE existence problem in parity games is PSPACE-easy. ▶Fixed-parameter tractability In [BHR18b], it has been shown that an Emerson-Lei gameGwith vertex space푆and objective휓 can be solved in time: 푂 2 2 |휓| |휓|+ 2 |휓|2 |휓| card푉 d 5 . In our case, the size of the objective휓depends only on푘and푛, whilecard푉 d =2 푛 card푉: our problem is therefore fixed-parameter tractable with 푘 and 푛 as parameters. ▶PSPACE-hardness (in co-Büchi games) We reduce the quantified satisfiability problem (QSat) to the SSE fixed-payoff existence problem (with휑 =⊤) in co-Büchi games (or more precisely to its complement). This will prove that our problem cannot be fixed-parameter tractable with 푘 as a single parameter, unless we have P = PSPACE. Reduction. Let: 휑 =∃푥 1 ∀푥 2 ·∃푥 푛−1 ∀푥 푛 푝 Û 푖=1 푚 푖 Ü 푗=1 ℓ 푖푗 9.3. PARITY GAMES139 be an instance ofQSatin conjunctive normal form where for all푖, 푗, we defineℓ 푖푗 as a literal of the form푥 푘 or¬푥 푘 for some푘 ∈ 1, . . .,푛. We write휓 ′ (푥 1 , . . .,푥 푛 ) = Ó 푝 푖=1 Ô 푚 푖 푗=1 ℓ 푖푗 for the non-quantified part of the formula. We construct a game with 2푛+2 players: one for each literal player and two additional players Solver (written픖) and Opponent (written픒) as depicted in Figure 32. Round vertices belong to Solver, hexagonal vertices belong to Opponent, rectangular vertices belong to the literal indicated in the bottom right corner (if any: some vertices only have a single outgoing edge and therefore could belong to anyone without loss of generality). Red clouds indicate the list of players for which a given vertex is in the co-Büchi set: infinitely many visits to these vertices means losing (payoff 0) while only finitely many visits means winning (payoff 1). The game is composed of three parts: in the first part, the setting module, a valuation can be chosen by visiting the vertex푥 s 푖 (setting푥 푖 to true) or the vertex¬푥 s 푖 (setting푥 푖 to false), the choice belonging to Solver if푥 푖 is existentially quantified and to Opponent if it is universally quantified. When the value of a variable is chosen, the corresponding literal player can choose to go to a sink vertex▽(duplicated in the figure for clarity) or continue the game. If all choose to continue, then the game moves into the second part, the checking module, where the formula is “checked”: each disjunctive clause퐶 푖 has a vertex owned by Solver who must choose a literalℓ 푖푗 from the clause. Vertexℓ 푖푗 is in the co-Büchi set of player ̄ ℓ 푖푗 , who loses if said vertex was visited infinitely often. When this “checking” is done, the last part, called punishing module, visits the co-Büchi set of one literal per variable, the choice of which being left to Solver. The game continues back to the checking module. As the last vertex of these modules is in the co-Büchi set of Opponent, Opponent will lose if the game enters these modules. Remark that there are no co-Büchi vertices for Solver who therefore always wins. We will now show that the formula휑holds if and only if there is no SSE with payoff vector(1) 푖∈Π . The intuition is that a coalition made of Solver and all true literals can deviate to make Opponent and all false literals lose while not penalizing the players in the coalition. If the formula휓does not hold, then there exists an SSE generating the payoff vector (1) 푖∈Π .Let us assume휑does not hold. Let ̄ 휎be the following profile of strategies: each literal goes to ▽whenever possible in the setting module, Solver plays any strategy, and Opponent plays the optimal strategy in response to Solver’s choice of valuation so that휓 ′ (푥 1 , . . .,푥 푛 )is false. The outcome of this profile is either path ?푥 1 · 푥 s 1 · ▽ 휔 or ?푥 1 ·¬푥 s 1 · ▽ 휔 , both with payoff vector(1) 푖∈Π . Let us prove that it is an SSE. Let퐶be a coalition of players and assume that there exists a harmful deviation for this coalition. As Solver always wins, she can be assumed to be in the coalition. And since no one loses in the setting module or▽, any harmful deviation must reach the checking module. Because that means that Opponent would then lose, Opponent cannot be in the coalition. For each variable푥 푖 , players푥 푖 and¬푥 푖 cannot be both in the coalition as (at least) one of them loses through the infinite visits in the punishing module. To actually reach the checking module, all literal players whose vertex was visited in the setting module must have deviated from ̄ 휎 , and therefore are in the coalition. As a result the coalition is made of Solver and exactly one literal per variable, defining a valuation. Because Opponent played optimally, that valuation does not satisfy휓 ′ (푥 1 , . . .,푥 푛 ), so there is a clause퐶 푖 where all literalsℓ 푖1 , . . .,ℓ 푖푚 푖 are false, meaning that the player corresponding to their negations are all in the coalition. The choice by Solver of any of these literals therefore enforces a visit 140CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS ?푥 1 푥 s 1 푥 1 ¬푥 s 1 ¬푥 1 ?푥 2 푥 s 2 푥 2 ¬푥 s 2 ¬푥 2 · 푥 s 푛 푥 푛 ¬푥 s 푛 ¬푥 푛 퐶 1 ℓ 1,1 ¬ℓ 1,1 ℓ 1,푚 1 ¬ℓ 1,푚 1 . . . 퐶 2 . . . . . . . . . . . . . . . . . . · · · 푥 1 ? 푥 p 1 ¬푥 1 ¬푥 p 1 푥 1 푥 2 ? 푥 p 2 ¬푥 2 ¬푥 p 2 푥 2 · 푥 p 푛 ¬푥 푛 ¬푥 p 푛 푥 푛 픒 ▽ ▽ Figure 32: Game to encodeQSatinstance∃푥 1 ∀푥 2 ·∃푥 푛−1 ∀푥 푛 Ó 푝 푖=1 Ô 푚 푖 푗=1 ℓ 푖푗 into the existence of an SSE with payoff(1, . . .,1,1,1). Round vertices belong to Solver, hexagonal vertices belong to Opponent, rectangular vertices belong to the literal indicated in the bottom right corner (if any). The red cloud lists the set of players for which this vertex is in the co-Büchi condition. to the co-Büchi set of (at least) one player in the coalition. As this choice is infinitely repeated, that makes this player lose and therefore it should not be part of the coalition, which is a contradiction. So ̄ 휎 is an SSE. If the formula휓holds, then there is no such SSE. Let us assume that휑holds. Let ̄ 휎be a strategy profile with payoff(1) 푖∈Π . As noted above, that payoff requires the outcome of ̄ 휎to end in the vertex▽. Consider the following strategy for Solver: in the setting module, choose the literal that ensures satisfaction of the formula휓 ′ (푥 1 , . . .,푥 푛 ), based on the already known choices of Opponent. This is possible since the formula휑holds. Let퐶be the coalition made of Solver and all literal players corresponding to vertices푥 s 푖 or¬푥 s 푖 visited in the setting module. These players deviate by reaching the vertex ?푥 푖+1 (or퐶 1 if푖 = 푛) instead of going to▽. Since the checking module is reached, Opponent (at least) will lose and be harmed by the deviation. The valuation thus built ensures that휓 ′ (푥 1 , . . .,푥 푛 ) is satisfied, so in the checking module, for every clause there is a true literal. The strategy of Solver consists in consistently choosing the vertex for these literals. Note that this means visiting a co-Büchi set for a player that is not in the coalition. In the punishing module, Solver visits vertices corresponding to players in the coalition (i.e. the exact same one that were visited in the setting module: if푥 s 푖 was visited, consistently choose푥 p 푖 , if¬푥 s 푖 was visited, consistently choose¬푥 p 푖 . This ensures that the co-Büchi set of no player in the coalition퐶is ever visited, hence all players in this coalition still win, hence the deviation is harmful and ̄ 휎 is not an SSE. 9.4. IN 휔-REGULAR GAMES141 ▶ Parity-hardness (with two players) We show here that when푛is fixed to 2 (and when휑 =⊤), the fixed-payoff SSE problem is at least as hard as solving a two-player zero-sum parity game. This will show that our problem cannot be fixed-parameter tractable with only푛as a parameter, unless parity games can be solved in polynomial time. Let us recall that this problem is known to belong to NP∩ coNP. LetGbe a parity gameG ↾푣 0 with two players, Adélaïde (written프) and Barthélémy (written픅), and such that휅 픅 =휅 프 +1 (i.e., Barthélémy’s winning condition is the complement of Adélaïde’s one). Let us now consider the gameG ′ ↾푣 0 , with the same arena, and with color mappings휅 ′ 프 =휅 프 and휅 픅 constantly equal to 2. Let us assume Adélaïde has a winning strategy휎 프 in the gameG ↾푣 0 : then, for any strategy휎 픅 , the strategy profile ̄ 휎is, in the gameG ′ ↾푣 0 , an SSE where both Adélaïde and Barthélémy get the payoff 1. Indeed, Barthélémy gets the payoff 1 as he always does, Adélaïde gets the payoff 1 because휎 프 wins against휎 픅 , and that strategy profile is an SSE because no player can deviate to harm the other one: Barthélémy wins every play, and Adélaïde is playing a winning strategy. Conversely, if the strategy profile ̄ 휎is an SSE inG ′ ↾푣 0 in which both Adélaïde and Barthélémy get the payoff 1, then consider some alternative strategy휎 ′ 픅 for Barthélémy: like every play, the play ⟨휎 프 ,휎 ′ 픅 ⟩is won by Barthélémy inG ′ . Consequently, it is also won by Adélaïde, otherwise it would constitute a harmful deviation for Barthélémy. It is therefore also won by Adélaïde inG. This proves that 휎 프 is a winning strategy inG ↾푣 0 . As a consequence, the problem of solving a two-player zero-sum parity game reduces to the SSE fixed-payoff existence problem in parity games with only two players; and therefore, the SSE fixed-payoff existence problem in parity games cannot be solved in polynomial time if only푛is fixed (unless parity games are solvable in P).□ 9.4IN 휔 -REGULAR GAMES When the payoff functions are defined by parity automata, those ones must be incorporated in the arena of the game, entailing an exponential blowup. However, there is no exponential blowup in the formula that defines the winning conditions of Prover and Challenger, hence the game can be solved by using algorithms that exists in the literature, in exponential time. Theorem 26. In휔-regular games, the fixed-payoff SSE problem isEXPTIME-complete. Hardness holds even if 휑 =⊤ and if all automata are Büchi or co-Büchi automata. When both the numbers of players and colors are fixed, the problem becomes P-easy. Proof.▶EXPTIME-easiness, and P-easiness with fixed numbers of colors and players. LetG ↾푣 0 be an휔-regular game, and let ̄ 푥 ∈ Q Π . Let푊 = 푖 ∈Π | 푥 푖 =1be the set of players winning according to payoff vector ̄ 푥. Let ( A 푖 ) 푖∈Π∪휑 be the parity automata defining the payoff of player푖and correctness constraint overG ↾푣 0 . Let푄 푖 (resp.훿 푖 ,휅 푖 ,푞 0푖 ) be the set of states (resp. transition function, coloring mapping, initial state) ofA 푖 . We assume all these automata use at most푘 142CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS colors, i.e. the co-domain of every color function 휅 푖 is included in0, . . .,푘. Let: 푚 = max 2,|푉|, max 푖∈Π∪휑 | 푄 푖 | be the maximal size of the graph and automata provided as input (assumed to be at least 2 for complexity purposes). We build the extended deviator game as the product of the deviator game with all automataA 푖 . Formally, the extended deviator game is the initialized game: edev ̄ 푥 (G) ↾푣 e 0 = 픓,ℭ,푉 e ,퐸 e , 푉 e 픓 ,푉 e ℭ , 휇 e ↾푣 e 0 , where: •the players are Prover and Challenger. •Prover controls the set푉 e 픓 and Challenger the set푉 e ℭ , where: 푉 e 픓 =푉 × 2 Π × Ö 푖∈Π∪휑 푄 푖 and푉 e ℭ = 퐸× 2 Π × Ö 푖∈Π∪휑 푄 푖 In the sequel, we write ̄ 푞 =(푞 푖 ) 푖∈Π∪휑 for elements of Î 푖∈Π∪휑 푄 푖 . •The initial vertex is 푣 e 0 =(푣 0 ,∅,(푞 0푖 ) 푖∈Π∪휑 ) =(푣 0 ,∅, ̄ 푞 0 ). • The edge set퐸 e is defined as follows: from the vertex(푢,퐷, ̄ 푞), Prover can go to every vertex of the form(푢푣,퐷, ̄ 푞)with푢푣 ∈ 퐸(she proposes the edge푢푣). From the vertex(푢푣,퐷, ̄ 푞), Challenger can go to the vertex(푣,퐷, ̄ 푞 ′ )where푞 ′ 푖 = 훿 푖 (푞 푖 ,푣)for each푖(he accepts the edge푢푣), or to every vertex(푤,퐷 ∪푗, ̄ 푞 ′ ), with푤 ≠ 푣,푢푤 ∈ 퐸, and푞 ′ 푖 = 훿 푖 (푞 푖 ,푤)for each푖, and where푗is the player controlling푢 (he deviates from Prover’s proposal). •We define the coloring functions 휅 ′ 푖 for 푖 ∈Π as: 휅 ′ 푖 (푢,퐷, ̄ 푞) =휅 푖 (푞 푖 ) and: 휅 ′ 푖 (푢푣,퐷, ̄ 푞) = 푘+ 1 (thus effectively ignoring proposal vertices). • Given a play휒 = (푢 0 ,퐷 0 , ̄ 푞 0 )(푢 0 푣 0 ,퐷 0 , ̄ 푞 0 )(푢 1 ,퐷 1 , ̄ 푞 1 )(푢 1 푣 1 ,퐷 1 , ̄ 푞 1 ) . . .in this game, we write D(휒) = Ð 푘 퐷 푘 , the set of players who deviated along the play 휒 . •Then, the Boolean payoff function휇 e is defined as follows. A play휒is won by Challenger if and only if either: –we have D(휒) ≠∅, and: 휒 ∈ Ø 푖∈푊 Parity(휅 ′ 푖 ) ! ∩ © « Ù 푖∈퐷(휒)∩푊 Parity(휅 ′ 푖 ) ª ® ¬ (i.e., the play휒gives to at least one player푖a payoff lower than푥 푖 , and to every player who deviated at least the same payoff as in ̄ 푥 ); 9.4. IN 휔-REGULAR GAMES143 –or we have D(휒) =∅, and: 휒 ∈ Ø 푖∈푊 Parity(휅 ′ 푖 ) ! ∪ Ø 푖∉푊 Parity(휅 ′ 푖 ) ! ∪ Parity(휅 ′ 휑 ) (i.e. the play휒gives to at least one player푖a payoff different from푥 푖 or the correction constraint is not satisfied). Note that we can convert all parity conditions into a chain of Rabin conditions: for푖 ∈Π∪휑and 푐 ∈ 0, . . .,푘, let: Ξ 푖 푐 = (푢,퐷, ̄ 푞) | 휅 푖 (푞 푖 ) ≤ 푐 . ThenParity(휅 ′ 푖 ) = Ð 0≤푐≤푘 푐 is even B(Ξ 푖 푐 )∩B(Ξ 푖 푐−1 ) is the set of runs whereΞ 푖 푐 appears infinitely often and Ξ 푖 푐−1 only finitely often for an even color푐. Remark that here we can (and do) ignore the color푘+1 added to the game. We add extra sets: Δ ∅ = (푢,∅, ̄ 푞) | 푢 ∈ 푉, ̄ 푞 ∈ Ö 푖∈Π∪휑 푄 푖 andΔ 푖 = (푢,퐷, ̄ 푞) | 푢 ∈ 푉,푖 ∈ 퐷, ̄ 푞 ∈ Ö 푖∈Π∪휑 푄 푖 that tracks whether no one (resp. player푖) deviated. The winning condition for Challenger therefore becomes: © « B(Δ ∅ )∩ © « Ø 푖∈푊 0≤푐≤푘 푐 is odd B(Ξ 푖 푐 )∩B(Ξ 푖 푐−1 ) ª ® ® ® ® ¬ ∩ © « Ù 푖∩푊 B(Δ 푖 )∪ Ø 0≤푐≤푘 푐 is even B(Ξ 푖 푐 )∩ B(Ξ 푖 푐−1 ) ª ® ® ¬ ª ® ® ® ® ¬ ∪ © « B(Δ ∅ )∩ © « © « Ø 푖∈푊 0≤푐≤푘 푐 is odd B(Ξ 푖 푐 )∩B(Ξ 푖 푐−1 ) ª ® ® ® ® ¬ ∪ © « Ø 푖∉푊 0≤푐≤푘 푐 is even B(Ξ 푖 푐 )∩B(Ξ 푖 푐−1 ) ª ® ® ® ® ¬ ∪ © « Ø 0≤푐≤푘 푐 is odd B(Ξ 휑 푐 )∩ B(Ξ 휑 푐−1 ) ª ® ® ¬ ª ® ® ¬ ª ® ® ¬ The extended deviator game therefore has 2푚2 푛+1 푚 푛 =푂(푚 푛 )vertices and a winning condition that is provided by an Emerson-Lei condition with(푛+1)푘colors, and a formula of size polynomial in푛and푘. Solving this game with thePSPACEalgorithm used in the proof of Theorem 25 would provide anEXPSPACEalgorithm. However, using the algorithm from [HLP23, Corollary 24], this game can be solved in푂((푛푘)!(푚 푛 ) 푛푘+2 ) =푂((푛푘)!푚 푛 2 푘+2푛 ). As a result the fixed-payoff SSE problem for 휔 -regular games is in EXPTIME. When both 푛 and 푘 are fixed, this complexity boils down to P. ▶EXPTIME-hardness. We first prove the following lemma. Lemma 25. Deciding a two-player zero-sum game where one player’s winning condition is defined as the intersection of the languages of Büchi automata isEXPTIME-hard, and similarly with co-Büchi automata. Proof.We proceed by reduction from the emptiness problem for the intersection of deterministic top-down tree automata. 144CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS Tree automata. We call here tree automata what is called more precisely deterministic top-down tree automata in the literature. Definition 42 (Tree automata). A tree automaton is a tupleT =(푄,(퐹 푘 ) 푘∈0,...,푘 max ,푞 0 ,Δ), with: •a finite set 푄 of states; • a finite family(퐹 푘 ) 푘∈0,...,푘 max , where each set퐹 푘 is a finite set of symbols of arity푘(and the sets 퐹 푘 are pairwise disjoint); •an initial state 푞 0 ; • and a setΔ ⊆ Ð 푘 푄 × 퐹 푘 × 푄 푘 of transitions, such that for each arity푘and each pair (푞, 푓) ∈ 푄× 퐹 푘 , there exists at most one transition(푞, 푓,푞 1 , . . .,푞 푘 ) ∈Δ. A(퐹 푘 ) 푘 -tree, or simply tree when the context is clear, is a tuple푇 =(푉,퐸,⪯,푟,휆) with: •a graph(푉,퐸) where for every two vertices푢 ≠ 푣 ∈ 푉 , there is at most one path from푢 to 푣 ; •a root 푟 ∈ 푉 ; •a total order⪯ on the set푉 ; •a mapping휆:푉 ! Ð 푘 퐹 푘 that labels each vertex푣with a symbol휆(푣)that has aritycard퐸(푣). A tree푇is recognized by the tree automatonTif and only if there exists a mapping푣 7! 푞 푣 such that: •we have 푞 푟 =푞 0 , • and for every푣, we have(푞 푣 ,휆(푣),푞 푤 1 , . . .,푞 푤 푘 ) ∈Δ , where푤 1 ⪯·⪯ 푤 푘 are the elements of the set 퐸(푣). A natural decision problem about tree automata is the following one. Problem 10 (Emptiness problem for the intersection of tree automata). Given푛tree automata T 1 , . . .,T 푛 with the same family(퐹 푘 ) 푘 , is there a tree푇 that is recognized by all of them? That problem is known to be EXPTIME-complete. Sublemma 6 ([CDG + 07]). The emptiness problem for the intersection of tree automata isEXPTIME- complete. Reduction. LetT 1 , . . .,T 푛 be푛tree automata with the same family(퐹 푘 ) 푘 . We construct in polynomial time a Boolean zero-sum gameG ↾푣 0 with two players, Adélaïde and Barthélémy, where Adélaïde’s winning condition is defined as the intersection of the languages of푛Büchi automata (or푛co-Büchi automata), and where Adélaïde has a winning strategy if and only if there is a tree recognized by all tree automata. Let 퐹 =∪ 푘 퐹 푘 . We define the set of Barthélémy’s vertices as: 푉 픅 = 퐹 ∪⊥, and Adélaïde’s vertices as: 푉 프 = (푓,ℓ) 푘 ∈ N, 푓 ∈ 퐹 푘 , ℓ ∈ 1, . . .,푘 ∪⊤. 9.4. IN 휔-REGULAR GAMES145 From each vertex푓where the symbol푓has positive arity, there is an edge to each vertex(푓,ℓ), and from each vertex(푓,ℓ), there is an edge to each vertex푓 ′ ∈ 퐹. From each vertex푓where the symbol푓has arity 0, there is an edge to the vertex⊥. From the vertex⊤, there is an edge to every vertex푓 ∈ 퐹, and the vertex⊤is the initial vertex. Intuitively, Adélaïde and Barthélémy define together a branch of a tree, from the root to a leaf: for each vertex, Adélaïde chooses a label, and Barthélémy chooses the index of the child that will be considered. Let us now construct the automata that will define Adélaïde’s winning condition. For each tree automatonT 푖 , we define an automatonA 푖 . The states of the automatonA 푖 are the same as the tree automatonT 푖 , plus two sink states▽and△. The initial state is the same. As for transitions, when reading a vertex of the form푓or⊤, the automatonA 푖 remains in the same state. When it reads a vertex of the form(푓,ℓ)from state푞, it switches to state푞 ℓ such that there is a transition (푞, 푓,푞 1 , . . .,푞 푘 ) ∈Δ 푖 , if such a transition exists—and to state▽otherwise. When it reads the vertex ⊥, it switches to state△. From the states△and▽, reading any vertex, the automaton remains in the same state. Colors are defined so that the automatonA 푖 accepts a play if and only if the state△is reached. Note that this can be defined with the colors 0 and 1 as well as with the colors 1 and 2, i.e., the automaton can be defined as a Büchi automaton or as a co-Büchi automaton. A play is therefore accepted by the automatonA 푖 if and only if along the branch that Adélaïde and Barthélémy have constructed together, the labeling푣 7! 푞 푣 compatible with the tree automaton T 푖 can be properly defined. Equivalence. Each strategy휎 프 for Adélaïde defines a tree푇, up to renaming vertices; and conversely, each tree푇defines a strategy for Adélaïde, up to ignoring what Adélaïde does in subgames where she has herself deviated. Thus, if there exists a tree푇that is accepted by each tree automatonT 푖 , then consider the corresponding strategy for Adélaïde: against that strategy, Barthélémy chooses a branch of푇. For each automatonA 푖 , the mapping푣 7! 푞 푣 testifying of the acceptance of the tree푇by the tree automatonT 푖 is well defined, and in particular is well defined along the branch chosen by Barthélémy. Hence the play that is generated is accepted by the automatonA 푖 , and is therefore won by Adélaïde. If now there is no such tree, then for every strategy휎 프 for Adélaïde, let us consider the corresponding tree푇. There exists, then, a tree automatonT 푖 that does not recognize the tree푇, and a vertex푣of푇on which the corresponding mapping푣 7! 푞 푣 cannot be defined. If Barthélémy chooses a branch that traverses that vertex, the automatonA 푖 will then switch to the state▽, making Adélaïde lose.□ Now that we know that this problem isEXPTIME-hard, we can prove our result by reducing it to the SSE fixed-payoff problem with휑 =⊤. LetG ↾푣 0 be a two-player zero-sum Boolean game with players Barthélémy and Adélaïde, where Adélaïde’s winning condition is defined as the intersection of the languages of 푛 parity automataA 1 , . . .,A 푛 . Reduction. We build the(푛+2)-player gameG † ↾푣 1 that has an SSE with payoff(1) 푖∈Π if and only if Barthélémy has a winning strategy inG. The푛+2 players areΠ =1, . . .,푛,픅,프. The game arena consists of the original gameG ↾푣 0 with a prepended path where each player from 1 to푛can 146CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS 푣 1 1 푣 2 2 · 푣 푛 푛 G ↾푣 0 ▽ Figure 33: The gameG † ↾푣 1 0 1 A 푖 푞 1 , . . .,푞 푛 ,▽ 푣 0 푣 ∈ 푉 \푣 0 푞 1 , . . .,푞 푛 ,▽ ∗ (a) Parity automataA † 푖 for player 푖 to win 0 1 푞 1 , . . .,푞 푛 ,▽,푣 ∈ 푉 \푣 0 푣 0 ∗ (b) Parity automatonA † 픅 . 1 0 푞 1 , . . .,푞 푛 ,▽,푣 ∈ 푉 \푣 0 푣 0 ∗ (c) Parity automatonA † 프 . Figure 34: Automata associated to the gameG † either choose to continue or reach a sink state▽, as depicted by Figure 33. Player푖 ∈ 1, . . .,푛wins if eitherA 푖 accepts inG ↾푣 0 or the play ends up in▽; otherwise he loses. This is ensured in the automata A † 푖 by adding two states toA 푖 to take into account this prefix (see Figure 34—the numbers indicated in each state is the color of that state). Barthélémy gets payoff 1 if the play ends up in▽and gets 0 if it reachesG ↾푣 0 . Adélaïde always wins in this game. If Barthélémy has a winning strategy inG ↾푣 0 , then there is an SSE inG 푞 1 where all players get payoff 1.Assume that Barthélémy has a winning strategy inG ↾푣 0 . Let ̄ 휎be the following profile inG † : for푖 ∈ 1, . . .,푛, from푞 푖 go to▽; Barthélémy plays his winning strategy inG ↾푣 0 , and Adélaïde plays an arbitrary strategy. Since▽is reached in the play induced by ̄ 휎, the payoff vector is(1) 푖∈Π . We will show that ̄ 휎 is an SSE. Let퐶be a coalition of players and assume ̄ 휎 ′ is a harmful deviation from players in퐶. First, since Adélaïde cannot be harmed by a deviation (she always wins), we can assume that프 ∈ 퐶. In addition, if one of the players푖 ∈ 1, . . .,푛does not effectively deviate, then the play ends up in▽and the payoff vector does not change; therefore all of these players must be in퐶. On the other hand, if the play changes the payoff by not ending in▽, then Barthélémy strictly decreases his payoff, and then cannot be part of the coalition. Therefore, any harmful deviation must be from the coalition퐶 =1, . . .,푛,프. As the deviation reachesG ↾푣 0 , Barthélémy is indeed harmed; for the deviation to be deemed harmful it remains to show that no player in the coalition is themself harmed. However, since Barthélémy has a winning strategy and plays it inG ↾푣 0 , the play is not accepted by (at least) one automatonA 푖 , so the corresponding player푖loses and thus decreases their payoff. As a result, player푖cannot be part of the 9.4. IN 휔-REGULAR GAMES147 coalition, which is a contradiction, and ̄ 휎 is an SSE. If Adélaïde has a winning strategy, then there is no such SSE.Now assume that Adélaïde has a winning strategy inG ↾푣 0 . Let ̄ 휎 be a profile with payoff vector(1) 푖∈Π . We will show this profile is not an SSE. For this payoff to occur, that means the play ends up in▽. Consider the coalition 퐶 =1, . . .,푛,프 and the deviation (which may not be actually effective for some players) ̄ 휎 ′ defined as follows: for푖 ∈ 1, . . .,푛, from푞 푖 go to푞 푖+1 (or푣 0 if푖 = 푛); then, Adélaïde plays her winning strategy. The play thus built will reachG ↾푣 0 where the created play is accepted by allA 푖 , so the payoff for player푖remains 1 while the payoff for Barthélémy strictly decreases from 1 to 0. Therefore ̄ 휎 ′ is a harmful deviation and ̄ 휎 is not an SSE. Conclusion. As a result there is an SSE with payoff(1) 푖∈Π if and only if Barthélémy has a winning strategy inG ↾푣 0 for the winning condition¬ Ñ 푛 푖=1 L(A 푖 ). Since solving a two-player game with such a winning condition isEXPTIME-hard, by Lemma 25, even if all automata are Büchi automata or co-Büchi automata, our hardness result follows.□ 148CHAPTER 9. STRONG SECURE EQUILIBRIA AND THEIR APPLICATIONS CHAPTER 10:RISK-SENSITIVE EQUILIBRIA 10.1INTRODUCTION 10.1.1 About randomness in games It is now time to highlight the fact that the formalism we have considered up to this point implicitly rules out any form of randomness. However, it is common in game theory to consider games where chance plays a role. This can be done in two ways: first, by introducing randomness into the game itself, for example, with stochastic vertices, i.e. vertices not controlled by any player, where the outgoing edge is chosen randomly according to a fixed probability distribution; and second, by allowing players to randomize their strategies, for instance, by tossing a coin when undecided between two edges. In the real world, especially in the context of computer-related systems (quantum computers being set aside), generating true randomness is exceedingly difficult, and the question of whether it can even exist is highly debated. However, randomness in games can be understood more broadly as an abstraction for the absence of information, in line with Andrey Kolmogorov’s conception of probabilities: if player푖 makes a randomized choice between actions푎and푏, each with probability 1 2 , this simply means that the other players have no way of knowing which action player푖will actually choose, or even which one is more likely. It is a well-known fact (a proof can be derived from Martin’s proof of determinacy in Blackwell games [Mar98]) that such randomization is unnecessary in (Boolean) two-player zero-sum games within our formalism: if players are allowed to randomize their strategies, pure (i.e., non-randomized) strategies are just as effective as randomized ones. However, this result no longer holds when the formalism is extended to concurrent games (where players make simultaneous choices) [Eve57]. In this case, a player (say Adélaïde) may have an incentive to randomize her strategy so that her opponent (say Barthélémy) cannot adapt to the action she plans. A canonical example is the rock-paper-scissors game: if Adélaïde plays a given action deterministically (e.g., rock), Barthélémy can easily counter it (in this case, by playing paper). To prevent this, an optimal strategy for Adélaïde is to play each action with probability 1 3 —which might be interpreted as playing a given action, but ensuring Barthélémy has no information about it. Another interesting example arises when analyzing penalty kicks in soccer: the kicker must choose between sending the ball to the left or right side of the goal, while the goalkeeper simultaneously (and ideally quickly) chooses to defend one of those sides. The probability of scoring is much higher when the goalkeeper chooses the wrong side. Data analysis from the German Bundesliga appears to show that players’ choices closely align with those suggested by theory [Sch07]. Similar phenomena arise in turn-based (i.e., non-concurrent) games with imperfect information [RCDH07]. To return to our framework, randomization might also be useful for constructing an equilibrium in a 149 150CHAPTER 10. RISK-SENSITIVE EQUILIBRIA 푎 푏 푐 푑 푒 푓 푔 ◦ 1 □ 1 ⋄ 1 ◦ 2 ⋄ 1 □ 2 Figure 35: A game where randomization might be useful 푎푐 푏푑 ◦ 0 □ 3 ◦ 0 □ 3 ◦ 2 □ 2 ◦ 1 □ 1 Figure 36: A game without SPE multiplayer game. In this case, its purpose is no longer to prevent a player, viewed as an opponent, from using some information, but rather to use randomness to distribute payoffs among several players in such a way that they are satisfied with the expected payoff they receive. Example 16. Consider the mean-payoff game depicted in Figure 35 (with all non-specified rewards being zero), and let us examine whether there exists a Nash equilibrium in which player^receives a payoff of 1. If randomization is not allowed, the answer is clearly no: to give player^a payoff of 1, the play must necessarily reach either vertex푓or vertex푔. In both cases, player ◦ (in the first case) or player□(in the second case) can deviate profitably by moving to vertex푑or푒, respectively. However, if randomization is allowed, player^, starting from vertex푐, can take both edges푐푓and푐푔with probability 1 2 . In this case, the expected payoffs for both player ◦ and player□are equal to 1, and they have no incentive to deviate—assuming that the measure used to define a profitable deviation is the expected payoff. Finally, randomization may also be used to punish infinite deviations, up to relaxing our equilibrium notions. Example 17. Consider the mean-payoff game depicted by Figure 36: we have shown in the proof of Theorem 13 that, when the players are not allowed to play randomized strategies, this game contains no SPE, and even no휀-SPE with휀<1. On the other hand, consider the following randomized stationary strategy profile: from the vertex푎, player ◦ goes to the vertex푐with probability휀>0, and to the vertex푏 with probability 1− 휀. From the vertex푏, player□always goes to the vertex푑. This strategy profile is an 휀-SPE (in every subgame, no player can improve their expected payoff by more than휀), and is defined for every 휀> 0. In line with this last example, it is already known [KFSV21] that randomized휀-SPEs always exist in deterministic mean-payoff games where all rewards are zero, except in self-loops on sink vertices (which corresponds to what we call later in this chapter simple stochastic games, but with no stochastic vertex). We conjecture that this remains true in all mean-payoff games. 10.1. INTRODUCTION151 Conjecture 1. LetG ↾푣 0 be a mean-payoff game (with no stochastic vertices). Then, for every휀>0, there exists a randomized 휀-SPE inG ↾푣 0 . 10.1.2 Constrained existence of randomized NEs (and SPEs) The constrained existence problem in stochastic games with randomized strategies has been studied by Michael Ummels and Dominik Wojtczak, who showed that it is undecidable in all the settings we have considered, as soon as randomness is introduced. More precisely, in [UW11a], they prove the undecidability of the constrained existence problem for randomized Nash equilibria in simple quantitative games, i.e., games with terminal vertices, from which no other vertex is accessible, and where each player’s payoff depends solely on the terminal vertex eventually reached (each player receives a payoff of 0 if no terminal vertex is ever reached). Clearly, simple games can be seen as a subclass of mean-payoff games. Notably, this result does not require the presence of stochastic vertices. On the other hand, in [UW11b], they show that undecidability still holds even when payoffs are restricted to 0 and 1, i.e., in Boolean simple games, provided stochastic vertices are introduced. Moreover, in that setting, undecidability persists even when players are not allowed to randomize their strategies. Boolean simple games can be viewed as a subclass of both energy and parity games (and even Büchi and co-Büchi games). Thus, the work of Michael Ummels and Dominik Wojtczak already covers all the game classes studied in this document, except for discounted-sum games, where decidability remains an open problem even in the absence of randomness (see Theorem 3). Furthermore, their proofs rely on variations of a single reduction from the halting problem of two- counter machines. With careful refinements, the same reduction can also be used to establish the undecid- ability of the constrained existence problem for SPEs in these game classes. This existing body of work thus provides a computational argument for exploring alternative equilib- rium notions when randomness is involved, in the hope of identifying problems that are decidable with reasonable complexity. Another motivation comes from the limitations of Nash equilibria, defined using the classical expected payoff, in capturing intuitively rational behavior. 10.1.3 Randomness and risk Motivating example.Let us consider a 1-player game where a protagonist is proposed two options: (a) earning€1; (b) playing a lottery in which, with probability 1 40 , she gets€40, and with probability 39 40 , she does not earn anything. Classically, rational strategies would be maximizing the expected payoff. From this perspective, both options yield an expected payoff of€1, making them equivalent. This approach is particularly justified when the game represents a scenario that can be repeated many times: the law of large numbers ensures that, in the long run, the average payoff will converge to the expected payoff. However, when the game is played only once, the protagonist may prioritize immediate needs. If she urgently requires€1, the guaranteed option (a) becomes preferable. Conversely, if she is a risk-taker, or finds herself in a situation where only the€40 can make a significant difference, she may prefer the high-risk option (b). Although this choice might appear irrational, it mirrors the behavior of millions of people who participate everyday in games that even have a negative expected payoff, driven by the allure of a potentially life-changing win, and generating an annual turnover of USD 536 billions [Cap23] for the gambling industry. That industry, on the other hand, operates on a large scale where expected payoff becomes the key metric. This contrast underscores the importance of alternative measures to expected payoff that account for each agent’s risk tolerance. 152CHAPTER 10. RISK-SENSITIVE EQUILIBRIA 푎 푏 푐 푡 1 : ◦ 40 푡 2 : ◦ 0 푡 3 : ◦ 1 1 2 1 2 1 40 39 40 Figure 37: A stochastic MDP Risk measures.A risk measure captures the perception that a player has of what their payoff will be. In that sense, they generalize the notion of expected payoff. Various risk measures exist in the literature, and have been used extensively in the field of economics and finance. Some of these risk measures include expected shortfall (ES), value at risk (VaR) [Aue18], variance [Bra99], entropic risk measure (ER) [FS02]. A lot of work has been done in considering these risk measures over Markov decision processes (MDPs) which use variance (along with mean) as a risk-measure [FK89,PSB22,MT11], ES [RRS15,KM18,Meg22] (also referred to as conditional value at risk (CVaR), average value at risk (AVaR), expected tail loss (ETL), and superquantile in literature) and ER [HM72,BR14,BCMP24]. Studying the entropic risk measure in MDPs appears more practical compared to expected shortfall or using variance-penalized risk-measures. This impracticability of ES and variance-penalized measure in particular is due to the intractable exponential memory [HK15] and time required to compute optimal strategies [PSB22], even for the one agent system of Markov decision processes (MDPs). On the other hand, when the risk measure used is ER, players have optimal positional strategies in MDPs [HM72], which makes it a prime candidate for consideration in multi-agent settings. Entropic risk measure. The entropic risk measure is computed by assigning to each agent a risk parameter, i.e., a value휌 ∈ R. The entropic risk measure of a random variable푋is then defined as M 휌 [푋] =− 1 휌 log 푒 E 푒 −휌푋 . Assume the random variable푋is a player’s payoff. If the risk parameter 휌is positive, then more weight will be given to the bad payoffs: the corresponding player can then be considered as risk-averse. Conversely, players with a negative휌are more risk-loving. When휌tends to 0, the entropic risk measure converges to the classical expectation E[푋]. The game depicted by Figure 37 extends the lottery example we discussed earlier. Black vertices are stochastic, and the circle vertex is controlled by player ◦ . A play can be seen as a sequence of moves of a token along the edges of the graph, starting from푎: from a stochastic vertex, it takes one of the outgoing edges with the probabilities indicated on those, and from a vertex controlled by the player, she chooses which edge it takes. The payoff 40, 0, or 1 is obtained when the terminal vertex푡 1 ,푡 2 , or푡 3 is reached, respectively. If no terminal vertex is reached, then the payoff is 0. Taking the red edge corresponds to option (a): then, her risk entropy is always 1, for every risk parameter휌. But if she chooses option (b), that is, if she takes the blue edge, her risk entropy isM 휌 [휇 ◦ ] =− 1 휌 log ( E [ 푒 −휌휇 ◦ ]) =− 1 휌 log 1 40 푒 −40휌 + 39 40 . Both cases are illustrated with red and blue curves in Figure 38. The curves cross at abscissa휌 =0, where the entropic risk measure corresponds to the expectation. Note that other strategies are possible if randomization is allowed—the player could, for example, toss a coin and participate in the lottery if the outcome is heads. The perceived reward of randomizing between outermost red and blue edges are illustrated in the intermediate cases with mixtures of red and blue in Figure 38. Unfortunately, even for two player zero-sum stochastic games with total-reward objectives (payoff is 10.2. DEFINITIONS153 −6−4−20246810 0 10 20 30 40 Risk parameter 휌 M ( Outcome ) Figure 38: Each curve represents the perceived reward of a player choosing only blue strategy, only red, or randomizing between both strategies. the sum of the rewards seen along the way), computing optimal strategies can only be done inPSPACE, when the base푒is replaced by a rational number; and if푒is the base of the exponent, then it is only inEXPTIME[GHM25]. Solving the two-player zero-sum case is a specific case of finding equilibria in two-agent systems where the payoffs of the two agents are exactly the negation of each others and so are the risk parameters of each of the agents. Therefore, reasoning about multi-agent systems with ER also has potential to be computationally intractable. 10.2DEFINITIONS We provide here additional definitions that will be used in this chapter and the following. 10.2.1 Probabilities Given a (finite or infinite) set of outcomesΩand a probability measurePoverΩ, let푋be a random variable overΩ, that is, a mapping푋:Ω ! R. We then writeE P [푋], or simplyE[푋], for the expectation of푋, when it is defined. Given a finite set푆, a probability distribution over푆is a mapping푑:푆 ! [0,1] that satisfies the equality Í 푥∈푆 푑(푥) =1. We writeSupp(푑)for the support of the distribution푑, that is, the set of elements 푥 ∈ 푆 such that 푑(푥)> 0. 10.2.2 Risk measures Given a setΩof outcomes, a risk measure overΩis a mapping푀which maps a probability measureP overΩ and a random variable 푋 to a real value 푀 P [푋]. 154CHAPTER 10. RISK-SENSITIVE EQUILIBRIA Sometimes, in the literature, risk measures are expected to have the following three properties: (1) they are normalized, i.e., we have푀 P [0] =0; (2) they are monotone, i.e., the pointwise inequality푋 ≤ 푌 implies푀 P [푋] ≤ 푀 P [푌]; and (3) they are translative, i.e., we have푀 P [푋 + 푐] = 푀 P [푋] + 푐for every constant푐. In particular, the expectation of a random variableEsatisfies those properties. Sometimes also, the word translative refers to the opposite of this definition, i.e., to the condition푀 P [푋 +푐] = 푀 P [푋]−푐 for every constant푐. We will not need any of those properties in the sequel—we simply note that they are satisfied (with the first definition of translativity) by the risk measures we consider. 10.2.3 Simple stochastic games Since we now want to consider games in which players must deal with stochastic phenomena, we need a new definition that allows for stochastic vertices. At the same time, since the equilibrium notions we study in this chapter and the next are new, we do not want this study to be influenced by the properties of sophisticated payoff functions, such as those examined in the previous chapters. Therefore, we restrict our work to simple stochastic games, i.e., quantitative games in which each player’s payoff is entirely determined by the terminal vertex that is eventually reached (or not). Definition 43 (Simple stochastic game). A simple stochastic game is a tupleG =(푉,퐸,Π,(푉 푖 ) 푖∈Π ,p, 휇), where we have: • a directed graph(푉,퐸), called the underlying graph ofG; • a finite setΠ of players; • a partition(푉 푖 ) 푖∈Π∪? of the set푉, where푉 푖 denotes the set of vertices controlled by player푖, and the vertices in푉 ? are called stochastic vertices; •a probability function p :퐸(푉 ? ) ! [0,1], such that for each stochastic vertex푢, the restriction of p to 퐸(푢) is a probability distribution; •a mapping휇:푇 ! R Π called payoff function, where푇is the set of terminal vertices, i.e. vertices of the graph(푉,퐸)that have no outgoing edges. We also write휇 푖 , for each player푖, for the function that maps a terminal vertex 푡 to the 푖 th coordinate of the tuple 휇(푡). The mapping 휇 extends to the set(푉 \푇) 휔 ∪(푉 \푇) ∗ 푇by defining휇(푣 1 . . .푣 푘 푡) = 휇(푡), and휇(푣 1 푣 2 . . .) = (0) 푖∈Π (if no terminal vertex is reached, everyone gets the payoff 0). Note that we dropped the assumption according to which every vertex has an outgoing edge: this was necessary to have terminal vertices, defining the players’ payoff. On the other hand, for computational reasons, we will often assume that every vertex, except possibly the initial one when defined, has an ingoing edge: thus, we always have card퐸 ≥ card푉 − 1. In this chapter and the following, we often use the word game for simple stochastic games. Moreover, we use the vocabulary and the notations that we have introduced for games as defined in Definition 1, for simple stochastic games, with analogous definitions. For instance, we will sometimes consider initialized simple stochastic games (without always mentioning that they are initialized). We often assume that we are given a simple stochastic gameG ↾푣 0 and implicitly use the same notations as in the definition above. In such games, histories are still defined as finite paths in the graph(푉,퐸). Plays are now paths that are either infinite, or end in a terminal vertex. 10.2. DEFINITIONS155 Finally, note that simple stochastic games can also be seen as a particular case of mean-payoff games (with stochastic vertices), up to adding a self-loop on terminal vertices—those loops are then the only edges with possibly nonzero rewards. Markov decision processes and Markov chains can be defined as particular cases of stochastic games. Definition 44 (Markov decision process, Markov chain). A (initialized or not) Markov decision process is a simple stochastic game with one player. A Markov chain is a simple stochastic game with zero player. 10.2.4 Strategies, and strategy profiles We now give a new definition for strategies, allowing for randomization. This definition will always be the one used in the sequel. Definition 45. In a gameG ↾푣 0 , a strategy for player푖is a mapping휎 푖 that maps each historyℎ푢 ∈ Hist 푖 G ↾푣 0 to a probability distribution over 퐸(푢). Again, the vocabulary and the notations that have been defined around strategies in the sense of Definition 4 are naturally adapted. For instance, a strategy profile is still a tuple of strategies indexed by players. A path훼 0 훼 1 . . .(be it a history or a play) is now said to be compatible with the strategy휎 푖 if for each 푘such that훼 푘 ∈ 푉 푖 , the probability that the strategy휎 푖 outputs the vertex훼 푘+1 after the history훼 ≤푘 is positive, that is, if we have 휎 푖 (훼 ≤푘 )(훼 푘+1 )> 0. That definition naturally extends to strategy profiles. A strategy profile ̄ 휎 −푖 in the gameG ↾푣 0 defines an initialized Markov decision processG ↾푣 0 [ ̄ 휎 −푖 ] , where the vertices of the (infinite) underlying graph are the histories ofG ↾푣 0 and the edges are added fromℎ푢 to each the historyℎ푢푣if and only if we have푢푣 ∈ 퐸. Similarly, a strategy profile ̄ 휎 forΠdefines an initialized Markov chainG ↾푣 0 [ ̄ 휎]. Thus, it also defines a probability measureP ̄ 휎 over plays—which turns the payoff functions 휇 푖 into random variables. Remark 10. A historyℎis compatible with ̄ 휎if and only if it has positive probability of being generated. This equivalence fails for plays: a play휋may be compatible with ̄ 휎but generated with probability 0, if it is infinite (and therefore never reaches any terminal vertex) and if along휋, infinitely many randomized choices occur, be it by the players or by stochastic vertices. We say that a strategy휎 푖 is pure when for each historyℎ푢, there is a vertex푣such that휎 푖 (ℎ푢)(푣) =1. Then, we often just write휎 푖 (ℎ푢) = 푣. The strategy휎 푖 is positional when it is pure and stationary. Those concepts are naturally generalized to strategy profiles. Remark 11. Pure strategies can be seen as strategies in the sense of Definition 4. 10.2.5 Risk-sensitive equilibria Let us first give a new definition of Nash equilibria, taking into account the existence of stochastic vertices and of randomized strategies. For the sake of readability, we write E( ̄ 휎) for E P ̄ 휎 . Definition 46 (Nash equilibrium). LetG ↾푣 0 be a game. Let ̄ 휎be a strategy profile inG ↾푣 0 , let푖be a player, and let휎 ′ 푖 be a deviation for player푖. The deviation휎 ′ 푖 is profitable if we haveE( ̄ 휎 −푖 ,휎 ′ 푖 )[휇 푖 ]> E( ̄ 휎)[휇 푖 ] The strategy profile ̄ 휎 is a Nash equilibrium, or NE, if no player has a profitable deviation from ̄ 휎 . As we saw in the introduction to this chapter, Nash equilibria in stochastic games fail to capture notions 156CHAPTER 10. RISK-SENSITIVE EQUILIBRIA of risk aversion (or tolerance). This is why we aim to study generalizations in which the expectation is replaced by other risk measures. Such generalizations are known as risk-sensitive equilibria [Now05]. We define them here for games played on graphs. Similarly to what we did with the expectation, when푀is a risk measure and ̄ 휎 is a strategy profile, we write 푀( ̄ 휎) to denote 푀 P ̄ 휎 . Definition 47 (Risk-sensitive equilibrium). LetG ↾푣 0 be a game, and let ̄ 푀 =(푀 푖 ) 푖∈Π be a profile of risk measures. Let ̄ 휎 be a strategy profile inG ↾푣 0 , let푖be a player, and let휎 ′ 푖 be a strategy for player푖, called deviation of player푖from ̄ 휎. The deviation휎 ′ 푖 is profitable with regards to the risk measure푀 푖 if we have푀 푖 ( ̄ 휎 −푖 ,휎 ′ 푖 )[휇 푖 ]> 푀 푖 ( ̄ 휎)[휇 푖 ] . The strategy profile ̄ 휎is a ̄ 푀-risk-sensitive equilibrium, or ̄ 푀-RSE, if no player 푖 has a profitable deviation from ̄ 휎 with regards to 푀 푖 . We can now define the following variant of the constrained existence problem, where each player’s payoff is replaced by a given risk measure of that payoff. Problem 11 (Constrained existence of risk-sensitive equilibria). Given a gameG ↾푣 0 , a profile of risk measures ̄ 푀, and two payoff vectors ̄ 푥, ̄ 푦 ∈ Q Π , does there exist a ̄ 푀-RSE ̄ 휎 inG ↾푣 0 such that for each푖 ∈Π, we have 푥 푖 ≤ 푀 푖 ( ̄ 휎)[휇 푖 ] ≤ 푦 푖 ? To turn this problem into an algorithmic decision problem, we still need to restrict it to some specific sets of risk measures that can be finitely encoded. That is what we do in the sequel of this chapter, with the entropic risk measure, and in the next one, with the extreme risk measure. 10.3ENTROPIC RISK MEASURE The entropic risk measure is a measure of the perceived payoff, which depends on the aversion or inclination of the player toward risk through the exponential utility function. It is defined using a risk parameter, i.e. a real value휌 ∈ R\0: large positive values indicate risk-averseness, large negative values risk-inclination. To see a visual representation of the entropic risk measure, see Figure 38 in the introduction of this chapter. Definition 48 (Entropic risk measure). Given a risk parameter휌, the entropic risk measure is defined, for every probability measure P and random variable 푋 , by: M P 휌 [푋] =− 1 휌 log 푒 E P 푒 −휌푋 . We generalize this definition by allowing every base훽>1 instead of Euler’s number. The entropic risk measure with base 훽 is then defined by: M P 훽휌 [푋] =− 1 휌 log 훽 E P 훽 −휌푋 . The probability measureP, the parameter휌and the base훽may be omitted when they are clear from the context. Remark 12.• For every훽and휌, the entropic risk measureM 훽휌 is a monotone, normalized and translative risk measure. •By enabling any base훽, we obtain a definition that is more general only on a computational level, since handling Euler’s number may not be equivalent to handling rational values. Baring 10.4. THE EXISTENCE OF ERSES157 computational concerns, these definitions with different bases are equivalent, since for every훽, we have M 훽휌 = M 푒휌 ′ , where 휌 ′ = 휌log 푒 (훽). •The above definition implies that for휌 =0, the function is not defined. However, it is known that for allP,훽and푋, the quantityM 휌 [푋]converges toE[푋]when휌tends to 0 (see e.g. [PDM20]). Therefore, we henceforth assume thatM 0 [푋] = E[푋]to make risk entropy defined for all finite risk parameters 휌 . When we are given a profile ̄ 휌 =(휌 푖 ) 푖∈Π of risk parameters, we will sometimes writeM 훽 ̄ 휌 [휇]for the tuple M 훽휌 푖 [휇 푖 ] 푖∈Π . Risk entropy defines a family of RSEs, namely the(M 훽휌 푖 ) 푖 -RSEs, that we also call (훽, ̄ 휌)-entropic risk-sensitive equilibria, or(훽, ̄ 휌)-ERSEs. As we will see, entropic risk-sensitive equilibria do not resolve the undecidability issue that affects Nash equilibria, for a simple reason: they are merely a generalization of the latter. However, they represent a first step toward our next notion of equilibrium, extreme risk-sensitive equilibria, which we will study in the next chapter. This chapter must therefore be understood as a warm-up for the next one, with a few first results. In Section 10.4, we give a sufficient condition for ERSEs to exist, and in Section 10.5, we give some complexity results. 10.4THE EXISTENCE OF ERSES The following theorem states the existence of such an RSE that uses no randomness in its strategy profile, in cases where all payoffs are non-negative. Theorem 27 (Existence of ERSE). LetG ↾푣 0 be a simple stochastic game with only non-negative payoffs. Then, there exists a (pure)(훽,휌)-ERSE inG ↾푣 0 . Proof.Pure Nash equilibria always exist in a stochastic multi-player game with prefix-closed Boolean objectives [Umm10, Theorem 3.10] (a correction of an existing proof [CMJ04]). It is known that simple stochastic games where rewards are all positive (or all negative) can be converted into a game with reachability objectives such that if there is an NE in the latter, then there is an NE in the converted game with the reachability objective. Indeed, if all the rewards are positive, we can always scale the rewards for each player of a stochastic game to ensure they are in the unit interval[0,1]. Then, we can replace terminal vertices by the gadget given by Figure 39 (we assume푥 1 ≤·≤ 푥 푛 , and in the reachability game, each vertex푟 푗 belongs to the set player푖 푗 wishes to reach). Therefore, with the same result, Nash equilibria always exist in simple stochastic games with non-negative rewards on the terminals. Then, we can conclude our theorem using the following lemma. Lemma 26. Given a gameG ↾푣 0 and a tuple ̄ 휌 ∈ R Π , there exists a gameG ′ ↾푣 0 with the same underlying graph, player set, and probability function (but possibly different payoff function), such that the(훽,휌)- ERSEs inG ↾푣 0 are exactly the Nash equilibria inG ′ ↾푣 0 . Proof.Consider the simple stochastic gameG ↾푣 0 = ( 푉,퐸,Π,(푉 푖 ) 푖∈Π , p,휇 ) . We will define a payoff function휇 ′ over the same set of terminals for gameG ′ ↾푣 0 such that the gameG ′ ↾푣 0 = ( 푉,퐸,Π,(푉 푖 ) 푖∈Π , p,휇 ′ ) has a Nash equilibrium if and only if the gameGhas a(훽, ̄ 휌)-ERSE. For a terminal vertex푡, we simply define 휇 ′ 푖 (푡) = 1− 훽 −휌 푖 휇 푖 (푡) if 휌> 0, and 휇 ′ 푖 (푡) = 훽 −휌 푖 휇 푖 (푡) − 1 if 휌< 0. 158CHAPTER 10. RISK-SENSITIVE EQUILIBRIA 푡 : 푖 1 푥 1 푖 2 푥 2 . . . 푖 푛 푥 푛 (a) A terminal vertex 푡 푟 1 푟 2 . . . 푟 푛 푥 1 푥 2 − 푥 1 푥 푛 − 푥 푛−1 (b) An equivalent gadget Figure 39: Converting simple quantitative games into reachability games Consider the function: 푓 훽휌 : 푥 7! 1− 훽 −휌푥 if 휌> 0 훽 −휌푥 − 1 if 휌< 0 as the modified reward function. This function is similar to the negative utility function defined in the work of Baier et al. [BCMP24], where they replace terminal rewards with the negative value of 훽 −휌휇 푖 (푡) (as they assume휌>0), in order to compute the winner in a two-player zero-sum game with risk-averse players. We additionally add or subtract 1 from their value to ensure that besides monotonicity, this function also maps the play that does not reach a terminal in the original game to the payoff 0, and therefore in the modified game to preserve that such plays are still mapped to 0. Since we have M 훽휌 [ 푋 ] = −1 휌 log 훽 E[훽 −휌푋 ] , we immediately obtain the following result. Sublemma 7. For any random variable 푋 and constant 푟, for a value 휌> 0, we have: M 훽휌 [ 푋 ] ≥ 푟 if and only if E 푓 훽휌 (푋) ≥ 1− 훽 −휌푟 Therefore, any Nash equilibrium in the gameG ′ ↾푣 0 implies that there is a strategy profile ̄ 휎such that, for all players푖 ∈Π, in the MDP induced by ̄ 휎 −푖 , the strategy휎 푖 of player푖is an optimal strategy. We first consider the case of a player푖where휌 푖 >0. The case휌 푖 <0 is analogous, so we omit it. For every strategy휏 푖 of player푖, where휌 푖 >0, if we write ̄ 휏 = ( ̄ 휎 −푖 ,휏 푖 ), we have E( ̄ 휎)[휇 ′ 푖 ] ≥ E( ̄ 휏)[휇 ′ 푖 ], since ̄ 휎is a Nash equilibrium. Since the payoffs ofG ′ at a terminal푡is just 푓 훽휌 (휇 푖 ()), we therefore have E( ̄ 휎)[푓 훽휌 푖 (휇 푖 )] ≥ E( ̄ 휏)[푓 훽휌 푖 (휇 푖 )]. From Sublemma 7, we have: E( ̄ 휎) 푓 훽휌 푖 (휇 푖 ) ≥ E( ̄ 휏) 푓 훽휌 푖 (휇 푖 ) if and only if we have: E( ̄ 휎) [ 1− 훽 −휌 푖 휇 푖 ] ≥ E( ̄ 휏) [ 1− 훽 휌 푖 휇 푖 ] , 10.5. CONSTRAINED EXISTENCE PROBLEM159 that is if and only if we have: E( ̄ 휎) [ −훽 −휌 푖 휇 푖 ] ≥ E( ̄ 휏) [ −훽 −휌 푖 휇 푖 ] . Taking 1 휌 log 훽 on both sides, we get the above is true if and only if: M 훽휌 푖 ( ̄ 휎)[휇 푖 ] ≥− 1 휌 푖 log 훽 −E 푓 훽휌 푖 ( ̄ 휏)[휇 푖 ] , i.e., if and only if M 훽휌 푖 [ ̄ 휎](휇 푖 ) ≥ M 훽휌 푖 [ ̄ 휏](휇 푖 ). Therefore, the strategy profile ̄ 휎 is an ERSE. □ □ We conjecture that this result remains true when we remove the guarantee that rewards are non- negative. Conjecture 2. LetG ↾푣 0 be a simple stochastic game. Then, there exists a (pure)(훽,휌)-ERSE inG ↾푣 0 . 10.5CONSTRAINED EXISTENCE PROBLEM We now turn to the constrained existence problem of ( 훽, ̄ 휌 ) -ERSEs. Unfortunately, it is undecidable in the general case, but we will see that it becomes decidable if we consider restricted classes of strategies. 10.5.1 Undecidability in the general case Theorem 28. The constrained existence problem of ( 훽, ̄ 휌 ) -ERSEs with ̄ 휌 ∈ Q Π is undecidable, even for any fixed value of 훽, for ̄ 휌 = ̄ 0, and with only nonnegative payoffs. Proof. The undecidability of the constrained existence problem follows from the work of Ummels and Wojtczak [UW11b, Theorem 4.9] where they show the undecidability of the constrained existence problem for Nash equilibria in the setting with 10 or more players. Since Nash equilibria constitute a specific instance of the setting of ERSEs where ̄ 휌 = ̄ 0, the undecidability of our setting follows. □ 10.5.2 Restrictions on strategies We therefore turn our attention to the constrained existence problem when the class of strategies considered is restricted. Before we state our theorem, we need to define the existential theory of the reals (with exponentiation in the second case), and the associated complexity class. Definition 49 (Existential theory of the reals (with exponentiation)). A formula of the existential theory of the reals, or ETR for short, is a formula of the form∃푥 1 . . .∃푥 푘 휑(푥 1 , . . .,푥 푘 ), where휑is a formula written with the symbols 0, 1,=,≤,<,+,−,×,∧,⇔,¬, and parentheses, with the expected syntactic rules and semantics. The existential theory of the reals with exponentiation, or ETRE for short, is defined similarly with one additional symbol: Euler’s number 푒. We write∃Rfor the complexity class of problems that can be reduced in polynomial time to the problem of deciding the validity of an ETR formula. That class is known to be included inPSPACE, by the following lemma. Lemma 27 ([Can88]). Deciding the validity of a formula in ETR is PSPACE-easy. 160CHAPTER 10. RISK-SENSITIVE EQUILIBRIA Similarly, we write∃R(푒)for the class of problems that can be reduced in polynomial time to the validity problem of an ETRE formula. That class is known to be included in 3EXPTIME. Lemma 28 ([GHM25]). Deciding the validity of a formula in ETRE is 3EXPTIME-easy. We can now state our theorem. Theorem 29. The constrained existence problem of(훽, ̄ 휌)-ERSEs, in quantitative simple stochastic games: 1. remains undecidable when players are restricted to pure strategies; 2. is decidable when players are restricted to stationary strategies, or to positional strategies: (a) in 3EXPTIME if 훽 = 푒 and the risk-parameters 휌 푖 are rational; (b) and is∃R-easy if the risk parameters and the base훽are rational. In the stationary case, the problem is∃R-complete, and in the positional case, it is NP-hard. Proof.▶Undecidability result The undecidability of the pure case is inherited from Nash equilibria [UW11b, Theorem 4.9], since the reduction for undecidability uses only pure strategies. ▶Decidable subcases Decidability lowerbounds.NP-hardness has been shown in the case of positional NEs, i.e. when each player’s risk parameter is 0, by Ummels and Wojtczak [UW11b, Theorem 4.4]. Similarly, ∃R-hardness has been proved for the constrained existence problem of stationary NEs by Kristoffer Arnsfelt Hansen and Steffan Christ Sølvsten [HS20]. Decidability upperbounds.Our proof is similar to the one by Ummels and Wojtczak [UW11b, Theorem 4.5]. However, we need to do slightly more work to encode the payoff expressed by the entropic risk measure. First, observe that it is enough to verify if there is a stationary Nash equilibrium in the modified game obtained where all the terminal rewards휇 푖 (푣)are replaced instead with 1−훽 휌 푖 휇 푖 (푣) . This follows from Lemma 26. Since the players are restricted to strategies that are stationary, we give a non-deterministic algorithm that uses the solution to sentences in∃Rif the values of훽and휌 푖 s are rational. Since we have NPSPACE = PSPACE by Savitch’s theorem, this does not change the complexity. For a gameG ↾푣 0 = ( 푉,퐸,Π,(푣 푖 ) 푖∈Π , p,휇 ) and two payoff vectors ̄ 푥and ̄ 푦, our algorithm guesses, first, the support푆 ⊆ 퐸of the strategies that will be considered; that is, the set of edges that will be used with positive probability. Lemma 29. For any푧, which requiresℓbits to encode, there is a formula in ETR that uses only polynomially many variables inℓto encode푓 훽휌 푖 (푧) =1− 훽 −휌 푖 푧 , where훽and휌 푖 can also be represented in ETR using a polynomial formula. If훽 = 푒and휌 푖 is rational, then푓 푒휌 (푧) =1− 훽 −휌푧 can be expressed in ETRE using a formula of most polynomial length. 10.5. CONSTRAINED EXISTENCE PROBLEM161 Proof.In the first case, we assume without loss of generality that훽is a natural number. If훽is rational instead, and is represented by a value 푎 푏 , then we can individually find푎 1 = 푎 휌 푖 푧 and푏 1 =푏 휌 푖 푧 , and just find 푏 1 푎 1 , which is written in ETR with the formula∃푡 푟 ∃푎 ′ 1 (푎 ′ 1 ×푎 1 =1)∧(푡 푟 = 푎 ′ 1 ×푏 1 ). Now, under this assumption, we deal with fixed finite exponentiation with rational values. Similarly, we can assume without loss of generality that휌 푖 푧is a natural number. Otherwise, we can write 푟 푎 푏 = 푟 푎 ×푧 1 푏 and 푧 1 푏 can be defined by the formula∃푦 : 푦 푏 = 푧. It now suffices to show that for two values푏,푎, both natural numbers, the quantity푏 푎 can be expressed in ETR succinctly, using only a formula that has length that is not more a poly-log of푏 or 푎. Let 푎 = Í log 2 푎 푖=0 푎 푖 2 푖 , where 푎 푖 ∈ 0, 1. This follows from the following observations. •First, we can write 푏 푎 = Î log 2 푎 푖=1 푏 푎 푖 푏 2 푖 . •Second, the quantity 2 푖 can be expressed in a formula with at most 푖+ 1 many variables. •Third, the quantity푏 2 푖 requires at most푖many variables to express, because if푏 푖 represents 푏 2 푖 , then we have 푏 2 푖+1 =푏 2 푖 ×푏 2 푖 . •Finally, using a similar trick, the quantity푏itself can be represented using at mostlog 2 푏-many variables. Finally, if훽 = 푒and휌 푖 is rational, the quantity푒 −휌푧 can be expressed following the same reasoning by a succinct formula that uses the symbol 푒.□ Lemma 29 ensures that we can efficiently represent the variables used for the payoffs of the modified game. We now can write an equation assuming that all terminal rewards are available to us as constants. This will write the equation in three parts. Since we have guessed the support, we first ensure that, in fact, there are variables corresponding to the probabilities of the strategy that only take positive values on the edges corresponding to the support set that we guessed. Then, we write equations using variables that compute the values of the induced Markov chain from this strategies. Finally, we also have a formula whose solution corresponds to the values of the MDP obtained for each player when playing against the strategies of all other players. Then we compare if the value of the MDP is at least as large as the underlying Markov chain for each player, to ensure that it is indeed an equilibrium. To write all of this in ETR, we introduce the following variables: •one variable푝 푣푤 for each pair of vertices푣푤, which corresponds to the probabilities associated with the strategy profile; •a variable푟 푖 푣 which corresponds to the entropic risk measure of player푖from vertex푣if they follow the strategy defined by the probabilities above; • a variable푚 푖 푣 which corresponds to the value obtained by player푖if the game is treated as an MDP against other players. We can now define our formulae. Lemma 30. Given a gameG ↾푣 0 , two payoff vectors ̄ 푥 and ̄ 푦, and a subset 푆 ⊆ 퐸: • if훽and all휌 푖 are rational values: there exists a formula in ETR, computable in polynomial time, that is satisfied if and only if there exists a stationary ERSE ̄ 휎with푥 푖 ≤ M 휌 푖 ( ̄ 휎)[휇 푖 ] ≤ 푦 푖 for each 푖 ∈Π that uses exactly (with positive probability) the edges in 푆; • if훽 = 푒, and if all휌 푖 are rational values: there exists a formula in ETRE, computable in polynomial time, that is satisfied if and only if there exists a stationary ERSE ̄ 휎with푥 푖 ≤ M 휌 푖 ( ̄ 휎)[휇 푖 ] ≤ 푦 푖 for each 푖 ∈Π that uses exactly (with positive probability) the edges in 푆. 162CHAPTER 10. RISK-SENSITIVE EQUILIBRIA •Those results remain true if the restriction to stationary XRSEs is replaced by a restriction to positional XRSEs. Proof.This proof is similar to the one found in Ummels and Wojtczak [UW11b, Theorem 4.5], but we provide it to suit our setting, for the sake of completeness. First, we have a formula that states that the values푝 푣푤 indeed describe a strategy. We further ensure that for stochastic vertices, the value푝 푣푤 encodes exactly the value dictated by the probability function p by the stochastic vertex: Φ 푆 ( ̄ 푝) = Û 푣푤∈푆 (( 푝 푣푤 > 0 ) ∧ ( 푝 푣푤 ≤ 1 )) ∧ Û 푣푤∈퐸\푆 ( 푝 푣푤 = 0 ) ∧ Û 푖∈Π Û 푣∈푉 푖 © « ∑︁ 푤∈퐸(푣) 푝 푣푤 = 1 ª ® ¬ ∧ Û 푣∈푉 ? ( 푝 푣푤 = p(푣푤) ) . For a fixed support푆of a strategy ̄ 휎 , it is possible to compute the set푇 푆 of terminal vertices that have non-zero probability of being reached in the underlying Markov chain that is formed, and the vertices푉 푆 from which such terminals can be reached with non-zero probability. For a given terminal vertex푡and player푖, we write휇 ′ 푖 (푡) for the quantity 1− 훽 휌 푖 휇(푡) . We can now write a statement that define the values of the variables 푟 푖 푣 . Ω 푖 푆 ( ̄ 푝, ̄ 푟 푖 ) = Û 푡∈푇 푆 푟 푖 푡 = 휇 ′ 푖 (푡) ∧ Û 푣∉푉 푆 푟 푖 푣 = 1 ∧ Û 푣∈푉 푆 \푇 푆 © « 푟 푖 푣 = ∑︁ 푤∈퐸(푣) 푝 푣푤 푟 푖 푣 ª ® ¬ . To compute the values of the variables푚 푖 푣 , we construct a similar first-order statement: Ψ 푖 푆 ( ̄ 푝, ̄ 푚 푖 ) = Û 푡∈푇 푚 푖 = 휇 ′ 푖 (푡) ∧ Û 푣∈푉 푖 ,푤∈퐸(푣) 푚 푖 푣 ≥ 푚 푖 푤 ∧ Û 푣∉푉\푉 푖 © « 푚 푖 푤 = ∑︁ 푤∈퐸(푣) 푝 푣푤 푚 푖 푣 ª ® ¬ . Then, our statement is: ∃ ̄ 푝∃ ̄ 푟 ∃ ̄ 푚Φ( ̄ 푝)∧ Û 푖∈Π 푥 푖 푣 0 ≤ 푟 푖 푣 0 ∧ 푟 푖 푣 0 ≤ 푦 푖 푣 0 ∧Ω 푖 푆 ( ̄ 푝, ̄ 푟 푖 )∧Ψ 푖 푆 ( ̄ 푝, ̄ 푚 푖 )∧ 푚 푖 푣 0 ≤ 푟 푖 푣 0 . Using Lemma 29, this formula belongs to ETR if 훽 is rational, and to ETRE if 훽 = 푒. Finally, this result can be extended to the positional case by slightly changing the definition of the formulaΦ 푆 ( ̄ 푝): Φ 푆 ( ̄ 푝) = Û 푣푤∈푆 ( 푝 푣푤 = 1 ) ∧ Û 푣푤∈퐸\푆 ( 푝 푣푤 = 0 ) ∧ Û 푖∈Π Û 푣∈푉 푖 © « ∑︁ 푤∈퐸(푣) p 푣푤 = 1 ª ® ¬ ∧ Û 푣∈푉 ? ( 푝 푣푤 = p(푣푤) ) □ The conclusion follows from Lemmas 27 and 28.□ CHAPTER 11:EXTREME RISK-SENSITIVE EQUI- LIBRIA In Chapter 10, we introduced the formalism of simple stochastic games, randomized strategies, and the notion of risk-sensitive equilibria for an abstract profile of risk measures. We studied a first risk measure: the entropic risk measure, which is widely used in economics and finance. We showed that, unless strategies are restricted to stationary or positional ones, the constrained existence problem remains undecidable in this setting. A natural question then arises: can alternative risk measures be found that make this problem decidable? For this reason, in this chapter, we introduce a risk measure derived from the limit cases of the entropic risk measure: the extreme risk measure. Let us first examine what such a measure means in a two-player zero-sum game. Consider a pirate attempting to hack a computer system, such as a bank server. The interactions between these two agents may involve randomized actions—for instance, when the bank server generates a security key or when the hacker tries to guess a user’s password. They may also include external stochastic processes, such as the number of users attempting to connect to the bank server at any given time. To safeguard users’ funds, the bank server must be programmed in a pessimistic manner, assuming that any security breach with a nonzero probability is a security breach that will inevitably occur. Conversely, the pirate can be seen as an optimistic agent, content with even a small probability of success—he may be able to retry his attempts indefinitely, or to execute them in parallel using a large network of computers. Such interactions are often modeled using non-stochastic games, by merging the pirate and random events into a single environment player. However, this approach is not entirely equivalent to a probabilistic one, as the latter allows discarding events that have probability zero—such as consistently winning heads-or-tails-like interactions over an infinite horizon. More importantly, the adversarial approach cannot be generalized to multiplayer settings, in which it may be necessary to account for the existence of stochastic phenomena that are perceived differently by each agent. Consider, for example, two secured computer systems interacting with each other—say, a bank server and an online selling platform. Their behavior cannot be properly modeled without accounting for the fact that both are pessimistic: each has a specification to uphold and cannot tolerate even a minimal probability of being compromised. If a pirate discovers a method that grants a nonzero probability of breaching either the bank server or the selling platform (but never both simultaneously), then each system will anticipate its own failure. The bank server will assume that it will be hacked, while the selling platform will assume the same for itself—even though these two scenarios are mutually exclusive. This example highlights the importance of studying equilibria where all players are either extreme optimists or extreme pessimists—that is, players who, when faced with a random experiment, always expect either the best or the worst possible outcome. In other words, we consider risk-sensitive equilibria defined by risk measures corresponding to the extreme cases of risk entropy—what we refer to as the 163 164CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA extreme risk measure. 11.1THE EXTREME RISK MEASURE 11.1.1 Definition Let us consider a random variable푋that ranges overR. The pessimistic risk measure of푋is the highest value푥such that푋almost-surely takes a value above푥. When푋takes finitely many values, that corresponds to the least value that it takes with positive probability. In probability theory, that measure is sometimes referred to as essential infimum, writteness inf. The definition of optimistic risk measure is symmetric. Definition 50 (Optimistic, pessimistic risk measure). The pessimistic risk measure of a random variable푋 is defined byPM P [푋] = ess inf(푋) = sup푥 ∈ 푋 | P(푋 ≥ 푥) =1. Analogously, the optimistic risk measure of 푋 is OM P [푋] = ess sup(푋) = inf푥 ∈ 푋 | P(푋 ≤ 푥) = 1. When we are given a gameG ↾푣 0 , we can assign a risk measure to each player by defining a partition (푃,푂)ofΠ, where the set푃represents the set of players that are pessimists, whose perceived payoffs are defined by the pessimistic risk measure, while푂represents the optimists, who intend to maximize their optimistic risk measure. For convenience, we group both measures under the umbrella term extreme risk measure (XR), and often assume that the pair(푃,푂)is given; then, we writeX 푖 forPMwhen푖 ∈ 푃, and forOMwhen푖 ∈ 푂. Since each player푖is usually interested only in the risk measure of their own payoff, we will also writeX 푖 ( ̄ 휎)for the quantityX 푖 ( ̄ 휎)[휇 푖 ]. We define extreme risk-sensitive equilibria, or XRSEs for short, as(X 푖 ) 푖 -RSEs. 11.1.2 Link with the entropic risk measure We show that our definition of extreme risk measure corresponds to the limit cases of entropic risk measure. Observe that in Figure 37, following the blue strategy, the only payoffs that are obtained with positive probability were 40 and 0, which are also the limits of the risk entropy when휌tends to infinite values. On the other hand, in the red strategy, the only payoff obtained with positive probability is the payoff 1. Although the payoff 0 is possible since the play푎 휔 is compatible with every strategy, this outcome must be ignored since it is realized with probability 0. Theorem 30. Let 푋 be a random variable that ranges over R, and let 훽> 1. • The limit risk entropy of푋when휌tends to+∞exists and is equal to the pessimistic risk measure, that is, we have lim 휌!+∞ M 훽휌 [푋] = PM[푋]. •Similarly, the limit risk entropy of푋when휌tends to−∞exists and is equal to the optimistic risk measure, that is, we have lim 휌!−∞ M 훽휌 [푋] = OM[푋]. Proof.▶Limit in+∞ First, let us note that for every휌, we always haveM 훽휌 [푋] ≥ PM[푋]. Let now휀>0. We want to prove that there exists 휌 0 ∈ R such that for every 휌 ≥ 휌 0 , we have M 훽휌 [푋] ≤ PM[푋]+ 휀. 11.1. THE EXTREME RISK MEASURE165 Let us first notice that we have: M 훽휌 [푋] =− 1 휌 log 훽 ∫ 푥∈R 훽 −휌푥 dP(푋 = 푥) =− 1 휌 log 훽 ∫ 푥∈R 훽 −휌PM[푋] 훽 −휌(푥−PM[푋]) dP(푋 = 푥) = PM[푋]− 1 휌 log 훽 ∫ 푥∈R 훽 −휌(푥−PM[푋]) dP(푋 = 푥) = PM[푋]− 1 휌 log 훽 ∫ 푥≤PM[푋]+ 휀 2 훽 −휌(푥−PM[푋]) dP(푋 = 푥) + ∫ 푥≥PM[푋]+ 휀 2 훽 −휌(푥−PM[푋]) dP(푋 = 푥) ! ≤ PM[푋]− 1 휌 log 훽 ∫ 푥≤PM[푋]+ 휀 2 훽 −휌 휀 2 dP(푋 = 푥)+ 0 ! = PM[푋]− 1 휌 log 훽 P 푋 ≤ PM[푋]+ 휀 2 훽 −휌 휀 2 = PM[푋]− 1 휌 log 훽 P 푋 ≤ PM[푋]+ 휀 2 + 휀 2 . For 휌 large enough, this quantity is indeed smaller than PM[푋]+ 휀. ▶Limit in−∞ Let us first notice that for every훽,휌,푋, we have the equalityM 훽휌 [푋] =−M 훽(−휌) [−푋]. Thus, we can apply the previous result, and find: lim 휌!−∞ M 훽휌 [푋] = lim 휌!−∞ −M 훽(−휌) [−푋] =− lim 휌!+∞ M 훽휌 [−푋] =−PM[−푋] =− inf푥 ∈ R| P(−푋 ≤ 푥)> 0 =− inf푥 ∈ R| P(푋 ≥−푥)> 0 = sup푥 ∈ R| P(푋 ≥ 푥)> 0 = OM[푋]. □ 11.1.3 A technical lemma Before moving to our result, we give the following lemma, that simply rephrases a well-known result adapted to our context, but that will be extensively used in the sequel. Lemma 31. LetG ↾푣 0 be a game with two players, called푖and푗. We assume given a partition(푃,푂)of푖, 푗. Then, the quantity: inf 휎 푗 ∈Strat 푗 G ↾푣 0 sup 휎 푖 ∈Strat 푖 G ↾푣 0 X 푖 ( ̄ 휎) 166CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA can be computed in time푂(푚), where푚is the number of edges inG. Moreover, the infimum is reached with a positional strategy of player푗; and there is a positional strategy of player푖that realizes the supremum for every strategy of player 푗. Consequently, the optimality of positional strategies and the푂(푚)upper bound also hold in Markov decision processes, and in Markov chains; and, on the other hand, it holds when푗is a fictional player that represents a coalition of players who all have as unique objective to minimize player 푖’s risk measure. Proof.The푂(푚)upper bound holds by a slight adaptation of the classical attractor algorithm [AG11, Chapter 5.3]. Note that those algorithms run in time푂(푚+푛), where푛is the number of vertices; but here, we assumed that each vertex (except possibly푣 0 ) has at least one ingoing edge, hence푛 ≤ 푚+1 and푚+푛 =푂(푚). That algorithm immediately induces positional optimal strategies. Another way to obtain that second result, however, is the following: once the quantity푥 = inf 휎 푗 sup 휎 푖 X 푖 ( ̄ 휎)is known, strategies that realize the infimum and the supremum can be seen as optimal strategies in the Boolean zero-sum game in which player푖wants with positive probability (if they are optimist) or with probability 1 (if they are pessimist) to reach the set of terminals yielding them at least payoff푥(if푥>0) or to avoid the set of terminals yielding them less than payoff푥(if푥 ≤0). This is then a reachability game (seen either from player푖’s of from player푗’s perspective), and it is well-known [AG11] that in such a game, for both players, positional strategies suffice to maximize the probability of winning. In particular, if one has a strategy to win that game with positive probability, or with probability 1, there is also such a strategy that is positional.□ 11.2THE EXISTENCE OF XRSES We now answer a fundamental question about every notion of equilibrium, which is the condition under which it exists. We show that (stationary) XRSEs are guaranteed to exist in games with only non-negative rewards, similarly as ERSEs. But our proof does not rely on the same arguments, and we instead give a proof that goes along with an algorithm. Theorem 31. LetG ↾푣 0 be a game with only non-negative rewards, and let(푃,푂) be a partition ofΠ. Then, there exists a stationary XRSE inG ↾푣 0 . Moreover, there exists an algorithm that, given such a game, outputs the representation of such an XRSE in time푂(푚 2 푝), where푚is the number of edges, and푝the number of pessimistic players. Proof.▶Example Our algorithm generates an XRSE by constructing a decreasing sequence퐸 = 퐸 0 ,퐸 1 , . . .of sets of edges, and considering, for each푘, the stationary strategy profile that randomizes between all the outgoing edges in 퐸 푘 from all vertices. Let us illustrate it with the game depicted by Figure 40a, which involves two pessimists, player ◦ and player□. In that game, both players want to leave the cycle, but each of them would prefer the other player to leave. If we first consider the strategy profile that always randomizes between all the available edges, then both terminal vertices are reached with positive probability, and it is almost sure that one of them is reached: each player, as a pessimist, will therefore expect that they will leave the cycle first, and get payoff 1. Thus, their risk measure is 1. Then, player□(and symmetrically player ◦ ) has a profitable deviation by refusing to leave the cycle, and by always going back to the vertex푎: 11.2. THE EXISTENCE OF XRSES167 then, it is almost sure that player ◦ will eventually leave the cycle, and player□’s risk measure is now 2. Note that player ◦ cannot detect such a deviation of strategy, since she does not have access to the internal coins tossed by player □. Then, we remove the edge푏푡 2 (or푎푡 1 ). This results in a set of edges such that, if we consider again the strategy profile where both players always randomize, from each vertex, between the remaining outgoing edge, then player□gets risk measure 2, and player ◦ cannot get more than 1 by deviating. That new strategy profile is a (stationary) XRSE. ▶Algorithm Throughout this proof, for a given set of edges퐹 ⊆ 퐸, we writeG 퐹 for the game obtained from Gby removing all the edges that do not belong to퐹. In that game, we define ̄ 휎 퐹 as the stationary strategy profile that maps each vertex푣to some probability distribution whose support is퐹(푣). Note that the probabilities do not matter here: we are only interested in the support of the distribution of the strategy profile. 푎 푏 푡 1 : ◦ 1 □ 2 푡 2 : ◦ 2 □ 1 (a) Two pessimists 푐 푎 푏 푡 1 : ◦ 1 □ 2 푡 2 : ◦ 2 □ 1 (b) Two pessimists with a common coin 푐 푑 푒 푎 푏 푡 1 : ◦ 1 □ 2 푡 2 : ◦ 2 □ 1 푡 3 : ◦ 2 □ 2 (c) Two pessimists with a common coin and some temptation Figure 40: Some games involving two pessimistic players We proceed by presenting the algorithm, Algorithm 1, that takes as an input the gameG ↾푣 0 and the partition(푃,푂), and returns a subset퐹 ⊆ 퐸such that, as we will show, the strategy profile ̄ 휎 퐹 is always an XRSE. That algorithm defines a decreasing sequence퐸 0 ,퐸 1 , . . .of subsets of퐸, where 퐸 0 = 퐸. At each step푘, for each pessimist푖, it computes the risk measure푧 푘 푖 of player푖in ̄ 휎 퐸 푘 , and then the set푊 푘 푖 of vertices푣such that, from푣, whatever player푖does, that player almost surely gets a payoff smaller than or equal to푧 푘 푖 . If we have푣 0 ∈ 푊 푘 푖 for each푖, then the algorithm stops there and returns the set퐸 푘 (and we will show below that it means that ̄ 휎 퐸 푘 is an XRSE). Otherwise, we pick player푖such that푣 0 ∉푊 푘 푖 (a player who provably has a profitable deviation), and define퐸 푘+1 by removing all the edges accessible from 푣 0 leading from푉 \푊 푘 푖 to푊 푘 푖 . ▶Correctness A first quick invariant that we need to prove is the following one, which will guarantee that the gamesG 퐸 푘 and the strategies ̄ 휎 퐸 푘 are well-defined. 168CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA Algorithm 1 Exhibition of one stationary XRSE procedure Existence(G ↾푣 0 ,푃,푂 ) 푘 0 퐸 푘 퐸 while⊤ do Compute 퐴 푘 =푣 ∈ 푉 | 푣 is accessible from 푣 0 in(푉,퐸 푘 ) for all 푖 ∈ 푃 do Compute 푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푘 ) Compute푊 푘 푖 = n 푣 ∈ 푉 ∀휏 푖 ∈ Strat 푖 G 퐸 푘 ↾푣 0 , we have P ̄ 휎 퐸 푘 −푖 ,휏 푖 (휇 푖 ≤ 푧 푘 푖 )> 0 o end for if ∃푖 such that 푣 0 ∉푊 푘 푖 then Pick one such 푖 퐸 푘+1 퐸 푘 \((퐴 푘 \푊 푘 푖 )×푊 푘 푖 ) 푘 푘+ 1 else return 퐸 푘 end if end while end procedure Invariant 1. For each푘, each stochastic vertex푣, we have퐸(푣) ⊆ 퐸 푘 , and for each non-stochastic vertex 푣, we have 퐸(푣)∩ 퐸 푘 ≠∅. Proof.The set 퐸 0 = 퐸 trivially satisfies the invariant. Now, let us assume that퐸 푘 satisfies the invariant. At step푘, an edge is removed if and only if it goes from a vertex푢 ∈ 퐴 푘 \푊 푘 푖 to a vertex푣 ∈ 푊 푘 푖 . Consider a stochastic vertex푢 ∈ 퐴 푘 : if it has an edge that leads to vertex푣 ∈ 푊 푘 푖 , then whatever player푖plays from푢, with positive probability, the vertex푣is reached; and then, if the other players play the strategy profile ̄ 휎 퐸 푘 , then with positive probability, player푖gets the payoff푧 푘 푖 or less. Hence푢 ∈ 푊 푘 푖 , and the edge푢푣 is not removed, and remains in the set퐸 푘+1 . Similarly, if푢is not a stochastic vertex, but all its outgoing edges lead to a vertex that belong to푊 푘 푖 , then the vertex푢itself belongs to푊 푘 푖 , hence the outgoing edges of푢will not all be removed. The invariant is therefore still true at step푘+1, and by induction, is true for all 푘 .□ Each step of the algorithm is then also properly defined. Moreover, we have termination. Proposition 2. Algorithm 1 terminates. Proof.With Invariant 1, we now know that Algorithm 1 successfully constructs a sequence퐸 0 ,퐸 1 , . . . of sets of edges until it stops and returns the last of those sets. Termination is an immediate consequence of the fact that this sequence is decreasing. Indeed, for each step푘at which nothing is returned, there exists a player푖with푣 0 ∉푊 푘 푖 . On the other hand, the set푊 푘 푖 is necessarily accessible from 푣 0 : Lemma 32. The set푊 푘 푖 is nonempty, and accessible from 푣 0 in the graph(푉,퐸 푘 ). Proof. If푧 푘 푖 is obtained by reaching a terminal vertex푡, then we have푡 ∈ 푊 푘 푖 , and푡is accessible from푣 0 . If now푧 푘 푖 =0 is obtained by reaching no terminal vertex, then when following ̄ 휎 푘 , with positive probability, no terminal is reached. Then, there is in particular a vertex푢that has positive probability of being visited infinitely often. And when playing ̄ 휎 푘 from푢, the 11.2. THE EXISTENCE OF XRSES169 probability that some terminal is ever reached is actually 0, since if it was some constant푞>0, then the probability of visiting푢infinitely often would belim ℓ (1−푞) ℓ =0. In other words, no terminal vertex is accessible from푢 in(푉,퐸 푘 ), and then, we have푢 ∈ 푊 푘 푖 .□ Now, along a play that starts from푣 0 ∉푊 푘 푖 and visits푊 푘 푖 , there exists at least one edge that goes from a vertex that does not belong to푊 푘 푖 , to a vertex that does. Such an edge is then removed in the set퐸 푘+1 , which is therefore strictly included in the set퐸 푘 . This holds for every푘, ensuring termination.□ We now know that the algorithm terminates, i.e., constructs a finite decreasing sequence퐸 = 퐸 0 ,퐸 1 , . . .,퐸 푛 , and then returns the set퐸 푛 , as a succinct representation of the stationary strategy profile ̄ 휎 퐸 푛 . What remains to be proven is that this strategy profile is an XRSE. Before proving that it is an XRSE in the gameG ↾푣 0 , we first prove that it is one in the gameG 퐸 푛 ↾푣 0 , i.e., when the edges that have been removed cannot be used to deviate. Proposition 3. The strategy profile ̄ 휎 퐸 푛 is an XRSE in the gameG 퐸 푛 ↾푣 0 . Proof. Consider a player푖, and a deviation휎 ′ 푖 of player푖from the strategy profile ̄ 휎 퐸 푛 in the game G 퐸 푛 ↾푣 0 . Let 푥 = X 푖 ( ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 ). If player푖is an optimist. If푥 =0, then since all rewards are non-negative, we have 푥 ≤ X 푖 ( ̄ 휎 퐸 푛 ). If푥>0, then the payoff푥is obtained by reaching a terminal vertex푡. But then, that terminal vertex is accessible from푣 0 in the graph(푉,퐸 푛 ), and is therefore also reached with positive probability when all players follow the strategy profile ̄ 휎 퐸 푛 . Hence, again, the inequality 푥 ≤ X 푖 ( ̄ 휎 퐸 푛 ). If player푖is a pessimist. Then, since the algorithm terminated at step푛, player푖is such that푣 0 ∈ 푊 푘 푖 . The strategy휎 ′ 푖 , like every strategy휏 푖 for player푖, satisfies therefore the inequality P ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 (휇 푖 ≤ 푧 푛 푖 )> 0. Consequently, we have 푥 ≤ 푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푛 ). In both cases, the deviation 휎 ′ 푖 is not profitable, hence the conclusion.□ Let us now prove that putting back the removed edges does not change that result, and therefore conclude the correctness proof. Proposition 4. The strategy profile ̄ 휎 퐸 푛 is an XRSE in the gameG ↾푣 0 . Proof. Let푖be a player, and let휎 ′ 푖 be a deviation from ̄ 휎 퐸 푛 for player푖inG ↾푣 0 . Since ̄ 휎 푛 is stationary, we can assume that휎 ′ 푖 is positional by Lemma 31. If the strategy휎 ′ 푖 uses only edges of퐸 푛 , then it can be considered as a deviation from ̄ 휎 퐸 푛 in the gameG 퐸 푛 , hence by Proposition 3, it is not a profitable deviation. Let us now assume that the strategy휎 ′ 푖 uses an edge that does not belong to퐸 푛 , i.e. there exists a vertex푣that is visited with positive probability in the strategy profile( ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 ) and an edge 푣푤 ∈ 퐸\ 퐸 푛 such that푤 = 휎 ′ 푖 (푣) . Since only edges controlled by pessimists have been removed, we can immediately deduce that player 푖 is a pessimist. Now, among such edges, let us choose one whose removal occurred the earliest, that is, let us choose it in order to minimize the index푘such that푣푤 ∈ 퐸 푘 \ 퐸 푘+1 . Thus, in the strategy profile ( ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 ), it is almost sure that only edges of 퐸 푘 are used. 170CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA The fact that the edge푢푣has been removed at step푘means that we had푢 ∉푊 푘 푖 and푣 ∈ 푊 푘 푖 . Thus, from the vertex푣, if player푖uses only edges of퐸 푘 (which is the case when they follow휎 ′ 푖 ), and if the other players follow the strategy profile ̄ 휎 퐸 푘 −푖 , player푖gets the payoff푧 푘 푖 or less with positive probability. Let us show that it is also the case when the other players follow the strategy profile ̄ 휎 퐸 푛 instead of ̄ 휎 퐸 푘 . Lemma 33. From the vertex 푣, we have P ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 (휇 푖 ≤ 푧 푘 푖 )> 0. Proof. We proceed by proving that when the players follow, from the vertex푣, the strategy profile( ̄ 휎 퐸 푘 ,휎 ′ 푖 ) , there is a positive probability that player푖gets the payoff푧 푘 푖 or less and that the set푊 푘 푖 is never left. Indeed, if in that strategy profile there is a positive probability that player푖gets a payoff smaller than or equal to푧 푘 푖 by reaching a terminal vertex푡that yields such a payoff, then we have푡 ∈ 푊 푘 푖 , and with positive probability the terminal vertex푡is reached without leaving푊 푘 푖 . Similarly, if such a payoff is obtained by reaching no terminal vertex, and therefore getting payoff 0, then, using a reasoning that has already been used above, with positive probability a vertex푤is reached without leaving푊 푘 푖 , such that from푤, no terminal vertex is accessible anymore in(푉,퐸 푘 ); and, therefore, all the vertices accessible from푤in that graph belong to 푊 푘 푖 , hence once 푤 is reached it is almost sure that푊 푘 푖 is never left. Then, the set퐸 푘+1 was defined so that푊 푘 푖 is no longer accessible from푣 0 in the graph (푉,퐸 푘+1 ). Therefore, those vertices are not accessible at any stepℓ> 푘, and therefore no outgoing edge of a vertex of푊 푘 푖 is ever removed in the sequel, i.e.퐸 푛 ∩(푊 푘 푖 ×푉) = 퐸 푘 ∩(푊 푘 푖 ×푉). Consequently, since휎 ′ 푖 uses only edges of퐸 푘 , when the strategy profile( ̄ 휎 푛 ,휎 ′ 푖 ) is played from 푣 , it is also true that with positive probability player 푖 gets the payoff 푧 푘 푖 , or less.□ This lemma proves that we haveX 푖 ( ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 ) ≤ 푧 푘 푖 . To conclude that the deviation ̄ 휎 ′ 푖 is not profitable, we still need to prove that this quantity푧 푘 푖 is smaller than or equal to (actually strictly smaller) the risk measureX 푖 ( ̄ 휎 푛 ). That inequality is an immediate consequence of the following lemma. Lemma 34. For every pessimistic player푗, the sequence(푧 ℓ 푗 )of player푗’s risk measures is non- decreasing. Proof.Letℓbe a step index, and let us prove that we have푧 ℓ 푗 < 푧 ℓ+1 푗 . The quantity푧 ℓ+1 푗 is the pessimistic risk measure of player푗in the strategy profile ̄ 휎 퐸 ℓ+1 : there is therefore a positive probability that player푗gets the payoff푧 ℓ+1 푗 when that strategy profile is followed. Player푗 obtains that payoff either by reaching a terminal vertex to which that payoff is assigned, or by reaching no terminal vertex at all. Let us first show that the second case is actually impossible. If player푗gets the payoff 푧 ℓ+1 푗 =0 by reaching no terminal vertex, with the same reasoning as above, there is, in particular, a vertex푣that has positive probability of being visited infinitely often when ̄ 휎 ℓ+1 is played from푣 0 , and therefore such that no terminal vertex is accessible from푣in(푉,퐸 ℓ+1 ). But then, let푗 ′ be the player that was controlling the edges that were removed at stepℓ. Let us consider a strategy휏 푗 ′ of player푗 ′ that uses only edges of퐸 ℓ . Then, when the strategy profile( ̄ 휎 ℓ −푗 ′ ,휏 푗 ′ ) is played from the vertex푣, it will almost surely be true that either no terminal vertex is reached, leading to the payoff 0, or an edge of퐸 ℓ \퐸 ℓ+1 is taken, leading therefore to a vertex of푊 ℓ 푖 , and to a risk measure of푧 ℓ 푗 ′ or less. Thus, since all rewards are non-negative and therefore푧 ℓ 푗 ′ ≥0, 11.3. CONSTRAINED EXISTENCE PROBLEM171 the vertex푣belongs to the set푊 ℓ 푗 ′ , which is impossible since it should then have been made unaccessible in the graph(푉,퐸 ℓ+1 ). Therefore, player푗gets payoff푧 ℓ+1 푗 by reaching a terminal giving them that payoff, which means that such a terminal is accessible from푣 0 in the graph(푉,퐸 ℓ+1 ). Then, it is also accessible from푣 0 in the graph(푉,퐸 ℓ ), and therefore it is reached with positive probability when following the strategy profile ̄ 휎 퐸 ℓ . Consequently, we have 푧 ℓ 푗 ≤ 푧 ℓ+1 푗 .□ Consequently, with푗 = 푖, we have푧 푘 푖 ≤ 푧 푛 푖 , and thereforeX 푖 ( ̄ 휎 퐸 푛 −푖 ,휎 ′ 푖 ) ≤ 푧 푘 푖 ≤ 푧 푛 푖 = X 푖 ( ̄ 휎 퐸 푛 ) . The strategy휎 ′ 푖 is not a profitable deviation from the strategy profile ̄ 휎 퐸 푛 , which is therefore a (stationary) XRSE.□ ▶Complexity We finally show that Algorithm 1 runs with time푂(푚 2 푝). At each iteration of the while loop, at least one edge is removed; we therefore have at most푚 iterations of said loop. Now, during the푘 th iteration, the algorithm computes the set퐴 푘 , and for each pessimistic player푖, the algorithm also computes the quantity푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푘 ), and then the set푊 푘 푖 . All of those computations can be done in time푂(푚)using Lemma 31. Since there are푝players, this step therefore takes푂(푚푝) time. Finally, checking whether푣 0 ∈ 푊 푘 푖 for each player푖takes time푂(푝), and removing all the edges leading to푊 푘 푖 to define the set퐸 푘+1 takes time푂(푚). Hence, the complexity of the algorithm is 푂(푚 2 푝).□ Like in the case of ERSEs, we conjecture that existence, and even existence of a stationary strategy profile, remain true in the general case. Conjecture 3. LetG ↾푣 0 be a simple stochastic game, and a partition(푃,푂)of the player setΠ. Then, there exists a (stationary) XRSE inG ↾푣 0 . 11.3CONSTRAINED EXISTENCE PROBLEM We now study the computational complexity of the constrained existence problem of XRSEs. The main result of this section is the following theorem, which proves that, contrary to the same problem with ERSEs, it is a decidable fragment of the constrained existence of RSEs. Theorem 32. The constrained existence problem for XRSEs isNP-complete and isNP-hard even when all players are pessimistic and all rewards are non-negative. To prove Theorem 32, we first prove Theorem 35, which shows that if there is an XRSE then there is one that uses finite memory. We show that this finite-memory strategy profile can be described only using polynomial size, which in turn proves NP membership (Lemma 36). Later, we consider the problem of XRSEs when the players are restricted to pure, stationary or positional strategies. We show that in all the above cases, the problem remainsNP-complete. The upperbound is similar to the general case, but the lower bound is shown in Lemma 38 by showing a reduction from the problem 3Sat to the constrained existence problem. Theorem 33. The constrained existence problem of XRSEs is alsoNP-complete when the players are restricted to positional, stationary, or pure strategies. 172CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA Finally, we show in the following theorem that when all players are optimistic, the problem can even be solved in polynomial time. Theorem 34. The constrained existence problem of XRSE is P-complete when all players are optimists, that is, when we have 푃 =∅. We dedicate the rest of the section to proving these three results. 11.3.1 Membership in NP NP membership is a consequence of the fact that when an XRSE exists, there also exists one with the same extreme risk measures that uses finite memory, with a number of states that is polynomial in the size of the game. Let us therefore illustrate, with examples, how and why memory is required in such XRSEs. We consider the following constrained existence question and analyse the same question on three example graphs. (∗) Is there an XRSE in the game in which both players have exactly the risk measure 1? Game in Figure 40a. Let us consider the game in Figure 40a again. The answer to Question(∗) here is no. Intuitively, such an XRSE would require at least two plays of positive probability: one that ends in푡 1 , and one that ends in푡 2 . For those two plays to occur with positive probability, the strategy profile must proceed to a randomized action at vertex푎or푏: i.e., one of the players, at some point of time, must toss a coin to give the payoff 1 to player ◦ in one case and to□in the other. But then, since that player is the only one that can see that coin, they have a profitable deviation by lying about the outcome, and always choose the option that gives them the best payoff. More randomization will not help: as long as one of the players randomizes, be it once, several times, or infinitely often, they have an incentive to deviate and stay in the cycle and wait until the other player leaves. Game in Figure 40b. Consider now a slight modification, as shown in Figure 40b. There, the first player that plays is determined at random by the edge that is taken from an initial stochastic vertex. The answer to Question(∗)for this game is yes. The random choice on which player gets payoff 1 is decided by the stochastic vertex. Since both players can see which edge is taken from there, this serves as a source of unbiased randomness based on which they act. For example, it can be decided that if the play visits the vertex푎immediately after푐, then player ◦ must visit the terminal푡 1 , and similarly, if it visits the vertex푏, then player□must visit the terminal푡 2 . If the edge푎푏is taken, player□punishes player ◦ by always going back to푎, and vice versa. In other words, the stochastic vertex provides the players with a common coin. Game in Figure 40c.Finally, consider the game depicted in Figure 40c. Here, both players ◦ and□ have the possibility of deviating to a terminal with payoff 2 in one play. The stationary strategy profile in which from vertex푑, player ◦ goes from푑to푎and then to푡 1 , and in which player□goes from푒to푏 and then to푡 2 , is therefore not an XRSE: both players have a profitable deviation that goes to terminal 푡 3 . But the answer to Question(∗)still remains yes! If, from vertex푑, player ◦ goes from vertex푑to 푎and then to푏, from which player□leaves to푡 2 , and symmetrically, from vertex푒, player□goes to푎 through푏from which player ◦ goes to푡 1 , then that strategy profile is an XRSE, in which everyone gets the risk measure 1. This is because a player has a profitable deviation only if they can play in a way that guarantees them a risk measure better than 1, i.e., that guarantees them almost surely a payoff greater 11.3. CONSTRAINED EXISTENCE PROBLEM173 than 1. If there remains a play that occurs with nonzero probability and offers a lower reward, then the player does not increase their risk measure. Therefore, an XRSE where player ◦ gets the extreme risk measure 1 only needs to have one play with positive probability in which she gets the payoff 1, and in which she cannot increase her payoff by deviating. We say that such a play anchors that player. In our example, the play 푐푒푏푎푡 1 anchors player ◦ . We see in this last example that memory is required to remember either the subset of players that are being anchored, or if a player has deviated from the strategy and must be punished. Given one or more players that are being anchored, the memory state of any of the players does not change unless either a player deviates or, more importantly, randomization occurs. When randomization occurs, the set of players that are anchored in each of the plays is a subset of the set of players anchored before this play split. In our examples, the set of players that are anchored at푐is both ◦ and□, and it immediately splits. After the splits, when we have only one player to anchor, the players can follow a positional strategy profile; and similarly when one player deviates and must be punished, the players can follow a positional strategy profile. Following this intuition, we prove a theorem that bounds the amount of memory required by a strategy to a polynomial in the number of players and vertices in the game. Theorem 35. Let ̄ 휎be an XRSE in the gameG ↾푣 0 with푛vertices and푝players, and a partition(푃,푂)of the setΠ. Then, there exists a finite-memory XRSE ̄ 휎 ★ with at most 3푛푝−2푛+ 푝+1 many memory states, and such that X( ̄ 휎 ★ ) = X( ̄ 휎). Furthermore, if ̄ 휎 is pure, then there is such a strategy profile ̄ 휎 ★ that is pure. Proof.The XRSE ̄ 휎being given, we first define the labelingΛwhich, intuitively, defines which players are being anchored after every given history. ▶Definition ofΛ. Lemma 35 (The labelingΛ). There exists a labelingΛthat maps each historyℎ ∈ HistG ↾푣 0 compatible with ̄ 휎to a setΛ(ℎ) ⊆Π, such that for each suchℎ, if we write푣 1 , . . .,푣 푘 = Supp( ̄ 휎(ℎ)), we have the following properties. 1. If the vertexlast(ℎ)is stochastic, or belongs to some player푖 ∉Λ(ℎ), then the setsΛ(ℎ푣 1 ), . . .,Λ(ℎ푣 푘 ) form a partition of the setΛ(ℎ). 2.If the vertexlast(ℎ)belongs to some player푖 ∈Λ(ℎ), then the setsΛ(ℎ푣 1 )\푖, . . .,Λ(ℎ푣 푘 )\푖 form a partition of the setΛ(ℎ)\푖, and player 푖 belongs to all setsΛ(ℎ푣 1 ), . . .,Λ(ℎ푣 푘 ). 3. For each optimistic player 푖 ∈Λ(ℎ), we have X 푖 ( ̄ 휎 ↾ℎ ) = 푧 푖 . 4. For each pessimistic 푖 ∈Λ(ℎ), for all strategies 휏 푖 of player 푖, we have X 푖 ( ̄ 휎 −푖↾ℎ ,휏 푖 ) ≤ 푧 푖 . 5. If there is a successor푣 ℓ such thatΛ(ℎ푣 ℓ ) =Λ(ℎ), then all other successors푣 ℓ ′ are such that X 푖 ( ̄ 휎 ↾ℎ푣 ℓ ′ )< 푧 푖 for each optimist푖 ∈Λ(ℎ), and there exists휏 푖 withX 푖 ( ̄ 휎 −푖↾ℎ푣 ℓ ′ ,휏 푖 )> 푧 푖 for each pessimist 푖 ∈Λ(ℎ). Proof.We define the labelingΛ inductively. Base case. First, on the one-vertex history푣 0 , we defineΛ(푣 0 ) =Π: at the start, all players must be anchored. Let us notice that the history푣 0 satisfies Property 3, which states that the optimists get the optimistic expectation they are supposed to get, and Property 4, which states that 174CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA the pessimists have no profitable deviations. The other properties will be checked in the inductive case. Inductive case. SupposeΛ(ℎ푣)has already been defined, whereℎ푣is a history compatible with the strategy profile ̄ 휎, and that the five properties are satisfied byΛon all histories on which it is already defined. Let푤 1 , . . .,푤 푘 be the successors of푣that are chosen by the strategy ̄ 휎(ℎ푣) with non-zero probability, that is, the support of ̄ 휎(ℎ푣). If푘 =1, then we defineΛ(ℎ푣푤 1 ) =Λ(ℎ푣). Note that Properties 1, 2, and 5 are immediately satisfied, and that Properties 3 and 4 are satisfied by induction hypothesis. If푘>1, we need to partition the setΛ(ℎ)between the푘successors. To do so, we will use the following result. Sublemma 8. For each player 푖 ∈Λ(ℎ푣) that does not control the vertex 푣, the following holds. • If player푖is an optimist, andX 푖 ( ̄ 휎 ↾ℎ푣 ) = 푧 푖 , then there is at least one successor푤 ℓ such that we have X 푖 ( ̄ 휎 ↾ℎ푣푤 ℓ ) = 푧 푖 . •If player푖is a pessimist, then there is a successor푤 ℓ ∈ Supp( ̄ 휎(ℎ푣)), such that for every strategy 휏 ℓ 푖 from 푤 ℓ , we have X 푖 ( ̄ 휎 −푖↾ℎ푣푤 ℓ ,휏 푖 ) ≤ 푧 푖 . Proof. The first case follows from Property 3 in the induction hypothesis. As for the second case, we proceed by contradiction. Let us assume that for each푤 ℓ , there exists a strategy휏 ℓ 푖 such thatX 푖 ( ̄ 휎 −푖↾ℎ푣푤 ℓ ,휏 ℓ 푖 )> 푧 푖 . Then, the strategy휏 푖 defined by휏 푖↾푣푤 ℓ = 휏 ℓ for eachℓis such that X 푖 ( ̄ 휎 −푖↾ℎ푣 ,휏 푗 )> 푧 푖 , which is impossible sinceΛ(ℎ푣) is assumed to satisfy Property 4. □ We define each setΛ(ℎ푣푤 ℓ ) by iterating through each element ofΛ(ℎ푣) as follows: • Initialization. For all 푤 ℓ , declareΛ(ℎ푣푤 ℓ ) =∅. • Iteration over players. Consider each player 푖 ∈Λ(ℎ푣) sequentially and proceed as follows: –if player 푖 controls the vertex 푣 , then add 푖 to every setΛ(ℎ푣푤). –If player푖does not control푣, then add푖to the setΛ(ℎ푣푤 ℓ )where푤 ℓ is defined by Sublemma 8. Not that the first four properties are thus guaranteed to be satisfied. Moreover, when there are several successors푤 ℓ possible, we always favour those such that, at the moment where the decision is taken, the setsΛ(ℎ푣푤 ℓ ) are the smallest. This suffices to guarantee Property 5. □ ▶Construction of the strategy profile ̄ 휎 ★ Based onΛas in Lemma 35, we construct a finite-memory strategy profile. Formally, the definition of the strategy profile ̄ 휎 ★ will be done by defining its memory structure. The memory states are the following: •for each player 푖, the state punish 푖 ; • for each vertex푣and each subset퐴 ⊆Πof players such that there existsℎwith퐴 =Λ(ℎ), the state anchor 퐴푣 ; •the state anchor Π⊥ . 11.3. CONSTRAINED EXISTENCE PROBLEM175 We now define the transitions from each of those states and from each vertex. Observe that as long as the memory state does not change, the memory structure follows a positional strategy profile. So, we describe such stationary strategy profiles and also describe when the memory state changes. For each set퐴 ⊆Π, we write푊 퐴 for the set of vertices that ̄ 휎may visit while anchoring the set퐴, i.e., the set of vertices 푣 such that there exists a history ℎ푣 withΛ(ℎ푣) = 퐴. Punishing memory statespunish 푖 . First, let us define what ̄ 휎 ★ does when in statepunish 푖 , for some player푖. Those memory states will correspond to the punishing strategies, followed when player푖deviates from the assigned strategy with the other memory states. By Lemma 31, there is a positional strategy profile ̄ 휏 †푖 −푖 that minimizes, from every vertex of the game, the payoff that player푖 can enforce. In addition, we pick an arbitrary positional strategy휎 †푖 푖 . Then, when the strategy profile ̄ 휎 ★ is in the memory statepunish 푖 and reads a vertex푣, the memory structure outputs the vertex ̄ 휏 †푖 (푣) and remains in the same memory state punish 푖 . Anchoring states with no player to anchor.Let us now define what happens in memory state anchor ∅푣 . Consider the objective of achieving a payoff vector that has positive probability of being achieved in ̄ 휎, while visiting only vertices of푊 ∅ . Using classical attractor-based proofs (or Lemma 31), there exists a positional strategy profile that achieves that objective with probability 1 from every vertex from which that is possible: let us call it ̄ 휏 ï∅ . Then, when in memory stateanchor ∅푣 and reading the vertex 푤 , we distinguish two cases. •If푣푤is an edge that is compatible with the strategy profile ̄ 휏 ï∅ , then the strategy profile ̄ 휎 ★ outputs the vertex ̄ 휏 †푖 (푤), where푖is the player controlling푣, and shifts to the memory state punish 푖 . •Otherwise, it outputs the vertex ̄ 휏 ï∅ (푤) and moves to the memory state anchor ∅푤 . Anchoring state with one player to anchor.We can now move to singletons, and define what happens in the states of the formanchor 푖푣 . In such a state, we define a strategy profile that gives player푖exactly the extreme risk measure푧 푖 . More precisely, we want player푖to receive payoff푧 푖 with positive probability, and never leave the set of vertices푊 푖 with probability 1. Using Lemma 31, there exists a positional strategy profile ̄ 휏 ï푖 that satisfies that property from every vertex from which it is possible. Note that this objective is in particular satisfiable, and therefore satisfied by ̄ 휏 ï푖 , from every vertex푣 ∈ 푊 푖 . Similar to the previous step, we define the strategy profile ̄ 휎 ★ in the states of the form anchor 푖푣 so that it follows the strategy profile ̄ 휏 ï푖 , remembers the last vertex that was visited and uses that memory to switch to the corresponding punishing state when some player 푗 deviates. Anchoring states with two or more players to anchor. Now, let us consider the states of the formanchor 퐴푣 , where퐴has cardinality at least 2. The existence ofanchor 퐴푣 implies that there is a historyℎsuch thatΛ(ℎ) = 퐴. Moreover, since each randomization splits the label of histories in sets that have at most one element in common (Properties 1 and 2), there is only one side of each split that can contain퐴, which implies that among such historiesℎ, we can choose one that is a prefix of all others. After historyℎ, the histories labelled by퐴form a sequenceℎ,ℎ푣 1 ,ℎ푣 1 푣 2 , . . .which may be infinite, end in a terminal vertex, or end with a new split. 176CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA • If that sequence ends with a split, then there is a longest history ℎ푣 1 . . .푣 푞 withΛ(ℎ푣 1 . . .푣 푞 ) = 퐴 and푘 ≥2 vertices푤 1 , . . .,푤 푘 ∈ Supp( ̄ 휎(ℎ푣 1 . . .푣 푞 ))such that we haveΛ(ℎ푣 1 . . .푣 푞 푤 ℓ ) ≠ ∅,퐴 for eachℓ. We can then define a simple historyℎ ′ 푣 푞 that also goes from the vertexlast(ℎ)to the vertex푣 푞 , withOcc(ℎ ′ 푣 푞 ) ⊆ Occ(last(ℎ)푣 1 . . .푣 푞 ). We then define the strategy profile ̄ 휎 ★ in each stateanchor 퐴푣 so that it follows the historyℎ ′ 푣 푞 and remembers the last vertex visited, and switches to the state punish 푖 and follows the strategy profile ̄ 휏 †푖 when a given player 푖 deviates and takes an edge that they are not supposed to take. Moreover, when an edge is taken that does belong to the historyℎ ′ 푣 푞 , but not because a player deviated (it is then necessarily because of a stochastic vertex), the memory switches to the stateanchor ∅푣 (where푣is the last vertex seen) and immediately follows the corresponding strategy. Finally, when the vertex푣 푞 is reached and the memory is in stateanchor 퐴last(ℎ ′ ) , the strategy profile ̄ 휎 ★ chooses randomly between the edges푣 푞 푤 1 , . . . , and푣 푞 푤 푘 , all with positive probability. Such action will often be referred to as a split. •If that sequence is infinite or ends in a terminal vertex, then휋 퐴 = 푣 1 푣 2 . . .is a play, and satisfies Λ(ℎ휋 퐴 <푘 ) = 퐴for each푘. We can then consider a play휋 퐴★ withOcc(휋 퐴★ ) ⊆ Occ(휋 퐴 )and Inf(휋 퐴★ ) ⊆ Inf(휋 퐴 )that is either a simple path from휋 퐴 0 to the terminal reached by휋 퐴 , or, if휋 퐴 is infinite, a simple lasso (let us recall that those are defined as plays of the formℎ ′ 푐 휔 , where the historyℎ ′ 푐is simple). We can moreover choose휋 퐴★ so that the set of vertices visited infinitely often (if there are any) in휋 퐴★ is included in the set of vertices visited infinitely often in휋 퐴 . Then, we can define ̄ 휎 ★ in the states of the formanchor 퐴푣 as following the play휋 퐴★ , and remembering the last vertex seen. When a player푖deviates and takes an edge that should not be taken, the memory switches to the statepunish 푖 and follows the strategy profile ̄ 휏 †푖 . Finally, when an edge is taken that does not belong to휋 퐴★ but does not correspond to a deviation either, we switch to the state anchor ∅푤 where 푤 is the last vertex seen, and to the corresponding strategy profile. Initialization. The strategy profile ̄ 휎 ★ has the stateanchor Π⊥ as the initial memory state. In this state, it behaves exactly as in any state of the formanchor Π푣 , but without having memorized a last visited vertex푣, since there is no such vertex. From that memory state therefore, it necessarily reads the vertex 푣 0 , and starts acting as described in the previous case. The pure case.In this construction, the vertices on which the strategy profile ̄ 휎 ★ proceeds to an actual randomization (i.e., the vertices푣such that there exists a historyℎ푣such that the support of the distribution ̄ 휎 ★ (ℎ푣)contains more than one element) are vertices on which ̄ 휎also proceeds to such a randomization. Therefore, if ̄ 휎is pure (i.e., if randomizations occur only on stochastic vertices), so is ̄ 휎 ★ . ▶A combinatorial break: counting states Now that the strategy ̄ 휎 ★ is defined, let us bound the memory it uses. There are, obviously, exactly 푝 states of the form punish 푖 , and one state anchor Π⊥ . To prove that there are at most 3푛푝− 2푛 states of the formanchor 퐴푣 , we need to prove that there are at most 3푝−2 sets퐴such that there is a history ℎ withΛ(ℎ) = 퐴. Let us callΛ-anchored all such sets퐴. By analogy with strategies, we writeΛ ↾ℎ푣 for the labeling that maps each historyℎ ′ ∈ HistG ↾푣 compatible with ̄ 휎 ↾ℎ푣 to the setΛ(ℎ ′ ), and we will also use the 11.3. CONSTRAINED EXISTENCE PROBLEM177 notion of anchoredness for each of those labelingsΛ ↾ℎ푣 . We proceed by proving the following stronger result. Proposition 5. For every historyℎcompatible with ̄ 휎 , ifΛ(ℎ)contains at least two elements, then there are at most 3cardΛ(ℎ)− 2 sets that areΛ ↾ℎ -anchored. Proof.For each historyℎ, we write푓(ℎ)for the number ofΛ ↾ℎ -anchored sets that have cardinal at least 2. There arecardΛ(ℎ)+1 subsets ofΛ(ℎ)that have cardinality 0 or 1: the result will therefore be proved if we prove 푓(ℎ) ≤ 2cardΛ(ℎ)− 3. The proof goes by induction on푚 =Λ(ℎ) ≥ 2. Base case. If푚 =2, the setΛ(ℎ)is a pairΛ(ℎ) =푖, 푗. Then, since theΛ ↾ℎ -anchored sets are all subsets ofΛ(ℎ), the only set of cardinality at most 2 that isΛ ↾ℎ -anchored is the pair푖, 푗 itself, hence we have 푓(ℎ) = 2× 2− 3 = 1, as desired. Inductive case. If푚>2, and if we assume that the result is true for every historyℎ ′ with 2≤ cardΛ(ℎ ′ ) ≤ 푚−1, then let푣 1 , . . .,푣 푘 ⊆ Supp( ̄ 휎(ℎ))be the set of possible next vertices푣 such that cardΛ(ℎ푣) ≥ 2. If푘 =1, i.e., ifΛ(ℎ푣 1 ) =Λ(ℎ), then we have푓(ℎ) = 푓(ℎ푣 ℓ )and the result forℎwill be proved if we prove it forℎ푣 1 . Following that reasoning, we can extend the historyℎuntil we are not in that case: if we always are, then the onlyΛ ↾ℎ -anchored sets areΛ(ℎ)itself, and possibly the empty set and some singletons, hence푓(ℎ) =1 and the result is immediate. We can therefore assume that 푘> 1. Let푖be the player controlling the vertexlast(ℎ); we set푖 =⊥iflast(ℎ)is a stochastic vertex. Then, by Properties 1 and 2 of Lemma 35 guaranteed during the construction ofΛ, the sets Λ(ℎ푣 1 )\푖, . . .,Λ(ℎ푣 푘 )\푖form a partition ofΛ(ℎ)\푖. Therefore, no set of cardinality 2 or more can be simultaneouslyΛ(ℎ푣 ℓ )-anchored andΛ(ℎ푣 ℓ ′ ) -anchored forℓ ≠ ℓ ′ , hence the equality 푓(ℎ) =1+ Í ℓ 푓(ℎ푣 ℓ ). Now, since each setΛ(ℎ푣 ℓ )has at least 2 and less than푚elements, we can apply the induction hypothesis to deduce: 푓(ℎ) ≤ 1+ 푘 ∑︁ ℓ=1 (2cardΛ(ℎ푣 ℓ )− 3). Moreover, we have Í ℓ cardΛ(ℎ푣 ℓ ) ≤ 푚+ 푘 − 1 (each element ofΛ(ℎ)occurs in one of the sets Λ(ℎ푣 ℓ ), except possibly one that would occur in all of them), hence the inequality above becomes: 푓(ℎ) ≤ 1+ 2(푚+푘− 1)− 3푘 = 2푚− 1−푘 and since we have assumed 푘 ≥ 2, we obtain 푓(ℎ) ≤ 2푚− 3.□ Let us recall thatΛ(푣 0 ) =Π. As a particular case of this result, we obtain, if푝 ≥2, that there are at most 3푝−2 sets that areΛ-anchored, as desired. In the case푝 =1, the gameG ↾푣 0 is an MDP, and using Lemma 31, we can immediately construct ̄ 휎 ★ as a positional strategy (which has therefore 1≤ 3푛× 1− 2푛+ 1+ 1 memory states) with the same risk measure as ̄ 휎 . 178CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA ▶The strategy profile ̄ 휎 ★ has the desired extreme risk measures. We now show that X( ̄ 휎 ★ ) = X( ̄ 휎). Let us recall that we defined ̄ 푧 = X( ̄ 휎). Proposition 6. The strategy profile ̄ 휎 ★ satisfies the equality X( ̄ 휎 ★ ) = ̄ 푧. Proof.Let푖be a player: we want to prove thatX( ̄ 휎 ★ ) = 푧 푖 . Let us first see how푧 푖 has positive probability of being obtained in the strategy profile ̄ 휎 , and we will then show that no larger (respectively smaller) payoff, if푖is optimistic (respectively pessimistic), has a positive probability of being obtained with the same strategy profile. Player푖gets payoff푧 푖 with positive probability. Let퐴 ⊆Πbe one of the smallest sets (for the inclusion relation) containing푖such that there exists a historyℎ푢withΛ(ℎ푢) = 퐴. Then, by construction ofΛ, there exists a finite sequence of setsΠ = 퐴 0 ,퐴 1 , . . .,퐴 푚 = 퐴and of histories ℎ 1 푣 1 푤 1 , . . .,ℎ 푚 푣 푚 푤 푚 where for each푘, the historyℎ 푘+1 starts from푤 푘 , the historyℎ 1 푣 1 . . .ℎ 푘 푣 푘 푤 푘 is compatible with ̄ 휎 , and we haveΛ(ℎ 1 푣 1 . . .ℎ 푘 푣 푘 ) = 퐴 푘−1 andΛ(ℎ 1 푣 1 . . .ℎ 푘 푣 푘 푤 푘 ) = 퐴 푘 . We can then write ℎ푢 =ℎ 1 푣 1 . . .ℎ 푚 푣 푚 푤 푚 . Consider the strategy profile ̄ 휎 ★ , which initially follows the positional strategy profile ̄ 휏 ïΠ . This strategy profile generates, with nonzero probability, a historyℎ ′ 1 푣 1 starting from vertex푣 0 to vertex 푣 1 , based on our construction. From that vertex푣 1 , it proceeds to a randomized action and, with positive probability, moves to the vertex푤 1 and switches to the positional strategy profile ̄ 휏 ï퐴 1 , and so on: there is, therefore, a historyℎ ′ 1 푣 1 ℎ ′ 2 푣 2 . . .ℎ ′ 푚 푣 푚 푤 푚 that is compatible with the strategy profile ̄ 휎 ★ and after which the collective memory is in the stateanchor 퐴푣 푚 , and plays accordingly. Since퐴 푚 = 퐴is the subset ofΠwhere푖 ∈ 퐴 =Λ(ℎ푢), we are in the case where the set퐴is no longer split further by our labeling. That is, there is a play휋from푤 푚 such thatΛ(ℎ휋 ≤푘 ) = 퐴 for every푘such that this is defined. Then, in the construction of the strategy profile ̄ 휎 ★ we have distinguished two cases: the one where퐴was a singleton, and the one where it had at least two elements (the empty case is excluded, since 퐴 contains player 푖). • If퐴is a singleton, then after the historyℎ푢, without any player deviating, all players are following the strategy profile ̄ 휏 ï푖 . By its definition, that strategy profile achieves the payoff푧 푖 for player 푖 with positive probability. •If퐴has at least two elements, then after that same history, all players are following the play 휋 퐴★ , which yields the same payoffs as휋 퐴 . However, we must still prove that player푖actually gets the payoff푧 푖 in휋 퐴 , and that the play휋 퐴★ is generated with positive probability (i.e. that it does not cross infinitely many stochastic vertices—which we must first show for휋 퐴 ). We do so in the following result, which we will use again later. Sublemma 9. The play휋 퐴 is (eventually) generated with positive probability when the players follow the strategy profile ̄ 휎. Similarly, the play휋 퐴★ is generated with positive probability when they follow ̄ 휎 ★ . Both plays yield to each player 푗 ∈ 퐴 the payoff 푧 푗 . Proof. Let푗 ∈ 퐴. Let us proceed by case disjunction according to the risk measure used by player 푗 . •If player푗is an optimist, then, by Property 3 of Lemma 35, we haveX 푗 ( ̄ 휎 ↾ℎ ) = 푧 푗 , and 11.3. CONSTRAINED EXISTENCE PROBLEM179 therefore P ̄ 휎 ↾ℎ (휇 푗 = 푧 푗 )> 0, i.e., by the law of total probability: P ̄ 휎 ↾ℎ푢 (휋 퐴 )P ̄ 휎 ↾ℎ (휇 푗 = 푧 푗 | 휋 퐴 )+ ∑︁ 푘 ∑︁ 푤∈퐸(휋 푘 )\휋 푘+1 P ̄ 휎 ↾ℎ푢 (휋 퐴 ≤푘 푤)P ̄ 휎 ↾ℎ푢 (휇 푗 = 푧 푗 | 휋 퐴 ≤푘 푤)> 0. But using Property 5, all the terms of the summation on the right are zero, hence the product: P ̄ 휎 ↾ℎ (휋 퐴 )P ̄ 휎 ↾ℎ (휇 푗 = 푧 푗 | 휋 퐴 ) is positive, i.e. the play 휋 퐴 has a positive probability of being generated and 휇 푗 (휋 퐴 ) = 푧 푗 . • If player푗is a pessimist, then because of Property 5 again, for every푘 ≥0 and each 푤 ∈ Supp ̄ 휎 ℎ휋 퐴 ≤푘 , there exists a strategy휏 푘푤 푗 such thatX 푗 ( ̄ 휎 −푗↾ℎ휋 퐴 ≤푘 푤 ,휏 푘푤 푗 )> 푧 푗 . By composing all those strategies, we obtain a deviation휏 푗 of the strategy휎 푗↾ℎ ; which, by Property 4, satisfies the inequality X 푗 ( ̄ 휎 −푗↾ℎ ,휏 푗 ) ≤ 푧 푗 . Therefore, either: –we have: min 푘 min 푤∈Supp ̄ 휎 ℎ휋 퐴 ≤푘 \휋 퐴 푘+1 X 푗 ( ̄ 휎 −푗↾ℎ휋 ≤푘 푤 ,휏 푘푤 푗 ) ≤ 푧 푗 , which is impossible by definition of the strategies 휏 푘푤 푗 ; – or we haveP ̄ 휎 −푗↾ℎ ,휏 푗 (휋 퐴 ) = P ̄ 휎 ↾ℎ (휋 퐴 ) ≠0 and휇 푗 (휋 퐴 ) ≤ 푧 푗 , and then actually휇 푗 (휋 퐴 ) = 푧 푗 . We have thus proven that player푗gets the payoff푧 푗 in휋 퐴 , and that the play휋 퐴 is generated with positive probability in ̄ 휎. The analogous results about휋 퐴★ follow using the equalities Occ(휋 퐴★ ) = Occ(휋 퐴 ) and Inf(휋 퐴★ ) = Inf(휋 퐴 ).□ In those two cases (if퐴is a singleton or has several elements), we obtain that the strategy profile ̄ 휎 ★ is such that, with some positive probability, player 푖 gets the payoff 푧 푖 . Player푖gets risk measure푧 푖 . We still have to prove that player푖has zero probability of getting a lower payoff (if they are a pessimist) or a higher payoff (if they are an optimist). To show both cases, we prove the following result: Sublemma 10. Every payoff vector that has a positive probability of being achieved in the strategy profile ̄ 휎 ★ also has a positive probability of being achieved in the strategy profile ̄ 휎. Proof.Let ̄ 푧 ′ be such a payoff vector. Then, there is a historyℎ푤compatible with ̄ 휎 ★ and a set 퐴 ⊆Πsuch that, after the historyℎ푤, the strategy profile ̄ 휎 ★ is in stateanchor 퐴last(ℎ) , and from that point it has a nonzero probability of achieving the payoff vector ̄ 푧 ′ while staying in states of the form anchor 퐴푣 . • If퐴is empty, then the strategy profile ̄ 휏 ï∅ has been defined as a strategy profile that almost surely generates a payoff vector that is generated with positive probability by ̄ 휎, from every vertex from which that is possible. That requirement is satisfiable, and therefore satisfied by ̄ 휏 ï∅ , from the vertex푤, since that vertex is itself reached with positive probability in the strategy profile ̄ 휎. Therefore, the payoff vector ̄ 푧 ′ is also achieved with positive probability in ̄ 휎 . • If퐴is a singleton, say퐴 =푗, then the strategy profile ̄ 휏 ï푗 has been defined so that from every vertex from which that is possible, on the one hand, it generates the payoff푧 푗 with positive probability, and on the other hand, it is almost sure that the payoff vector that will 180CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA be generated has also positive probability of being generated in ̄ 휎. Similarly as above, that requirement is satisfiable from the vertex푤, since ̄ 휎 ↾ℎ푤 satisfies it. Therefore, again, the payoff vector ̄ 푧 ′ is also achieved with positive probability in ̄ 휎 . •If퐴has at least two elements, then the strategy profile ̄ 휎 ★ stays in states of the form anchor 퐴푣 only along one play, namely휋 퐴★ , and that play generates a payoff vector that was also associated with the play휋 퐴 that by Sublemma 9, has positive probability of being generated in ̄ 휎 , hence the same conclusion.□ This proves the equality X 푖 ( ̄ 휎 ★ ) = 푧 푖 .□ ▶The strategy profile ̄ 휎 ★ is an XRSE. We have now constructed the finite-memory strategy profile ̄ 휎 ★ , showed that it had the expected number of memory states, and that it generates the expected risk measures. We must now give the final argument for our construction: that strategy profile is also an extreme risk-sensitive equilibrium. We will prove that result by showing separately that optimists have no profitable deviations, and then that neither do pessimists. Proposition 7. No optimist has a profitable deviation in ̄ 휎 ★ . Proof. Let푖be an optimist, and let us consider a deviation휎 ′ 푖 of that player from ̄ 휎 ★ . Let us write푧 ′ for the risk measure 푧 ′ = X 푖 ( ̄ 휎 ★ −푖 ,휎 ′ 푖 ). Let us notice that along every play compatible with ̄ 휎 ★ −푖 , the transitions that are possible in the memory structure of the strategy profile ̄ 휎 ★ can be classified as follows: •transitions among states of the form anchor 퐴푣 for a fixed 퐴; •transitions from a state of the form anchor 퐴푣 to a state of the form anchor 퐵푤 with 퐵 ⊂ 퐴; •transitions from a state of the form anchor 퐴푣 to the state punish 푖 ; •and transitions from punish 푖 to itself. Therefore, any such play stabilizes either in the statepunish 푖 , or among the states of the form anchor 퐴푣 for a fixed set퐴. Consequently, if in the strategy profile( ̄ 휎 ★ −푖 ,휎 ′ 푖 )player푖gets the payoff 푧 ′ with positive probability, then we can also say that either: •with positive probability, player 푖 gets the payoff 푧 ′ and the state punish 푖 is reached; • or there exists a set퐴 ⊆Πsuch that with positive probability, player푖gets the payoff푧 ′ , and the collective memory remains in states of the form anchor 퐴푣 . In the first case, let us consider a historyℎ푣compatible with ̄ 휎 ★ −푖 such that the collective memory is in an anchoring state afterℎand in statepunish 푖 afterℎ푣. If player푖can obtain the risk measure 푧 ′ by going to푣from that vertex against ̄ 휎 ★ −푖↾ℎ , and therefore, against the punishing strategy profile ̄ 휏 †푖 −푖 , it means that they can enforce that risk measure against every possible strategy profile from last(ℎ). On the other hand, if the collective memory is in an anchoring state afterℎ, it means that the vertexlast(ℎ)is also visited with positive probability in the strategy profile ̄ 휎 (otherwise we would have switched to a punishing state earlier). There is therefore a historyℎ ′ compatible with ̄ 휎such thatlast(ℎ) = last(ℎ ′ ); and after that history, against the strategy profile ̄ 휎 ↾ℎ ′ , player푖 also has the possibility of getting with positive probability the payoff푧 ′ . Since ̄ 휎is an XRSE, that implies 푧 ′ ≤ 푧 푖 . 11.3. CONSTRAINED EXISTENCE PROBLEM181 In the second case, let us notice that the strategy profiles of the form ̄ 휏 ï퐴 are pure, and therefore that any deviation of player푖is immediately detected and leads to a switch to statepunish 푖 . Therefore, if the collective memory remains in states of the formanchor 퐴푣 , it means that player푖 is actually following the strategy 휎 ★ 푖 . Thus, we also have 푧 ′ ≤ 푧 푖 . The strategy 휎 ′ 푖 is not a profitable deviation from ̄ 휎 ★ .□ We can now end the proof with the dual proposition. Proposition 8. No pessimist has a profitable deviation in ̄ 휎 ★ . Proof.Let푖be a pessimist, and consider a deviation휎 ′ 푖 of that player from ̄ 휎 ★ . We intend to prove that the deviation휎 ′ 푖 is not profitable, that is, when following the strategy profile( ̄ 휎 ★ −푖 ,휎 ′ 푖 ), there is still a positive probability that player푖receives a payoff smaller than or equal to푧 푖 . Using Lemma 31, we can assume without loss of generality that 휎 ′ 푖 is pure. First, we observe that for each historyℎ푣compatible with ̄ 휎 ★ such that, afterℎ푣, the collective memory is in stateanchor 퐴last(ℎ) with푖 ∈ 퐴, the vertex푣is such that there also exists a historyℎ ′ 푣 compatible with ̄ 휎withΛ(ℎ ′ 푣) = 퐴. By Property 4 of Lemma 35, we haveX 푖 ( ̄ 휎 −푖↾ℎ ′ 푣 ,휏 푖 ) ≤ 푧 푖 for every휏 푖 . Therefore, if player푖accepts to follow the historyℎ푣and, then, deviates and takes an edge that makes the collective memory switch to the statepunish 푖 , then with positive probability player 푖gets a payoff lesser than or equal to푧 푖 . If such an action is ever performed, then the deviation휎 ′ 푖 is not profitable. Let us now assume that휎 ′ 푖 performs no such action: after every historyℎ푣 ∈ Hist 푖 G ↾푣 0 , if the collective memory is in a state of the formanchor 퐴last(ℎ) with푖 ∈ 퐴, the vertex휎 ′ 푖 (ℎ푣)belongs to the setSupp(휎 ★ 푖 (ℎ푣)). Then, by Property 2 of Lemma 35, we also have푖 ∈Λ(ℎ푣휎 ′ 푖 (ℎ푣)). Thus, there still exists a set퐴with푖 ∈ 퐴such that, with positive probability, when following the strategy profile( ̄ 휎 ★ −푖 ,휎 ′ 푖 ), the strategy profile ̄ 휎 ★ −푖 stabilizes among memory states of the formanchor 퐴푣 ; and then, the strategy profile ̄ 휎 ★ only proceeds to pure actions, hence the strategy휎 ′ 푖 is actually following 휎 ★ 푖 . Using the same arguments as in the proof of Proposition 6 (definition of ̄ 휏 ï푖 in case퐴 =푖, and Sublemma 9 in case퐴has more elements), we can then conclude that player푖gets the payoff 푧 푖 with positive probability and therefore that the deviation 휎 ′ 푖 is not profitable.□ The strategy profile ̄ 휎 ★ is an XRSE, satisfies the equalityX( ̄ 휎 ★ ) = X( ̄ 휎), and uses the desired number of memory states. Furthermore, if ̄ 휎 is pure, so is ̄ 휎 ★ .□ Finally, using Theorem 35, we can show the following lemma. Lemma 36. The constrained existence problem of XRSEs is inNP. The same problem when players are restricted to pure strategies is still in NP. Proof.LetG ↾푣 0 be a simple quantitative stochastic game. Let(푃,푂)be a partition ofΠ, and let ̄ 푥and ̄ 푦be threshold vectors. By Theorem 35, if there exists a (pure) XRSE with ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦, then there exists one with at most 3푛푝−2푛+ 푝+1 memory states, where푝is the number of players and푛is the number of vertices. Such a strategy profile can be guessed in polynomial time. We now show that, once such a finite-memory strategy profile ̄ 휎is guessed, one can check in polynomial time whether it is an XRSE, and satisfies the constraint ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦. 182CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA •First, given ̄ 휎, for each player푖, the quantityX 푖 ( ̄ 휎)can be computed in polynomial time, since it reduces to computing player푖’s risk measure in the Markov chain induced by ̄ 휎 (which has polynomial size) (by Lemma 31). •Second, checking that ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦 can be done in polynomial time. • Third, for each player푖, one must check that player푖has no profitable deviation. This can also be done in polynomial time (by Lemma 31) by computing the best risk measure player푖can get in the MDP induced by ̄ 휎 −푖 (which has polynomial size).□ 11.3.2 Restrictions on strategies We now consider subcases where the space of a strategies is restricted. We show in Theorem 33 that restricting the memory or amount of randomness of the strategy still renders the problem only inNP. Later in this section, we prove that all these problems, including the general problem, are NP-hard. This subsection therefore completes the proof of Theorems 32 and 33. We restrict the set of strategies of each player to stationary, positional or pure. We show that the problem is in NP for each of these cases. Lemma 37. The constrained existence problem, when all the players are restricted to positional, stationary, or pure strategies, is in NP. Proof. We show that we can still guess a strategy profile, and verify in polynomial time if it is indeed an XRSE. For the cases of positional and stationary strategies, guessing a strategy profile is straightforward, since such a strategy profile ̄ 휎can be represented using polynomially many bits. We can then verify that a given strategy profile ̄ 휎gives risk measures within the constraints, and also is an XRSE in polynomial time (Lemma 31). However, for pure strategies, memory might be required. But we showed with Theorem 35 that if there is a pure strategy profile, then there is one that requires polynomial memory, and therefore our result follows.□ We now proveNP-hardness of the constrained existence problem for the general setting as well as the cases where strategies are restricted. Lemma 38. The constrained existence problem of XRSEs isNP-hard, even when all players are pessimists and all rewards are non-negative. It remainsNP-hard when the strategies are reduced to stationary, pure, or positional ones. Proof. We proveNP-hardness by reducing from the problem 3Sat. Consider a 3Satformula휑, over the variables푥 1 , . . .,푥 푛 , where휑 =퐶 1 ∧퐶 2 ∧·∧퐶 푚 , where for each푖we have퐶 푖 =(ℓ 푖1 ∨ ℓ 푖2 ∨ ℓ 푖3 )and for푗 =1,2,3, we haveℓ 푖푗 = 푥 푘 orℓ 푖푗 =¬푥 푘 for some푘 ∈ 1, . . .,푛. We construct a gameG 휑 with two players for each literalℓ, denoted by ◦ ℓand□ℓ. The game is depicted in Figure 41. For convenience, some terminal vertices have been represented several times. Each player ◦ ℓcontrols one vertex, the vertex ◦ ℓ, of circled shape, and symmetrically, each player□ℓcontrols the square-shaped vertex□ℓ. Further, we add a player퐶 푖 , who controls the vertex퐶 푖 , for each clause퐶 푖 . Finally, there is a player who does not control any vertex. There are also stochastic vertices, that are represented by the black circles. In each terminal vertex, the symbol∀ should be understood as "every (other) player". We assume all players are pessimistic, and ask if there is an XRSE where player^’s risk measure is exactly 2. We give the formal definition of the gameG 휑 below. 11.3. CONSTRAINED EXISTENCE PROBLEM183 ¬푥 1 푥 1 푡 † : ∀ 0 푠 ¬푥 1 푠 푥 1 푓 푥 1 : ◦푥 1 1 ∀ 2 푓 ¬푥 1 : ◦¬푥 1 1 ∀ 2 ¬푥 1 푥 1 푡 ⋄ : ⋄ 0 ∀ 2 푡 ⋄ : ⋄ 0 ∀ 2 ¬푥 2 푥 2 푠 푥 2 푠 ¬푥 2 푓 푥 2 : ◦푥 2 1 ∀ 2 푓 ¬푥 2 : ◦¬푥 2 1 ∀ 2 ¬푥 2 푥 2 푡 ⋄ : ⋄ 0 ∀ 2 푡 ⋄ : ⋄ 0 ∀ 2 푡 † : ∀ 0 . . . ¬푥 푛 푥 푛 푡 ⋄ : ⋄ 0 ∀ 2 푡 ⋄ : ⋄ 0 ∀ 2 푠 r 퐶 1 퐶 2 . . . 퐶 푚 푡 푥 2 : □푥 2 1 ∀ 2 푡 푥 4 : □푥 4 1 ∀ 2 푡 ¬푥 11 : □¬푥 11 1 ∀ 2 Figure 41: Construction of a gameG 휑 from a 3Sat formula 휑 ▶Construction of the gameG 휑 : vertices, edges and payoffs For each literalℓ, we define two players□ℓand ◦ ℓ. We add one other player퐶 푖 for each clause퐶 푖 , and an additional constraining player^. All players are pessimists. Each player owns at most one vertex in the game, and therefore, we will refer to the player and vertex interchangeably. There is one vertex for each of the players mentioned above other than^, who owns no vertices. Further, there are 2푛+1 many stochastic vertices: one for each literal푠 푥 1 ,푠 푥 2 , . . .,푠 푥 푛 , 푠 ¬푥 1 ,푠 ¬푥 2 , . . .,푠 ¬푥 푛 , and finally one clause-randomizer푠 r . There are also 2푛+2 terminal vertices, written 푓 ℓ and 푡 ℓ for each literal ℓ , and further the terminal vertices 푡 and 푡 † . We now define the edges between the vertices of the graph for all푖 ∈ 1, . . .,푛: there are edges from ◦ 푥 푖 to ◦ ¬푥 푖 , and edges from ◦ ¬푥 푖 to푡 † . Further, for every literalℓ = 푥 푖 or¬푥 푖 , there are edges: •from ◦ ℓ to 푠 ℓ ; •from 푠 ℓ to 푓 ℓ and to □ ̄ ℓ , where ̄ ℓ =¬푥 푖 if ℓ = 푥 푖 and ̄ ℓ = 푥 푖 if ℓ =¬푥 푖 ; •from □ℓ to 푡 ; •from □ℓ to ◦ 푥 푖+1 if 푖< 푛, and to 푠 r if 푖 = 푛. Finally, for all clauses퐶 푗 , there are edges from푠 r to퐶 푗 and from퐶 푗 to푡 ℓ such thatℓoccurs positively in the clause퐶 푗 . The terminal vertices yield the following payoffs. •In terminal 푡 ℓ , all players get payoff 2, except the player □ℓ who gets payoff 1. •In terminal 푓 ℓ , all players get payoff 2, except player ◦ ℓ who gets payoff 1. •In terminal 푡 † , all players get payoff 0. •In terminal 푡 , all players get payoff 2, except player who gets payoff 0. Finally, we let the constraints be that player^gets a risk measure of exactly 2. Equivalently, we define ̄ 푥 and ̄ 푦 by ̄ 푦 =(2) 푖∈Π , 푥 푖 = 0 for each 푖 ∈Π\^, and 푥 = 2. ▶If 휑 is satisfiable, then there is an XRSE satisfying the constraints. Consider a satisfying valuation 휈 of the 3Sat formula 휑 . For each푖, letℓ 푖 denote the literal, among푥 푖 and¬푥 푖 , which is set to true by the satisfying valuation 휈 . Let us define the (positional) strategy profile ̄ 휎 휈 . 184CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA •Player ◦ ℓ 푖 goes to 푠 ℓ 푖 . •Player ◦ 푥 푖 goes to 푠 푥 푖 if 휈(푥 푖 ) = 1, and to ◦ ¬푥 푖 otherwise. •Player ◦ ¬푥 푖 goes to 푠 ¬푥 푖 if 휈(푥 푖 ) = 1 and to 푡 † otherwise. • For each player□ℓ, the strategy is to chose the edge that does not lead to푡 . That is, the edge to ◦ 푥 푖+1 if ℓ = 푥 푖 or¬푥 푖 and 푖< 푛, and the edge to 푠 r if 푖 = 푛. •Each clause player퐶 푖 takes the edge to the vertexℓ 푗 such that the litteralℓ 푗 was set to true by the satisfying valuation 휈 . We now show that this is an XRSE that satisfies the constraint. First, we verify if the constraints are satisfied. Observe that following the strategy profile ̄ 휎 휈 , it is almost sure that none of the terminals where player has payoff less than 2 will be reached. Therefore this satisfies the constraints. We now argue that ̄ 휎 휈 is an XRSE, i.e. that no player can get a better risk measure by deviating. The result is immediate for player^and for the clause players, who all get risk measure 2, the best they could hope for. For each literalℓ, player ◦ ℓgets risk measure 1 ifℓis set to true, and risk measure 2 ifℓis set to false. The same argument as above holds therefore in the second case. In the first case, she gets risk measure 0, but she has no profitable deviation, since the only deviation available leads to푡 † and to the payoff 0. Player□ℓhas also risk measure 2 whenℓis set to false. Otherwise, he gets payoff 1. In that second case, the vertex owned by the player is not visited in any history of the game, hence he has no possibility of deviating. The (positional) strategy profile ̄ 휎 휈 is therefore an XRSE. ▶If there is an XRSE satisfying the constraints, then 휑 is satisfiable. Let us assume that there exists an XRSE ̄ 휎in the gameG 휑 , such that player^gets the risk measure 2. We prove, first, that we can assume that ̄ 휎is pure (and therefore positional, since there is then only one history leading to each vertex). Sublemma 11. There exists a positional XRSE ̄ 휎 ★ inG 휑↾푥 1 where player gets risk measure 2. Proof. Let us first focus on what happens in vertices that have positive probability of being reached. If the vertex ◦ ¬푥 푖 has a positive probability of being reached in ̄ 휎 , then any strategy of the player ◦ ¬푥 푖 that goes to푡 † with positive probability gives the player^the risk measure 0. Therefore, necessarily, the strategy휎 ◦¬푥 푖 consists of deterministically going to푠 ¬푥 푖 . The same argument holds for the vertices of the form □ℓ . If now the vertex ◦ 푥 푖 has a positive probability of being reached and if the player ◦ 푥 푖 random- izes between the two edges available, then she gets the risk measure 1, since the terminal vertex푓 푥 푖 is reached with positive probability and푡 † with probability zero. But then, if she deviates and goes to the vertex ◦ ¬푥 푖 with probability 1, she avoids the terminal vertex푓 푥 푖 , and the other players will not react since they do not detect the deviation. She therefore gets the risk measure 2, and the deviation is profitable. Consequently, the strategy휎 ◦푥 푖 can only deterministically select one of those two edges. 11.3. CONSTRAINED EXISTENCE PROBLEM185 At the end of the game, for each푗, the player퐶 푗 could play a randomized strategy. In such a case, her strategy can be replaced by a pure strategy that takes, deterministically, one of the edges that she was previously taking. Such a modification in her strategy can only increase the risk measure of some players (namely, those of the form□ℓ) without impacting player^’s risk measure or giving any player the possibility of profitably deviating. Finally, if one of those vertices is reached after a history that is not compatible with ̄ 휎, i.e. if one of those players deviates: it is necessarily due to a deviation of a player of the form ◦ 푥 푖 , since any other deviation would immediately lead to a terminal vertex. If she went to푠 푥 푖 instead of ◦ ¬푥 푖 , what the other players do afterwards does not matter, since such a deviation cannot be profitable: with positive probability, the terminal vertex푓 푥 푖 is reached, and she gets payoff 1. If she went to ◦ 푥 푖 instead of푠 푥 푖 , then we can assume that player ◦ ¬푥 푖 ’s strategy consists of going to the terminal vertex푡 † , giving her the payoff 0. Those modifications do not impact the fact that ̄ 휎 is an XRSE.□ We therefore assume that ̄ 휎is positional. Let us now define the valuation휈as follows: for each variable푥 푖 we have휈(푥 푖 ) =1 if휎 ◦푥 푖 ( ◦ 푥 푖 ) = 푠 푥 푖 , and휈(푥 푖 ) =0 if휎 ◦푥 푖 ( ◦ 푥 푖 ) = ◦ ¬푥 푖 . Let then퐶 푗 be a clause, and let us prove that it is satisfied by휈. Let푡 ℓ = 휎 퐶 푗 (퐶 푗 ). Then, the player□ℓgets risk measure 1 in the XRSE ̄ 휎. Consequently, the vertex□ℓis never reached: otherwise, the only play compatible with ̄ 휎in which player□ℓgets payoff 1 would traverse the vertex□ℓ, and player□ℓwould have a profitable deviation by going to the terminal vertex푡 . If that is the case, then the definition of휈given above implies that the literal ℓ is true. The valuation 휈 satisfies therefore the formula 휑 . ▶Conclusion. We have defined an instance of the constrained existence problem of XRSEs from an instance of 3Satand proved that one is a positive instance if and only if the other is. This proves theNP-hardness of the constrained existence problem of XRSEs, since the gameG 휑 can clearly be constructed in polynomial time. Moreover, the gameG 휑 is such that if an XRSE where player^gets risk measure 2 exists, then there also exists such an equilibrium that it positional, which proves alsoNP-hardness when the players are restricted to pure, stationary or positional strategies.□ This lemma, along with Lemma 36, proves Theorem 32; and along with Lemma 37, it proves Theorem 33. 11.3.3 Things get easier when everyone is optimistic Since ourNP-hardness results involved only pessimistic players, we now show that the constrained existence problem of XRSEs becomes P-complete when the perceived reward of each player is computed based on the risk measureOM, thus proving Theorem 34. We first show an upperbound by giving a polynomial-time algorithm. Lemma 39. If all players are optimists, then the constrained existence problem of XRSEs is in P, and there is an algorithm that decides it in time푂(푝푚 2 ), where푚is the number of edges inGand푝the number of players. Moreover, the algorithm can be modified to output an XRSE that satisfies the constraints, if one exists, in time 푂(푝푚 2 +푚 3 )—or in time 푂(푝푚 2 ) if all upper thresholds are non-negative. 186CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA Proof.▶Preliminary remarks We are given the gameG ↾푣 0 and two threshold vectors ̄ 푥, ̄ 푦 ∈ Q Π ; we wish to find an XRSE ̄ 휎 such that ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦. Throughout the proof, when푊 ⊆ 푉is a set of vertices and퐹 ⊆ 퐸is a set of edges, we write Attr(푊,퐹)for the positive probabilistic attractor of푊in the graph(푉,퐹), i.e. the set of vertices푣such that for every strategy profile ̄ 휎inG ↾푣 that uses only edges of퐹, there is a positive probability of reaching푊. As a consequence of Lemma 31 (replacing the vertices of푊with terminal vertices), we have the following. Sublemma 12. Given푊 , the set Attr(푊,퐹) can be computed in time 푂(푚). Similarly, Lemma 31 enables us to compute the adversarial values of each vertex, i.e., the best risk measure that the player controlling that vertex can ensure from that vertex when the other players are fully hostile. Sublemma 13. For each 푖 and 푣 ∈ 푉 푖 , the quantity: val(푣) =inf ̄ 휏 −푖 ∈Strat −푖 G ↾푣 sup 휏 푖 ∈Strat 푖 G ↾푣 X 푖 ( ̄ 휏) can be computed in time 푂(푚). Then, computing all those values can be done in time푂(푚 2 ). We can therefore assume that those quantities val(푣) are given with the input. ▶Cycle-friendly and cycle-averse cases We differentiate two types of instances. If there exists a player푖such that we have푦 푖 <0, then the requirement ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦implies that ̄ 휎must almost surely reach a terminal vertex: we call that case the cycle-averse case. If there is no such player, we are in the cycle-friendly case. Our algorithm will work slightly differently in those two cases. However, the fundamental idea is still the same in both cases: we prune iteratively the set of edges, and each of the subsets퐹 ⊆ 퐸which we obtain will induce a strategy profile ̄ 휎 퐹 , in which the profitable deviations will be detected and used to prune new edges. However, the definition of ̄ 휎 퐹 differs in the cycle-averse and the cycle-friendly case. ▶Algorithm in the cycle-friendly case In the cycle-friendly case, for a given set of edges퐹, the strategy profile ̄ 휎 퐹 in the gameG ↾푣 0 , is defined as follows: from each non-stochastic vertex푣, when푣is seen for the first time, the strategy profile randomizes uniformly between all the edges푣푤 ∈ 퐹. Later, when푣is visited again, it always repeats the same choice. Equivalently, each player initially chooses, at random, a positional strategy, and then follows it. If some player푖deviates and takes an edge that they are not supposed to take (be it an edge that does not belong to퐹or an outgoing edge of a vertex from which a different edge has already been taken), then all the players switch to the positional strategy profile ̄ 휏 †푖 , where ̄ 휏 †푖 −푖 minimizes the best risk measure that player푖can get (a positional such strategy profile exists by Lemma 31), and 휏 †푖 푖 is some positional strategy. 11.3. CONSTRAINED EXISTENCE PROBLEM187 Our algorithm in the cycle-friendly case is presented in Algorithm 2. Each step푘consists of identifying a new set of vertices푉 푘 / that must be avoided. At step푘 =0, it is the set of terminal vertices that give some player푖a payoff that is larger than푦 푖 , which would then make them have an off-constraints risk measure. At step푘 ≥1, it is the set of vertices푣whose adversarial valueval(푣)is greater than the risk푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푘 ) , where푖is the player controlling푣. In other words, the vertices from which that player can have a profitable deviation. Note that it that second case, the computation of푉 푘 / requires the computation of푧 푘 푖 , which can be done in time푂(푚)by computing the set of terminals that are accessible from푣 0 in(푉,퐸 푘 ), and by deciding whether the probability of reaching no terminal is positive: that will be the case if and only if there exists a positional strategy profile that uses only edges of퐸 푘 (and therefore that ̄ 휎 퐸 푘 is following with positive probability) such that with positive probability no terminal vertex is reached, which can be decided in time 푂(푚) using Lemma 31. Then, the positive probabilistic attractor퐴 푘 = Attr(푣 푘 / ,퐸 푘 )is computed. If푘 ≥1 and푣 0 ∈ 퐴 푘 , i.e., if it is not possible to avoid reaching the set푉 푘 / , the answerNois returned. Otherwise, the set퐸 푘+1 is defined from퐸 푘 by removing all the edges that lead from a vertex that does not belong to퐴 푘 to a vertex that does, thus making sure that푉 푘 / will never be reached. The algorithm stops when there is no more edge to remove. Then, if we have푧 푘 푖 ≥ 푥 푖 for each푖, the algorithm answersYesand outputs the set 퐸 푘 as a succinct representation of the strategy profile ̄ 휎 퐸 푘+1 . Otherwise, it answers No. Algorithm 2 Constrained existence problem with optimists in the cycle-friendly case procedure CycleFriendly(G, ̄ 푥, ̄ 푦) 푘 0 퐸 푘 퐸 푉 푘 / =푡 ∈ 푇 | 휇 푖 (푡)> 푦 푖 퐴 푘 Attr(푉 푘 / ,퐸 푘 ) if 푣 0 ∈ 퐴 푘 then return No else 퐸 푘+1 퐸 푘 \푢푣 ∈ 퐸 푘 | 푢 ∉ 퐴 푘 and 푣 ∈ 퐴 푘 while 푘 = 0 or 퐸 푘+1 ≠ 퐸 푘 do 푘 푘+ 1 Compute 푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푘 ) for each 푖 ∈Π 푉 푘 / 푣 | val(푣)> 푧 푘 푖 for 푖 ∈Π such that 푣 ∈ 푉 푖 퐴 푘 Attr(푉 푘 / ,퐸 푘 ) if 푣 0 ∈ 퐴 푘 then return No else 퐸 푘+1 퐸 푘 \푢푣 ∈ 퐸 푘 | 푢 ∉ 퐴 푘 and 푣 ∈ 퐴 푘 end if end while end if if 푧 푘 푖 ≥ 푥 푖 for all players 푖 then return(Yes,퐸 푘+1 ) else return No end if end procedure Correctness in the cycle-friendly case.To prove the correctness of Algorithm 2, we first need to prove that the edge removals are such that all vertices always keep at least one outgoing edge, and that the stochastic ones always keep all of them, so that the strategy ̄ 휎 퐸 푘 is always properly defined. Invariant 2. At each step푘, every vertex푣 ∉ 푉 ? is such that퐸 푘 (푣) ≠∅, and every vertex푣 ∈ 푉 ? is such 188CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA that 퐸 푘 (푣) = 퐸(푣). The proof is an immediate induction. We also need termination. Proposition 9. Algorithm 2 terminates. Proof.At each step푘 ≥1, we either have that the algorithm terminates, or that an edge is removed. The sequence퐸 1 ,퐸 2 , . . .is therefore strictly decreasing (note that we might have퐸 0 = 퐸 1 ), hence it cannot be infinite.□ Now, to prove correctness, we will first prove the following result. Sublemma 14. For each player푖and each index푘, every strategy profile ̄ 휎 ′ that uses only edges of퐸 푘 is such that X 푖 ( ̄ 휎 ′ ) ≤ 푧 푘 푖 . Proof.This result is a consequence of the fact that every payoff vector that can be obtained with positive probability in the strategy profile ̄ 휎 ′ is obtained with positive probability in the strategy profile ̄ 휎 퐸 푘 . Indeed, consider some payoff vector ̄ 푧 that has a positive probability of being generated in ̄ 휎 ′ . If ̄ 푧 is obtained by reaching a terminal vertex, then that terminal vertex is accessible from푣 0 in the graph(푉,퐸 푘 ), and it therefore has a positive probability of being reached in ̄ 휎 퐸 푘 . If ̄ 푧 = (0) 푖 is obtained by reaching no terminal vertex, then by Lemma 31, there exists a positional strategy profile ̄ 휏that uses only edges of퐸 푘 such that with positive probability, no terminal vertex is reached. Then, when following the strategy profile ̄ 휎 퐸 푘 , there is a positive probability that the players actually follow ̄ 휏. And therefore, there is also a positive probability to get the payoff vector ̄ 푧 =(0) 푖 in the strategy profile ̄ 휎 퐸 푘 . Since all players are optimists, the result follows.□ Note that this result implies that the sequence(푧 푘 푖 ) 푘 , for each 푖, is nondecreasing. We can now prove correctness. To do so, we need to prove two propositions: the algorithm recognizes only positive instances, and recognizes all of them. Proposition 10. The algorithm recognizes only positive instances. Proof.Let us assume that the algorithm answersYesat step푘: let us show that the strategy profile ̄ 휎 퐸 푘 is an XRSE that satisfies the desired constraints. Note that the algorithm does never switch to the final refinements at step 0, hence we necessarily have 푘 ≥ 1. The strategy profile ̄ 휎 퐸 푘 satisfies ̄ 푥 ≤ X( ̄ 휎 퐸 푘 ) ≤ ̄ 푦 . The lower bound is immediate since the algorithm answers Yes at step 푘 only if the strategy profile ̄ 휎 퐸 푘 satisfies that constraint. Regarding the upper bound, observe that the set퐸 1 has been defined so that the setAttr(푉 0 / ,퐸), and therefore the set푉 0 / , is not accessible from푣 0 in the graph(푉,퐸 1 ), and therefore not in the graph(푉,퐸 푘+1 ). Thus, it is almost sure in ̄ 휎 퐸 푘+1 that no vertex of푉 0 / will ever be reached. In other words, all terminals that have a positive probability of being reached give each player푖a lower payoff than푦 푖 . Now, if there is a positive probability that the play never reaches a terminal, that also does not give any player 푖 such a payoff, since we are in the cycle-friendly case. 11.3. CONSTRAINED EXISTENCE PROBLEM189 The strategy profile휎 퐸 푘 is an XRSE.Let푖be a player, and let휎 ′ 푖 be a deviation of player푖 from ̄ 휎 퐸 푘 . We can assume without loss of generality that휎 ′ 푖 is pure. Let푧 ′ = X 푖 ( ̄ 휎 퐸 푘 −푖 ,휎 ′ 푖 ) be the extreme risk measure obtained by player푖. We want to prove that the deviation휎 ′ 푖 is not profitable, that is, we have 푧 ′ ≤ 푧 푘 푖 . If the deviation휎 ′ 푖 uses only the edges of퐸 푘 , then it cannot be profitable by Sublemma 14. But if it does use more edges, let us show that it cannot be a profitable deviation either. Sublemma 15. If there is a historyℎ푣compatible with ̄ 휎 퐸 푘 such that푣휎 ′ 푖 (ℎ푣) ∉ 퐸 푘 , then we have X 푖 ( ̄ 휎 퐸 푘 ↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ) ≤ 푧 푘 푖 . Proof.After such a history, the strategy profile ̄ 휎 퐸 푘 −푖 follows the positional strategy profile ̄ 휏 †푖 −푖 . By the definition of that strategy profile, we haveX 푖 ( ̄ 휎 퐸 푘 ↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ) ≤ val(푣). On the other hand, the vertex푣is accessible from푣 0 in(푉,퐸 푘 ), since it is visited with a positive probability in ̄ 휎 퐸 푘 . Therefore, it does not belong to the set퐴 푘 , and in particular not to the set푉 푘 / , which means that we have val(푣) ≤ 푧 푘 푖 . Hence, the conclusion follows.□ In the general case, the payoffs that player푖obtains with positive probability in the strategy profile( ̄ 휎 퐸 푘 −푖 ,휎 ′ 푖 ) are obtained either by using only edges that belong to퐸 푘 , or by using an edge that does not. In both cases, we have shown that player푖cannot get a payoff greater than푧 푘 푖 , which proves that the strategy profile ̄ 휎 퐸 푘 is an XRSE. Algorithm 2 answersYesonly on positive instances, and outputs in that case a succinct representation of an XRSE matching the constraints.□ It now remains to prove the converse. Proposition 11. Algorithm 2 recognizes all positive instances. Proof. Let us assume that we have a positive instance, i.e., that there exists an XRSE ̄ 휎 with ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦 . Let us show that the algorithm will answerYes. To do so, we first prove the following invariant: if an edge is removed at some step, then it is never taken by the XRSE ̄ 휎 . Invariant 3. For each푘 ≥0, every edge that has positive probability of being eventually taken in ̄ 휎 belongs to 퐸 푘 . Proof.We prove the invariant by induction. Base cases. The case 푘 = 0 is immediate, since we have 퐸 0 = 퐸. Further, at step푘 =1, if the strategy profile ̄ 휎 uses eventually, with positive probability, an edge that does not belong to퐸 1 , then it goes with positive probability to a vertex푣 ∈ Attr(푉 0 / ,퐸 0 ) . Then, with positive probability, a terminal vertex will be reached that gives to some player푖a payoff greater than푦 푖 , which is impossible. Therefore, such an edge cannot be taken in ̄ 휎 . Induction step. Let us assume that the invariant is true until step푘 ≥1, and let us show that it holds at step푘+1. Let푢푣be an edge that is used with positive probability when following ̄ 휎, and let us assume toward contradiction that it does not belong to퐸 푘+1 . Since the invariant is true at each step until푘, we can assume that푢푣has been removed at step푘, i.e., that we have푢푣 ∈ 퐸 푘 \ 퐸 푘+1 . Then, we have푢 ∉ 퐴 푘 and푣 ∈ 퐴 푘 . The strategy profile ̄ 휎 has therefore positive probability of visiting the set퐴 푘 , and therefore the set푉 푘 / . Then, from a vertex of푉 푘 / , i.e., a vertex푣withval(푣)> 푧 푘 푖 , player푖can deviate and get a risk measure 190CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA strictly better than푧 푘 푖 . But since the invariant is true at step푘, the strategy profile ̄ 휎uses only vertices of퐸 푘 , and therefore, by Sublemma 14, we haveX 푖 ( ̄ 휎) ≤ 푧 푘 푖 : player푖has a profitable deviation in ̄ 휎 , which is impossible.□ We are now able to conclude. The answer No can be given in the two following cases: • If at step푘, we have푣 0 ∈ Attr(푉 푘 / ,퐸 푘 ). Then, the strategy profile ̄ 휎visits the set푉 푘 / with positive probability. With the same arguments that were used in the proof of Invariant 3, that is not possible. • If during step푘, no edge is removed, but we have푧 푘 푖 < 푥 푖 for some player푖. Since ̄ 휎uses only edges of퐸 푘 by Invariant 3, we can apply Sublemma 14, and obtainX 푖 ( ̄ 휎) ≤ 푧 푘 푖 , and therefore X 푖 ( ̄ 휎)< 푥 푖 : that case is therefore also impossible by definition of ̄ 휎 . None of those cases is possible, hence our algorithm will eventually answer Yes.□ ▶Algorithm in the cycle-averse case The algorithm and the structure of the proof will be similar. However, we need some significant modifications, especially in the definition of the strategy profiles ̄ 휎 퐹 . In the cycle-averse case, for a given set of edges퐹, the strategy profile ̄ 휎 퐹 , in the gameG ↾푣 0 , is defined as follows: from each vertex푣 ∉ 푉 ? , it randomizes uniformly between all the edges푣푤 ∈ 퐹. Contrary to the cycle-friendly case, the outcome of such a randomization has no influence on what will happen if 푣 is seen again. If some player 푖 deviates and takes an edge that they are not supposed to take (an edge that does not belong to퐸 푘 , then), then all the players switch to the positional strategy profile ̄ 휏 †푖 , where ̄ 휏 †푖 −푖 minimizes the best risk measure that player푖can get (a positional such strategy profile exists by Lemma 31), and 휏 †푖 푖 is some positional strategy. Our algorithm in the cycle-averse case is presented in Algorithm 3. Again, each step푘identifies a new set of vertices that must be avoided. Their definition depends now on the parity of푘. When푘is even, it is the same as in the cycle-friendly case: the set푉 푘 / is the set of vertices푣such thatval(푣)> 푧 푘 푖 , where푖is the player controlling푣, and퐴 푘 is the positive probabilistic attractor of푉 푘 / . When푘is odd, we define directly퐴 푘 as the set of vertices from which whatever the players play, there is a positive probability of reaching no terminal vertex. Again, the computation of푉 푘 / for an even step푘 ≥2 requires the computation of푧 푘 푖 , which can be done in time푂(푚)by computing the set of terminals that are accessible from푣 0 in(푉,퐸 푘 ), and by deciding whether the probability of reaching no terminal is positive: that will be the case, now, if and only if there exists a vertex from which no terminal vertex is accessible, which can also be decided in time 푂(푚). As for odd steps, the computation of 퐴 푘 can also be done in 푂(푚) using Lemma 31. If푘 ≥1 and푣 0 ∈ 퐴 푘 , i.e., if it is not possible to avoid reaching the set푉 푘 / , the answerNois returned. Otherwise, the set퐸 푘+1 is defined from퐸 푘 by removing all the edges that lead from a vertex that does not belong to 퐴 푘 to a vertex that does, thus making sure that푉 푘 / will never be reached. The loop stops when there is no more edge to remove, i.e., when we get퐸 푘+2 = 퐸 푘 . Then, the algorithm answersNoif we have푧 푘 푖 < 푥 푖 for some푖. Otherwise, it performs final refinements, defined as follows: first, it defines퐹 0 = 퐸 푘 . Then, once퐹 ℓ is defined for someℓ, it checks whether there exists an edge푢푣 that matches the following conditions in the graph(푉,퐸 ℓ ): 1.the vertex푢 is not stochastic and has several outgoing edges; 11.3. CONSTRAINED EXISTENCE PROBLEM191 Algorithm 3 Constrained existence problem with optimists in the cycle-averse case procedure CycleAverse(G, ̄ 푥, ̄ 푦) 푘 0 퐸 푘 퐸 푉 푘 / =푡 ∈ 푇 | 휇 푖 (푡)> 푦 푖 퐴 푘 Attr(푉 푘 / ,퐸 푘 ) if 푣 0 ∈ 퐴 푘 then return No else 퐸 푘+1 퐸 푘 \푢푣 ∈ 퐸 푘 | 푢 ∉ 퐴 푘 and 푣 ∈ 퐴 푘 while 퐸 푘+2 ≠ 퐸 푘 or 푘 ≤ 1 do 푘 푘+ 1 if 푘 is even then Compute 푧 푘 푖 = X 푖 ( ̄ 휎 퐸 푘 ) for each 푖 ∈Π 푉 푘 / 푣 | val(푣)> 푧 푘 푖 for 푖 ∈Π such that 푣 ∈ 푉 푖 퐴 푘 Attr(푉 푘 / ,퐸 푘 ) else 퐴 푘 / 푣 |∀ ̄ 휏 ∈ Strat Π G ↾푣 ,P ̄ 휏 (Occ∩푇 =∅)> 0 end if if 푣 0 ∈ 퐴 푘 then return No else 퐸 푘+1 퐸 푘 \푢푣 ∈ 퐸 푘 | 푢 ∉ 퐴 푘 and 푣 ∈ 퐴 푘 end if end while end if if 푧 푘 푖 < 푥 푖 for some player 푖 then return No else ℓ 0 퐹 ℓ 퐸 푘 ⊲ Final refinement steps while there exists푢푣 satisfying Conditions 1, 2, and 3 do ℓ ℓ+ 1 퐹 ℓ+1 퐹 ℓ \푢푣 end while return(Yes,퐹 ℓ ) end if end procedure 2.all the terminal vertices accessible from 푣 are also accessible from 푣 0 without using푢푣 ; 3.at least one terminal vertex is accessible from푢 without using푢푣 . In the following, we will refer to those conditions as Conditions 1, 2, and 3. If there exists such an edge, then we define퐹 ℓ+1 = 퐹 ℓ \푢푣. If there is no such edge, the algorithm stops there, answersYes, and returns 퐹 ℓ as a succinct representation of ̄ 휎 퐹 ℓ . ▶Correctness in the cycle-averse case The fact that ̄ 휎 퐸 푘 and ̄ 휎 퐹 ℓ are always correctly defined can be proved with arguments similar as those that were used for the cycle-friendly case. We now focus on correctness properly said. We first need the following properties. Invariant 4. For every even푘>0, and for everyℓ, the graph(푉,퐸 푘 ), or(푉,퐹 ℓ ), contains no vertex that is accessible from 푣 0 and from which no terminal vertex is accessible. 192CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA Proof.In the graph(푉,퐸 푘 )(for푘>0 even), no induction is required: the set퐸 푘 has been obtained after an odd step, in which the set퐴 푘−1 has been made inaccessible. Thus, if we have a vertex푣 from which no terminal vertex is accessible, it means that in the graph(푉,퐸 푘−1 ), all paths from푣 to a terminal vertex were traversing a vertex of퐴 푘−1 , which implies that푣itself belonged to퐴 푘−1 , and is therefore not accessible from 푣 0 in(푉,퐸 푘 ). This also proves that the invariant is true during the final refinements at stepℓ =0. If now we assume that it is true at some stepℓ, then Condition 3 guarantees that it remains true at step ℓ+ 1.□ Invariant 5. If the algorithm switches to final refinements after step푘, then for each stepℓof final refinements and for each player 푖, we have X 푖 ( ̄ 휎 퐹 ℓ ) = 푧 푘 푖 . Proof.The invariant is immediate forℓ =0, since we have퐹 ℓ = 퐸 푘 . Then, if it is true at stepℓ, it remains true at stepℓ+1. Indeed, Condition 2 guarantees that the set of terminal vertices accessible from푣 0 in(푉,퐹 ℓ )is the same as in(푉,퐹 ℓ+1 ). In other words, the terminal vertices that are reached with positive probability in ̄ 휎 퐹 ℓ and ̄ 휎 퐹 ℓ+1 are the same. Moreover, Invariant 4 guarantees that it is almost sure that some terminal vertex will be reached, in ̄ 휎 퐹 ℓ as well as in ̄ 휎 퐹 ℓ+1 . Therefore, the set of payoff vectors that have positive probability of being obtained is the same in both strategy profiles, hence the risk measures are the same.□ We can now prove correctness. To do so, we need to prove two propositions: the algorithm recognizes only positive instances, and recognizes all of them. Proposition 12. The algorithm recognizes only positive instances. Proof. Let us assume that the algorithm answersYesat stepℓof the final refinements, after having switched to the final refinements loop at step푘: let us show that the strategy profile ̄ 휎 퐹 ℓ is an XRSE that satisfies the desired constraints. Note that the algorithm does never answerYesat step 0, hence we necessarily have 푘 ≥ 1. The strategy profile ̄ 휎 퐹 ℓ satisfies ̄ 푥 ≤ X( ̄ 휎 퐹 ℓ ) ≤ ̄ 푦. The algorithm switches to the final refinements at step푘only if the strategy profile ̄ 휎 퐹 ℓ satisfiesX( ̄ 휎 퐸 푘 ) ≥ ̄ 푥 . Then, by Invariant 5, we also have X( ̄ 휎 퐹 ℓ ) ≥ ̄ 푥 . As for the upper bound, observe that the set퐸 1 has been defined so that the setAttr(푉 0 / ,퐸), and therefore the set푉 0 / , is not accessible from푣 0 in the graph(푉,퐸 1 ), and therefore not in the graph(푉,퐹 ℓ )either. Thus, it is almost sure in ̄ 휎 퐹 ℓ that no vertex of푉 0 / will ever be reached. In other words, all terminals that have positive probability of being reached give to each player푖a payoff smaller than푦 푖 . That is sufficient to prove the lower bound, because it is almost sure, when following ̄ 휎 퐹 ℓ , that some terminal vertex will eventually be reached, by Invariant 4. The strategy profile ̄ 휎 퐹 ℓ is an XRSE.Let푖be a player, and let휎 ′ 푖 be a deviation of player푖 from ̄ 휎 퐹 ℓ . Let푧 ′ = X 푖 ( ̄ 휎 퐹 ℓ −푖 ,휎 ′ 푖 )be the extreme risk measure obtained by player푖. We want to prove that the deviation휎 ′ 푖 is not profitable, i.e., that we have푧 ′ ≤ 푧 푘 푖 (since we haveX 푖 ( ̄ 휎 퐸 ℓ ) = 푧 푘 푖 by Invariant 5). To do so, we first show that player푖cannot obtain a payoff better than푧 푘 푖 after using an edge that does not belong to 퐹 ℓ . We first show that if the deviation 휎 ′ 푖 uses only edges of 퐹 ℓ , then it cannot be profitable. 11.3. CONSTRAINED EXISTENCE PROBLEM193 Sublemma 16. If it is almost sure, when following( ̄ 휎 퐹 ℓ −푖 ,휎 ′ 푖 ) , that only edges of퐹 ℓ will be used, then the deviation 휎 ′ 푖 is not profitable. Proof. First, let us note that as long as player푖uses only edges that belong to퐹 ℓ , the strategy profile ̄ 휎 퐹 ℓ behaves in a stationary way, and we can therefore assume without loss of generality that 휎 ′ 푖 is positional. The payoff푧 ′ may be obtained by reaching a terminal vertex: in that case, that terminal vertex is accessible from푣 0 in(푉,퐹 ℓ ), and therefore also reached with positive probability when following the strategy profile ̄ 휎 퐹 ℓ , hence 푧 ′ ≤ 푧 푘 푖 . Let us show that it cannot be obtained by reaching no terminal. We proceed by contradiction: if, in the strategy profile( ̄ 휎 퐹 ℓ ,휎 ′ 푖 ) , there is a positive probability of reaching no terminal when following that strategy profile, then there is a vertex that has positive probability of being visited infinitely often. We can then define the set푊of such vertices, i.e., the set 푊 =푣 ∈ 푉 | P ̄ 휎 퐹 ℓ ,휎 ′ 푖 (푣 ∈ Inf)>0. Thus, when the strategy profile( ̄ 휎 퐹 ℓ ,휎 ′ 푖 ) is followed from a vertex of푊, it is almost sure that no terminal vertex is reached, and that the set푊will never be left. We can then choose푤 ∈ 푊such that it has positive probability of being reached without visiting any other vertex of푊before, i.e., such that there exists a historyℎ푤from푣 0 with Occ(ℎ)∩푊 =∅ (note that ℎ can be empty). On the other hand, in the graph(푉,퐹 ℓ ), there is at least one terminal vertex accessible from 푤: all vertices from which no terminal is accessible are made themselves inaccessible at odd steps, the switch to final refinements loop happens only if there is no more edge to remove in that perspective, and Condition 3 guarantees that the final refinements loop leave at least one terminal vertex accessible from every vertex accessible from 푣 0 . From each terminal푡accessible from푤, we pick a simple pathℎ 푡 0 . . .ℎ 푡 푞 푡 fromℎ 푡 0 = 푤to ℎ 푡 푞 푡 = 푡 in the graph(푉,퐹 ℓ ). Those paths define a directed acyclic graph (DAG)퐷 =(푉 퐷 ,퐸 퐷 ) rooted at푤, where all non-terminal vertices have at least one outgoing edge, with푉 퐷 ⊆ 푉 and퐸 퐷 ⊆ 퐹 ℓ . Now, since the strategy휎 ′ 푖 guarantees that no terminal vertex will be reached, each branchℎ 푡 of that DAG is such that there exists a (smallest) index푗withℎ 푡 푗 ∈ 푊 ∩푉 푖 , and 휎 ′ 푖 (ℎ 푡 푗 ) ≠ ℎ 푡 푗+1 . It may be the case thatℎ 푡 푗 휎 ′ 푖 (ℎ 푡 푗 ) ∈ 퐸 퐷 , i.e., that fromℎ 푡 푗 , player푖proceeds to an undetectable deviation and takes another branch of the DAG. But that cannot be the case for all푡: otherwise, there would be a branchℎ 푡 that would be followed with positive probability when following( ̄ 휎 퐹 ℓ −푖 ,휎 ′ 푖 )from푤, and therefore a terminal vertex푡that would be reached with positive probability, which contradicts the definition of 푤 . There must therefore exist an edge푢푣 ∈ 퐹 ℓ \ 퐸 퐷 , with푢 ∈ 푉 퐷 ∩푉 푖 ∩푊. We will show that such an edge should have been removed during the final refinements loop. First, it immediately satisfies Condition 1. Moreover, the vertex푢is necessarily on a branchℎ 푡 of퐷that leads to a terminal vertex푡, hence it satisfies Condition 3. Finally, since푣is accessible from푤in(푉,퐹 ℓ ), the terminal vertices that are accessible from푣in(푉,퐹 ℓ )are all accessible from푤in that same graph, and therefore are accessible from푤in the DAG퐷. Since푤is accessible from푣 0 without visiting any vertex of푊, and it particular without visiting푢, it means that the terminal vertices accessible from푣are also accessible from푣 0 without using the edge푢푣. In other words, the edge푢푣 satisfies Condition 2, and should have been removed during the final refinements. This case is therefore impossible: when the players follow the strategy profile( ̄ 휎 퐹 ℓ ,휎 ′ 푖 ), it is almost sure that some terminal vertex will be reached, and that concludes the proof. □ But now, if the strategy휎 ′ 푖 does use edges that do not belong to퐹 ℓ , let us show that it cannot 194CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA be a profitable deviation either. Sublemma 17. If there is a historyℎ푣compatible with ̄ 휎 퐹 ℓ such that we have푣휎 ′ 푖 (ℎ푣) ∉ 퐹 ℓ , then we have X 푖 ( ̄ 휎 퐹 ℓ ↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ) ≤ 푧 푘 푖 . Proof. After such a history, the strategy profile ̄ 휎 퐹 ℓ −푖 follows the positional strategy profile ̄ 휏 †푖 −푖 . By definition of that strategy profile, we haveX 푖 ( ̄ 휎 퐹 ℓ ↾ℎ푣 ,휎 ′ 푖↾ℎ푣 ) ≤ val(푣) . On the other hand, the vertex푣is accessible from푣 0 in(푉,퐹 ℓ ), since it is visited with positive probability in ̄ 휎 퐹 ℓ . Therefore, it does not belong to the set 퐴 푘 (if 푘 is even) or 퐴 푘−1 (if 푘 is odd), and in particular not to the set푉 푘 / or푉 푘−1 / , which means that we have val(푣) ≤ 푧 푘 푖 . Hence the conclusion. □ In the general case, the payoffs that player푖obtains with positive probability in the strategy profile( ̄ 휎 퐹 ℓ −푖 ,휎 ′ 푖 ) are either obtained using only edges that belong to퐹 ℓ , or by using an edge that does not: in both cases, we have shown that player푖cannot get a payoff greater than푧 푘 푖 , which proves that the strategy profile ̄ 휎 퐹 ℓ is an XRSE. Algorithm 2 answersYesonly on positive instances, and outputs in that case a succinct representation of an XRSE matching the constraints.□ We will now prove the converse. Proposition 13. Algorithm 3 recognizes all positive instances. Proof. Let us assume that we have a positive instance, i.e., that there exists an XRSE ̄ 휎 with ̄ 푥 ≤ X( ̄ 휎) ≤ ̄ 푦 . Let us show that the algorithm will answerYes. To do so, we first prove the following invariant: if an edge is removed at some step before the final refinements, then it is never taken by the XRSE ̄ 휎 . Invariant 6. For each푘 ≥0, every edge that has positive probability of being eventually taken in ̄ 휎 belongs to 퐸 푘 . Proof.We prove the invariant by induction. Base case. The case 푘 = 0 is immediate, since we have 퐸 0 = 퐸. Further, at step푘 =1, if the strategy profile ̄ 휎uses eventually, with positive probability, an edge that does not belong to퐸 1 , then it goes with positive probability to a vertex푣 ∈ Attr(푉 0 / ,퐸 0 ). Then, with positive probability, a terminal vertex will be reached that gives to some player푖a payoff greater than푦 푖 , which is impossible. Therefore, such an edge cannot be taken in ̄ 휎 . Induction step. Let us assume that the invariant is true until step푘 ≥1, and let us show that it holds at step푘+1. Let푢푣be an edge that is used with positive probability when following ̄ 휎, and let us assume toward contradiction that it does not belong to퐸 푘+1 . Since the invariant is true at each step until푘, we can assume that푢푣has been removed at step푘, i.e., that we have푢푣 ∈ 퐸 푘 \ 퐸 푘+1 . Then, we have푢 ∉ 퐴 푘 and푣 ∈ 퐴 푘 . The strategy profile ̄ 휎has therefore positive probability of visiting the set 퐴 푘 . We must now distinguish the cases where푘is even or odd. If푘is even, then the strategy profile ̄ 휎has therefore positive probability of visiting a vertex푣 ∈ 푉 푘 / . If player푖is the player controlling푣, then that player has a deviation in which they get risk measure at least val(푣)> 푧 푘 푖 . Let us now note that when the players follow the strategy profile ̄ 휎, it is almost sure that a terminal vertex will eventually be reached, and that all the terminal vertices that 11.3. CONSTRAINED EXISTENCE PROBLEM195 have positive probability of being reached also have positive probability of being reached in ̄ 휎 퐸 푘 , since the invariant is true at step푘: therefore, we have푧 푘 푖 ≥ X 푖 ( ̄ 휎), and player푖has a profitable deviation after 푣 , which contradicts the fact that ̄ 휎 is an XRSE. If푘is odd, then, by definition of퐴 푘 , when the strategy profile ̄ 휎is followed, there is a positive probability of reaching no terminal. But that is impossible in the cycle-averse case. The invariant is therefore necessarily still true at step 푘+ 1.□ We are now able to conclude. The answer No can be given in the two following cases: •If at step푘, we have푣 0 ∈ Attr(푉 푘 / ,퐸 푘 ). Then, the strategy profile ̄ 휎 visits the set푉 푘 / with positive probability. With the same arguments that were used in the proof of Invariant 6, that is not possible. •If during steps푘−1 and푘, no edge is removed, but we have푧 푘 푖 < 푥 푖 for some player푖. Since ̄ 휎 uses only edges of퐸 푘 by Invariant 6 and reaches almost surely a terminal (since we are in the cycle-averse case), we haveX 푖 ( ̄ 휎) ≤ 푧 푘 푖 , and thereforeX 푖 ( ̄ 휎)< 푥 푖 : that case is therefore also impossible by definition of ̄ 휎 . None of those cases is possible, hence our algorithm will eventually answer Yes.□ ▶Complexities We consider here the complexity of the two algorithms. Since at least one edge is removed every two steps, there are푂(푚)steps. In each of them, we need푂(푝)calls to simple algorithms: computation of 푧 푘 푖 , of푉 푘 / , of 퐴 푘 . Hence the complexity 푂(푝푚 2 ). In the cycle-averse case, when an output is asked, we need to add the final refinements loop, which consist of푂(푚)additional steps in which we check, for each of the푂(푚)remaining edges, whether they satisfy Conditions 1, 2, and 3: that can be done in time푂(푚). Hence the complexity 푂(푝푚 2 +푚 3 ).□ Finally, we show that the problem is P-hard, even when there are only two players. Lemma 40. The constrained existence problem of XRSEs with optimistic players is P-hard even with only two players. Proof. Given a deterministic two-player (between players ◦ and□) zero-sum reachability gameG ↾푣 0 with target set of vertices푇, we construct a simple stochastic game (with no stochastic vertices) where there is an XRSE ̄ 휎 satisfying X ◦ ( ̄ 휎) = 1 and X □ ( ̄ 휎) =−1 if and only if player ◦ wins the game. The game is simply obtained by assigning rewards on the zero-sum two player game as follows: we make all nodes in the target set푇of the reachability game as a terminal node where player ◦ gets reward 1 and player□the reward−1. Recall that if no terminal is reached, both players get reward 0. If ◦ has a strategy to win the reachability game, then the same strategy for ◦ , along with any strategy for□, will be an XRSE in that new game, and then satisfies the constraint. Similarly, if on the other hand, player□has a strategy to avoid the states푇, then no strategy of ◦ that gives her payoff +1 and gives player□the payoff−1 will be an equilibrium, since□can always deviate to the winning strategy in the reachability game that offers him the better payoff of 0.□ 196CHAPTER 11. EXTREME RISK-SENSITIVE EQUILIBRIA DISCUSSION 197 DISCUSSION Throughout this document, we have introduced new techniques and complexity results concerning Nash equilibria and SPEs in parity, mean-payoff, discounted-sum, and energy games. We have also shown how these results extend to rational verification. We have studied strong secure equilibria in parity and 휔-regular games, demonstrating that this concept effectively meets key criteria for secure protocols. Regarding stochastic games and the use of randomized strategies, we considered the entropic risk measure, proving that the constrained existence problem for equilibria defined with this measure is undecidable. However, we established decidability results under certain restrictions on players’ strategies. We then introduced the extreme risk measure and demonstrated that, when equilibria are defined using this measure, the constrained existence problem becomes decidable. We believe these results provide novel insights into the study of multi-player games and their applica- tions in computer science. Several questions remain open, such as the complexity of the SPE constrained existence problem in Büchi games (Open Problem 1), its fixed-parameter tractability in mean-payoff games (Open Problem 2), its recursive enumerability in energy games (Open Problem 3), the rational synthesis problem in mean-payoff games (Open Problem 4), the existence of randomized휀-SPEs in mean-payoff games (Conjecture 1), or the existence of entropic or extreme risk-sensitive equilibria in simple stochastic games with negative rewards (Conjectures 2 and 3). More broadly, interesting research directions emerge from studying alternative or more sophisticated classes of games. For instance, multiplayer games where players pursue a mix of parity and mean-payoff objectives, as in [CHJ05b], or games where some players have parity objectives while others have mean-payoff ones. Additionally, the unconventional equilibrium notions we introduce, such as extreme risk-sensitive equilibria, could be explored within the game classes where we have studied NEs and SPEs: parity games, mean-payoff games, discounted-sum games, and energy games. Such equilibria could also be examined in concurrent or imperfect information games, or combined with other equilibrium properties, leading to concepts such as extreme risk-sensitive subgame-perfect equilibria or strong secure extreme risk-sensitive equilibria. The application of these results must, however, be approached with caution. Like all mathematical results, the proofs presented in this work offer the significant advantage of being non-ambiguous and fully verifiable, but they deal with abstract objects. Modeling real-life situations through such abstractions is inherently a simplification, carrying the risk of omitting crucial elements. Applications of game theory, particularly in economics, have faced criticism [Rub13,Gue00] for their ambition to describe—and even predict—human behavior, despite empirical observations often contradicting theoretical predictions. Indeed, when multiple individuals must make decisions based on their beliefs about others’ choices, neither empirical studies nor theoretical proofs guarantee that they will select strategies forming a Nash equilibrium or any other equilibrium notion. In such situations, the conceptual tools of game theory primarily serve to clarify the dilemmas that agents face rather than to 199 200 predict their behavior—though it is worth noting that designing game-theoretic notions that better align with human behavior is an active area of research (see, for instance, Herbert A. Simon’s work on the satisficing concept [Sim56]). In contrast, computer-science-related applications may represent some of the most relevant uses of game theory, as they often have a more normative than descriptive objective, in Bernard Guerrien’s terminology [Gue00]. In scenarios involving multiple agents, each with a set of possible strategies and well-defined objectives, the existence of a strong secure equilibrium provides a concrete guarantee: if the agents are given such an equilibrium as a protocol to follow, no coalition can disrupt the protocol to harm another agent without also harming one of its own members. Such a protocol could, therefore, be reasonably proposed as a safe protocol. However, the (legitimate) desire to explore all the potential applications of game theory in our field may also lead to interpretations that attribute to it more predictive power than it actually possesses. For instance, the underlying approaches of rational verification or rational synthesis may be more debatable. Is it reasonable to consider a system safe if it guarantees a safety property only against responses that form a Nash equilibrium or an SPE? Such claims should be carefully examined within specific application contexts, supported by both solid theoretical foundations and empirical validation. However, we believe that developing a deeper understanding of these concepts through algorithmic results can only contribute positively in that regard. BIBLIOGRAPHY [ADGH06]Ittai Abraham, Danny Dolev, Rica Gonen, and Joe Halpern. Distributed computing meets game theory: robust mechanisms for rational secret sharing and multiparty computation. In Proceedings of the 25th Annual ACM Symposium on Principles of Distributed Computing (PODC’06), page 53–62, New York, NY, USA, 2006. Association for Computing Machinery. doi:10.1145/1146381.1146393. [AG11]Krzysztof R. Apt and Erich Gradel. Lectures in Game Theory for Computer Scientists. Cam- bridge University Press, USA, 1st edition, 2011. [AL87]E. Allen Emerson and Chin-Laung Lei.Modalities for model checking: branching time logic strikes back. Science of Computer Programming, 8(3):275–306, 1987. URL: https://w.sciencedirect.com/science/article/pii/0167642387900360,doi:10. 1016/0167-6423(87)90036-0. [Aue18]M. Auer. Hands-On Value-at-Risk and Expected Shortfall: A Practical Primer. Management for Professionals. Springer International Publishing, 2018. URL:https://books.google.at/ books?id=4EFKDwAAQBAJ. [BBG + 19]Thomas Brihaye, Véronique Bruyère, Aline Goeminne, Jean-François Raskin, and Marie van den Bogaard. The complexity of subgame perfect equilibria in quantitative reachability games. In Wan J. Fokkink and Rob van Glabbeek, editors, 30th Int. Conf. on Concurrency Theory, CONCUR 2019, August 27-30, 2019, Amsterdam, the Netherlands, volume 140 of LIPIcs, pages 13:1–13:16. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2019.doi: 10.4230/LIPIcs.CONCUR.2019.13. [BBMR15]Thomas Brihaye, Véronique Bruyère, Noémie Meunier, and Jean-François Raskin. Weak subgame perfect equilibria and their application to quantitative reachability. In 24th EACSL Annual Conference on Computer Science Logic, CSL 2015, September 7-10, 2015, Berlin, Germany, volume 41 of LIPIcs, pages 504–518. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2015. [BBMU15]Patricia Bouyer, Romain Brenguier, Nicolas Markey, and Michael Ummels. Pure Nash equilibria in concurrent deterministic games. Log. Methods Comput. Sci., 11(2), 2015.doi: 10.2168/LMCS-11(2:9)2015. [BCH + 16]Romain Brenguier, Lorenzo Clemente, Paul Hunter, Guillermo A. Pérez, Mickael Randour, Jean-François Raskin, Ocan Sankur, and Mathieu Sassolas. Non-zero sum games for reactive synthesis. In Language and Automata Theory and Applications - 10th International Conference, LATA 2016, Prague, Czech Republic, March 14-18, 2016, Proceedings, volume 9618 of Lecture Notes in Computer Science, pages 3–23. Springer, 2016. 201 202BIBLIOGRAPHY [BCK10]Michael Backes, Oana Ciobotaru, and Anton Krohmer. Ratfish: A file sharing protocol provably secure against rational users. In Dimitris Gritzalis, Bart Preneel, and Marianthi Theoharidou, editors, Proceedings of the 5th European Symposium on Research in Computer Security (ESORICS’10), pages 607–625, Berlin, Heidelberg, 9 2010. Springer.doi:10.1007/ 978-3-642-15497-3_37. [BCMP24]Christel Baier, Krishnendu Chatterjee, Tobias Meggendorfer, and Jakob Piribauer. Entropic risk for turn-based stochastic games. Information and Computation, 301:105214, 2024.doi: 10.1016/j.ic.2024.105214. [BDS13]Thomas Brihaye, Julie De Pril, and Sven Schewe. Multiplayer cost games with simple Nash equilibria. In Sergei N. Artëmov and Anil Nerode, editors, Logical Foundations of Computer Science, Int. Symp., LFCS 2013, San Diego, CA, USA, January 6-8, 2013. Proc., volume 7734 of Lecture Notes in Computer Science, pages 59–73. Springer, 2013.doi:10.1007/ 978-3-642-35722-0\_5. [BFL + 08]Patricia Bouyer, Ulrich Fahrenberg, Kim Guldstrand Larsen, Nicolas Markey, and Jirí Srba. Infinite runs in weighted timed automata with energy constraints. In Franck Cassez and Claude Jard, editors, Formal Modeling and Analysis of Timed Systems, 6th Int. Conf., FORMATS 2008, Saint Malo, France, September 15-17, 2008. Proc., volume 5215 of Lecture Notes in Computer Science, pages 33–47. Springer, 2008. doi:10.1007/978-3-540-85778-5\_4. [BHC04]Levente Buttyán, Jean-Pierre Hubaux, and Srdjan Capkun. A formal model of rational exchange and its application to the analysis of syverson’s protocol. Journal of Computer Security, 12(3-4):551–587, 2004. doi:10.3233/JCS-2004-123-408. [BHO15]Udi Boker, Thomas A. Henzinger, and Jan Otop. The target discounted-sum problem. In 30th Annual ACM/IEEE Symp. on Logic in Computer Science, LICS 2015, Kyoto, Japan, July 6-10, 2015, pages 750–761. IEEE Computer Society, 2015. doi:10.1109/LICS.2015.74. [BHR18a]Véronique Bruyère, Quentin Hautem, and Jean-François Raskin. Parameterized complexity of games with monotonically ordered omega-regular objectives. In Sven Schewe and Lijun Zhang, editors, 29th International Conference on Concurrency Theory, CONCUR 2018, September 4-7, 2018, Beijing, China, volume 118 of LIPIcs, pages 29:1–29:16. Schloss Dagstuhl - Leibniz- Zentrum für Informatik, 2018. doi:10.4230/LIPIcs.CONCUR.2018.29. [BHR18b] Véronique Bruyère, Quentin Hautem, and Jean-François Raskin. Parameterized complexity of games with monotonically ordered omega-regular objectives. In Sven Schewe and Lijun Zhang, editors, 29th International Conference on Concurrency Theory, CONCUR 2018, September 4-7, 2018, Beijing, China, volume 118 of LIPIcs, pages 29:1–29:16. Schloss Dagstuhl - Leibniz- Zentrum für Informatik, 2018. URL:https://doi.org/10.4230/LIPIcs.CONCUR.2018.29, doi:10.4230/LIPICS.CONCUR.2018.29. [BHT25]Léonard Brice, Thomas Henzinger, and K. S. Thejaswini. Finding equilibria: simpler for pessimists, simplest for optimists, 2025. URL:https://arxiv.org/abs/2502.05316, arXiv:2502.05316. [Bis06] Christopher M. Bishop.Pattern Recognition and Machine Learning.Springer, 2006.URL:https://w.microsoft.com/en-us/research/publication/ pattern-recognition-machine-learning/. BIBLIOGRAPHY203 [BMR14]Véronique Bruyère, Noémie Meunier, and Jean-François Raskin. Secure equilibria in weighted games. In CSL-LICS, pages 26:1–26:26. ACM, 2014. [Bou19]Patricia Bouyer. On the computation of Nash equilibria in games on graphs. In Johann Gam- per, Sophie Pinchinat, and Guido Sciavicco, editors, 26th International Symposium on Temporal Representation and Reasoning (TIME 2019), volume 147 of Leibniz International Proceedings in Informatics (LIPIcs), pages 3:1–3:3, Dagstuhl, Germany, 2019. Schloss Dagstuhl – Leibniz- Zentrum für Informatik. URL:https://drops.dagstuhl.de/entities/document/10. 4230/LIPIcs.TIME.2019.3, doi:10.4230/LIPIcs.TIME.2019.3. [BR14]Nicole Bäuerle and Ulrich Rieder. More risk-sensitive Markov decision processes. Mathe- matics of Operations Research, 39(1):105–120, 2014. [BR15]Romain Brenguier and Jean-François Raskin. Pareto curves of multidimensional mean-payoff games. In Daniel Kroening and Corina S. Pasareanu, editors, Computer Aided Verification - 27th International Conference, CAV 2015, San Francisco, CA, USA, July 18-24, 2015, Proceedings, Part I, volume 9207 of Lecture Notes in Computer Science, pages 251–267. Springer, 2015. doi:10.1007/978-3-319-21668-3\_15. [Bra99]Hans Wolfgang Brachinger. From variance to value at risk: A unified perspective on standard- ized risk measures. In Wolfgang Gaul and Hermann Locarek-Junge, editors, Classification in the Information Age, pages 91–99, Berlin, Heidelberg, 1999. Springer Berlin Heidelberg. [Bre16] Romain Brenguier. Robust equilibria in mean-payoff games. In Bart Jacobs and Christof Löding, editors, Proceedings of the 19th International Conference on Foundations of Software Science and Computation Structures (FOSSACS’16), volume 9634 of Lecture Notes in Computer Science, pages 217–233. Springer, April 2016. doi:10.1007/978-3-662-49630-5\_13. [Brø83] Arne Brøndsted. An Introduction to Convex Polytopes, volume 90 of Graduate Texts in Mathematics. Springer-Verlag, 1983. doi:10.1007/978-1-4612-5460-8. [BRPR17]Véronique Bruyère, Stéphane Le Roux, Arno Pauly, and Jean-François Raskin. On the exis- tence of weak subgame perfect equilibria. In Foundations of Software Science and Computation Structures - 20th International Conference, FOSSACS 2017, Held as Part of the European Joint Conferences on Theory and Practice of Software, ETAPS 2017, Uppsala, Sweden, April 22-29, 2017, Proceedings, volume 10203 of Lecture Notes in Computer Science, pages 145–161, 2017. [BRRvdB24]Véronique Bruyère, Jean-François Raskin, Alexis Reynouard, and Marie van den Bogaard. The non-cooperative rational synthesis problem for subgame perfect equilibria and omega- regular objectives. CoRR, abs/2412.08547, 2024. URL:https://doi.org/10.48550/arXiv. 2412.08547, arXiv:2412.08547, doi:10.48550/ARXIV.2412.08547. [BRS + 24]Léonard Brice, Jean-François Raskin, Mathieu Sassolas, Guillaume Scerri, and Marie van den Bogaard. Pessimism of the will, optimism of the intellect: Fair protocols with malicious but rational agents. CoRR, abs/2405.18958, 2024. URL:https://doi.org/10.48550/arXiv. 2405.18958, arXiv:2405.18958, doi:10.48550/ARXIV.2405.18958. [Bru17]Véronique Bruyère. Computer aided synthesis: A game-theoretic approach. In Developments in Language Theory - 21st International Conference, DLT 2017, Liège, Belgium, August 7-11, 204BIBLIOGRAPHY 2017, Proceedings, volume 10396 of Lecture Notes in Computer Science, pages 3–35. Springer, 2017. [BRvdB21]Léonard Brice, Jean-François Raskin, and Marie van den Bogaard. Subgame-perfect equilibria in mean-payoff games. In Serge Haddad and Daniele Varacca, editors, 32nd International Conference on Concurrency Theory, CONCUR 2021, August 24-27, 2021, Virtual Conference, volume 203 of LIPIcs, pages 8:1–8:17. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2021. URL:https://doi.org/10.4230/LIPIcs.CONCUR.2021.8,doi:10.4230/LIPICS. CONCUR.2021.8. [BRvdB22a]Léonard Brice, Jean-François Raskin, and Marie van den Bogaard. The complexity of SPEs in mean-payoff games. In Mikolaj Bojanczyk, Emanuela Merelli, and David P. Woodruff, editors, 49th International Colloquium on Automata, Languages, and Programming, ICALP 2022, July 4-8, 2022, Paris, France, volume 229 of LIPIcs, pages 116:1–116:20. Schloss Dagstuhl - Leibniz- Zentrum für Informatik, 2022. URL:https://doi.org/10.4230/LIPIcs.ICALP.2022.116, doi:10.4230/LIPICS.ICALP.2022.116. [BRvdB22b]Léonard Brice, Jean-François Raskin, and Marie van den Bogaard. On the complexity of SPEs in parity games. In Florin Manea and Alex Simpson, editors, 30th EACSL Annual Conference on Computer Science Logic, CSL 2022, February 14-19, 2022, Göttingen, Germany (Virtual Conference), volume 216 of LIPIcs, pages 10:1–10:17. Schloss Dagstuhl - Leibniz- Zentrum für Informatik, 2022. URL:https://doi.org/10.4230/LIPIcs.CSL.2022.10, doi:10.4230/LIPICS.CSL.2022.10. [BRvdB23a]Léonard Brice, Jean-François Raskin, and Marie van den Bogaard. Rational verification for Nash and subgame-perfect equilibria in graph games. In Jérôme Leroux, Sylvain Lom- bardy, and David Peleg, editors, 48th International Symposium on Mathematical Foundations of Computer Science, MFCS 2023, August 28 to September 1, 2023, Bordeaux, France, vol- ume 272 of LIPIcs, pages 26:1–26:15. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2023. URL:https://doi.org/10.4230/LIPIcs.MFCS.2023.26,doi:10.4230/LIPICS. MFCS.2023.26. [BRvdB23b]Léonard Brice, Jean-François Raskin, and Marie van den Bogaard. Subgame-perfect equilibria in mean-payoff games (journal version). Log. Methods Comput. Sci., 19(4), 2023. URL: https://doi.org/10.46298/lmcs-19(4:6)2023, doi:10.46298/LMCS-19(4:6)2023. [Can88] John Canny. Some algebraic and geometric computations in pspace. In Proceedings of the Twentieth Annual ACM Symposium on Theory of Computing STOC, STOC ’88, page 460–467. Association for Computing Machinery, 1988. doi:10.1145/62212.62257. [Cap23]H2 Gambling Capital. Global gambling industry generates $536bn in 2023 with 7% growth expected in 2024, 2023. Accessed: 2025-01-15. URL:https://h2gc.com/news/general/ global-gambling-industry-generates-536bn-in-2023-with-7-growth-expected-in-2024 . [CDE + 10]Krishnendu Chatterjee, Laurent Doyen, Herbert Edelsbrunner, Thomas A. Henzinger, and Philippe Rannou. Mean-payoff automaton expressions. In Paul Gastin and François Laroussinie, editors, CONCUR 2010 - Concurrency Theory, 21th International Conference, CONCUR 2010, Paris, France, August 31-September 3, 2010. Proceedings, volume 6269 of Lecture Notes in Computer Science, pages 269–283. Springer, 2010. BIBLIOGRAPHY205 [CDG + 07]Hubert Comon-Lundh, Max Dauchet, Rémi Gilleron, Cristof Löding, Florent Jacquemard, Denis Lugiez, Sophie Tison, and Marc Tommasi. Tree Automata Techniques and Applications. Université de Lille, ENS Cachan, RWTH Aachen, Université Aix-Marseille, INRIA, CNRS, November 2007. URL: https://inria.hal.science/hal-03367725v1/file/tata.pdf. [CFGR16] Rodica Condurache, Emmanuel Filiot, Raffaella Gentilini, and Jean-François Raskin. The complexity of rational synthesis. In Ioannis Chatzigiannakis, Michael Mitzenmacher, Yuval Rabani, and Davide Sangiorgi, editors, 43rd International Colloquium on Automata, Languages, and Programming, ICALP 2016, July 11-15, 2016, Rome, Italy, volume 55 of LIPIcs, pages 121:1– 121:15. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2016. URL:https://doi.org/ 10.4230/LIPIcs.ICALP.2016.121, doi:10.4230/LIPICS.ICALP.2016.121. [CHJ05a]Krishnendu Chatterjee, Thomas A. Henzinger, and Marcin Jurdziński. Games with secure equilibria. In Frank S. de Boer, Marcello M. Bonsangue, Susanne Graf, and Willem-Paul de Roever, editors, Formal Methods for Components and Objects, pages 141–161, Berlin, Heidelberg, 2005. Springer Berlin Heidelberg. [CHJ05b]Krishnendu Chatterjee, Thomas A. Henzinger, and Marcin Jurdzinski. Mean-payoff parity games. In 20th IEEE Symposium on Logic in Computer Science (LICS 2005), 26-29 June 2005, Chicago, IL, USA, Proceedings, pages 178–187. IEEE Computer Society, 2005.doi:10.1109/ LICS.2005.26. [CHJ06]Krishnendu Chatterjee, Thomas A. Henzinger, and Marcin Jurdziński. Games with secure equilibria. Theoretical Computer Science, 365(1):67–82, 2006. Formal Methods for Com- ponents and Objects. URL:https://w.sciencedirect.com/science/article/pii/ S0304397506004816, doi:10.1016/j.tcs.2006.07.032. [CHP10] Krishnendu Chatterjee, Thomas A. Henzinger, and Nir Piterman. Strategy logic. Inf. Comput., 208(6):677–693, 2010. doi:10.1016/j.ic.2009.07.004. [CMJ04]Krishnendu Chatterjee, Rupak Majumdar, and Marcin Jurdzinski. On Nash equilibria in stochastic games. In Computer Science Logic, 18th International Workshop, CSL 2004, volume 3210 of Lecture Notes in Computer Science, pages 26–40. Springer, 2004.doi:10.1007/ 978-3-540-30124-0\_6. [Cou38] Augustin Cournot. Recherches sur les principes mathématiques de la théorie des richesses. Hachette, Paris, 1838. Edition originale. [CR12]Krishnendu Chatterjee and Vishwanath Raman. Synthesizing protocols for digital contract signing. In Viktor Kuncak and Andrey Rybalchenko, editors, Verification, Model Checking, and Abstract Interpretation, pages 152–168, Berlin, Heidelberg, 2012. Springer Berlin Heidelberg. [CS10]Michael R. Clarkson and Fred B. Schneider. Hyperproperties. Journal of Computer Security, 18(6):1157–1210, September 2010. doi:10.3233/JCS-2009-0393. [Eve57]Hugh Everett. Recursive games. In C. Berg and M. Dresher, editors, Contributions to the Theory of Games, Volume I, pages 47–78. Princeton University Press, 1957. URL:https:// w.degruyter.com/document/doi/10.1515/9781400829156-014/html,doi:10.1515/ 9781400829156-014. 206BIBLIOGRAPHY [FGR20]Emmanuel Filiot, Raffaella Gentilini, and Jean-François Raskin. The adversarial Stackel- berg value in quantitative games. In Artur Czumaj, Anuj Dawar, and Emanuela Merelli, editors, 47th International Colloquium on Automata, Languages, and Programming (ICALP 2020), volume 168 of Leibniz International Proceedings in Informatics (LIPIcs), pages 127:1– 127:18, Dagstuhl, Germany, 2020. Schloss Dagstuhl – Leibniz-Zentrum für Informatik. URL:https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ICALP.2020. 127, doi:10.4230/LIPIcs.ICALP.2020.127. [FK89]Jerzy Filar and Lodewijk Kallenberg. Variance-penalized Markov decision process. Mathe- matics of Operations Research, 14, 02 1989. doi:10.1287/moor.14.1.147. [FKL10]Dana Fisman, Orna Kupferman, and Yoad Lustig. Rational synthesis. In Javier Esparza and Rupak Majumdar, editors, Tools and Algorithms for the Construction and Analysis of Systems, pages 190–204, Berlin, Heidelberg, 2010. Springer Berlin Heidelberg. [FKM + 10] János Flesch, Jeroen Kuipers, Ayala Mashiah-Yaakovi, Gijs Schoenmakers, Eilon Solan, and Koos Vrieze. Perfect-information games with lower-semicontinuous payoffs. Math. Oper. Res., 35(4):742–755, 2010. [FP17]János Flesch and Arkadi Predtetchinski. A characterization of subgame-perfect equilibrium plays in Borel games of perfect information. Math. Oper. Res., 42(4):1162–1179, 2017.doi: 10.1287/moor.2016.0843. [FS02] Hans Föllmer and Alexander Schied. Convex measures of risk and trading constraints. Finance and Stochastics, 6(4):429–447, Oct 2002. doi:10.1007/s007800200072. [GHM25]Jorge Gallego-Hernández and Alessio Mansutti. On the Existential Theory of the Reals En- riched with Integer Powers of a Computable Number. In Olaf Beyersdorff, Michal Pilipczuk, Elaine Pimentel, and Nguy ̃ ên Kim Th ́ ăng, editors, 42nd International Symposium on Theoreti- cal Aspects of Computer Science (STACS 2025), volume 327 of Leibniz International Proceedings in Informatics (LIPIcs), pages 37:1–37:18, Dagstuhl, Germany, 2025. Schloss Dagstuhl – Leibniz- Zentrum für Informatik. URL:https://drops.dagstuhl.de/entities/document/10. 4230/LIPIcs.STACS.2025.37, doi:10.4230/LIPIcs.STACS.2025.37. [GNPW23] Julian Gutierrez, Muhammad Najib, Giuseppe Perelli, and Michael J. Wooldridge. On the complexity of rational verification. Ann. Math. Artif. Intell., 91(4):409–430, 2023. URL: https://doi.org/10.1007/s10472-022-09804-3, doi:10.1007/S10472-022-09804-3. [GU08]Erich Grädel and Michael Ummels. Solution concepts and algorithms for infinite multiplayer games. In Krzysztof Apt and Robert van Rooij, editors, New Perspectives on Games and Interaction, volume 4 of Texts in Logic and Games, pages 151–178. Amsterdam University Press, 2008. [Gue00] Bernard Guerrien. À quoi sert la théorie des jeux ? Blog post on autisme-economie.org, 2000. Available online: http://w.autisme-economie.org. [HD05]Paul Hunter and Anuj Dawar. Complexity bounds for regular games. In Joanna Je ̧drzejowicz and Andrzej Szepietowski, editors, Mathematical Foundations of Computer Science 2005, pages 495–506, Berlin, Heidelberg, 2005. Springer Berlin Heidelberg. BIBLIOGRAPHY207 [HK15]Christoph Haase and Stefan Kiefer. The odds of staying on budget. In Automata, Languages, and Programming, pages 234–246, Berlin, Heidelberg, 2015. Springer Berlin Heidelberg. [HLP23]Daniel Hausmann, Mathieu Lehaut, and Nir Pitermann. Symbolic reactive synthesis for the safety and EL-fragment of LTL, 2023. arXiv:2305.02793. [HM72] Ronald A. Howard and James E. Matheson. Risk-sensitive Markov decision processes. Management Science, 18(7):356–369, 1972. URL:http://w.jstor.org/stable/2629352. [HS20]Kristoffer Arnsfelt Hansen and Steffan Christ Sølvsten.∃∖-completeness of stationary nash equilibria in perfect information stochastic games. In Javier Esparza and Daniel Král’, editors, 45th International Symposium on Mathematical Foundations of Computer Science, MFCS 2020, August 24-28, 2020, Prague, Czech Republic, volume 170 of LIPIcs, pages 45:1–45:15. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2020. URL:https://doi.org/10.4230/ LIPIcs.MFCS.2020.45, doi:10.4230/LIPICS.MFCS.2020.45. [Kar78]Richard M. Karp. A characterization of the minimum cycle mean in a digraph. Discret. Math., 23(3):309–311, 1978. [KFSV21] Jeroen Kuipers, János Flesch, Gijs Schoenmakers, and Koos Vrieze. Subgame perfection in recursive perfect information games. Economic Theory, 71:603–662, 2021.doi:10.1007/ s00199-020-01260-6. [KM18]Jan Křetínský and Tobias Meggendorfer. Conditional value-at-risk for reachability and mean payoff in Markov decision processes. In Proceedings of the 33rd Annual ACM/IEEE Symposium on Logic in Computer Science, LICS ’18, page 609–618. Association for Computing Machinery, 2018. doi:10.1145/3209108.3209176. [Kop06]Eryk Kopczynski. Half-positional determinacy of infinite games. In Michele Bugliesi, Bart Preneel, Vladimiro Sassone, and Ingo Wegener, editors, Automata, Languages and Programming, 33rd International Colloquium, ICALP 2006, Venice, Italy, July 10-14, 2006, Proceedings, Part I, volume 4052 of Lecture Notes in Computer Science, pages 336–347. Springer, 2006. [KPV16] Orna Kupferman, Giuseppe Perelli, and Moshe Y. Vardi. Synthesis with rational envi- ronments. Ann. Math. Artif. Intell., 78(1):3–20, 2016. URL:https://doi.org/10.1007/ s10472-016-9508-8, doi:10.1007/S10472-016-9508-8. [KR03]Steve Kremer and Jean-François Raskin. A game-based verification of non-repudiation and fair exchange protocols. Journal of Computer Security, 11(3):399–429, 2003. URL: http://w.lsv.ens-cachan.fr/Publis/PAPERS/PS/Kremer-gameNRextended.ps. [Kuh53]Harold W. Kuhn. Extensive games and the problem of information. In Harold W. Kuhn and Albert W. Tucker, editors, Contributions to the Theory of Games, Volume I, volume 28 of Annals of Mathematics Studies, pages 193–216. Princeton University Press, 1953. [LRP14]Stéphane Le Roux and Arno Pauly. Infinite sequential games with real-valued payoffs. In Proceedings of the Joint Meeting of the Twenty-Third EACSL Annual Conference on Computer Science Logic (CSL) and the Twenty-Ninth Annual ACM/IEEE Symposium on Logic in Computer Science (LICS), CSL-LICS ’14, New York, NY, USA, 2014. Association for Computing Machinery. doi:10.1145/2603088.2603120. 208BIBLIOGRAPHY [Mar75]Donald A. Martin. Borel determinacy. Annals of Mathematics, pages 363–371, 1975. [Mar98]Donald A. Martin. The determinacy of blackwell games. J. Symb. Log., 63(4):1565–1581, 1998. doi:10.2307/2586667. [Meg22]Tobias Meggendorfer. Risk-aware stochastic shortest path. Proceedings of the AAAI Conference on Artificial Intelligence, pages 9858–9867, 2022. doi:10.1609/AAAI.V36I9.21222. [Meu16] Noémie Meunier. Multi-Player Quantitative Games: Equilibria and Algorithms. PhD thesis, Université de Mons, 2016. [Min61]Marvin L Minsky. Recursive unsolvability of post’s problem of "tag" and other topics in theory of turing machines. Annals of Mathematics, 74(3):437–455, 1961. [MT11]Shie Mannor and John N. Tsitsiklis. Mean-variance optimization in Markov decision processes. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, pages 177–184. Omnipress, 2011. URL: https://icml.c/2011/papers/156_icmlpaper.pdf. [Nas51]John Nash. Non-cooperative games. Annals of Mathematics, 54(2):286–295, 1951.doi: 10.2307/1969529. [Now05] Andrzej S. Nowak. Notes on Risk-Sensitive Nash Equilibria, pages 95–109. Birkhäuser Boston, Boston, MA, 2005. doi:10.1007/0-8176-4429-6_5. [PDM20]Bernardo K. Pagnoncelli, Oscar Dowson, and David P. Morton. Multistage stochastic pro- grams with the entropic risk measure, 2020. URL:https://optimization-online.org/ 2020/08/7984/. [PG99]Henning Pagnia and Felix C Gärtner. On the impossibility of fair exchange without a trusted third party. Technical Report TUD-BS-1999-02, Darmstadt University of Technology, 1999. [PSB22]Jakob Piribauer, Ocan Sankur, and Christel Baier. The variance-penalized stochastic shortest path problem. In 49th International Colloquium on Automata, Languages, and Programming, ICALP, volume 229 of LIPIcs, pages 129:1–129:19. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2022. doi:10.4230/LIPICS.ICALP.2022.129. [RBS24]Itsaka Rakotonirina, Gilles Barthe, and Clara Schneidewind. Decision and complexity of Dolev-Yao hyperproperties. Proceedings of ACM on Programming Languages, 8(POPL), January 2024. doi:10.1145/3632906. [RCDH07] Jean-François Raskin, Krishnendu Chatterjee, Laurent Doyen, and Thomas A. Henzinger. Algorithms for omega-regular games with imperfect information. Log. Methods Comput. Sci., 3(3), 2007. doi:10.2168/LMCS-3(3:4)2007. [RRS15]Mickael Randour, Jean-François Raskin, and Ocan Sankur. Variations on the stochastic shortest path problem. In Verification, Model Checking, and Abstract Interpretation, pages 1–18. Springer, 2015. [RSV05] J.-F. Raskin, M. Samuelides, and L. Van Begin. Games for counting abstractions. Electronic Notes in Theoretical Computer Science, 128(6):69–85, 2005. Proceedings of the Fouth Interna- tional Workshop on Automated Verification of Critical Systems (AVoCS 2004). URL:https: //w.sciencedirect.com/science/article/pii/S1571066105002379,doi:10.1016/ j.entcs.2005.04.005. BIBLIOGRAPHY209 [Rub13]Ariel Rubinstein. How game theory will solve the problems of the euro bloc and stop Iranian nukes. Frankfurter Allgemeine Zeitung, March 2013. [Sch07]Ricardo Schumann. Links oder rechts; das ist hier die Frage. Eine spieltheoretische Analyse von Elfmeterschüssen mit Bundesligadaten. Technical Report 47, Institut für Soziologie, Universität Leipzig, January 2007. Arbeitsberichte des Instituts für Soziologie der Univer- sität Leipzig. URL:https://archive.wikiwix.com/cache/index2.php?url=http%3A% 2F%2Fwww.uni-leipzig.de%2F~sozio%2Fcontent%2Fsite%2Fa_berichte%2F47.pdf. [Sel65]Reinhard Selten. Spieltheoretische Behandlung eines Oligopolmodells mit Nachfrageträgheit: Teil I: Bestimmung des dynamischen Preisgleichgewichts. Zeitschrift für die gesamte Staatswis- senschaft / Journal of Institutional and Theoretical Economics, 121(2):301–324, 1965. URL: http://w.jstor.org/stable/40748884. [Sim56]Herbert A. Simon. Rational choice and the structure of the environment. Psychological Review, 63(2):129–138, 1956. doi:10.1037/h0042769. [Sip97]Michael Sipser. Introduction to the theory of computation. PWS Publishing Company, 1997. [SV03]Eilon Solan and Nicolas Vieille. Deterministic multi-player Dynkin games. Journal of Mathematical Economics, 39(8):911–929, 2003. URL:https://w.sciencedirect.com/ science/article/pii/S0304406803000211, doi:10.1016/S0304-4068(03)00021-1. [Umm08]Michael Ummels. The complexity of Nash equilibria in infinite multiplayer games. In Roberto M. Amadio, editor, Foundations of Software Science and Computational Structures, 11th International Conference, FOSSACS 2008, Held as Part of the Joint European Conferences on Theory and Practice of Software, ETAPS 2008, Budapest, Hungary, March 29 - April 6, 2008. Proceedings, volume 4962 of Lecture Notes in Computer Science, pages 20–34. Springer, 2008. doi:10.1007/978-3-540-78499-9\_3. [Umm10]Michael Ummels. Stochastic multiplayer games: theory and algorithms. PhD thesis, RWTH Aachen, Germany, January 2010. [UW11a]Michael Ummels and Dominik Wojtczak. The complexity of Nash equilibria in limit-average games. In Joost-Pieter Katoen and Barbara König, editors, CONCUR 2011 – Concurrency Theory, volume 6901 of Lecture Notes in Computer Science, pages 482–496. Springer, 2011. doi:10.1007/978-3-642-23217-6_32. [UW11b] Michael Ummels and Dominik Wojtczak. The complexity of Nash equilibria in stochastic multiplayer games. Logical Methods in Computer Science, Volume 7, Issue 3, September 2011. URL: https://lmcs.episciences.org/1209, doi:10.2168/LMCS-7(3:20)2011. [VCD + 15]Yaron Velner, Krishnendu Chatterjee, Laurent Doyen, Thomas A. Henzinger, Alexan- der Moshe Rabinovich, and Jean-François Raskin. The complexity of multi-mean-payoff and multi-energy games. Inf. Comput., 241:177–196, 2015. [vS34]Heinrich von Stackelberg. Marktform und Gleichgewicht. Julius Springer, Wien und Berlin, 1934. 210BIBLIOGRAPHY [Zer13]Ernst Zermelo. Über eine Anwendung der Mengenlehre auf die Theorie des Schachspiels. Proceedings of the Fifth International Congress of Mathematicians, Cambridge 1912, 2:501–504, 1913. [ZG97]Jianying Zhou and Dieter Gollmann. An efficient non-repudiation protocol. In Proceedings of The 10th Computer Security Foundations Workshop, pages 126–132. IEEE Computer Society, June 1997. [ZP06]Uri Zwick and Mike Paterson. The complexity of mean payoff games, volume 959, pages 1–10. Springer, 04 2006. doi:10.1007/BFb0030814.