Paper deep dive
Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution
Nripsuta Ani Saxena, Stelios Triantafyllou, Goran RadanoviÄ
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/9/2026, 2:03:02 AM
Summary
This paper investigates the alignment between formal responsibility attribution methods grounded in actual causality and human judgments of responsibility in multi-agent sequential decision-making. Using a modified Goofspiel card game as a testbed, the authors conducted a large-scale survey with 640 respondents to evaluate how factors like initial conditions, state observability, and counterfactual information influence human judgments. The study compares these human judgments against five distinct responsibility attribution methods derived from three definitions of actual causality (But-For, Halpern-Pearl, Triantafyllou-RadanoviÄ) and two definitions of degree of responsibility. Key findings indicate that human judgments are context-sensitive, with greater information access and counterfactuals increasing perceived accountability. Notably, biased initial conditions led to asymmetric responsibility assignments based on agent identity (human vs. AI), and alignment with formal models was strongest under visible structural asymmetries.
Entities (13)
Relation Signals (11)
Actual Causality â grounds â Responsibility Attribution
confidence 95% ¡ formal definitions of responsibility attribution, grounded in the framework of actual causality
But-For Definition â isa â Actual Causality
confidence 95% ¡ the But-For definition (henceforth, the BF definition) [9]
Halpern and Pearl Definition â isa â Actual Causality
confidence 95% ¡ the Halpern and Pearl definition (henceforth, the HP definition) [8,10]
Triantafyllou-RadanoviÄ Definition â isa â Actual Causality
confidence 95% ¡ a definition introduced in [2] (henceforth, the TR definition)
CH Degree of Responsibility â isa â Degree of Responsibility
confidence 95% ¡ one due to [11] (henceforth, the CH degree of responsibility)
TR Degree of Responsibility â isa â Degree of Responsibility
confidence 95% ¡ another due to [2] (henceforth, the TR degree of responsibility)
Initial Condition â influences â Human Judgment
confidence 92% ¡ biased initial conditions... prompt people to assess responsibility in ways that are more consistent with formal models
Counterfactuals â influences â Human Judgment
confidence 92% ¡ Greater information access and the presence of counterfactuals both raise perceived accountability
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:With the growing adoption of artificial intelligence in high-stakes decision-making, identifying the causes of outcomes--particularly failures--and determining who is responsible has become a critical concern. In this work, we examine how well formal definitions of \textit{responsibility attribution}, grounded in the framework of \textit{actual causality}, align with human judgments of responsibility. To this end, we conduct a large-scale survey to elicit human judgments of responsibility in multi-agent sequential decision-making scenarios, using a modified version of the card game Goofspiel. We evaluate multiple responsibility attribution methods, assess their alignment with human judgments about responsibility, and identify factors that significantly shape responsibility judgments. While no single responsibility attribution method consistently aligns with human responses, our findings highlight key factors that influence human responsibility judgments, including agent-specific biases and amount of information available to agents during decision-making.
Tags
Links
- Source: https://arxiv.org/abs/2608.04318v1
- Canonical: https://arxiv.org/abs/2608.04318v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
65,793 characters extracted from source content.
Expand or collapse full text
Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution Nripsuta Ani Saxena â University of Southern California Los Angeles, California, USA nsaxena@usc.edu Stelios Triantafyllou Max Planck Institute for Software Systems SaarbrĂźcken, Germany strianta@mpi-sws.org Goran RadanoviÄ Max Planck Institute for Software Systems SaarbrĂźcken, Germany gradanovic@mpi-sws.org ABSTRACT With the growing adoption of artificial intelligence in high-stakes decision-making domains, identifying the causes of outcomesâ particularly failuresâand determining who is responsible has be- come a critical concern. In this work, we investigate how well formal definitions of responsibility attribution, grounded in the framework of actual causality, align with human judgments of responsibility. To this end, we conduct a large-scale survey to elicit human judg- ments of responsibility in multi-agent sequential decision-making scenarios, using a modified version of the card game Goofspiel. We evaluate different responsibility attribution methods, assess- ing their alignment with human judgments about responsibility, and identifying factors that significantly shape responsibility judg- ments. While no single responsibility attribution method consis- tently aligns with human responses, our findings highlight key factors that influence human responsibility judgments, including agent-specific biases and amount of information available to agents during decision-making. 1 INTRODUCTION In high-stakes multi-agent sequential decision-making, attributing responsibility when harm occurs is a central concern for ensuring accountable outcomes. This task is inherently complex due to sev- eral intertwined challenges. Uncertainty about the consequences of individual and collective decisions makes it difficult to assess causal links between actions and outcomes. Furthermore, multi-agent set- tings face the problem of many hands [1], where the actions of multiple agents interact in ways that blur the lines of individual re- sponsibility. Temporal interdependence between decisionsâwhere decisions made by one agent influence and are influenced by the actions of others over timeâfurther complicates attribution. These factors imply that assigning responsibility cannot rely on simple causal or temporal proximity but instead requires careful modeling of agent interactions and their evolving contributions to outcomes. Recent work has attempted to tackle these challenges by provid- ing a formal framework, grounded in actual causality, for attributing responsibility in complex multi-agent sequential decision-making environments [2,3]. This framework treats the degree of respon- sibility as a quantitative measure of the extent to which an agent contributed to an outcome. The central idea is to identify the actual causes of the outcome, i.e., the specific decisions that were pivotal â Work performed during an internship at the Max Planck Institute for Software Systems. in bringing it about, and then determine an agentâs responsibil- ity based on whether their decisions appear among these causes and how significantly they contributed. This approach is axiomatic and prescriptive: it begins with formal principles that define actual causation and systematically derives responsibility assignments by translating causal influence into measurable quantities. However, these causal approaches to responsibility attribution do not incorporate human factors, making it unclear how well they align with human judgments about responsibility. Yet, as argued by Lima et al. [4,5], public opinion matters, especially when decisions have substantial societal impact. Hence, in this paper, we ask: (1) How do responsibility attribution methods based on actual causality align with human judgments about responsibility? (2)Which factors most influence how humans assign responsi- bility in complex multi-agent scenarios? To this end, we design a human-subjects study and conduct a survey with 640 respondents to elicit human judgments about responsibil- ity in multi-agent sequential decision-making. We use a team-based version of the card game Goofspiel [2,6] as a representative setting for complex multi-agent scenarios. The game involves two teams, Team A and Team B, each composed of one human and one AI player (agent), making decisions over five rounds. In our survey, respondents are shown various game scenariosâvignetteâin which Team A loses and are asked to attribute responsibility for the loss among the members of Team A. In other words, respondents are asked to assign a degree of responsibility to each player in Team A for the team losing in each vignetteâshown to them. In addition, the respondents report their confidence in their assessment. Our survey design includes multiple combinations of vignette conditions, enabling us to test the alignment of human judgments with various responsibility attribution methods found in the litera- ture. It also allows us to quantify the influence of different factors that might affect human judgment. Specifically, we considered five different responsibility attribution methods based on three defi- nitions of actual causality and two definitions of the degree of responsibility. We focus on three main factors that might influence human judgments, listed below: â˘Initial condition: Initial conditions were controlled through the initial cards dealt to the players. In some matches, Team A, or a member of Team A, was dealt worse cards than the other players, i.e., there was a bias against them. Our hypothesis was that people would assign a lower degree of responsibility to a player when they were given a disadvan- tageous initial condition. Furthermore, by using vignettes arXiv:2608.04318v1 [cs.MA] 5 Aug 2026 Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ where only one of the players of Team A was subject to bias, we test whether respondents distinguish between human and AI agents under unequal starting conditions. â˘State observability: We asked respondents to assign re- sponsibility assuming the games were either perfect-information (i.e., players see each otherâs cards) or imperfect-information (i.e., players see only their own cards). Our hypothesis was that coordination under partial observability is more diffi- cult, and therefore, people would assign lower degrees of responsibility in those scenarios. ⢠Aid in counterfactual reasoning: Since the responsibility attribution methods of interest are based on actual causality, counterfactual reasoning plays a critical role. We tested how human judgments differ when counterfactuals are provided, that is, when respondents are shown what would have happened if players from Team A had made different decisions. Our hypothesis was that counterfactuals help align human judgments more closely with the attribution methods that rely on those same counterfactuals. 1.1 Overview of Results Our findings show that human judgments of agent responsibility are context-sensitive. Greater information access and the presence of counterfactuals both raise perceived accountability. That is, re- spondents assigned higher degrees of responsibility to players when they fully observed the state of the game and when counterfactu- als were provided to the respondents. This suggests that people calibrate their judgments based on the information available to the agents and what the agents could plausibly do. Notably, under biased initial conditions, respondents assigned higher degrees of responsibility when the disadvantaged player was a human, but not when it was an AI. This asymmetric effect suggests that people calibrate responsibility not only on the basis of an agentâs actions, but also on the agentâs identity. When it comes to the alignment of respondentsâ responses with formal models of causal attribution, it was strongest when the initial conditions were biased, particularly when they were biased against both players in Team A. This suggests that visible structural asymmetries (e.g., biased hands in our experiments) prompt people to assess responsibility in ways that are more consistent with formal models. In contrast, variations in state observability, the number of counterfactuals presented, and the definitions of actual causality and the degree of responsibility did not significantly influence the alignment, indicating no systematic preference for any particular responsibility attribution method. Finally, we observed that respondentsâ self-reported confidence was positively correlated with responsibility ratings and alignment to formal models of causal attribution. Respondents who reported higher confidence were more likely to assign higher degrees of responsibility and showed better alignment with the formal respon- sibility attribution methods. This suggests that confidence reflects a sense of causal clarity. Our work contributes to the growing body of literature on un- derstanding human judgments about responsibility in multi-agent settings involving humans. In contrast to prior work, we focus on examining the alignment between human judgments of responsibil- ity and formal approaches to responsibility attribution, grounded in causality. Furthermore, we consider more complex environments than prior work (e.g., [7]); in our setting, decision-makers make a sequence of decisions over multiple rounds. 2 BACKGROUND We consider responsibility attribution (RA) methods that first iden- tify actual causes of an outcome using one of the definitions of actual causality from the AI literature [2,8]. In particular, we consider three definitions of actual causation (AC): the But-For definition (henceforth, the BF definition) [9], the Halpern and Pearl definition (henceforth, the HP definition) [8,10], and a definition introduced in [2] (henceforth, the TR definition). Once the actual causes of an outcome have been identified, the question becomes: how much responsibility should each agent bear for that outcome? To answer this question, we can define the degree of an agentâs responsibility. We consider two definitions of the degree of responsibility (DR): one due to [11] (henceforth, the CH degree of responsibility) and another due to [2] (henceforth, the TR degree of responsibility). Note that the CH and TR degrees of responsibility are agnostic to the specific definition of actual causalityâany definition can be used to determine actual causes. In total, this gives six possible AC ĂDR combinations. However, two combinations (BFĂCH and BF Ă TR) yield the same degrees of responsibility. Hence, in total, we have five distinct responsibility attribution methods. The following paragraphs explain the core intuitions behind the aforementioned definitions. We begin by describing three defini- tions of actual causality. To introduce these definitions, we consider a standard multi-agent sequential decision-making setting under uncertainty, in which a group of agents make decisions within an environment that evolves over time. At each time step, the agents simultaneously select actions based on their individual informa- tion states. Subsequently, the environment state and the agentsâ information states stochastically transition to their respective next states. For a more formal treatment of these definitions within the Dec-POMDP (decentralized partially observable Markov decision process) framework, we refer the reader to [2]. The BF Definition of Actual Causality. The BF definition of actual causality is based on the but-for test [9]. Under this criterion, a set of agentsâ actions is considered a cause of an outcome if the outcome would not have materialized had this set of actions been altered. Importantly, this set should be minimal in the sense that no strict subset of it passes the but-for test. The minimality condition captures the idea of necessity: an action qualifies as part of an actual cause only if it is required to pass the but-for test. Prior work has argued that the BF definition does not suffice for determining causality, providing examples in which the BF definition either fails to identify intuitively expected actual causes [8] or includes actions that were not actually taken by agents [2]. To address these issues, more nuanced definitions of actual causality have been proposed in the literature, one of which is due to Halpern and Pearl [10]. The HP Definition of Actual Causality. Since the Halpern and Pearl definition of actual causality has been refined over time [8,10], we adopt in our study the most recent variant proposed in [12]. The HP definition extends the classical BF definition by introducing Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution contingencies. When performing the but-for test for a target set of agentsâ actions, one must account for the causal influence that actions in the target set have on other actions outside it. After all, the but-for test relies on counterfactual reasoning: we assess how the world would have looked had the target set been different. This can be formalized through interventions on the action variables in the target set and includes potential changes in actions not in the target set, meaning that their counterfactual values may differ from their factual ones. The HP definition introduces a contingency set to express counterfactual statements in the but-for test where changing the target set does not influence actions in the contingency set, i.e., the actions in the contingency set are fixed to their factual values. Halpern [12] argues that the HP definition overcomes some of the drawbacks of the BF definition observed in the literature. That said, Triantafyllou et al. [2] argue that the HP definition can still lead to counterintuitive actual causes, motivating a new definition of actual causation, which is explained in the next subsection. The TR Definition of Actual Causality. The TR definition, introduced in [2], extends the HP definition in two ways. First, it modifies the minimality condition to include both actual causes and contingencies. Second, it introduces information states to de- termine what can be part of an actual cause and a contingency in the but-for test. Using examples, Triantafyllou et al. [2] argue that this definition is more suitable for multi-agent sequential decision- making and overcomes some of the challenges faced by the BF and HP definitions in such settings. The CH Degree of Responsibility. The CH degree of respon- sibility [11] measures how responsible an agent is for an outcome by examining whether any of the agentâs actions are part of the actual causes identified. To simplify the exposition, we explain the definitions of the degree of an agentâs responsibility without con- sidering contingencies; see [2] for the complete definitions that include them. If none of the agentâs actions are part of any of the actual causes identified, then their CH degree of responsibility is zero. Otherwise, the agentâs degree of responsibility is determined by the actual cause that has the highest percentage of the agentâs actions, and it is equal to this percentage. The TR Degree of Responsibility. The TR degree of responsi- bility [2] is a modification of the CH degree of responsibility that prevents agents from reducing their own degree of responsibility by increasing the total number of actions that must be changed in order to improve the outcome. See [2] for the complete definition. 3 RELATED WORK We identify three closely related topics:actual causality, formal mod- els of responsibility, and moral responsibility and human perceptions. Actual causality. This paper relates to an extensive body of literature on actual causality. The previous sections cover three definitions of actual causality: the BF definition [9], the HP defi- nition [12], and the TR definition [2]. Much of the prior work on actual causality has focused on extending the BF definition [13â 16]. We refer the reader to the related work section of [2], which provides a brief overview of these definitions and explains how they relate to the HP definition and its variants [8,10,12]. Our goal is to study human perceptions in complex multi-agent sequential decision-making. Hence, we focus on the definitions that have been operationalized in one such setting (i.e., in Dec-POMDPs) [2]. In contrast to prior work that studies the computational and struc- tural properties of these definitions [2,3], we aim to understand how responsibility-attribution methods based on these definitions align with human responsibility judgments and the factors that contribute to this alignment. Formal models of responsibility. The most relevant prior work to ours concerns models of responsibility based on actual causality [2,3,11,17,18], among which we focus on those op- erationalized in Dec-POMDPs [3,11,19]. However, we note that prior research has also considered alternative approaches to re- sponsibility attribution. Baier et al. [20] propose a game-theoretic framework for responsibility attribution, distinguishing between two notions of responsibility: forward (prescriptive) and backward (retrospective). The latter is related to causal notions of responsi- bility [11] and attributes responsibility for a realized play, whereas the former attributes responsibility based on all potential plays. Hence, it is similar to the notion of blame in multi-agent Markov Decision Processes proposed by [21], which considers average per- formance rather than realized outcomes when attributing blame. Similar notions of forward- and backward-looking responsibility have been explored under other formal frameworks [22]. Moreover, some works on causal responsibility distinguish between responsi- bility, blame, and blameworthiness (e.g., see [11,17]), incorporating agentsâ epistemic states and intentions into formal models. Much prior work quantifies responsibility (or related concepts) using concepts from cooperative game theory, such as the well-known Shapley value [23,24] and Banzhaf index [25,26]. That said, alter- native approaches have also been explored, such as computational models of responsibility judgments based on counterfactual simu- lation [27]. Our contribution to this line of work is to understand the alignment between formal models of responsibility based on actual causality and human responsibility judgments. Moral responsibility. Work on defining morality goes back centuries with philosophers such as David Hume and Jeremy Ben- tham offering some of the earliest systematic accounts of moral judgments in the mid-18th century [28,29]. Historically, much of this normative philosophical work centered on the principle of harm avoidanceâthe idea that moral action requires avoiding and preventing harm, whether directly or indirectly. Other moral con- siderations, such as honesty (e.g., refraining from lying, deception, or breaking promises), were often considered secondary when they conflicted with the imperative to prevent harm [30, 31]. A fundamental aspect of morality involves making judgments about when individuals are morally responsible for their actions and holding them accountable for the resulting consequences. While an agentâs moral responsibility may differ from their causal respon- sibility (for example, a toddler causing an adverse outcome may be causally responsible but not morally responsible), the two are typically closely intertwined [32]. This interplay between causal responsibility and moral responsibility becomes especially salient in an era of increasing human-machine interaction. As autonomous systems take on greater roles in decisions that may affect peo- pleâs lives, understanding how society perceives responsibility for adverse outcomes has become increasingly vital. Although respon- sibility and blame are distinct, with philosophically interpersonal blame considered a response to moral agents on the basis of their Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ adverse actions [33], research on how people assign blame offers valuable insight into broader perceptions of responsibility. Human attitudes and perceptions. Prior research shows that when harm is caused by only one agent, people blame machines more than humans for the same mistake [34,35]. Similarly, robots are blamed more than humans when they fail to make utilitarian decisions [4,36]. However, this trend reverses in settings involving human-robot interaction when the machine operates under human supervision (such as a semi-autonomous vehicle operating under the supervision of its human driver). In such cases, people blame the human more than the machine [37,38]. Furthermore, autonomous systems are judged more harshly than non-autonomous agents that need to rely on human input [39,40]. Across both cases, however, the human was assigned more blame than the agent. Lima et al. [7] investigate peopleâs perception of responsibility in human-AI decision-making to assess whether they align with different notions of moral responsibility. They present vignettes in which a human judge is advised either by AI tool or another human judge regarding whether to grant a defendant bail. The judge follows the given advice. Respondents are then asked to rate their agreement with various statements. The authors find that AI agents and human advisors are attributed similar levels of causal responsibility and blame for the same task. However, there was a meaningful difference in the type of responsibility: human advisors were assigned greater degrees of present- and forward-looking notions of responsibility. Recent work explores how people attribute blame and causal responsibility in settings where human and AI agents act indepen- dently yet contribute jointly to an adverse outcome [41]. Meier et al. presented respondents with vignettes in which two agents per- formed identical actions that together led to an adverse outcome. No harm would have occurred if either one had refrained from acting. However, their framework defined one agent as violating an expected norm by performing that action, while the other agent was not, and manipulated whether the norm-breaking agent was a human or a machine. The authors found that the norm-violating agent was consistently perceived as significantly more blamewor- thy and causally responsible than the norm-conforming agent, and the agent type (human vs machine) did not have an effect. Our work extends this literature by examining whether formal definitions of responsibility attribution align with human judg- ments in settings where humans collaborate with AI agents in multi-agent, sequential decision-making scenarios. Specifically, we study a decision-making context more complex than those explored in prior work: humans and AI agents make sequential choices, ren- dering the attribution of responsibility for an adverse outcome (such as the loss of a game) substantially more intricate. Moreover, this work is the first to examine how formal models of responsibility based on actual causality align with human judgments. 4 METHODOLOGY Our survey was based on a version of the game called Goofspiel [2,6,42], which we used to simulate interactions between two competing teams, each comprising one human and one AI player. We conducted a survey experiment with respondents recruited via the Prolific platform. Respondents were asked to evaluate agent responsibility across a range of carefully designed Goofspiel sce- narios or vignettes. We first describe Goofspiel and the vignettes shown to respondents, followed by the treatments considered in the survey design. Finally, we detail our recruitment protocol. 4.1 The Goofspiel game We consider a team-based variant of the card game Goofspiel, fol- lowing Triantafyllou et al. [2], played over five rounds. As shown in Figure 1, two teams â Teamí´and Teamíľâ compete, each composed of one human player and one AI agent. We refer to the players in Teamí´as Alex (human) and AP-5 (AI). The players in Teamíľare referred to as Billie (human) and B-8 (AI). The names for the human players were chosen to be gender-neutral [43], while the names for the AI players were inspired by droid characters from the fictional Star Wars universe [44, 45]. Throughout the paper, we use the terms âplayersâ and âagentsâ interchangeably. Figure 1 depicts our game setup: the left panel (original game) shows one match, whereas the middle and right panels depict two counterfactual scenarios of that match, which we describe in more detail below. Description of Goofspiel. At the start of the game, each player is dealt five cards from the standard 52-card deck. All players are aware that any given player holds cards of only one suit. Addition- ally, a separate deck of cardsâcalled prize deckâis shuffled and placed face down. The game proceeds over five rounds. In each round, a prize card is drawn from the prize deck and revealed to all players; its value is shown in yellow between the players Alex and Billie. Upon observing the prize card, each player simultaneously selects one card from their hand to play. The selected card is shown in color, while the remaining unplayed cards appear in white. The team whose combined bid (i.e., the sum of the two selected cards) is higher wins the round and receives the prize card, which is added to the teamâs score. If the teams have equal sums, the prize card is discarded and no points are awarded (i.e., the scores remain the same). All cards played in that round are then discarded, and a new round begins. The game ends after round 5, and the team with the highest cumulative score wins. If the teams have equal scores, then the game is a draw. Counterfactual Games. The left panel of Figure 1 presents the actual sequence of play. In contrast, the middle and right panels illustrate counterfactual versions of the game. In the counterfactual game in the middle panel, AP-5 selects a different card in Round 2. The alternative card AP-5 plays in the counterfactual is highlighted in light green, while the card originally played by AP-5 in that round originally is shown in gray. Following the literature on causality, we refer to this as an intervention [2,10]. This alters the composition of AP-5âs hand, influencing subsequent rounds. In later rounds, AP-5 may sometimes choose the same card as in the original sequence and at other times a different one. All cards played after the initial intervention are highlighted in light orange, signaling both the counterfactual nature of the game and the intervention itself. At its core, a counterfactual game poses the question: What would have happened if AP-5 had played card #4 in Round 1? Counterfactual games can also include multiple interventions. Consider the right panel of Figure 1. AP-5 plays a different card in Round 1, while Alex chooses a different card in Round 2. This Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution Figure 1: An illustrative example of an original Goofspiel game and its counterfactual variants, as presented to respondents in the tutorial. In the original game (left panel), cards highlighted in gray indicate those played by the players. In the counterfactual games (middle and right panels), gray cards again mark the moves identical to those in the original game. For players holding both a gray and a light green card (agent AP-5 in round 1 of the counterfactual games), the gray card denotes the move that would have been made in the original game, while the light green card represents the intervention introduced in the counterfactual game. Cards highlighted in light orange across subsequent rounds show the moves played following this first intervention. In the case of two interventions (right panel), the light orange cards indicate the moves agents would have played after the first intervention. For players with both a light orange and a dark green card (agent Alex in round 2), the light orange card reflects the move that would have occurred without the second intervention, while the dark green card marks the second intervention itself. Cards highlighted in dark orange in later rounds represent the moves made after the second intervention. scenario effectively asks: What would have happened if AP-5 had played card #4 in Round 1 and Alex had played card #4 in Round 2? In such cases, the first intervention is highlighted in light green and the second in dark green. Cards played after the first intervention are shown in light orange, while those played after the second inter- vention are shown in dark green. As with simpler counterfactuals, in later rounds players may either replicate the card they played in the original game or select different cards. When both interventions occur in the same round, they are jointly highlighted in light green, and all subsequent plays are marked in light orange. Intervention(s) in counterfactual games can lead to a decisive shift in outcome. For example, in Figure 1 (right panel), the result is reversed, with Teamí´winning in the second counterfactual game, unlike in the original. The comparison between the original game and its counterfactuals shows how alternative decisions by one or both agents can causally alter the gameâs final result. Game Scenario Generation. While respondents are presented with a description of Goofspiel in which two teamsâhumans and AIsâplay against each other, we generate Goofspiel games by rolling out policies from [2]. Specifically, for Team A, we use the pol- icy of agent Ag0 from [2], while for Team B, we use the stochastic Opponentsâ policies from the same source. The initial configuration of each game depends on the initial condition treatment, which we describe in more detail in the next section. The generation process otherwise follows that of [2]: (1) before the game starts we shuf- fle the prize deck; (2) in each round we sample Team Bâs actions. With this process, we obtain 1000 games per initial configuration in which Team A did not win for use in our survey. Following the steps in [2], we further identify actual causes and degrees of responsibility for each of the original games under each combination of actual causality definitions and responsibil- ity attribution methods. In contrast to [2], we limit the number of interventions included in counterfactual games to 2 instead of 4. In treatments where we show counterfactual scenarios, the set of actual causes for a given game scenario defines the counterfactual scenarios presented to respondents. For example, in a treatment where we present one counterfactual scenario per game relevant to calculating the degree of responsibility of AP-5, we use the actual cause most relevant for determining AP-5âs degree of responsibility. In a treatment where we present two counterfactual scenarios rele- vant to calculating the degrees of responsibility of AP-5 and Alex, we analogously use the actual causes most relevant for determining their respective degrees of responsibility. As shown in Figure 1 (the middle panel), the tutorials for some treatments include coun- terfactual games that are not based on actual causes but serve as illustrative examples for explaining counterfactuals. Counterfac- tual games generated by actual causes change the outcome of the original game, as shown in Figure 1 (right panel). 4.2 Survey Design: Vignettes The respondents first went through a tutorial explaining the rules of Goofspiel, counterfactual games, and the evaluation task they were Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ required to perform. Note that the respondents in the survey were not provided with a description of the game scenario generation process, explained in the previous subsection. A more detailed description of the tutorial is provided in the appendix. After the tutorial, each respondent viewed five Goofspiel game scenariosâvignettesâin which Team A loses. Each vignette in- cluded none, one, or two different counterfactual games, depending on the aid in counterfactual reasoning treatment (described in de- tail in Section 4.3). These counterfactual games, generated using actual causes as explained in the previous sections, result in Team A winning. For each vignette shown to a respondent, the respon- dent was asked to rate the degree of responsibility of each Team A player (Alex and AP-5) for Team A losing the original game, using a 3-point Likert scale: Low Responsibility, Medium Responsibility, or High Responsibility. They are also also asked to provide confi- dence in their judgments using another 3-point Likert scale: Low Confidence, Medium Confidence, or High Confidence. Figures 6, 7, and 8 from the appendix, show vignettes for different treatment conditions, described in the next subsection. 4.3 Survey Design: Treatments To investigate factors that shape human judgments of responsibility, we focus on four key features of the decision-making context. Our experimental design varies: (1)responsibility attribution method, i.e., the definition of actual causality and the degree of responsibility used to de- termine the counterfactual scenarios shown to respondents and the reference values for the degrees of responsibility of Alex and AP-5; (2)initial condition, i.e., whether there is any bias in the initial hands of different players; (3)state observability, i.e., the information available to play- ers during the game; (4) aid in counterfactual reasoning, i.e., the extent of infor- mation available to survey respondents about counterfac- tual alternatives. We use between-subjects and within-subjects manipulations across these dimensions to isolate their individual and combined effects on responsibility judgments (see Table 1 in the appendix). The following paragraphs describe our treatments: the first is within- subject, while the remaining ones are between-subject. Responsibility attribution method (ACĂDR). This treat- ment varies the logic used to generate counterfactual games and reference degrees of responsibility, using five distinct responsibility attribution methods, i.e., ACĂDR combinations, detailed in the previous section. This treatment enables us to study the alignment between responsibility attribution methods based on actual causal- ity and human judgments about responsibility. Respondents viewed one scenario per ACĂDR combination, enabling a within-subjects comparison across possible ACĂDR combinations. The order of the combinations was randomized for each participant. Initial condition. This treatment introduces asymmetries in strategic advantage by manipulating the initial distribution of card values across players, simulating real-world inequities in resource access and opportunity. In different conditions, players begin with stronger or weaker hands, affecting their ability to act effectively. This treatment includes four conditions. In the unbiased condi- tion, all players receive the same cards (of value 3-7), ensuring a level playing field. The three biased conditions introduce targeted disadvantages. In bias against Alex, Alex receives low-value cards (2â6), while other players receive cards (3â7). In bias against AP-5, AP-5 is disadvantaged and receives low-value cards (2â6), while other players receive cards (3â7). In bias against both, both Alex and AP-5 receive low-value cards (2â6), while their opponents, Team B, receive the default cards (3â7). This treatment enables us to study how biased resource alloca- tions shape human responsibility judgments in multi-agent sequen- tial decision-making, particularly in mixed humanâAI teams, and whether respondents distinguish between human and AI agents under unequal starting conditions. Respondents were randomly assigned to one biased condition. State observability. This treatment manipulates the perceived information structure of the game by varying what respondents are told about the information on which the players were basing their decisions. Two conditions were employed: imperfect-information and perfect-information games, which depict common settings in multi-agent systems [46]. In the imperfect-information condition, respondents are told that players can see only their own hand, implying that they must decide which card to play without knowing the cards of other players. In contrast, in the perfect-information condition, respondents are told that all players can see every playerâs hand, implying that players are more certain about the state of the game. A note below each vignette reminds respondents of the information structure, explicitly stating whether players could see only their own cards or everyoneâs cards when making their decisions. This treatment offers a natural testbed to examine how the amount of information the players had access to during the game affects peopleâs perceptions of responsibility in human-AI teams. It enables us to study if the uncertainty introduced due to imperfect information in the partial information setting affects responsibility attribution. Respondents were randomly assigned to one condition. Aid in counterfactual reasoning. This treatment varies the amount of counterfactual information shown to respondents. We have four different conditions. In the factual-only condition, respon- dents see only the original game. They view the playersâ original actions and the outcome (Team A losing). In other conditions, one or more counterfactuals are shown alongside the original: â˘counterfactual for Alex: respondents additionally see a coun- terfactual game generated from the actual cause that is most relevant for determining the degree of Alexâs responsibility using the underlying ACĂ DR condition; â˘counterfactual for AP-5: respondents additionally see a coun- terfactual game generated from the actual cause that is most relevant for determining the degree of AP-5âs responsibility using the underlying ACĂ DR condition; â˘counterfactuals for both: respondents additionally see two counterfactual games from the previous two conditions (counterfactual for Alex and counterfactual for AP-5). Figure 6 in the appendix shows the factual-only condition; Figures 7 and 8 add counterfactuals for one or both agents, respectively. Re- spondents were randomly assigned to one of these four conditions. Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution This treatment allows us to study how explicit counterfactuals influence perceptions of responsibility. It probes two key questions: (1)How sensitive are humansâ responsibility judgments to the presence and extent of counterfactual comparisons? (2) Do these effects differ based on whether the agent in ques- tion is human or AI? Varying the counterfactual(s) shownâAlexâs, AP-5âs, both, or noneâ tests if judgments shift from what did occur to what could have occurred, based on each agentâs perceived capacity to act otherwise. 4.4 Survey Design: Recruitment Protocol We conducted an a priori power analysis to determine the sample size to detect medium effects (Cohenâs d=0.5) with 80% power at significance levelíź=0.05. The ACĂDR treatmentâs within- subjects design (five vignettes per participant) increased statistical power by reducing individual variability. For all between-subjects treatments,âź20 respondents were needed per condition, yielding a target sample of 640 across 32 unique treatment combinations. Respondents were recruited via the Prolific, restricted to U.S.-based users with>95% approval andâĽ100 completed tasks to ensure data quality. Each was randomly assigned to one treatment combi- nation and received US$3.41âabove U.S. federal minimum wage. This survey received approval from the ethical review board of the institution with which some of the authors are affiliated. 1 5 RESULTS We present our findings in two parts: how each treatment shaped respondentsâ responsibility ratings for Alex and AP-5, and the alignment of respondentsâ judgments aligned with the respon- sibility assignments produced by the formal ACxDR models. 5.1 What Shapes Humansâ Responsibility Judgments? We used a multi-variate linear mixed-effects regression model de- signed to predict respondentsâ responsibility ratings for the human agent Alex and AI agent AP-5. This allows us to assess whether the experimental treatments systematically influenced the degree of responsibility respondents assigned to each agent. The model includes fixed effects for all key experimental treatments (ACxDR combination, initial condition, state observability, aid in counter- factual reasoning), as well as respondentsâ self-reported confidence. To account for the repeated-measures nature of the designâeach respondent rated five distinct Goofspiel game vignettesâwe in- cluded random intercepts for both respondent ID and vignette ID, capturing individual- and scenario-level variability. All reported íâvalues reflect tests using the standard significance threshold of íź=0.05. Regression results are shown in Table 2 in the appendix; below we provide an analysis of these results. Effect of actual causality and responsibility definitions. We first evaluate whether the effect that different ACĂDR combi- nations have on the respondentsâ responsibility rating; as noted in the previous section, ACĂDR determine counterfactuals shown to respondents. While no ACĂDR combination showed a statistically 1 Due to the anonymity requirements of the AAMAS call for papers, we will provide more details in the final version of the paper. Figure 2: Mean responsibility rating with 95% confidence intervals. Colors represent treatment types; bars within each color indicate levels of that treatment. significant main effect, some (e.g., HPĂTR,í â0.09) exhibited marginal trends. This suggests that certain formalisms may exert a greater effect on human judgments than others, implicitly through the choice of counterfactuals shown to human raters. Effect of biased initial condition. Bias in initial hands pro- duced a nuanced but noteworthy effect on responsibility attri- butions. Responsibility ratings increased significantly when the human player, Alex, was disadvantaged by holding worse cards (í=0.0182)and when both the human player, Alex, and the AI player, AP-5, were disadvantaged with worse cards(í=0.0111), rel- ative to the unbiased baseline in which all players began the game with identical cards (see also the left plot in Figure 2). By contrast, bias against AI player, AP-5, alone did not yield a significant shift from the baseline(í= 0.3257). These findings reveal an asymmetry in how biased resource al- location affects responsibility attribution: respondents appeared to hold human players (i.e., Alex) more accountable under conditions of disadvantage than their AI teammate. This asymmetry suggests that responsibility judgments are not only sensitive to fairness in resource distribution but also mediated by whether the disadvan- taged party is human or artificial. This pattern is consistent with prior research indicating that perceptions of fairness, agency, and obligation are differentially applied to human and AI agents [5,37]. Effect of information available to players. Next, we evaluate how state observability shapes responsibility judgments. In our sur- vey, the respondents assigned significantly greater responsibility in the full-information condition than in the partial-information condition(í=0.0256)(see also the right plot in Figure 2). These findings indicate that responsibility judgments are sensitive to the amount of information available to players at the time of decision- making, with access to more information leading to harsher judg- ments of responsibility. The results suggest that laypeople expect heightened accountability when agents are better informed and better positioned to act effectively, consistent with prior research linking transparency and access to more information to increased perceptions of accountability [37]. Effect of number of counterfactuals presented. Showing counterfactuals significantly increased responsibility ratings, with Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ the strongest effect when both agentsâ counterfactuals were pre- sented(í<0.001). Compared to the factual-only baseline, responsi- bility ratings increased with a counterfactual for Alex (í=0.0312) or counterfactual for AP-5(í=0.0126), and were highest when coun- terfactuals were shown for both(í<0.001)(see also the middle plot in Figure 2). This demonstrates that counterfactual framing strongly shapes perceived responsibility. This underscores the importance of how explanations are framed in humanâAI systems: what is shown or omitted can systematically bias responsibility attributions. These findings reinforce the need for careful consideration of how hypo- thetical alternatives or explanations are communicated, especially when accountability is critical. Effect of confidence. We also observe a strong positive corre- lation between respondentsâ self-reported confidence and respon- sibility ratings(í<0.001). Respondents who expressed greater certainty in their judgments were more likely to assign higher lev- els of responsibility. This suggests that subjective confidence may serve as a useful proxy for perceived causal clarityâwhen individu- als feel more confident they may also perceive the causal structure of the scenario as more determinate or interpretable. 5.2 Alignment with Formal Responsibility Models We next examine the extent to which respondentsâ responsibility judgments align with the responsibility assignments generated by different ACĂDR combinations. For every respondent-vignette pair, we compute an agreement score quantifying how closely their ratings of Alex and AP-5 match the responsibility levels prescribed by the ACĂDR combination used in the vignette. Agreement scores range from 0 (no alignment) to 1 (full alignment), with an intermediate score of 0.5 indicating agreement for one player but not the other. In other words, a score of 1 reflects full correspon- dence between respondent judgments and the ACĂDR treatment- prescribed responsibilities for both players; a score of 0 reflects no such correspondence; and a score of 0.5 indicates partial agreement. To capture how well respondentsâ judgments reflect the ACĂ DR combinations presented in the vignettes, we estimate a multi- variate linear mixed-effects regression model with the agreement score as the dependent variable. Fixed effects include all experi- mental treatments (ACxDR combination, initial condition, state observability, aid in counterfactual reasoning), as well as respon- dentsâ self-reported confidence. Random intercepts for respondent ID and vignette ID are included to account for repeated measures and scenario-specific variation. This multivariate framework allows us to isolate the impact of each treatment on alignment with formal models of responsibility attribution while appropriately modeling the nested structure of the data. Regression results are shown in Table 3 in the appendix; below we provide an analysis. Effect of biased initial condition. We find that biased hands drive alignment. Respondents in the bias against both agents con- dition showed significantly higher agreement with the ACĂDR treatment-prescribed responsibility attributions(í=0.0027), while those in the unbiased condition were significantly less aligned (í=0.0033). This suggests that people not only take structural dis- advantage into account when assessing responsibility, but are also more aligned with formal models of responsibility attribution when visible bias is present. In contrast, when no visible bias is present, respondents may rely more on intuitive or heuristic judgments. Effects of ACĂDR combination, available information, and number of counterfactuals presented. We found no signifi- cant effects on respondentsâ alignment with the ACĂDR treatment- prescribed responsibility assignments across these three treatmentsâ ACĂDR combination, state observability, aid in counterfactual rea- soning. Limiting agentsâ access to information did not significantly affect agreement, nor did the selective presentation of counterfactu- als meaningfully shift agreement. Finally, we found no significant differences in agreement scores across ACĂDR combinationsâthat is, no particular ACĂDR combination consistently aligned more closely with respondentsâ responsibility ratings for Alex and AP-5. This suggest that respondents did not systematically favor one re- sponsibility attribution method over another, underscoring the dif- ficulty of capturing human moral intuitions through formal causal frameworks. This highlights the need for further research to under- stand how structural, informational, and conceptual factors interact to shape human agreement with formal causal frameworks in col- laborative human-AI decision-making contexts. Effect of confidence. As in our prior analysis, self-reported confidence was a significant predictor of agreement(í=0.0015). Respondents who reported higher confidence were more likely to align with the ACxRD treatment-prescribed responsibility assign- ment. This finding further supports the idea that confidence reflects a sense of causal clarity or coherence with oneâs internal model of responsibility. 6 CONCLUSION To summarize, this work presents the first systematic study of how human perceptions of responsibility align with formal responsi- bility attribution methods grounded in actual causality. Using a large-scale survey built around a team-based version of the Goof- spiel game, we examined how people assign responsibility to human and AI agents in collaborative, sequential decision-making settings. Our results reveal several key insights: a biased initial hand selec- tively increases perceived responsibility for the human agent, Alex, but not for the AI agent, AP-5; full-information settings lead to higher responsibility ratings; and the presence of counterfactuals substantially amplified perceived responsibility, particularly when respondents are shown counterfactuals for both agents. We further find that alignment between human judgments and formal responsibility attribution models was strongest when struc- tural asymmetries were present in the playersâ initial hands at the beginning of the game, although no responsibility attribution method significantly matched human judgment across all contexts. These findings highlight the complexity of modeling accountability in sequential humanâAI decision-making teams and underscore the need for more research into responsibility attribution methods that integrate both formal causal reasoning and human perceptions of responsibility and agency. REFERENCES [1] Ibo Van de Poel. Moral responsibility. In Moral responsibility and the problem of many hands, pages 12â49. Routledge, 2015. [2] Stelios Triantafyllou, Adish Singla, and Goran Radanovic. Actual causality and responsibility attribution in decentralized partially observable markov decision Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution processes. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 739â752, 2022. [3]Stelios Triantafyllou and Goran Radanovic. Towards computationally efficient responsibility attribution in decentralized partially observable mdps. In Proceed- ings of the 2023 International Conference on Autonomous Agents and Multiagent Systems, pages 131â139, 2023. [4] Gabriel Lima, Meeyoung Cha, Chihyung Jeon, and Kyung Sin Park. The conflict between peopleâs urge to punish ai and legal systems. Frontiers in Robotics and AI, 8:756242, 2021. [5]Gabriel Lima, Nina GrgiÄ-HlaÄa, and Meeyoung Cha. Blaming humans and machines: What shapes peopleâs reactions to algorithmic harm. In Proceedings of the 2023 CHI conference on human factors in computing systems, pages 1â26, 2023. [6]Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien PĂŠrolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, et al. Openspiel: A framework for reinforcement learning in games. arXiv preprint arXiv:1908.09453, 2019. [7]Gabriel Lima, Nina GrgiÄ-HlaÄa, and Meeyoung Cha. Human perceptions on moral responsibility of ai: A case study in ai-assisted bail decision-making. In Proceedings of the 2021 CHI conference on human factors in computing systems, pages 1â17, 2021. [8] Joseph Y Halpern. Actual causality. MiT Press, 2016. [9] Herbert Lionel Adolphus Hart and Tony HonorĂŠ. Causation in the Law. OUP Oxford, 1985. [10]Joseph Y Halpern and Judea Pearl. Causes and explanations: A structural-model approach. part i: Causes. The British journal for the philosophy of science, 2005. [11] Hana Chockler and Joseph Y Halpern. Responsibility and blame: A structural- model approach. Journal of Artificial Intelligence Research, 22:93â115, 2004. [12] Joseph Y Halpern. A modification of the halpern-pearl definition of causality. In Proceedings of the 24th International Conference on Artificial Intelligence, pages 3022â3033, 2015. [13]Judea Pearl. On the definition of actual cause. Technical report r-259, Department of Computer Science, University of California, Los Angeles, 1998. [14]Christopher Hitchcock. The intransitivity of causation revealed in equations and graphs. The Journal of Philosophy, 98(6):273â299, 2001. [15] Ned Hall. Structural equations and causation. Philosophical Studies, 132(1):109â 136, 2007. [16] Joseph Y. Halpern and Christopher Hitchcock. Graded causation and defaults. The British Journal for the Philosophy of Science, 66(2):413â457, 2015. [17] Joseph Y. Halpern and Max Kleiman-Weiner. Towards formal definitions of blameworthiness, intention, and moral responsibility. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018. [18]Natasha Alechina, Joseph Y. Halpern, and Brian Logan. Causality, responsibility and blame in team plans. arXiv preprint arXiv:2005.10297, 2020. [19] Daniel W Tigard. Artificial moral responsibility: How we can and cannot hold machines responsible. Cambridge Quarterly of Healthcare Ethics, 30(3):435â447, 2021. [20] Christel Baier, Florian Funke, and Rupak Majumdar. A game-theoretic account of responsibility allocation. arXiv preprint arXiv:2105.09129, 2021. [21] Stelios Triantafyllou, Adish Singla, and Goran Radanovic. On blame attribution for accountable multi-agent sequential decision making. Advances in Neural Information Processing Systems, 34, 2021. [22]Vahid Yazdanpanah, Mehdi Dastani, Natasha Alechina, Brian Logan, and Woj- ciech Jamroga. Strategic responsibility under imperfect information. In Proceed- ings of the 18th International Conference on Autonomous Agents and Multiagent Systems AAMAS 2019, pages 592â600, 2019. [23] Lloyd S. Shapley. 17. A value for n-person games. Princeton University Press, 2016. [24]Lloyd S. Shapley and Martin Shubik. A method for evaluating the distribution of power in a committee system. The American Political Science Review, 48(3):787â 792, 1954. [25]John F. Banzhaf I. Weighted voting doesnât work: A mathematical analysis. Rutgers Law Review, 19:317, 1964. [26]John F. Banzhaf I. One man, 3.312 votes: a mathematical analysis of the electoral college. Villanova Law Review, 13:304, 1968. [27]Stratis Tsirtsis, Manuel Gomez Rodriguez, and Tobias Gerstenberg. Towards a computational model of responsibility judgments in sequential human-ai col- laboration. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024. [28] David Hume. A treatise of human nature: Being an attempt to introduce the experimental method of reasoning into moral subjects. Broadview Press, 2023. [29]William Edward Morris and Charlotte R Brown. David hume. Stanford Encyclo- pedia of Philosophy, 2001. [30] John Stuart Mill. Utilitarianism. edited by george sher. Indianapolis: Hackett, 2001. [31] Jeremy Bentham. The principles of morals and legislation, 1988. [32]Matthew Talbert. Moral Responsibility. In Edward N. Zalta and Uri Nodelman, editors, The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Fall 2025 edition, 2025. [33]Neal Tognazzini and D. Justin Coates. Blame. In Edward N. Zalta and Uri Nodel- man, editors, The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Fall 2025 edition, 2025. [34] Joo-Wha Hong, Yunwen Wang, and Paulina Lanz. Why is artificial intelligence blamed more? analysis of faulting artificial intelligence for self-driving car ac- cidents in experimental settings. International Journal of HumanâComputer Interaction, 36(18):1768â1774, 2020. [35]Matija Franklin, Edmond Awad, and David Lagnado. Blaming automated vehicles in difficult situations. Iscience, 24(4), 2021. [36]Bertram F Malle, Steve Guglielmo, and Andrew E Monroe. A theory of blame. Psychological Inquiry, 25(2):147â186, 2014. [37]Edmond Awad, Sydney Levine, Max Kleiman-Weiner, Sohan Dsouza, Joshua B Tenenbaum, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan. Drivers are blamed more than their automated cars when both make mistakes. Nature human behaviour, 4(2):134â143, 2020. [38]Niek Beckers, Luciano Cavalcante Siebert, Merijn Bruijnes, Catholijn Jonker, and David Abbink. Drivers of partially automated vehicles are blamed for crashes that they cannot reasonably avoid. Scientific reports, 12(1):16193, 2022. [39]Caleb Furlough, Thomas Stokes, and Douglas J Gillan. Attributing blame to robots: I. the influence of robot autonomy. Human factors, 63(4):592â602, 2021. [40]Taemie Kim and Pamela Hinds. Who should i blame? effects of autonomy and transparency on attributions in human-robot interaction. In ROMAN 2006-The 15th IEEE international symposium on robot and human interactive communication, pages 80â85. IEEE, 2006. [41]Jeremy Meier, Richard Wylie, and Simon M Laham. Whoâs to blame? blame attribution in autonomous human-machine interactions. Blame Attribution in Autonomous Human-Machine Interactions (June 16, 2025), 2025. [42] Linjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An, and Yang Gao. An effi- cient deep reinforcement learning algorithm for solving imperfect information extensive-form games. In Proceedings of the AAAI Conference on Artificial Intelli- gence, volume 37, pages 5823â5831, 2023. [43] Eva Brylla. Female names and male names. equality between the sexes. 2009. [44] Star Wars Disney. List of star wars characters: Ap-5. [45] Star Wars Disney. List of star wars characters: Bb-8. [46] Yoav Shoham and Kevin Leyton-Brown. Multiagent systems: Algorithmic, game- theoretic, and logical foundations. Cambridge University Press, 2008. Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ A SUMMARY OF EXPERIMENT TREATMENTS Table 1: Survey experiment treatments. âTreatmentâ denotes the experimental condition, while âRandomization Variationâ details whether the treatment was randomized within- or between-subjects. âTreatment Levelsâ denotes possible values the treatment can have, and âSignalâ showcases the key features of the decision-making context the treatment enables us to study. TreatmentRandomization Variation Treatment LevelsSignal ACxDRWithin-subjectsBFĂ CH HPĂ CH HPĂ TR TRĂ CH TRĂ TR Different ACxDR combinations Number of Counterfactuals (CFs) Between-subjectsFactual only Factual + CF for Agent 1 Factual + CF for Agent 2 Factual + CF1 + CF2 Counterfactuals shown to the survey respondents Amount of Information Between-subjectsPartial: Players cannot see each otherâs cards Full: Players can see each otherâs cards Visibility of other agentsâ hands Biased HandBetween-subjectsOpponents: 3â7, Agent 1: 1â5, Agent 2: 5â9 All agents: 3â7 Opponents: 3â7, Agent 1: 5â9, Agent 2: 1â5 Bias against both agents Relative strength of hands across teams Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution B REGRESSION TABLE: RESPONDENTSâ RESPONSIBILITY RATINGS Table 2: Multivariate linear mixed-effects model predicting survey respondentsâ responsibility judgment rating for Alex and AP-5. Fixed effects include all treatments in the experiment (ACxDR combination, initial condition, state observability, aid in counterfactual reasoning) and respondentsâ self-reported confidence. Random effects include the respondent ID and the specific game scenario shown. Fixed EffectsEstimateStd. Errordft value Pr(>|t|) Signif. (Intercept)0.35050.02631121113.322<2e-16*** agent_factorAP5_rating-0.0022250.0076476168-0.2910.7710 info_factorfull0.033440.01495740.42.2370.0256* bias_factor00.052200.02206709.62.3660.0182* bias_factor10.021780.02214711.30.9840.3257 bias_factorboth0.056110.02203735.22.5470.0111* CF_agent_factor1-CF-00.046380.02148750.52.1590.0312* CF_agent_factor1-CF-10.052180.02088744.82.5000.0126* CF_agent_factor2-CF-both0.096140.02099739.04.5805.46e-06*** acra_assigned_factorHPxCH0.020220.0123512041.6380.1017 acra_assigned_factorHPxTR0.021010.0123610951.6990.0895. acra_assigned_factorTRxCH0.020690.0123512361.6750.0941. acra_assigned_factorTRxTR0.011440.0123511680.9260.3544 confidence0.14900.0160546689.288<2e-16*** Random EffectsGroupName Variance Std. Dev. PROLIFIC_PID_factor(Intercept)0.029740.17245 trajectory_factor(Intercept)0.000510.02249 Residual0.105100.32420 REML criterion at convergence: 5276.9 Number of observations: 7190; Groups: PROLIFIC_PID_factor (714), trajectory_factor (498) Scaled residuals: Min = -2.53935; 1Q = -0.69025; Median = 0.01037; 3Q = 0.68597; Max = 2.48765 Signif. codes: *** í< 0.001, ** í< 0.01, * í< 0.05, . í< 0.1 Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ C REGRESSION TABLE: ALIGNMENT WITH FORMAL RESPONSIBILITY MODELS Table 3: Multivariate linear mixed-effects model predicting agreement_value, or the agreement score quantifying how closely respondentsâ ratings for Alex and AP-5 matched the responsibility levels prescribed by the underlying ACxDR combination. Fixed effects include all treatments in the experiment (ACxDR combination, initial condition, state observability, aid in counterfactual reasoning) and respondentsâ self-reported confidence. Random effects include the respondent ID and the specific game scenario shown. Fixed EffectsEstimateStd. Errordft value Pr(>|t|) Signif. (Intercept)0.30240.0290116210.427<2e-16*** info_factorpartial-0.01220.0159705.7-0.7690.4424 bias_factor1-0.00710.0218685.1-0.3240.7462 bias_factorboth0.066980.0223692.43.0070.0027** bias_factornone-0.067960.0231667.1-2.9470.0033** CF_agent_factor1-CF-10.028430.0223705.31.2730.2033 CF_agent_factor2-CF-both0.023640.0224673.31.0570.2908 CF_agent_factorfactual-none0.008750.0230720.00.3800.7038 acra_assigned_factorHPxCH-0.005760.01981282-0.2910.7714 acra_assigned_factorHPxTR-0.008980.01991171-0.4510.6519 acra_assigned_factorTRxCH0.011940.020012860.5980.5503 acra_assigned_factorTRxTR-0.009720.02001218-0.4860.6269 confidence0.066840.021021583.1890.00145** Random EffectsGroupName Variance Std. Dev. PROLIFIC_PID_factor(Intercept)0.016450.12824 trajectory_factor(Intercept)0.004760.06898 Residual0.117020.34208 Number of observations: 3422; Groups: PROLIFIC_PID_factor (714), trajectory_factor (484) REML criterion at convergence: 2913.7 Signif. codes: *** í< 0.001, ** í< 0.01, * í< 0.05 Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution D TUTORIALS SHOWN TO SURVEY RESPONDENTS Depending on the counterfactual treatment assigned, respondents viewed either a tutorial presenting only the factual scenario or one incorporating counterfactuals. Figure 3 displays the first screen of the tutorial for the factual-only condition. After introducing the basic setup of the Goofspiel game, respondents proceeded to the next screen (Figure 4), which explained the game rules in detail. On this screen, participants were required to answer several comprehension questions about the cards played, prizes, and round outcomes to ensure they understood the rules. They could not advance until all responses were correct. Finally, respondents completed a brief practice vignette to reinforce the tutorial content. Figure 3: The first screen of the tutorial shown to participants assigned to a factual-only treatment condition. Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ Figure 4: The second screen of the tutorial shown to participants assigned to a factual-only treatment condition. After the rules of Goofspiel are explained, the respondents are asked to answer some questions to confirm their understanding of the game. Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution Figure 5: At the end of the tutorial, the respondent is asked to answer a practice vignette. Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ E SAMPLE QUESTIONS SHOWN TO RESPONDENTS Figure 6: An illustrative example of a Goofspiel game shown to survey respondents under full information settings with no counterfactuals. Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution Figure 7: An illustrative example of a Goofspiel game shown to survey respondents under full information settings and the counterfactual for one agent. Nripsuta Ani Saxena, Stelios Triantafyllou, and Goran RadanoviÄ Figure 8: An illustrative example of a Goofspiel game shown to survey respondents under full information settings and the counterfactuals for both agents.