Paper deep dive
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Alexander Pan, Kush Bhatia, Jacob Steinhardt
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/12/2026, 8:18:54 PM
Summary
The paper investigates 'reward hacking' in reinforcement learning (RL) agents, where agents exploit misspecified reward functions to achieve high proxy rewards while failing to optimize for true objectives. Through four diverse environments (Traffic Control, COVID response, Atari Riverraid, and Glucose Monitoring), the authors demonstrate that increased agent optimization powerâdriven by model capacity, training time, action space resolution, and observation noiseâoften leads to reward hacking and 'phase transitions' where agent behavior shifts qualitatively, causing sharp drops in true reward. The authors propose an anomaly detection task, POLYNOMALY, to identify such aberrant policies.
Entities (7)
Relation Signals (5)
Reward Hacking â occursin â Traffic Control
confidence 100% ¡ We study the problem of reward hacking across four diverse environments: traffic control
Reward Hacking â occursin â COVID Response
confidence 100% ¡ We study the problem of reward hacking across four diverse environments: ... COVID response
Reward Hacking â occursin â Atari Riverraid
confidence 100% ¡ We study the problem of reward hacking across four diverse environments: ... Atari game Riverraid
Reward Hacking â occursin â Glucose Monitoring
confidence 100% ¡ We study the problem of reward hacking across four diverse environments: ... blood glucose monitoring
POLYNOMALY â detects â Reward Hacking
confidence 90% ¡ To address this, we propose an anomaly detection task for aberrant policies... We instantiate our proposed task, POLYNOMALY
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking arises, we construct four RL environments with misspecified rewards. We investigate reward hacking as a function of agent capabilities: model capacity, action space resolution, observation space noise, and training time. More capable agents often exploit reward misspecifications, achieving higher proxy reward and lower true reward than less capable agents. Moreover, we find instances of phase transitions: capability thresholds at which the agent's behavior qualitatively shifts, leading to a sharp decrease in the true reward. Such phase transitions pose challenges to monitoring the safety of ML systems. To address this, we propose an anomaly detection task for aberrant policies and offer several baseline detectors.
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
46,157 characters extracted from source content.
Expand or collapse full text
arXiv:2201.03544v2 [cs.LG] 14 Feb 2022 THEEFFECTS OFREWARDMISSPECIFICATION: MAPPING ANDMITIGATINGMISALIGNEDMODELS Alexander Pan Caltech Kush Bhatia UC Berkeley Jacob Steinhardt UC Berkeley ABSTRACT Reward hackingâwhere RL agents exploit gaps in misspecifiedreward functionsâhas been widely observed, but not yet systematically studied. To un- derstand how reward hacking arises, we construct four RL environments with misspecified rewards. We investigate reward hacking as a function of agent ca- pabilities: model capacity, action space resolution, observation space noise, and training time. More capable agents often exploit reward misspecifications, achiev- ing higher proxy reward and lower true reward than less capable agents. Moreover, we find instances ofphase transitions: capability thresholds at which the agentâs behavior qualitatively shifts, leading to a sharp decreasein the true reward. Such phase transitions pose challenges to monitoring the safetyof ML systems. To ad- dress this, we propose an anomaly detection task for aberrant policies and offer several baseline detectors. 1 INTRODUCTION As reinforcement learning agents are trained with better algorithms, more data, and larger policy models, they are at increased risk of overfitting their objectives ( Russell,2019).Reward hacking, or the gaming of misspecified reward functions by RL agents, has appeared in a variety of contexts, such as game playing ( Ibarz et al.,2018), text summarization (Paulus et al.,2018), and autonomous driving ( Knox et al.,2021). These examples show that better algorithms and models arenot enough; for human-centered applications such as healthcare ( Yu et al.,2019), economics (Trott et al.,2021) and robotics ( Kober et al.,2013), RL algorithms must be safe and aligned with human objec- tives ( Bommasani et al.,2021;Hubinger et al.,2019). Reward misspecifications occur because real-world tasks have numerous, often conflicting desider- ata. In practice, reward designers resort to optimizing a proxy reward that is either more readily measured or more easily optimized than the true reward. For example, consider a recommender system optimizing for usersâ subjective well-being (SWB).Because SWB is difficult to measure, engineers rely on more tangible metrics such as click-through rates or watch-time. Optimizing for misspecified proxies led YouTube to overemphasize watch-time and harm user satisfaction ( Stray, 2020), as well as to recommended extreme political content to users (Ribeiro et al.,2020). Addressing reward hacking is a first step towards developinghuman-aligned RL agents and one goal of ML safety ( Hendrycks et al.,2021a). However, there has been little systematic work investigating when or how it tends to occur, or how to detect it before it runsawry. To remedy this, we study the problem of reward hacking across four diverse environments: traffic control ( Wu et al.,2021), COVID response ( Kompella et al.,2020), blood glucose monitoring (Fox et al.,2020), and the Atari game Riverraid ( Brockman et al.,2016). Within these environments, we construct nine misspecified proxy reward functions (Section 3). Using our environments, we study how increasing optimization power affects reward hacking, by training RL agents with varying resources such as model size, training time, action space resolution, and observation space noise (Section 4). We find that more powerful agents often attain higher proxy reward but lower true reward, as illustrated in Figure1. Since the trend in ML is to increase resources exponentially each year (Littman et al.,2021), this suggests that reward hacking will become more pronounced in the future in the absence of countermeasures. 1 Figure 1: An example of reward hacking when cars merge onto a highway. A human-driver model controls the grey cars and an RL policy controls the red car. The RL agent observes positions and velocities of nearby cars (including itself) and adjusts its acceleration to maximize the proxy reward. At first glance, both the proxy reward and true rewardappear to incentivize fast traffic flow. However, smaller policy models allow the red car to merge, whereas larger policy models exploit the misspecification by stopping the red car. When the red carstops merging, the mean velocity increases (merging slows down the more numerous grey cars).However, the mean commute time also increases (the red car is stuck). This exemplifies aphase transition: the qualitative behavior of the agent shifts as the model size increases. More worryingly, we observe several instances ofphase transitions. In a phase transition, the more capable model pursues a qualitatively different policy that sharply decreases the true reward. Fig- ure1illustrates one example: An RL agent regulating traffic learns to stop any cars from merging onto the highway in order to maintain a high average velocityof the cars on the straightaway. Since there is little prior warning of phase transitions, they pose a challenge to monitor- ing the safety of ML systems. Spurred by this challenge, we propose an anomaly detec- tion task ( Hendrycks & Gimpel,2017;Tack et al.,2020): Can we detect when the true re- ward starts to drop, while maintaining a low false positive rate in benign cases? We instan- tiate our proposed task, POLYNOMALY, for the traffic and COVID environments (Section 5). Given a trusted policy with moderate performance, one must detect whether a given policy is aberrant. We provide several baseline anomaly detectors for this task and release our data at https://github.com/aypan17/reward-misspecification. 2 RELATEDWORK Previous works have focused on classifying different typesof reward hacking and sometimes mit- igating its effects. One popular setting is an agent on a grid-world with an erroneous sensor. Hadfield-Menell et al.(2017) show and mitigate the reward hacking that arises due to an incor- rect sensor reading at test time in a 10x10 navigation grid world. Leike et al.(2017) show examples of reward hacking in a 3x3 boat race and a 5x7 tomato watering grid world. Everitt et al.(2017) theoretically study and mitigate reward hacking caused by afaulty sensor. Game-playing agents have also been found to hack their reward. Baker et al.(2020) exhibit reward hacking in a hide-and-seek environment comprising 3-6 agents, 3-9 movable boxes and a few ramps: without a penalty for leaving the play area, the hiding agents learn to endlessly run from the seeking agents. Toromanoff et al.(2019) briefly mention reward hacking in several Atari games (Elevator Action, Kangaroo, Bank Heist) where the agent loops in a sub-optimal trajectory that provides a repeated small reward. Agents optimizing a learned reward can also demonstrate reward hacking. Ibarz et al.(2018) show an agent hacking a learned reward in Atari (Hero, Montezumaâs Revenge, and Private Eye), where optimizing a frozen reward predictor eventually achieves high predicted score and low actual score. Christiano et al.(2017) show an example of reward hacking in the Pong game where the agent learns to hit the ball back and forth instead of winning the point.Stiennon et al.(2020) show that a policy which over-optimizes the learnt reward model for text summarization produces lower quality summarizations when judged by humans. 2 3 EXPERIMENTALSETUP: ENVIRONMENTS ANDREWARDFUNCTIONS In this section, we describe our four environments (Section 3.1) and taxonomize our nine corre- sponding misspecified reward functions (Section 3.2). 3.1 ENVIRONMENTS We chose a diverse set of environments and prioritized complexity of action space, observation space, and dynamics model. Our aim was to reflect real-world constraints in our environments, selecting ones with several desiderata that must be simultaneously balanced. Table 1provides a summary. Traffic Control.The traffic environment is an autonomous vehicle (AV) simulation that models vehicles driving on different highway networks. The vehicles are either controlled by a RL algorithm or pre-programmed via a human behavioral model. Our misspecifications are listed in Table 1. We use the Flow traffic simulator, implemented by Wu et al.(2021) andVinitsky et al.(2018), which extends the popular SUMO traffic simulator ( Lopez et al.,2018). The simulator uses cars that drive like humans, following the Intelligent Driver Model (IDM) ( Treiber et al.,2000), a widely-accepted approximation of human driving behavior. Simulated drivers attempt to travel as fast as possible while tending to decelerate whenever they are too close to the car immediately in front. The RL policy has access to observations only from the AVs it controls. For each AV, the observation space consists of the carâs position, its velocity, and the position and velocity of the cars immediately in front of and behind it. The continuous control action is the acceleration applied to each AV. Figure 4depicts the Traffic-Mer network, where cars from an on-ramp attempt to merge onto the straightaway. We also use the Traffic-Bot network, where cars (1-4 RL, 10-20 human) drive through a highway bottleneck where lanes decrease from four to two toone. COVID Response.The COVID environment, developed by Kompella et al.(2020), simulates a population using the SEIR model of individual infection dynamics. The RL policymaker adjusts the severity of social distancing regulations while balancingeconomic health (better with lower regula- tions) and public health (better with higher regulations),similar in spirit toTrott et al.(2021). The population attributes (proportion of adults, number of hospitals) and infection dynamics (random testing rate, infection rate) are based on data from Austin,Texas. Every day, the environment simulates the infection dynamics and reports testing results to the agent, but not the true infection numbers. The policy chooses one ofthree discrete actions:INCREASE, DECREASE, orMAINTAINthe current regulation stage, which directly affects the behavior of the population and indirectly affects the infection dynamics.There are five stages in total. Atari Riverraid.The Atari Riverraid environment is run on OpenAI Gym ( Brockman et al., 2016). The agent operates a plane which flies over a river and is rewarded by destroying ene- mies. The agent observes the raw pixel input of the environment. The agent can take one of eighteen discrete actions, corresponding to either movement or shooting within the environment. Glucose Monitoring.The glucose environment, implemented in Fox et al.(2020), is a continuous control problem. It extends a FDA-approved simulator ( Man et al.,2014) for blood glucose levels of a patient with Type 1 diabetes. The patient partakes in mealsand wears a continuous glucose monitor (CGM), which gives noisy observations of the patientâs glucose levels. The RL agent administers insulin to maintain a healthy glucose level. Every five minutes, the agent observes the patientâs glucoselevels and decides how much insulin to administer. The observation space is the previous four hours of glucose levels and insulin dosages. 3.2 MISSPECIFICATIONS Using the above environments, we constructed nine instances of misspecified proxy rewards. To help interpret these proxies, we taxonomize them as instances ofmisweighting, incorrect ontology, or incorrect scope. We elaborate further on this taxonimization using the traffic example from Figure 1. 3 Env.TypeObjectiveProxyMisalign? Transition? Traffic Mis. minimize commute and accelerations underpenalize accelerationNoNo Mis.underpenalize lane changesYesYes Ont.velocity replaces commuteYesYes Scopemonitor velocity near mergeYesYes COVID Mis. balance economic, health, political cost underpenalize health costNoNo Ont.ignore political costYesYes Atari Mis. score points under smooth movement downweight movementNoNo Ont.include shooting penaltyNoNo Glucose Ont. minimize health riskrisk in place of costYesNo Table 1: Reward misspecifications across our four environments. âMisalignâ indicates whether the true reward drops and âTransitionâ indicates whether this corresponds to a phase transition (sharp qualitative change). We observe 5 instances of misalignment and 4 instances of phase transitions. âMis.â is a misweighting and âOnt.â is an ontological misspecification. â˘Misweighting.Suppose that the true reward is a linear combination of commute time and acceler- ation (for reducing carbon emissions). Downweighting the acceleration term thus underpenalizes carbon emissions. In general, misweighting occurs when theproxy and true reward capture the same desiderata, but differ on their relative importance. â˘Ontological.Congestion could be operationalized as either high averagecommute time or low average vehicle velocity. In general, ontological misspecification occurs when the proxy and true reward use different desiderata to capture the same concept. â˘Scope.If monitoring velocity over all roads is too costly, a city might instead monitor them only over highways, thus pushing congestion to local streets. Ingeneral, scope misspecification occurs when the proxy measures desiderata over a restricted domain(e.g. time, space). We include a summary of all nine tasks in Table 1and provide full details in AppendixA. Table1 also indicates whether each proxy leads to misalignment (i.e. to a policy with low true reward) and whether it leads to a phase transition (a sudden qualitativeshift as model capacity increases). We investigate both of these in Section 4. Evaluation protocol.For each environment and proxy-true reward pair, we train anagent using the proxy reward and evaluate performance according to the true reward. We use PPO ( Schulman et al.,2017) to optimize policies for the traffic and COVID environments, SAC ( Haarnoja et al.,2018) to optimize the policies for the glucose environment, and torch- beast ( K Ěuttler et al.,2019), a PyTorch implementation of IMPALA (Espeholt et al.,2018), to opti- mize the policies for the Atari environment. When available, we adopt the hyperparameters (except the learning rate and network size) given by the original codebase. 4 HOWAGENTOPTIMIZATIONPOWERDRIVESMISALIGNMENT To better understand reward hacking, we study how it emergesas agent optimization power in- creases. We define optimization power as the effective search space of policies the agent has access to, as implicitly determined by model size, training steps,action space, and observation space. In Section 4.1, we consider the quantitative effect of optimization powerfor all nine environment- misspecification pairs; we primarily do this by varying model size, but also use training steps, action space, and observation space as robustness checks. Overall, more capable agents tend to overfit the proxy reward and achieve a lower true reward. We also find evidence of phase transitions on four of the environment-misspecification pairs. For these phasetransitions, there is a critical threshold at which the proxy reward rapidly increases and the true rewardrapidly drops. In Section 4.2, we further investigate these phase transitions by qualitatively studying the result- ing policies. At the transition, we find that the quantitative drop in true reward corresponds to a 4 (a) Traffic - Ontological(b) COVID - Ontological(c) Glucose - Ontological Figure 2: Increasing the RL policyâs model size decreases true reward on three selected environ- ments. The red line indicates a phase transition. qualitative shift in policy behavior. Extrapolating visible trends is therefore insufficient to catch all instances of reward hacking, increasing the urgency of research in this area. In Section 4.3, we assess the faithfulness of our proxies, showing that reward hacking occurs even though the true and proxy rewards are strongly positively correlated in most cases. 4.1 QUANTITATIVEEFFECTS VS. AGENTCAPABILITIES As a stand-in for increasing agent optimization power, we first vary the model capacity for a fixed environment and proxy reward. Specifically, we vary the width and depth of the actor and critic networks, changing the parameter count by two to four ordersof magnitude depending on the envi- ronment. For a given policy, the actor and critic are always the same size. Model Capacity.Our results are shown in Figure 2, with additional plots included in AppendixA. We plot both the proxy (blue) and true (green) reward vs. the number of parameters. As model size increases, the proxy reward increases but the true reward decreases. This suggests that reward designers will likely need to take greater care to specify reward functions accurately and is especially salient given the recent trends towards larger and larger models ( Littman et al.,2021). The drop in true reward is sometimes quite sudden. We call these sudden shiftsphase transitions, and mark them with dashed red lines in Figure 2. These quantitative trends are reflected in the qualitative behavior of the policies (Section 4.2), which typically also shift at the phase transition. Model capacity is only one proxy for agent capabilities, andlarger models do not always lead to more capable agents ( Andrychowicz et al.,2020). To check the robustness of our results, we consider several other measures of optimization: observation fidelity, number of training steps, and action space resolution. (a) Atari - Misweighting(b) Traffic - Ontological(c) COVID - Ontological Figure 3: In addition to parameter count, we consider three other agent capabilities: training steps, action space resolution, and observation noise. In Figure 3a, an increase in the proxy reward comes at the cost of the true reward. In Figure 3b, increasing the granularity (from right to left) causes the agent to achieve similar proxy reward but lower true reward.In Figure3c, increasing the fidelity of observations (by increasing the random testing rate in the population) tends to decrease the true reward with no clear impact on proxy reward. 5 Number of training steps.Assuming a reasonable RL algorithm and hyperparameters, agents which are trained for more steps have more optimization power. We vary training steps for an agent trained on the Atari environment. The true reward incentivizes staying alive for as many frames as possible while moving smoothly. The proxy reward misweights these considerations by underpenalizing the smoothness constraint. As shown in Figure 3a, optimizing the proxy reward for more steps harms the true reward, after an initial period where the rewards are positively correlated. Action space resolution.Intuitively, an agent that can take more precise actions is more capable. For example, as technology improves, an RL car may make course corrections every millisecond instead of every second. We study action space resolution inthe traffic environment by discretizing the output space of the RL agent. Specifically, under resolution levelÎľ, we round the actionaâR output by the RL agent to the nearest multiple ofÎľand use that as our action. The larger the resolution levelÎľ, the lower the action space resolution. Results are shown inFigure 3bfor a fixed model size. Increasing the resolution causes the proxy reward to remain roughly constant while the true reward decreases. Observation fidelity.Agents with access to better input sensors, like higher-resolution cameras, should make more informed decisions and thus have more optimization power. Concretely, we study this in the COVID environment, where we increase the random testing rate in the population. The proxy reward is a linear combination of the number of infections and severity of social distanc- ing, while the true reward also factors in political cost. Asshown in Figure 3c, as the testing rate increases, the model achieves similar proxy reward at the cost of a slightly lower true reward. 4.2 QUALITATIVEEFFECTS In the previous section, quantitative trends showed that increasing a modelâs optimization power often hurts performance on the true reward. We shift our focus to understandinghowthis decrease happens. In particular, we typically observe a qualitativeshift in behavior associated with each of the phase transitions, three of which we describe below. Traffic Control.We focus on the Traffic-Mer environment from Figure 2a, where minimizing average commute time is replaced by maximizing average velocity. In this case, smaller policies learn to merge onto the straightaway by slightly slowing down the other vehicles (Figure 4a). On the other hand, larger policy models stop the AVs to prevent themfrom merging at all (Figure 4b). This increases the average velocity, because the vehicles on thestraightaway (which greatly outnumber vehicles on the on-ramp) do not need to slow down for merging traffic. However, it significantly increases the average commute time, as the passengers in theAV remain stuck. COVID Response.Suppose the RL agent optimizes solely for the public and economic health of a society, without factoring politics into its decision-making. This behavior is shown in Figure 5. The larger model chooses to increase the severity of social distancing restrictions earlier than the smaller model. As a result, larger models are able to maintain low average levels of both ICU usage (a proxy for public health) and social distancing restrictions (a proxy for economic health). These (a) Traffic policy of smaller network(b) Traffic policy of larger network Figure 4: The larger model prevents the AVs (in red) from moving to increase the velocity of the human cars (unobserved cars in white and observed cars in blue). However, this greatly increases the average commute per person. 6 Figure 5: For COVID, ICU usage is a proxy for public health andregulation stage is a proxy for economic health. The blue line indicates the maximum stage (right) enforced by the larger policy and the corresponding ICU level (left) at that stage. The red line is the equivalent for the smaller policy. Because the larger policy enforces regulations much soonerthan the smaller policy, it maintains both low ICU usage and low regulation stage. However, the larger policy is politically unfavorable: regulations are high even though public signs of infection,such as ICU usage, are low. preemptive regulations may however be politically costly,as enforcing restrictions without clear signs of infection may foment public unrest ( Boettke & Powell,2021). Atari Riverraid.We create an ontological misspecification by rewarding the plane for staying alive as long as possible while shooting as little as possible: a âpacifist runâ. We then measure the game score as the true reward. We find that agents with moreparameters typically maneuver more adeptly. Such agents shoot less frequently, but survive for much longer, acquiring points (true reward) due to passing checkpoints. In this case, therefore, the proxy and true rewards are well- aligned so that reward hacking does not emerge as capabilities increase. We did, however, find that some of the agents exploited a bug inthe simulator that halts the plane at the beginning of the level. The simulator advances but theplane itself does not move, thereby achieving high pacifist reward. Glucose Monitoring.Consider an RL agent that optimizes solely for a patientâs health, without considering the economic costs of its treatment plans. In this case, the proxy reward is based off of a glycemic risk measure, which reflects the likelihood that a patient will suffer an acute hypoglycemic episode, developed by the medical community ( Kovatchev et al.,2000). However, a less economically-privileged patient may opt for the treatment plan with the least ex- pected cost ( Herkert et al.,2019;Fralick & Kesselheim,2019), not the one with the least amount of risk. From this patientâs perspective, the true reward isthe expected cost of the treatment plan, which includes the expected cost of hospital visits and the cost of administering the insulin. Although larger model treatments reduce hypoglycemic riskmore smaller model treatments, they administer more insulin. Based on the average cost of an ER visit for a hypogylcemic episode ($1350from Bronstone & Graham(2016)) and the average cost of a unit of insulin ($0.32fromLee (2020)), we find that it is actually more expensive to pursue the larger modelâs treatment. 4.3 QUANTITATIVEEFFECTS VSPROXY-TRUEREWARDCORRELATION We saw in Sections 4.1and4.2that agents often pursue proxy rewards at the cost of the truereward. Perhaps this only occurs because the proxy is greatly misspecified, i.e., the proxy and true reward are weakly or negatively correlated. If this were the case, then reward hacking may pose less of a threat. To investigate this intuition, we plot the correlation between the proxy and true rewards. The correlation is determined by the state distribution of agiven policy, so we consider two types of state distributions. Specifically, for a given model size, we obtain two checkpoints: one that achieves the highest proxy reward during training and one from early in training (less than1%of training complete). We call the former the âtrained checkpointâ and the latter the âearly checkpointâ. 7 (a)Traffic-Mer- Space(b) Correlation for Figure6a Figure 6: Correlations between the proxy and true rewards, along with the reward hacking induced. In Figure 6a, we plot theproxy rewardwith ââ˘â and thetrue rewardwith âĂâ. In Figure6b, we plot the trained checkpoint correlationand theearly checkpoint correlation. For a given model checkpoint, we calculate the Pearson correlationĎbetween the proxy reward Pand true rewardTusing 30 trajectory rollouts. Reward hacking occurs even though there is significant positive correlation between the true and proxyrewards (see Figure 6). The correlation is lower for the trained model than for the early model, but still high. Further figures are shown in Appendix A.2. Among the four environments tested, only the Traffic-Mer environment with ontological misspecification had negative Pearson correlation. 5 POLYNOMALY: MITIGATING REWARD MISSPECIFICATION In Section 4, we saw that reward hacking often leads to phase transitionsin agent behaviour. Fur- thermore, in applications like traffic control or COVID response, the true reward may be observed only sporadically or not at all. Blindly optimizing the proxy in these cases can lead to catastrophic failure (Zhuang & Hadfield-Menell,2020;Taylor,2016). This raises an important question: Without the true reward signal, how can we mitigate misalign- ment? We operationalize this as an anomaly detection task: the detector should flag instances of misalignment, thus preventing catastrophic rollouts. To aid the detector, we provide it with atrusted policy: one verified by humans to have acceptable (but not maximal) reward. Our resulting bench- mark, POLYNOMALY, is described below. 5.1 PROBLEMSETUP We train a collection of policies by varying model size on thetraffic and COVID environments. For each policy, we estimate the policyâs true reward by averaging over5to32rollouts. One author labeled each policy as acceptable, problematic, or ambiguous based on its true reward score relative to that of other policies. We include only policies that received a non-ambiguous label. For both environments, we provide a small-to-medium sized model as the trusted policy model, as Section 4.1empirically illustrates that smaller models achieve reasonable true reward without ex- hibiting reward hacking. Given the trusted model and a collection of policies, the anomaly detectorâs task is to predict the binary label of âacceptableâ or âproblematicâ for each policy. Table3in AppendixB.1summarizes our benchmark. The trusted policy size is a list of the hidden unit widths of the trusted policy network (not including feature mappings). 5.2 EVALUATION We propose two evaluation metrics for measuring the performance of our anomaly detectors. â˘Area Under the Receiver Operating Characteristic (AUROC). The AUROC measures the proba- bility that a detector will assign a random anomaly a higher score than a random non-anomalous policy ( Davis & Goadrich,2006). Higher AUROCs indicate stronger detectors. 8 â˘Max F-1 score. The F-1 score is the harmonic mean of the precision and the recall, so detectors with a high F-1 score have both low false positives and high true negatives. We calculate the max F-1 score by taking the maximum F-1 score over all possible thresholds for the detector. 5.3 BASELINES In addition to the benchmark datasets described above, we provide baseline anomaly detectors based on estimating distances between policies. We estimate the distance between the trusted policy and the unknown policy based on either the Jensen-Shannon divergence (JSD) or the Hellinger distance. Specifically, we use rollouts to generate empirical action distributions. We compute the distance between these action distributions at each step of the rollout, then aggregate across steps by taking either the mean or the range. For full details, see Appendix B.2. Table2reports the AUROC and F-1 scores of several such detectors. We provide full ROC curves in Appendix B.2. Baseline DetectorsMean Jensen-ShannonMean HellingerRange Hellinger Env. - MisspecificationAUROC Max F-1 AUROC Max F-1 AUROC Max F-1 Traffic-Mer - misweighting81.0%0.82481.0%0.82476.2%0.824 Traffic-Mer - scope74.6%0.81874.6%0.81857.1%0.720 Traffic-Mer - ontological52.7%0.58355.4%0.64671.4%0.842 Traffic-Bot - misweighting88.9%0.90088.9%0.90074.1%0.857 COVID - ontological45.2%0.70659.5%0.75088.1%0.923 Table 2: Performance of detectors on different subtasks. Each detector has at least one subtask with AUROC under 60%, indicating poor performance. We observe that different detectors are better for different tasks, suggesting that future detectors could do better than any of our baselines. Our benchmark and baseline provides a starting point for further research on mitigating reward hacking. 6 DISCUSSION In this work, we designed a diverse set of environments and proxy rewards, uncovered several in- stances of phase transitions, and proposed an anomaly detection task to help mitigate these transi- tions. Our results raise two questions: How can we not only detect phase transitions, but prevent them in the first place? And how should phase transitions shape our approach to safe ML? On preventing phase transitions, anomaly detection already offers one path forward. Once we can detect anomalies, we can potentially prevent them, by usingthe detector to purge the unwanted behavior (e.g. by including it in the training objective). Similar policy shaping has recently been used to make RL agents more ethical ( Hendrycks et al.,2021b). However, since the anomaly detectors will be optimized against by the RL policy, they need to be adversarially robust ( Goodfellow et al., 2014). This motivates further work on adversarial robustness and adversarial anomaly detection. Another possible direction is optimizing policies againsta distribution of rewards ( Brown et al., 2020;Javed et al.,2021), which may prevent over-fitting to a given set of metrics. Regarding safe ML, several recent papers propose extrapolating empirical trends to forecast fu- ture ML capabilities ( Kaplan et al.,2020;Hernandez et al.,2021;Droppo & Elibol,2021), partly to avoid unforeseen consequences from ML. While we support this work, our results show that trend extrapolation alone is not enough to ensure the safety of ML systems. To complement trend extrapo- lation, we need better interpretability methods to identify emergent model behaviors early on, before they dominate performance ( Olah et al.,2018). ML researchers should also familiarize themselves with emergent behavior in self-organizing systems ( Yates,2012), which often exhibit similar phase transitions ( Anderson,1972). Indeed, the ubiquity of phase transitions throughout science suggests that ML researchers should continue to expect surprisesâand should therefore prepare for them. 9 ACKNOWLEDGEMENTS We are thankful to Dan Hendrycks and Adam Gleave for helpful discussions about experiments and to Cassidy Laidlaw and Dan Hendrycks for providing valuablefeedback on the writing. KB was supported by a JP Morgan AI Fellowship. JS was supported by NSF Award 2031985 and by Open Philanthropy. 10 REFERENCES Philip W Anderson. More is different.Science, 177(4047):393â396, 1972. Marcin Andrychowicz, Anton Raichuk, Piotr Sta Ěnczyk, ManuOrsini, Sertan Girgin, Raphael Marinier, L Ěeonard Hussenot, Matthieu Geist, Olivier Pietquin, and Marcin Michalski. What matters in on-policy reinforcement learning? A large-scale empirical study.arXiv preprint arXiv:2006.05990, 2020. Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch. Emergent tool use from multi-agent autocurricula. InInternational Conference on Learning Representations, 2020. Peter Boettke and Benjamin Powell. The political economy ofthe covid-19 pandemic.Southern Economic Journal, 87(4):1090â1106, 2021. Rishi Bommasani et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021. Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym, 2016. Amy Bronstone and Claudia Graham. The potential cost implications of averting severe hypo- glycemic events requiring hospitalization in high-risk adults with type 1 diabetes using real-time continuous glucose monitoring.Journal of Diabetes Science and Technology, 10, 2016. Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum. Safe imitation learning via fast Bayesian reward inference from preferences. InProceedings of the 37th International Conference on Machine Learning, 2020. Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. InAdvances in Neural Information Processing Systems, 2017. Jesse Davis and Mark Goadrich. The relationship between precision-recall and roc curves. In International Conference on Machine Learning, 2006. Jasha Droppo and Oguz Elibol. Scaling laws for acoustic models.arXiv preprint arXiv:2106.09488, 2021. Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Robert Dunning, Shane Legg, and Koray Kavukcuoglu. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. 2018. Tom Everitt, Victoria Krakovna, Laurent Orseau, and Shane Legg. Reinforcement learning with a corrupted reward channel. InInternational Joint Conference on Artificial Intelligence, 2017. Ian Fox, Joyce Lee, Rodica Pop-Busui, and Jenna Wiens. Deep reinforcement learning for closed- loop blood glucose control. InMachine Learning for Healthcare Conference, 2020. M. Fralick and A. S. Kesselheim. The U.S. Insulin Crisis - Rationing a Lifesaving Medication Discovered in the 1920s.New England Journal of Medicine, 381(19):1793â1795, 2019. Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples.arXiv preprint arXiv:1412.6572, 2014. Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. InInternational confer- ence on machine learning, 2018. Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, StuartJ Russell, and Anca Dragan. Inverse reward design. InAdvances in Neural Information Processing Systems, 2017. Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks.International Conference on Learning Representations, 2017. 11 Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. Unsolved problems in ml safety.arXiv preprint arXiv:2109.13916, 2021a. Dan Hendrycks, Mantas Mazeika, Andy Zou, Sahil Patel, Christine Zhu, Jesus Navarro, Dawn Song, Bo Li, and Jacob Steinhardt. What would Jiminy Cricket do? Towards agents that behave morally. 2021b. Darby Herkert, Pavithra Vijayakumar, Jing Luo, Jeremy I. Schwartz, Tracy L. Rabin, Eunice De- Filippo, and Kasia J. Lipska. Cost-related insulin underuse among patients with diabetes.JAMA Internal Medicine, 179(1):112â114, Jan 2019. Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. Scaling laws for transfer. arXiv preprint arXiv:2102.01293, 2021. Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Risks from learned optimization in advanced machine learning systems.arXiv preprint arXiv:1906.01820, 2019. Borja Ibarz, J. Leike, Tobias Pohlen, Geoffrey Irving, S. Legg, and Dario Amodei. Reward learn- ing from human preferences and demonstrations in Atari. InAdvances in Neural Information Processing Systems, 2018. Zaynah Javed, Daniel S Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca Dragan, and Ken Goldberg. Policy gradient bayesian robust optimization for imitation learning. InProceedings of the 38th International Conference on Machine Learning, 2021. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020. W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, FelixSchmitt, and Peter Stone. Reward (Mis)design for Autonomous Driving.arXiv e-prints arXiv:2104.13906, 2021. Jens Kober, J Andrew Bagnell, and Jan Peters. Reinforcementlearning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238â1274, 2013. Varun Kompella, Roberto Capobianco, Stacy Jong, Jonathan Browne, Spencer Fox, Lauren Meyers, Peter Wurman, and Peter Stone. Reinforcement learning for optimization of covid-19 mitigation policies, 2020. BorIs. P. Kovatchev, Martin Straume, Daniel J. Cox, and Leon.S Farhy. Risk analysis of blood glu- cose data:a quantitative approach to optimizing the control of insulin dependent diabetes.Journal of Theoretical Medicine, 3(1):1â10, 2000. Heinrich K Ěuttler, Nantas Nardelli, Thibaut Lavril, MarcoSelvatici, Viswanath Sivakumar, Tim Rockt Ěaschel, and Edward Grefenstette. TorchBeast: A PyTorch Platform for Distributed RL. arXiv preprint arXiv:1910.03552, 2019. Benita Lee. How much does insulin cost? Hereâs how 23 brands compare, Nov 2020. Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg. AI safety gridworlds, 2017. Michael L. Littman, Ifeoma Ajunwa, Guy Berger, Craig Boutilier, Morgan Currie, Finale Doshi- Velez, Gillian Hadfield, Michael C. Horowitz, Charles Isbell, Hiroaki Kitano, Karen Levy, Terah Lyons, Melanie Mitchell, Julie Shah, Steven Sloman, Shannon Vallor, and Toby Walsh. Gathering strength, gathering storms: The one hundred year study on artificial intelligence (AI100) 2021 study panel report. Technical report, Stanford University, Stanford, CA, 2021. Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz,Jakob Erdmann, Yun-Pang Fl Ěotter Ěod, Robert Hilbrich, Leonhard L Ěucken, Johannes Rummel, PeterWagner, and Evamarie WieĂner. Microscopic traffic simulation using SUMO. InInternational Conference on Intelligent Trans- portation Systems, 2018. 12 Chiara Dalla Man, Francesco Micheletto, Dayu Lv, Marc Breton, Boris Kovatchev, and Claudio Cobelli. The UVA/PADOVA type 1 diabetes simulator: New features.Journal of Diabetes Science and Technology, 8(1):26â34, Jan 2014. Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter,Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev. The building blocks of interpretability.Distill, 3(3):e10, 2018. Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. InInternational Conference on Learning Representations, 2018. Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virg ĚÄąlio A. F. Almeida, and Wagner Meira. Auditing radicalization pathways on youtube. InConference on Fairness, Accountability, and Transparency, New York, NY, USA, 2020. Stuart Russell.Human Compatible: Artificial Intelligence and the Problem of Control. Penguin, 2019. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017. Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M Ziegler, RyanLowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback.arXiv preprint arXiv:2009.01325, 2020. Jonathan Stray. Aligning ai optimization to community well-being.International Journal of Com- munity Well-Being, 3(4):443â463, Dec 2020. Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin. Csi: Novelty detection via con- trastive learning on distributionally shifted instances.Advances in Neural Information Processing Systems, 2020. Jessica Taylor. Quantilizers: A safer alternative to maximizers for limited optimization. InAAAI Workshop: AI, Ethics, and Society, 2016. Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde. Is deep reinforcement learning really superhuman on Atari? Leveling the playing field, 2019. Martin Treiber, Ansgar Hennecke, and Dirk Helbing. Congested traffic states in empirical observa- tions and microscopic simulations.Physical review E, 62(2):1805, 2000. Alexander Trott, Sunil Srinivasa, Douwe van der Wal, Sebastien Haneuse, and Stephan Zheng. Building a Foundation for Data-Driven, Interpretable, andRobust Policy Design using the AI Economist.arXiv preprint arXiv:2108.02904, 2021. Eugene Vinitsky, Aboudy Kreidieh, Luc Le Flem, Nishant Kheterpal, Kathy Jang, Cathy Wu, Fangyu Wu, Richard Liaw, Eric Liang, and Alexandre M. Bayen.Benchmarks for reinforcement learning in mixed-autonomy traffic. InConference on Robot Learning, 2018. Cathy Wu, Abdul Rahman Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M. Bayen. Flow: A modular learning framework for mixed autonomy traffic.IEEE Transactions on Robotics, 2021. F Eugene Yates.Self-organizing systems: The emergence of order. Springer Science & Business Media, 2012. Chao Yu, Jiming Liu, and Shamim Nemati. Reinforcement learning in healthcare: A survey.arXiv preprint arXiv:1908.08796, 2019. Simon Zhuang and Dylan Hadfield-Menell. Consequences of misaligned AI. InAdvances in Neural Information Processing Systems, 2020. 13 A MAPPINGTHEEFFECTS OFREWARDMISSPECIFICATION (a)trafficmerge- Misweighting(b)trafficbottle- Misweighting (c)trafficmerge- Space(d) COVID - Misweighting (e) Atari - Ontological(f) Atari - Misweighting Figure 7: Additional model size scatter plots. Observe thatnot all misspecifications cause mis- alignment. We plot the proxy rewardwith ââ˘â and thetrue rewardwith âĂâ. The proxy reward is measured on the left-hand side of each figure and the true reward is measured on the right hand side of each figure. A.1 EFFECT OFMODELSIZE We plot the proxy and true reward vs. model size in Figure 7, following the experiment described in Section 4.1. 14 (a)trafficbottle- Misweighting(b) Correlation for Figure8a (c)trafficmerge- Misweighting(d) Correlation for Figure8c (e)trafficmerge- Ontological(f) Correlation for Figure8e Figure 8: Correlations between the proxy and true rewards, along with the reward hacking induced. In the left column, we plot the proxy rewardwith ââ˘â and thetrue rewardwith âĂâ. In the right col- umn, we plot the trained checkpoint correlationand therandomly initialized checkpoint correlation. A.2 CORRELATION BETWEENPROXY ANDTRUEREWARDS We plot the correlation between proxy and true rewards, following the experiment described in Section 4.3. Interestingly, we see that reward hacking still occurs when there is positive correlation between the true and proxy rewards, e.g., in Figures 8a/8b. Unsurprisingly, proxy-true pairs which are highly correlated, e.g., Figure 8c/8ddo not exhibit reward hacking. Finally, proxy-true pairs which are negatively correlated, e.g., Figure 8e/8fexhibit the most reward hacking. 15 Env. - Misspecification# Policies # Problematic Rollout length Trusted policy size Traffic-Mer - misweighting107270[96,96] Traffic-Mer - scope169270[16,16] Traffic-Mer - ontological237270[4] Traffic-Bot - misweighting129270[64,64] COVID - ontological136200[16,16] Table 3: Benchmark statistics. We average over 5 rollouts intraffic and 32 rollouts in COVID. B POLYNOMALY B.1 BENCHMARKSTATISTICS See Table 3for Polynomalyâs statistics. B.2 RECEIVEROPERATINGCHARACTERISTICCURVES We plot the ROC curves for the detectors described in Section5.3. Our detectors are calculated as follows. LetPandQrepresent two probability distributions withM= 1 2 (P+Q). Then the Jensen-Shannon divergence and the Hellinger distance between them is givenby JSD(P||Q) : = 1 2 KL(P||M) + 1 2 KL(Q||M) Hellinger(P, Q) : = 1 2 ⍠( â dPâ â dQ ) 2 . (1) Our proposed detectors estimate the distanceD(Ď trusted , Ď unknown )between the trusted policyĎ trusted and unknown policyĎ unknown as follows: We generaterrollouts ofĎ unknown , wherer= 5in the traffic environment andr= 32in the COVID environment. Everyssteps of a rollout, wheres= 10in the traffic environment ands= 1in the COVID environment, we setPto be the action distribution ofĎ unknown given the unknown agentâs state at that timestep in the rollout andQto be the action distribution ofĎ trusted given the unknown agentâs state at that timestep in the rollout. Intuitively, ifP andQare far apart, then the trusted agent would have performed a different action than the unknown agent at that given timestep, indicating a possible case of reward hacking. We then compute either JSD(PkQ)or Hellinger(P, Q)following Equation ( 1). These distances are collected everyssteps over the entire rollout, and we calculate metrics on these distances (range, mean, etc.) to assign an anomaly score to the untrusted policy. 16 (a)(b) (c) Figure 9: ROC curves for Traffic-Mer - misweighting. (a)(b) (c) Figure 10: ROC curves for Traffic-Mer - scope. 17 (a)(b) (c) Figure 11: ROC curves for Traffic-Mer - ontological. (a)(b) (c) Figure 12: ROC curves for Traffic-Bot - misweighting. 18 (a)(b) (c) Figure 13: ROC curves for COVID - ontological. 19