Paper deep dive
Reasons to Doubt the Impact of AI Risk Evaluations
Gabriel Mukobi
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/12/2026, 6:14:21 PM
Summary
This paper critically examines the core value proposition of AI risk evaluations, arguing that current evaluation practices often fail to improve understanding of AI risks or lead to effective mitigation. The author identifies six ways evaluations fail to improve understanding, four ways they fail to improve mitigation, and five potential harms, including weaponization and opportunity costs. The paper concludes with 12 recommendations for AI labs, regulators, and researchers to shift toward a more strategic, impact-oriented approach to AI safety.
Entities (6)
Relation Signals (3)
AI Labs â investin â AI Risk Evaluations
confidence 95% · AI safety practitioners currently invest significant talent and resources in AI evaluations
AI Risk Evaluations â aimstoimprove â Understanding of AI Risks
confidence 90% · The core value proposition for AI system evaluations is that Evaluations improve our Understanding of the risks of advanced AI systems
AI Risk Evaluations â maycause â Harm
confidence 90% · Evaluations could even be harmful, for example, by triggering the weaponization of dual-use capabilities
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI safety practitioners invest considerable resources in AI system evaluations, but these investments may be wasted if evaluations fail to realize their impact. This paper questions the core value proposition of evaluations: that they significantly improve our understanding of AI risks and, consequently, our ability to mitigate those risks. Evaluations may fail to improve understanding in six ways, such as risks manifesting beyond the AI system or insignificant returns from evaluations compared to real-world observations. Improved understanding may also not lead to better risk mitigation in four ways, including challenges in upholding and enforcing commitments. Evaluations could even be harmful, for example, by triggering the weaponization of dual-use capabilities or invoking high opportunity costs for AI safety. This paper concludes with considerations for improving evaluation practices and 12 recommendations for AI labs, external evaluators, regulators, and academic researchers to encourage a more strategic and impactful approach to AI risk assessment and mitigation.
Tags
Links
- Source: https://arxiv.org/abs/2408.02565
- Canonical: https://arxiv.org/abs/2408.02565
Trouble viewing inline? Open PDF directly â
Full Text
42,269 characters extracted from source content.
Expand or collapse full text
ReasonstoDoubttheImpactofAIRiskEvaluations GabrielMukobi UCBerkeley gmukobi@berkeley.edu Abstract AIsafetypractitionersinvestconsiderableresourcesinAIsystemevaluations,buttheseinvestmentsmay bewastedifevaluationsfailtorealizetheirimpact.Thispaperquestionsthecorevaluepropositionof evaluations:thattheysignificantlyimproveourunderstandingofAIrisksand,consequently,ourabilityto mitigatethoserisks.Evaluationsmayfailtoimproveunderstandinginsixways,suchasrisksmanifesting beyondtheAIsystemorinsignificantreturnsfromevaluationscomparedtoreal-worldobservations. Improvedunderstandingmayalsonotleadtobetterriskmitigationinfourways,includingchallengesin upholdingandenforcingcommitments.Evaluationscouldevenbeharmful,forexample,bytriggeringthe weaponizationofdual-usecapabilitiesorinvokinghighopportunitycostsforAIsafety.Thispaper concludeswithconsiderationsforimprovingevaluationpracticesand12recommendationsforAIlabs, externalevaluators,regulators,andacademicresearcherstoencourageamorestrategicandimpactful approachtoAIriskassessmentandmitigation. Figure1:Thestructureofthispaper.EvaluationsmayfailtoimproveAIriskUnderstandingor Mitigationandcouldevencauseharm,butevaluationscouldstillbevaluablewithafewconsiderations. Contents 1.TheCoreValuePropositionofEvaluations...........................................................................................3 2.EvaluationsMayFailtoImproveUnderstanding.................................................................................3 2.1RisksManifestBeyondtheAISystem..............................................................................................3 2.2TheRealWorldRevealsRisks..........................................................................................................3 2.3DiminishingWarningReturnsBeyondScaryDemos.......................................................................4 2.4Measurement-DeploymentGap.........................................................................................................4 2.5GeneralCapabilitiesEntanglement...................................................................................................4 2.6ThresholdsAreNotUnderstanding...................................................................................................4 3.UnderstandingMayFailtoImproveMitigation...................................................................................4 3.1CannotTrustVoluntaryLabCommitments.......................................................................................4 3.2GovernmentsMayNotWanttoRestrictAI......................................................................................5 3.3EvaluationsMayNotBuyMuchTime..............................................................................................5 3.4EvaluationsAloneDoNotImproveSafetyCulture..........................................................................5 4.HarmfromEvaluations...........................................................................................................................5 4.1WeaponizationofDual-UseCapabilities...........................................................................................5 4.2HighOpportunityCosts.....................................................................................................................6 4.3HarmfulSafety-Washing...................................................................................................................6 4.4AccidentalLabLeaks........................................................................................................................6 4.5DelayingImpactsUntilCatastrophe..................................................................................................6 5.ConsiderationsforImprovement............................................................................................................7 5.1RecognizePossibleSourcesofEvaluationHype..............................................................................7 5.2EvaluationsareNecessaryButNotSufficient...................................................................................7 5.3PropensityEvaluationsareUnderrated.............................................................................................7 5.4EvaluatorsNeedBetterResourcesandAccess.................................................................................8 5.5TheScienceofEvaluationNeedsProgress.......................................................................................8 5.6ProposalforaLimitedEvaluationRegime.......................................................................................8 6.Recommendations....................................................................................................................................8 6.1AIDevelopmentLabs........................................................................................................................9 6.2GovernmentandThird-PartyEvaluators...........................................................................................9 6.4AIRegulators...................................................................................................................................10 6.5AcademicResearchers.....................................................................................................................10 7.Conclusion...............................................................................................................................................10 Acknowledgments......................................................................................................................................10 References...................................................................................................................................................11 1.TheCoreValuePropositionofEvaluations AIsafetypractitionerscurrentlyinvestsignificanttalentandresourcesinAIevaluations,meaning technicalmethodstotestandassessthecapabilitiesorpropensitiesofadvancedAIsystems. 1 For example,theU.S. 2 andUK 3 AISafetyInstitutesincludeevaluationsasoneoftheirtoppriorities,a significantportionofAIsafetytechnicalstaffwithinindustrylabsareallocatedtoevaluations,and ApolloResearch 4 andMETR, 5 whichareamongthelargestAIsafetynonprofits,bothfocuson evaluations. ThecorevaluepropositionforAIsystemevaluations(Figure1)isthatEvaluationsimproveour UnderstandingoftherisksofadvancedAIsystems,andinturn,thatimprovedUnderstandingenables ustobetterMitigatethoserisks. 6 However,Ifearthatthisvaluepropositionmaybedrivenbyatechno-solutionisthopethatfailsto adequatelyaccountforasocialmodelofimpact. 7 Thispaperanalyzesthispossiblefailurethroughthe linksbetweenthesethreesteps:whetherEvaluationsimproveUnderstanding(Section2)andwhether UnderstandingimprovesMitigation(Section3).Additionally,Idiscusspossibleharms(Section4), considerationsforimprovingevaluations(Section5),andrecommendationsforevaluators(Section6). 2.EvaluationsMayFailtoImproveUnderstanding IcritiquethefirstlinkbydescribingsixwaysinwhichAIsystemEvaluationsmayfailtosignificantly improveourUnderstandingofasystemâsrisks: 2.1RisksManifestBeyondtheAISystem AIevaluationsimplicitlyfocusonriskslocatedwithinanAIsystem,suchasthesystemknowingandnot refusingrequestsforinstructionsonhowtobuildabioweapon.However,manyAIrisksmanifestthrough anAIsystemâsinteractionswithcomplexsystemsintherealworld,makingitespeciallyhardtoassess sociotechnical, 8 systemic, 9 orunknown 10 riskswithevaluations. 2.2TheRealWorldRevealsRisks MuchofourunderstandingofAIriskscomesfromlearningaboutAIâsimpactsintherealworld. Evaluationsmaynotsignificantlybeatthesimplebaselineofincidentreportingfromreal-worldmodel deployments. 11 WecanalsolearnalotaboutrisksbycollectinginformationfromAIlabs,suchasthrough 11 PreventingRepeatedRealWorldAIFailuresbyCatalogingIncidents:TheAIIncidentDatabase|Proceedingsof theAAAIConferenceonArtificialIntelligence 10 EmergentAbilitiesinLargeLanguageModels:AnExplainer|CenterforSecurityandEmergingTechnology 9 [2401.07836]TwoTypesofAIExistentialRisk:DecisiveandAccumulative 8 [2310.11986]SociotechnicalSafetyEvaluationofGenerativeAISystems 7 Safetyisnâtsafetywithoutasocialmodel(or:dispellingthemythofpersetechnicalsafety)âAIAlignment Forum 6 TheoriesofChangeforAIAuditingâApolloResearch , ClarifyingMETR'sAuditingRole , [2305.15324]Model evaluationforextremerisks , AISafetyInstituteapproachtoevaluations-GOV.UK 5 METR 4 AnnouncingApolloResearchâApolloResearch 3 AISafetyInstituteapproachtoevaluations-GOV.UK 2 TheUnitedStatesArtificialIntelligenceSafetyInstitute:Vision,Mission,andStrategicGoals|NIST 1 ByâEvaluation,âImeantheevaluationofAIsoftwaresystems,notauditsoflabpracticesorotherartifacts,which aresometimesalsocalledâevaluations.âThispaperdoesdiscussrelatedtopicssuchasredteamingorscary demonstrations. reportingrequirements, 12 orembeddedgovernment-vettedauditors,andcollectingsubtle,early-warning signsthatexternaloverseerscanpiecetogetheracrosstheindustry. 2.3DiminishingWarningReturnsBeyondScaryDemos Toinformdecision-makersaboutforthcomingAIrisks,rigorousandcomplicatedevaluationsdonotadd muchmorethananotherbaselineofbuildingâscaryâdemonstrationsofspecificthreatmodels,butthey costmuchmore.ThesesmalldemonstrationsseemedusefulforconvincingpolicymakersattheUKAI SafetySummitoftheimportanceofAIsafety, 13 thoughdemonstratorsshouldtakecarenotto misrepresentrisks. 2.4Measurement-DeploymentGap Fortheforeseeablefuture,wewilllikelyhavealargegapbetweenwhatevaluationscanmeasureandan AIsystemâstruerisksoncedeployedduetoelicitationchallenges. 14 RapidchangestoAIsystems,suchas iffuturesystemsundergocontinuallearningorareconnectedtotheevolvingInternet,widenthisgap. 2.5GeneralCapabilitiesEntanglement Mostdangerouscapabilitiesarestronglycorrelatedwithgeneralcapabilities,someasuringgeneral capabilitiesâastheMLresearchcommunityisalreadyincentivizedtodoâtellsmostofthestory. Further,SuperintelligentAIrisksmainlystemfromrawgeneralintelligence,notnichedomain-specific capabilities,soweshouldbemoreconcernedwithunderstandingintelligenceastimegoeson. 2.6ThresholdsAreNotUnderstanding Thecurrentevaluationparadigmisheavilysituatedwithinriskmanagementframeworkssuchas ResponsibleScalingPolicies(RSPs) 15 thatonlyseektodetectwhenAIcapabilitiespassarbitrary capabilitythresholds. 16 ThisischieflydifferentfromamechanisticunderstandingofanAIsystemâsrisks. 3.UnderstandingMayFailtoImproveMitigation Next,Idiscussfourcritiquesofthesecondlinkinthevaluepropositionbetweenincreased UnderstandingofanAIsystemâsrisksascreatedbyevaluationsandbetterMitigationofthoserisks: 3.1CannotTrustVoluntaryLabCommitments ThoughAIlabsaremakingvoluntarycommitmentsnow, 17 thereisnogoodreasontoexpectthemto upholdthosecommitmentswhentheysignificantlyconflictwithcorporateinterests. 18 EspeciallyasAI getsmorepowerfulandracedynamics 19 strengthen,theleadinglabwillfaceincreasedpressuretorenege onitspromisesifitcansignificantlybenefitfromdeployingitsnextAIsystem. 19 [2306.12001]AnOverviewofCatastrophicAIRisks 18 RSPsarepausesdoneright(Comments)âAIAlignmentForum 17 AIcompaniesmakefreshsafetypromiseatSeoulsummit,nationsagreetoalignworkonrisks|APNews 16 [2406.14713]RiskthresholdsforfrontierAI 15 Anthropic'sResponsibleScalingPolicy , FrontierSafetyFramework-GoogleDeepMind , PreparednessFramework|OpenAI , ResponsibleScalingPolicies(RSPs)-METR,RSPsarepausesdonerightâ AIAlignmentForum 14 [2312.07413]AIcapabilitiescanbesignificantlyimprovedwithoutexpensiveretraining,Guidelinesforcapability elicitation|METRâsAutonomyEvaluationResources 13 TheUKAISafetySummit-ourrecommendationsâApolloResearch 12 [2404.02675]ResponsibleReportingforFrontierAIDevelopment 3.2GovernmentsMayNotWanttoRestrictAI GovernmentsmaynotbewillingtorestrictAIsystemsthatevaluationsindicatearedangerousdueto financialincentives,suchasifadvancedAIsystemshavebeencreatingsignificanteconomicvalue. Additionally,somegovernmentsmayhaveageneralaversiontoslowinginnovation,particularlyif politicalpowersinchargehaveestablishedapro-innovationpolicystance. Thislackofpoliticalwilltoactonevaluationsmayespeciallyrevealitselfifpre-deploymentevaluations indicateanuncertainpossibilityofAIrisk,buttherealworldhasyettorealizethatAIrisk(2.2). 3.3EvaluationsMayNotBuyMuchTime IfevaluationsindicateanAIsystemisdangerous,anddecision-makersdecidetorestrictitsdeployment,it isonlyamatteroftimebeforeAIlabscanpatchtheparticulardiscoveredissuesandundothatdecision. Further,itisnotclearifgovernmentscanadequatelymonitorâletaloneregulateâinternaldeployments, suchasanAIlabusingafrontiersystemtoautomateitsAIR&D. ItisevenmorechallengingtorestrictAIdevelopment,whichinvolvesmanydiffusefactorslikeplanning newdatacentersortestingresearchideas,sorestrictingonedangerousmodellikelydoesnotsignificantly affectthetimingofsubsequentgenerationsofmoredangerousmodels. 20 3.4EvaluationsAloneDoNotImproveSafetyCulture SomehopethatrequiringevaluationsinAIlabswillimprovethesafetyculture 21 ofthoselabs. 22 However, thishopeseemsunfounded,asorganizational-wideshiftsarerequiredtochangesafetyculture, 23 evaluationteamstendtobesiloedwithinlabs,andevaluationrequirementscouldevenbackfireby creatingresentmentforsafetypractices. 24 4.HarmfromEvaluations BeyondthesereasonsthatEvaluationsmaybeineffectiveatimprovingUnderstandingandMitigation, EvaluationscouldbeharmfulandincreaseAIrisksinatleastfivecases: 4.1WeaponizationofDual-UseCapabilities Dual-usecapabilitiessuchascyber-offense,persuasion,andautomatedAIR&D 25 arenotpurely riskyâtheyarealsohighlydesirableforcertainactorslikenationalsecurityorganizationsandAIlabs. Evaluationsmayactasprogressmeasuresandtriggersforthoseactorstoco-optdangerousAIsystemsfor theirownmeans. Evenifcontrollingactors,suchastheAIlabthatdevelopsasystemandthegovernmentofthenationthat labislocatedin,areresponsibleenoughtonotweaponizeanAIsystem,evaluationsofthesedual-use 25 Exclusive:OpenAIworkingonnewreasoningtechnologyundercodenameâStrawberryâ|Reuters,Examplesof AIImprovingAI(CAIS) 24 Whensafetyculturebackfires:Unintendedconsequencesofhalf-sharedgovernanceinahightechworkplace:The SocialScienceJournal:Vol46,No4 23 StrategyforCultureChange(COS) 22 ThismayhavehappenedwithAnthropic,butthatmaybeduetoAnthropicâsstrongpreexistingfocusonsafety. 21 BuildingaCultureofSafetyforAI:PerspectivesandChallengesbyDavidManheim::SSRN , ComplexSystems forAISafety[PragmaticAISafety#3]âAIAlignmentForum 20 Theexceptionmightbesecond-orderresourceeffectswheretherevenueandhypefromonesystemdeployment helpstofundthenextAIsystemâsdevelopment. capabilitiesmaystillalertnon-responsibleactorstothevalueofstealingthatAIsystemâsweightsand otherartifactsordevelopingtheirownsimilarlypowerfulAIsystem. 4.2HighOpportunityCosts SignificantAIsafetytalentandresourcesareinvestedinAIevaluations,asdescribedinSection1. Furthermore,thecurrentevaluationlandscapeishighlyredundant,withmanyorganizationsbuildingtests forthesamekindsofrisksontheirowninfrastructure.Focusingtoomuchonevaluationsmaybetoo costlyofadistractionfortheAIsafetyfield,especiallyifevaluationsdonotbuyusalotmoretime (3.3). 26 AllthesepeopleandresourcescouldinsteadbeappliedtoadvancingMLsafetyandAI governance,makingprogressonproblemsthatactuallymitigateAIrisks. 4.3HarmfulSafety-Washing Evaluationsmightcontributetosafety-washing, 27 wherenon-expertdecision-makersaremisledinto believingthatanAIsystemissafe.Thiscouldcreateafalsesenseofsecurity, 28 leadingtoharmfulmodels beingdeployed. 29 Relatedly,evaluationscouldalsoderiskAIinvestments, 30 leadingtoincreasedinvestmentsandAI capabilitiesacceleration. 31 4.4AccidentalLabLeaks Gain-of-function-likecapabilitieselicitationorintentionallyinducingmisalignment 32 forscary demonstrationscouldcreateunnecessarydangers.AsAIsystemsbecomemorecapable,thisworkcould increasethetheharmifdangerousmodelsareexfiltrateintotheworld,eitherbyescapingcontrolontheir ownorwiththeaidofhumaninsiders. 33 Thisrisk,analogoustolableaksfrombiosecuritylabsintendingtostudymoredangerouspathogens,is madeespeciallysignificantifAIlabsandexternalevaluatorscontinuetohaveunderdevelopedsecurity practices. 34 4.5DelayingImpactsUntilCatastrophe Relatedly,ifevaluationsdonotcatchallAIrisksbutratherfindmorelessperniciousorsevererisks,then theycounterintuitivelymightincreasetheharmofthefirstAIcatastrophe.Thatis,insteadofthefirst pointofsignificantAIharmbeingaminorincidentthatwouldhavegarneredsocietalresponse, incompleteevaluationsmaycatchandpreventthoseminorincidentsbutpushoutthefirstrealizedharm tothepointofalargercatastrophethatsocietymaybelesspreparedfor. 35 35 [2405.19832]AISafety:AClimbToArmageddon? 34 SecuringAIModelWeights:PreventingTheftandMisuseofFrontierModels|RAND 33 ImprovingthesafetyofAIevalsâLessWrong 32 ModelOrganismsofMisalignment:TheCaseforaNewPillarofAlignmentResearchâAIAlignmentForum 31 Thismayormaynotbeharmful,partiallydependingonyourviewofcapabilityoverhangs. 30 TheoriesofChangeforAIAuditing 29 RSPsarepausesdoneright(Comments)âAIAlignmentForum 28 WhenSafetyProvesDangerous 27 SafetywashingâAIAlignmentForum 26 Anexceptionisiftimelaterismuchmorevaluablethantimenowsuchthat,forexample,itisworthittospend3 yearsofAIsafetycommunityeffortin2024tobuy1yearoftimein2028.However,Iamskepticalthatthetime valuesaresoasymmetricandthetimewegainlaterissolongthatthesekindsofdealsmaybeworthit. Thisriskdepends,however,onboththeassumptionthatAIriskevaluationswillfailtodetectsomeofthe mostsevereAIrisksandtheassumptionthatsocietaldefensebenefitsfromincrementalexposureand adaptationtoAIsocietalimpacts. 5.ConsiderationsforImprovement Thatsaid,Idonotthinkallevaluationsarebadorthatsomeevaluations-likeworkhasnoplaceintheAI safetyportfolio.Idiscusssixadditionalconsiderationsformakingthemostofevaluations: 5.1RecognizePossibleSourcesofEvaluationHype Itisfirstimportanttorecognizehowevaluationsbecamesopopular:Evaluationsareeasytomake continuousprogresson,provideostensiblyprecisenumberstonon-technicaldecisionmakers, 36 parallel riskassessmenttechniquesfromotherdisciplines, 37 andareusefulforMLcapabilitiesdevelopment. 38 However,theseareallseparatefromanyreasonsthatevaluationswouldbeusefulforAIriskmitigation. IespeciallyworryaboutevaluationhypecomingattheexpenseofAIsafetyprogress.Forexample,by focusingonevaluations,theUKAISafetyInstitutemadeitselfmorecredibleandappealingtotherestof theUKgovernment 39 andmayhaveprotecteditssurvival. 40 However,theappearanceofvaluedoesnot necessarilymeanthisworksignificantlyreducedAIriskbetterthanalternativefocuses. 5.2EvaluationsareNecessaryButNotSufficient Evaluationscanrevealthepresenceofrisksbutnottheirabsence. 41 Someevaluationscouldbeusefulfor evaluatingbroadcapabilitylevelstotriggertieredsecurityrequirements,butwemightnotexpectorgain frommuchmorethanthecurrenteffortlabsareputtingintoevaluations. Instead,wemayneedtofliptheburdenofprooffromevaluatingwhetherassumed-safeAIsystemsare dangerousintomakingacaseforwhetherassumed-dangerousAIsystemsaresafe.Thesesafetycases 42 thendemandconsiderableeffortintootherAIsafetyinterventionsthatactuallyreducerisk,suchas propensityevaluations,mitigationevaluations, 43 novelmitigationmethods,andassurancevalidation techniques. 5.3PropensityEvaluationsareUnderrated Currently,mostevaluationsfocusoncapabilities,butwemaywanttoinvestmoreinevaluatingthe propensities,dispositions,oralignmentofAIsystems. 44 Propensityevaluationsaddacrucialcomponent oflikelihoodtoriskassessments,incentivizesomeprogressinalignment,andarehardertoweaponize. 44 [2305.15324]Modelevaluationforextremerisks 43 MitigationevaluationstesthowwellamitigationtechniquereducessomekindsofAIrisk.TheWMDP Benchmarkisanearlyexampleofevaluatingunlearningmitigations. 42 [2403.10462]SafetyCases:HowtoJustifytheSafetyofAdvancedAISystems,AffirmativeSafety:AnApproach toRiskManagementforAdvancedAIbyAkashWasil,JoshuaClymer,DavidKrueger,EmilyDardaman,Simeon Campos,EvanMurphy::SSRN 41 [2309.01933]Provablysafesystems:theonlypathtocontrollableAGI 40 WhatWeKnowAbouttheNewU.K.GovernmentâsApproachtoAI|TIME 39 RishiSunakonX:"AIisthedefiningtechnologyofourtimeandwehaveaclearstrategytodevelopitinasafe waythatwillbenefiteveryoneintheUK.Hereâswhatthatlookslike ï 38 Let'stalkaboutLLMevaluation 37 Riskassessment-Wikipedia 36 RichardNgoonX:"IâmworriedthatalotofworkonAIsafetyevalsisprimarilymotivatedbyâSomethingmust bedone.Thisissomething.Thereforethismustbedone.âOr,toputitanotherway:Ijudgeevalideason4criteria, andIoftenseeproposalswhichfailall4.Thecriteria:" PropensityevaluationsmoredirectlygetatourunderstandingofanAIsystemanditsrisks, 45 andasa result,theymayneedmoreinterpretabilityprogresstostartworking. 46 5.4EvaluatorsNeedBetterResourcesandAccess Industryracedynamicscreatelimitations,suchasOpenAIresearchersonlyhavingoneweektoevaluate GPT-4o 47 orAIlabswithholdingpre-deploymentaccesstofinalmodelsfromexternalauditors, 48 making rigorousevaluationschallenging.Thiscanbeaddressed,butonlywithstrongenoughforcestoovercome adversarialcorporateincentives. Idealaccessmightincludepre-deployment,transparent-box 49 accesstoAIsystemswithboththefinal versionsofmodels 50 andhelpful-only/non-refusalmodelsthatareeasiertoelicitcapabilitiesfrom,along withlab-developedelicitationandagenttools. 5.5TheScienceofEvaluationNeedsProgress Insofarasevaluationsarebeneficial,wecanmakethemmoreeffectivebydevelopingrigorousand reproducibleevaluationpractices.METRhasmadethisitstoppriority. 51 However,thisisoneofthemoreobviousareastodivertmoreAIsafetyresearchfundingto,soitmaynot beespeciallyneglectedsoon. 5.6ProposalforaLimitedEvaluationRegime OnesimplifiedâbutpossiblystillvaluableâevaluationregimenottoofarfromourcurrentRSP-heavy statecouldinvolvealightsetofevaluationsassessingthehigh-levelgeneralcapabilitytier 52,53 ofanAI systemtotriggerheightenedsecurityrequirementssuchassafetycases(5.2)andSecurityLevels. 54 These evaluationscouldmostlyfocusongeneralcapabilitiesratherthanmanycorrelated,narrowrisks(2.5). AIlabswoulddomostoftheevaluationwork,withexternalpartiesonlyreallyinvestinginoversight mechanismssuchassendinginredteamerstosubjectivelycheckthecapabilitytierorreviewingthelabâs evaluationpracticesandinfrastructure.AIlabscouldbeincentivizedtododecentcapabilitytier evaluationsiftheincreasedsecurityrequirementsmitigateactualbusinessrisks, 55 ifoversightandlegal penaltiesarestrongenoughthatlabswanttheirinternalassessmentstobeaccurate,orifadditional securityrequirementsalsocomewithhelpfromgovernmentsecurityorganizationsforachievingthose requirements. 6.Recommendations Iendwith12recommendationsforpossibleactionsdifferentAIlabs,externalevaluators,regulators,and academicscouldtaketoreducethefailingsofandimproveAIevaluations. 55 Forexample,considertheFinancialimpactoftheBoeing737MAXgroundings. 54 SecuringAIModelWeights:PreventingTheftandMisuseofFrontierModels|RAND 53 [2406.14713]RiskthresholdsforfrontierAI 52 Forexample,AISafetyLevels(ASL)inAnthropicâsRSP. 51 ClarifyingMETR'sAuditingRoleâAIAlignmentForum 50 GPT-4SystemCard(OpenAI) 49 [2401.14446]Black-BoxAccessisInsufficientforRigorousAIAudits 48 AIcompaniesaren'treallyusingexternalevaluators-AILabWatch 47 OpenAIemployeessayitâfailedâitsfirsttesttomakeitsAIsafe-TheWashingtonPost 46 AtransparencyandinterpretabilitytechtreeâAIAlignmentForum 45 Towardsunderstanding-basedsafetyevaluationsâAIAlignmentForum 6.1AIDevelopmentLabs 1.CrediblyCommit:MuchofthevalueofevaluationsdependsonwhetherwhicheverAIlabisin theleadwillupholditscommitments(3.1).Labscouldestablishbetterinternalgovernance mechanismsthatcouldmorecrediblyincentivizethemtohonortheircommitments. 56 2.ProvideResourcesandAccess:Theycanalsoprovideexternalevaluatorswithappropriate resourcesandaccesstoAIsystems(5.4).LabscoulddirectlyprovideevaluationAPIstolet governmentandthird-partyevaluatorssecurelyusesomeofthelabsâinternaltools,suchasagent scaffolding,capabilityelicitation,andgradingtools. 57 Labscanalsousetheirconsiderable financialandcomputationalresourcestosupporttheevaluationecosystem. 58 3.ShareEvaluationInfrastructure:Finally,labscanreduceredundancyandincrease transparencybysharingmuchmoreoftheirevaluationinfrastructure. 59 Forexample,theycould open-sourcemostoftheirinfrastructure 60 orprivatelyshareitwithtrustedgovernmentand third-partyevaluators.ThisalsoincreasesthetransparencyofAIlabevaluationmethods, allowingforgreaterscrutinyandtrust.Sharinghaslowdownsides,asinternalevaluationsare alreadysusceptibletoexploitationbymodeldeveloperswithinAIlabs,unlesslabshavehigh levelsofsiloingbetweenevaluationanddevelopmentteams. 6.2GovernmentandThird-PartyEvaluators 4.Specialize:Toreduceredundancy,AISafetyInstitutesandsimilarorganizationsmightspecialize inevaluationsthatneedspecialgovernmentresources,suchasnationalsecurityrisksrequiring securityclearances. 61 Third-partyevaluatorsmightspecializeinnovelrisksthatlabsareless incentivizedtoworkon. 62 5.CooperateonStandardsandSharing:Externalevaluatorscanalsofocusonplanning internationalcooperationswitheachothertoformgloballyconsistentAIsafetystandardsthat reducethecostsandincreasethelikelihoodofcomplianceandsharetheirworktoreduce redundancy. 63 6.CreateExternalOversight:ThesegroupscouldoverseeAIlabevaluationsandscrutinizelab commitments 64 toincentivizegreateraccountability.Government-vettedexpertscouldbeinternal auditorsembeddedwithinAIlabstoleverageheightenedtransparency. 7.BuildScaryDemos:Tocommunicateriskstodecision-makers,scarydemosmaybetteruse resourcesthanexpensiverigorousevaluations(2.3). 8.AdvancetheScienceofAISafetyBeyondEvaluations:Last,externalevaluatorscouldadvance thescienceofAIsafety. 65 Bythis,Imeannotjustthescienceofevaluationsbutalsomaking 65 StrategicVision|NIST 64 Commitments-AILabWatch 63 TheglobalnetworkofAISafetyInstitutes,theU.S.-UKAISIpartnership,andtheUKAISIopen-sourcingits Inspectevaluationsinfrastructureareallpositivesignsofthis.However,muchworkremainstofigureoutthedetails ofefficientresourcesharingpartnerships. 62 METRstartedtodothiswithautonomousadaptationandreplication,thoughthatseemslessofapriorityfor METRandmorerepresentedwithinthelabsnow.Apollohasbeenspecializingindeceptionevaluationsthatmay useinterpretabilitytools. 61 RecommendationsforthenextstagesoftheFrontierAITaskforceâApolloResearch 60 OpenAIandDeepMindhavemadeprogressinopen-sourcingalimitedamountofinfrastructure,thoughIexpect theycouldshareevenmore. 59 [2404.14068]HolisticSafetyandResponsibilityEvaluationsofAdvancedAIModels 58 Anewinitiativefordevelopingthird-partymodelevaluations 57 [2401.14446]Black-BoxAccessisInsufficientforRigorousAIAudits 56 TheWindfallClauseisanearlybuthighlynon-bindingattemptatsimilargoals. fundamentalscientificprogressonothertechnical,governance,andsociotechnicalAIsafety problems. 66 6.4AIRegulators 9.RequireLabCooperation:Governmentregulatorscanhelpenforcealltherecommendationsfor AIlabsinSection6.1,increasingthelikelihoodthatAIlabsupholdtheircommitments,provide appropriateresourcesandaccess,andshareevaluationinfrastructure. 10.ClarifyProtectionsforLabCooperation:Complementarily,regulatorscouldworkwithAIlabs toclarifylegalgreyzonesandcarveoutprotectionsforspecificinstancesoflabcooperation.This mightincludeanti-trustlaws,protectionsforsharinginformation,orprotectionsforgivingaccess toexternalevaluators. 67 6.5AcademicResearchers 11.AdvancetheScienceofEvaluation:Researcherscantargetpropensityevaluations(5.3), predictingpropertiesoffutureAIsystemsbeforetheyaredeveloped, 68 automatedevaluations, howtoupdatedynamicevaluations, 69 andotherscienceofevaluationquestions(5.5). 70 12.DevelopBetterThreatModels:Similarly,manyevaluationdevelopmentchallengesare bottleneckedbysolidthreatmodeling. 71 Evenonepersoncouldmakesignificantprogresshere, improvingtheefficiencyandeffectivenessofevaluationsbuiltonbetterthreatmodels. 7.Conclusion TherearemanyreasonstosuspectAIEvaluationsmayfailtoimproveourUnderstandingofAIrisks, thattheUnderstandingwegetfromevaluationsmayfailtoimproveourMitigationofthoserisks,and thatevaluationscouldevenbeharmful.However,evaluationsdonotseementirelydoomedâwejustneed tocarefullyconsiderhowtheyshouldfitintoahealthyanddiverseAIsafetyportfolio. AIriskevaluationisanewandrapidlyevolvingfield,sosomeofthesepointsmaybelessaccuratein differentcontextsorovertime.Ultimately,IaimforthispapertoinspireAIsafetypractitionerstothink aboutthevaluepropositionofevalsmoredeliberatelyanddecidewhetherandhowAIsafetytalentand resourcesmightbebetterspentonothermeanstoreduceAIrisk. Acknowledgments ManythankstoDanHendrycks,DavidKrueger,andPatriciaPaskovforhelpfulcommentsand discussionsthatinformedthispaper.Mistakesaremyown. 71 ThreatModels-AIAlignmentForum 70 SeeSection3.3of[2404.09932]FoundationalChallengesinAssuringAlignmentandSafetyofLargeLanguage Modelsformoreopenproblems. 69 [2405.10632]BeyondstaticAIevaluations:advancinghumaninteractionevaluationsforLLMharmsandrisks 68 [2406.04391]WhyHasPredictingDownstreamCapabilitiesofFrontierAIModelswithScaleRemainedElusive?, [2405.10938]ObservationalScalingLawsandthePredictabilityofLanguageModelPerformance 67 [2403.04893]ASafeHarborforAIEvaluationandRedTeaming 66 [2404.09932]FoundationalChallengesinAssuringAlignmentandSafetyofLargeLanguageModels, [2109.13916]UnsolvedProblemsinMLSafety,[2407.14981]OpenProblemsinTechnicalAIGovernance References Anewinitiativefordevelopingthird-partymodelevaluations.(n.d.).RetrievedJuly22,2024,from https://w.anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations AIcompaniesarenâtreallyusingexternalevaluators.(n.d.).RetrievedJuly22,2024,from https://ailabwatch.org/blog/external-evaluation/ AIcompaniesmakefreshsafetypromiseatSeoulsummit,nationsagreetoalignworkonrisks.(2024, May21).APNews.https://apnews.com/article/south-korea-seoul-ai-summit-uk-2c2b297872d86 0edc60545d5a5cf598 AISafetyInstituteapproachtoevaluations.(n.d.).GOV.UK.RetrievedJuly22,2024,fromhttps://w .gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-ap proach-to-evaluations AISafetyInstitutereleasesnewAIsafetyevaluationsplatform.(n.d.).GOV.UK.RetrievedJuly22,2024, fromhttps://w.gov.uk/government/news/ai-safety-institute-releases-new-ai-safety-evaluations- platform AnnouncingApolloResearch.(n.d.).ApolloResearch.RetrievedJuly22,2024,fromhttps://w.apollo research.ai/blog/announcing-apollo-research AnthropicâsResponsibleScalingPolicy.(n.d.).RetrievedJuly22,2024,fromhttps://w.anthropic.com /news/anthropics-responsible-scaling-policy Anwar,U.,Saparov,A.,Rando,J.,Paleka,D.,Turpin,M.,Hase,P.,Lubana,E.S.,Jenner,E.,Casper,S., Sourbut,O.,Edelman,B.L.,Zhang,Z.,GĂŒnther,M.,Korinek,A.,Hernandez-Orallo,J., Hammond,L.,Bigelow,E.,Pan,A.,Langosco,L.,...Krueger,D.(2024).Foundational ChallengesinAssuringAlignmentandSafetyofLargeLanguageModels(arXiv:2404.09932). arXiv.https://doi.org/10.48550/arXiv.2404.09932 Autonomousreplicationthreatmodels[draftinprogress].(n.d.).GoogleDocs.RetrievedJuly22,2024, fromhttps://docs.google.com/document/d/1be2HNkxPoH0P-q8SDsGrq4xMblPxlpYo9md_zEeS zl0/edit?usp=embed_facebook Barnes,B.(2024).ClarifyingMETRâsAuditingRole.https://w.alignmentforum.org/posts/yHFhWmu3 DmvXZ5Fsm/clarifying-metr-s-auditing-role Booth,H.(2024,July12).WhatWeKnowAbouttheNewU.K.GovernmentâsApproachtoAI.TIME. https://time.com/6997876/uk-labour-ai-kyle-starmer/ Cappelen,H.,Dever,J.,&Hawthorne,J.(2024).AISafety:AClimbToArmageddon? (arXiv:2405.19832).arXiv.https://doi.org/10.48550/arXiv.2405.19832 Casper,S.,Ezell,C.,Siegmann,C.,Kolt,N.,Curtis,T.L.,Bucknall,B.,Haupt,A.,Wei,K.,Scheurer,J., Hobbhahn,M.,Sharkey,L.,Krishna,S.,VonHagen,M.,Alberti,S.,Chan,A.,Sun,Q., Gerovitch,M.,Bau,D.,Tegmark,M.,...Hadfield-Menell,D.(2024).Black-BoxAccessis InsufficientforRigorousAIAudits.The2024ACMConferenceonFairness,Accountability,and Transparency,2254â2272.https://doi.org/10.1145/3630106.3659037 Commitments.(n.d.).RetrievedJuly22,2024,fromhttps://ailabwatch.org/resources/commitments/ Critch,A.(2024).Safetyisnâtsafetywithoutasocialmodel(or:Dispellingthemythofpersetechnical safety).https://w.alignmentforum.org/posts/F2voF4pr3BfejJawL/safety-isn-t-safety-without-a -social-model-or-dispelling-the Davidson,T.,Denain,J.-S.,Villalobos,P.,&Bas,G.(2023).AIcapabilitiescanbesignificantlyimproved withoutexpensiveretraining(arXiv:2312.07413).arXiv.https://doi.org/10.48550/arXiv.2312. 07413 Edwards,M.,&Jabs,L.B.(2009).Whensafetyculturebackfires:Unintendedconsequencesof half-sharedgovernanceinahightechworkplace.TheSocialScienceJournal,46(4),707â723. https://doi.org/10.1016/j.soscij.2009.05.007 Hubinger,E.(2022).Atransparencyandinterpretabilitytechtree.https://w.alignmentforum.org/posts /nbq2bWLcYmSGup9aF/a-transparency-and-interpretability-tech-tree Hubinger,E.(2023).Towardsunderstanding-basedsafetyevaluations.https://w.alignmentforum.org/ posts/uqAdqrvxqGqeBHjTP/towards-understanding-based-safety-evaluations Hubinger,E.RSPsarepausesdonerightâAIAlignmentForum.RetrievedJuly22,2024,from https://w.alignmentforum.org/posts/mcnWZBnbeDz7KKtjJ/rsps-are-pauses-done-right#comm ents Hubinger,E.,Schiefer,N.,Denison,C.,&Perez,E.(2023,August7).Modelorganismsofmisalignment: Thecaseforanewpillarofalignmentresearch-aialignmentforum.AIAlignmentForum. https://w.alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignmen t-the-case-for-a-new-pillar-of-1 ExamplesofAIImprovingAI.(n.d.).RetrievedJuly22,2024,fromhttps://ai-improving-ai.safe.ai/ FHI,F.ofH.I.-.(2020,January30).TheWindfallClause:DistributingtheBenefitsofAI.TheFutureof HumanityInstitute.http://w.fhi.ox.ac.uk/ FinancialimpactoftheBoeing737MAXgroundings.(2024).InWikipedia. https://en.wikipedia.org/w/index.php?title=Financial_impact_of_the_Boeing_737_MAX_groundi ngs&oldid=1228175994 Google-deepmind/dangerous-capability-evaluations.(2024).[Python].GoogleDeepMind. https://github.com/google-deepmind/dangerous-capability-evaluations(Originalworkpublished 2024) Guidelinesforcapabilityelicitation.(n.d.).METRâsAutonomyEvaluationResources.RetrievedJuly22, 2024,fromhttps://metr.github.io/autonomy-evals-guide/elicitation-protocol/ H,D.,&ThomasW.(2022).ComplexSystemsforAISafety[PragmaticAISafety#3]. https://w.alignmentforum.org/posts/n767Q8HqbrteaPA25/complex-systems-for-ai-safety-prag matic-ai-safety-3 Hendrycks,D.,Carlini,N.,Schulman,J.,&Steinhardt,J.(2022).UnsolvedProblemsinMLSafety (arXiv:2109.13916).arXiv.https://doi.org/10.48550/arXiv.2109.13916 Hendrycks,D.,Mazeika,M.,&Woodside,T.(2023).AnOverviewofCatastrophicAIRisks (arXiv:2306.12001).arXiv.https://doi.org/10.48550/arXiv.2306.12001 Ibrahim,L.,Huang,S.,Ahmad,L.,&Anderljung,M.(2024).BeyondstaticAIevaluations:Advancing humaninteractionevaluationsforLLMharmsandrisks(arXiv:2405.10632).arXiv. https://doi.org/10.48550/arXiv.2405.10632 IntroducingtheFrontierSafetyFramework.(2024,May14).GoogleDeepMind.https://deepmind.google/ discover/blog/introducing-the-frontier-safety-framework/ Jones,A.(2020).AreweinanAIoverhang?https://w.alignmentforum.org/posts/N6vZEnCn6A95Xn3 9p/are-we-in-an-ai-overhang JustinShovelain,&Elliot_Mckernon.(2023).ImprovingthesafetyofAIevals.https://w.lesswrong .com/posts/XCRsg2ZnHBNAN862T/improving-the-safety-of-ai-evals Kasirzadeh,A.(2024).TwoTypesofAIExistentialRisk:DecisiveandAccumulative(arXiv:2401.07836). arXiv.https://doi.org/10.48550/arXiv.2401.07836 Koessler,L.,Schuett,J.,&Anderljung,M.(2024).RiskthresholdsforfrontierAI(arXiv:2406.14713). arXiv.https://doi.org/10.48550/arXiv.2406.14713 Kolt,N.,Anderljung,M.,Barnhart,J.,Brass,A.,Esvelt,K.,Hadfield,G.K.,Heim,L.,Rodriguez,M., Sandbrink,J.B.,&Woodside,T.(2024).ResponsibleReportingforFrontierAIDevelopment (arXiv:2404.02675).arXiv.https://doi.org/10.48550/arXiv.2404.02675 LetâstalkaboutLLMevaluation.(n.d.).RetrievedJuly22,2024,fromhttps://huggingface.co/blog/ clefourrier/llm-evaluation Li,N.,Pan,A.,Gopal,A.,Yue,S.,Berrios,D.,Gatti,A.,Li,J.D.,Dombrowski,A.-K.,Goel,S.,Phan, L.,Mukobi,G.,Helm-Burger,N.,Lababidi,R.,Justen,L.,Liu,A.B.,Chen,M.,Barrass,I., Zhang,O.,Zhu,X.,...Hendrycks,D.(2024).TheWMDPBenchmark:MeasuringandReducing MaliciousUseWithUnlearning(arXiv:2403.03218).arXiv.https://doi.org/10.48550/arXiv.2403. 03218 Manheim,D.(2023).BuildingaCultureofSafetyforAI:PerspectivesandChallenges(SSRNScholarly Paper4491421).https://doi.org/10.2139/ssrn.4491421 McGregor,S.(2021).PreventingRepeatedRealWorldAIFailuresbyCatalogingIncidents:TheAI IncidentDatabase.ProceedingsoftheAAAIConferenceonArtificialIntelligence,35(17),Article 17.https://doi.org/10.1609/aaai.v35i17.17817 METR.(n.d.).RetrievedJuly22,2024,fromhttps://metr.org/ Nevo,S.,Lahav,D.,Karpur,A.,Bar-On,Y.,Bradley,H.A.,&Alstott,J.(2024).SecuringAIModel Weights:PreventingTheftandMisuseofFrontierModels.RANDCorporation.https://w.rand .org/pubs/research_reports/RRA2849-1.html Nosek,B.(n.d.).StrategyforCultureChange.RetrievedJuly22,2024,fromhttps://w.cos.io/blog/ strategy-for-culture-change OpenProblemsinTechnicalAIGovernance|GovAI.(n.d.).RetrievedJuly22,2024,from https://w.governance.ai/research-paper/open-problems-in-technical-ai-governance OpenAI,Achiam,J.,Adler,S.,Agarwal,S.,Ahmad,L.,Akkaya,I.,Aleman,F.L.,Almeida,D., Altenschmidt,J.,Altman,S.,Anadkat,S.,Avila,R.,Babuschkin,I.,Balaji,S.,Balcom,V., Baltescu,P.,Bao,H.,Bavarian,M.,Belgum,J.,...Zoph,B.(2024).GPT-4TechnicalReport (arXiv:2303.08774).arXiv.https://doi.org/10.48550/arXiv.2303.08774 Openai/evals.(2024).[Python].OpenAI.https://github.com/openai/evals(Originalworkpublished2023) Phuong,M.,Aitchison,M.,Catt,E.,Cogan,S.,Kaskasoli,A.,Krakovna,V.,Lindner,D.,Rahtz,M., Assael,Y.,Hodkinson,S.,Howard,H.,Lieberum,T.,Kumar,R.,Raad,M.A.,Webson,A.,Ho, L.,Lin,S.,Farquhar,S.,Hutter,M.,...Shevlane,T.(2024).EvaluatingFrontierModelsfor DangerousCapabilities(arXiv:2403.13793).arXiv.https://doi.org/10.48550/arXiv.2403.13793 Preparedness.(n.d.).RetrievedJuly22,2024,fromhttps://openai.com/preparedness/ RecommendationsforthenextstagesoftheFrontierAITaskforce.(n.d.).ApolloResearch.RetrievedJuly 22,2024,fromhttps://w.apolloresearch.ai/blog/recommendations-for-the-next-stages-of-the- frontier-ai-taskforce ReflectionsonourResponsibleScalingPolicy.(n.d.).RetrievedJuly22,2024,from https://w.anthropic.com/news/reflections-on-our-responsible-scaling-policy ResponsibleScalingPolicies(RSPs).(n.d.).RetrievedJuly22,2024,fromhttps://metr.org/blog/2023-09 -26-rsp/ RichardNgo[@RichardMCNgo].(2024,July18).IâmworriedthatalotofworkonAIsafetyevalsis primarilymotivatedbyâSomethingmustbedone.Thisissomething.Thereforethismustbe done.âOr,toputitanotherway:Ijudgeevalideason4criteria,andIoftenseeproposalswhich failall4.Thecriteria:[Tweet].Twitter.https://x.com/RichardMCNgo/status/ 1814049093393723609 RishiSunak[@RishiSunak].(2023,June12).AIisthedefiningtechnologyofourtimeandwehavea clearstrategytodevelopitinasafewaythatwillbenefiteveryoneintheUK.Hereâswhatthat lookslike ï https://t.co/jur7TCS84U[Tweet].Twitter.https://x.com/RishiSunak/status/ 1668169727552765954 Riskassessment.(2024).InWikipedia.https://en.wikipedia.org/w/index.php?title=Risk_assessment &oldid=1235318307 Ruan,Y.,Maddison,C.J.,&Hashimoto,T.(2024).ObservationalScalingLawsandthePredictabilityof LanguageModelPerformance(arXiv:2405.10938).arXiv.https://doi.org/10.48550/arXiv.2405. 10938 Schaeffer,R.,Schoelkopf,H.,Miranda,B.,Mukobi,G.,Madan,V.,Ibrahim,A.,Bradley,H.,Biderman, S.,&Koyejo,S.(2024).WhyHasPredictingDownstreamCapabilitiesofFrontierAIModels withScaleRemainedElusive?(arXiv:2406.04391).arXiv.https://doi.org/10.48550/arXiv.2406. 04391 Scholl,A.(2022).Safetywashing.https://w.alignmentforum.org/posts/xhD6SHAAE9ghKZ9HS/ safetywashing Shevlane,T.,Farquhar,S.,Garfinkel,B.,Phuong,M.,Whittlestone,J.,Leung,J.,Kokotajlo,D.,Marchal, N.,Anderljung,M.,Kolt,N.,Ho,L.,Siddarth,D.,Avin,S.,Hawkins,W.,Kim,B.,Gabriel,I., Bolina,V.,Clark,J.,Bengio,Y.,...Dafoe,A.(2023).Modelevaluationforextremerisks (arXiv:2305.15324).arXiv.https://doi.org/10.48550/arXiv.2305.15324 StrategicVision.(2024).NIST.https://w.nist.gov/aisi/strategic-vision Street,F.(2020,May25).WhenSafetyProvesDangerous.FarnamStreet.https://fs.blog/safety-proves- dangerous/ Tegmark,M.,&Omohundro,S.(2023).Provablysafesystems:TheonlypathtocontrollableAGI (arXiv:2309.01933).arXiv.https://doi.org/10.48550/arXiv.2309.01933 TheUKAISafetySummitâOurrecommendations.(n.d.).ApolloResearch.RetrievedJuly22,2024,from https://w.apolloresearch.ai/blog/the-uk-ai-safety-summit-our-recommendations TheoriesofChangeforAIAuditing.(n.d.).ApolloResearch.RetrievedJuly22,2024,from https://w.apolloresearch.ai/blog/theories-of-change-for-ai-auditing ThreatModelsâAIAlignmentForum.(2021,April19).https://w.alignmentforum.org/tag/threat- models Tong,A.,Paul,K.,&Tong,A.(2024,July15).Exclusive:OpenAIworkingonnewreasoningtechnology undercodenameâStrawberry.âReuters.https://w.reuters.com/technology/artificial-intelligenc e/openai-working-new-reasoning-technology-under-code-name-strawberry-2024-07-12/ U.S.andUKAnnouncePartnershiponScienceofAISafety.(2024,April1).U.S.Departmentof Commerce.https://w.commerce.gov/news/press-releases/2024/04/us-and-uk-announce- partnership-science-ai-safety U.S.SecretaryofCommerceGinaRaimondoReleasesStrategicVisiononAISafety,AnnouncesPlanfor GlobalCooperationAmongAISafetyInstitutes.(2024,May21).U.S.DepartmentofCommerce. https://w.commerce.gov/news/press-releases/2024/05/us-secretary-commerce-gina-raimondo-r eleases-strategic-vision-ai-safety Verma,P.,Tiku,N.,&Zakrzewski,C.(2024,July12).OpenAIpromisedtomakeitsAIsafe.Employees sayitâfailedâitsfirsttest.WashingtonPost.https://w.washingtonpost.com/technology/2024/ 07/12/openai-ai-safety-regulation-gpt4/ Wasil,A.,Clymer,J.,Krueger,D.,Dardaman,E.,Campos,S.,&Murphy,E.(2024).AffirmativeSafety: AnApproachtoRiskManagementforAdvancedAI(SSRNScholarlyPaper4806274). https://doi.org/10.2139/ssrn.4806274 Weidinger,L.,Barnhart,J.,Brennan,J.,Butterfield,C.,Young,S.,Hawkins,W.,Hendricks,L.A., Comanescu,R.,Chang,O.,Rodriguez,M.,Beroshi,J.,Bloxwich,D.,Proleev,L.,Chen,J., Farquhar,S.,Ho,L.,Gabriel,I.,Dafoe,A.,&Isaac,W.(2024).HolisticSafetyandResponsibility EvaluationsofAdvancedAIModels(arXiv:2404.14068).arXiv.https://doi.org/10.48550/arXiv. 2404.14068 Weidinger,L.,Rauh,M.,Marchal,N.,Manzini,A.,Hendricks,L.A.,Mateos-Garcia,J.,Bergman,S., Kay,J.,Griffin,C.,Bariach,B.,Gabriel,I.,Rieser,V.,&Isaac,W.(2023).SociotechnicalSafety EvaluationofGenerativeAISystems(arXiv:2310.11986).arXiv.https://doi.org/10.48550/arXiv. 2310.11986 Woodside,T.(2024,April16).EmergentAbilitiesinLargeLanguageModels:AnExplainer.Centerfor SecurityandEmergingTechnology.https://cset.georgetown.edu/article/emergent-abilities-in-large -language-models-an-explainer/