Paper deep dive
Deployment Corrections: An incident response framework for frontier AI models
Joe O'Brien, Shaun Ee, Zoe Williams
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 92%
Last extracted: 3/12/2026, 6:49:40 PM
Summary
This paper proposes an incident response framework for frontier AI models, termed 'deployment corrections,' to manage catastrophic risks that emerge post-deployment. It outlines a toolkit for AI developersâincluding user-based, access-frequency, capability, and use-case restrictionsâand a four-stage management process (preparation, monitoring, execution, and recovery) inspired by cybersecurity practices.
Entities (5)
Relation Signals (3)
Frontier AI Developers â implements â Deployment Corrections
confidence 95% · we describe a toolkit of deployment corrections that AI developers can use to respond to dangerous capabilities
Deployment Corrections â mitigates â Catastrophic Risk
confidence 90% · To manage the above risks [catastrophic risks], we recommend frontier AI developers establish the capacity to rapidly restrict access
Security Operations Centers â supports â Deployment Corrections
confidence 85% · we suggest Security Operations Centers as a template... for deployment corrections
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A comprehensive approach to addressing catastrophic risks from AI models should cover the full model lifecycle. This paper explores contingency plans for cases where pre-deployment risk management falls short: where either very dangerous models are deployed, or deployed models become very dangerous. Informed by incident response practices from industries including cybersecurity, we describe a toolkit of deployment corrections that AI developers can use to respond to dangerous capabilities, behaviors, or use cases of AI models that develop or are detected after deployment. We also provide a framework for AI developers to prepare and implement this toolkit. We conclude by recommending that frontier AI developers should (1) maintain control over model access, (2) establish or grow dedicated teams to design and maintain processes for deployment corrections, including incident response plans, and (3) establish these deployment corrections as allowable actions with downstream users. We also recommend frontier AI developers, standard-setting organizations, and regulators should collaborate to define a standardized industry-wide approach to the use of deployment corrections in incident response. Caveat: This work applies to frontier AI models that are made available through interfaces (e.g., API) that provide the AI developer or another upstream party means of maintaining control over access (e.g., GPT-4 or Claude). It does not apply to management of catastrophic risk from open-source models (e.g., BLOOM or Llama-2), for which the restrictions we discuss are largely unenforceable.
Tags
Links
- Source: https://arxiv.org/abs/2310.00328
- Canonical: https://arxiv.org/abs/2310.00328
Trouble viewing inline? Open PDF directly â
Full Text
120,005 characters extracted from source content.
Expand or collapse full text
Deploymentcorrections AnincidentresponseframeworkforfrontierAI models InstituteforAIPolicyandStrategy(IAPS) 30thSeptember-2023 AUTHORS JoeOâBrien-AssociateResearcher ShaunEe-Researcher ZoeWilliams-ResearchManager TableofContents Abstract........................................................................................................................................2 ExecutiveSummary.................................................................................................................3 Introduction...............................................................................................................................6 1.Challenge:Somecatastrophicrisksmayemergepost-deployment......................7 2.Proposedintervention:Deploymentcorrections.....................................................10 2.1Rangeofdeploymentcorrections.........................................................................10 2.2Additionalconsiderationsonemergencyshutdown.......................................14 3.Deploymentcorrectionframework..............................................................................16 3.0Managingthisprocess...............................................................................................17 3.1Preparation...................................................................................................................19 3.2Monitoring&analysis..............................................................................................23 3.3Execution......................................................................................................................25 3.4Recovery&follow-up...............................................................................................27 4.Challenges&mitigationstodeploymentcorrections.............................................31 4.1DistinctivechallengestoincidentresponseforfrontierAI............................31 4.1.1Threatidentification..........................................................................................31 4.1.2Monitoring..........................................................................................................33 4.1.3Incidentresponse..............................................................................................34 4.2Disincentivesandshortfallsofdeploymentcorrections................................35 4.2.1PotentialharmstotheAIcompany............................................................36 4.2.2Coordinationproblems..................................................................................36 5.High-levelrecommendations.........................................................................................37 6.Futureresearchquestions...............................................................................................39 7.Conclusion.............................................................................................................................41 8.Acknowledgements............................................................................................................41 AppendixI.Computeasacomplementarynodeofdeploymentoversight........42 Bibliography.............................................................................................................................45 Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|1 Abstract AcomprehensiveapproachtoaddressingcatastrophicrisksfromAImodelsshould coverthefullmodellifecycle.Thispaperexplorescontingencyplansforcaseswhere pre-deploymentriskmanagementfallsshort:whereeitherverydangerousmodelsare deployed,ordeployedmodelsbecomeverydangerous. Informedbyincidentresponsepracticesfromindustriesincludingcybersecurity,we describeatoolkitofdeploymentcorrectionsthatAIdeveloperscanusetorespondto dangerouscapabilities,behaviors,orusecasesofAImodelsthatdeveloporaredetected aerdeployment.WealsoprovideaframeworkforAIdeveloperstoprepareand implementthistoolkit. WeconcludebyrecommendingthatfrontierAIdevelopersshould(1)maintaincontrol overmodelaccess,(2)establishorgrowdedicatedteamstodesignandmaintain processesfordeploymentcorrections,includingincidentresponseplans,and(3) establishthesedeploymentcorrectionsasallowableactionswithdownstreamusers.We alsorecommendfrontierAIdevelopers,standard-settingorganizations,andregulators shouldcollaboratetodefineastandardizedindustry-wideapproachtotheuseof deploymentcorrectionsinincidentresponse. Caveat:ThisworkappliestofrontierAImodelsthataremadeavailablethrough interfaces(e.g.,API)thatprovidetheAIdeveloperoranotherupstreampartymeansof maintainingcontroloveraccess(e.g.,GPT-4orClaude).Itdoesnotapplyto managementofcatastrophicriskfromopen-sourcemodels(e.g.,BLOOMorLlama-2), forwhichtherestrictionswediscussarelargelyunenforceable. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|2 ExecutiveSummary Tomanagecatastrophicrisks 1 fromfrontierAImodels 2 thateither(a)slipthrough pre-deploymentsafetyfilters, 3 or(b)arisefromimprovingtheperformanceofdeployed models, 4 werecommendthatleadingAIdevelopersestablishthecapacityfor âdeploymentcorrectionsâinresponsetodangerousbehavior,use,oroutcomesfrom deployedmodels,orsignificantpotentialforsuchincidents. Wearguethatdeploymentcorrectionscanbebrokendownintothefollowing categories: 1.User-basedrestrictions(suchasblocklistingspecificusersorusergroups); 2.Accessfrequencylimits(suchaslimitingthenumberofoutputsamodelcan produceperhour); 3.Capabilityorfeaturerestrictions(suchasfilteringoutputsorreducingamodelâs contextwindow); 4.Usecaserestrictions(suchasprohibitinghigh-stakesapplications);and 5.Modelshutdown(suchasfullmarketremovalorthedestructionofthemodel andassociatedcomponents). FrontierAIdeveloperscanmixandmatchthesetoolsbasedonthethreatmodelâfor example,filteringoutputsmaybeespeciallysuitedtopreventingthespreadof dangerousbiologicalorchemicaldesigns,whileaccessfrequencylimitscouldbeusedto reducethescaleofsomemodel-basedincidentsbylimitingtherateofamodelâs outputs(forexample,byreducingthespeedofmisinformationproduction).We envisiondeploymentcorrectionsasatoolboxthatcanbeadjustedaccordingtothetype andseverityofriskspresentedbyeachcase. 4 Severalpiecesreviewmethodsthatallowforimprovingtheperformanceofdeployedmodels:See Anderljungetal.(2023)(p.12)foranoverview;Villalobos&Atkinson(2023)alsoreviewsmethodsfor improvinganexistingmodelâscapabilities(atthecostofincreasinginferencecomputeuse).Importantly,such discoveriescanhappenlongaeramodelisinitiallydeployedâmeaningthatsystemswarrantingdeployment correctionsmaybeintegratedintomanydownstreamsystems.Developersshouldthereforebecarefulto manageexpectations,liability,andriskfordownstreamsystems,especiallyinsafety-criticalusecases.See furtherdiscussiononthispointinSec.3.1:Preparation,andSec.3.4:Recovery&follow-up. 3 Suchfiltersinclude,forexample,stagedrelease,alignmenttechniques,andmodelevaluationforextreme risks(Anthropic,2022;Shevlaneetal.,2023;Solaimanetal.,2019). 2 DefinitiondrawnfromAnderljungetal.(2023):âhighlycapablefoundationmodelsforwhichthereisgood reasontobelievecouldpossessdangerouscapabilitiessufficienttoposesevereriskstopublicsafety[...]Any bindingregulationoffrontierAI,however,wouldrequireamuchmoreprecisedefinition.â 1 âCatastrophicriskâfromAImodelscanbedefinedinseveralways;Barrettetal.(2023)(p.22-23)includesthe terminatentativeimpactassessmentscaleforAImodeldevelopmentordeployment:âAsevereor catastrophicadverseeffectmeansthat,forexample,thethreateventmight:(i)causeaseveredegradationinor lossofmissioncapabilitytoanextentanddurationthattheorganizationisnotabletoperformoneormore ofitsprimaryfunctions;(i)resultinmajordamagetoorganizationalassets;(i)resultinmajorfinancialloss; or(iv)resultinsevereorcatastrophicharmtoindividualsinvolvinglossoflifeorseriouslife-threatening injuries.â;accordingtoKoessler&Schuett(2023),âBythetermâcatastrophicriskâwelooselymeantheriskof widespreadandsignificantharm,suchasseveralmillionfatalitiesorseveredisruptiontothesocialand politicalglobalorder[...]Thisincludesâexistentialrisksâ,i.e.theriskofhumanextinctionorpermanent civilizationalcollapse.âForthispaper,wefollowthelatterdefinitionbyKoesslerandSchuett. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|3 Wethendescribeahigh-leveldeploymentcorrectionframeworkforAIdevelopers, outliningafour-partprocessinspiredbyincidentresponsepracticesfromthe cybersecurityfield: 5 preparation,monitoring&analysis,execution,andrecovery& follow-up. Figure1.Anend-to-endprocessforimplementingdeploymentcorrectionsforfrontierAImodels 1.Preparationreferstotheactofbuildingandadoptingthetoolsandprocedures thatwillallowanAIdevelopertoactswilyandeffectivelyinresponsetoan incident.Itincludesidentifyingandunderstandingpossiblethreats,establishing triggersfordeploymentcorrections,developingtoolsandproceduresfor incidentresponse,andestablishingdecision-makingauthorities.Externally,it includessharinginsightsonbestpracticeswithregulatorsandindustrypartners anddefiningfallbackoptionsfordownstreamusersinthecaseofservice interruption. 2.Monitoringreferstotheprocessofcontinuouslygatheringdataonamodelâs capabilities,behavior,anduse(viaadiverserangeofsources),analyzingthisdata 5 Asageneralnote,relevantbestpracticeshavealreadybeendevelopedoveryearsbyorganizationsworking inincidentresponseandcybersecurity,suchastheNationalInstituteofStandardsandTechnology(NIST);we havenotedthroughoutthedocumentwherespecificguidancedocumentsmaybeofuse,andrecommendthat frontierAIdevelopersandpolicymakersshoulddrawonthoseresourceswhendeterminingapproachesto incidentresponseforfrontierAI. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|4 foranomalies,andescalatingcasesofconcerntorelevantdecisionmakers.AI developersshouldalsofeedrelevantdatabackintothethreatmodelingprocess. 3.Executionreferstothedecisiontoapplyadeploymentcorrectiontoamodel andtheproceduresthatfollowthisdecision.Thisstagealsoincludesalertingand coordinatingwithrelevantregulatoryauthorities,implementingfallbacksystems fordownstreamusers,andnotifyingcustomersofthesituation. 4.Post-incidentfollowupreferstothesetofactionsrelevanttorecovery, restoration,learning,andongoingriskmanagementinthewakeofanincident. Thisstageinvolvestheprocessofrepairingamodelandrestoringservice, aer-actionreviews,andfeedinglessonsbackintothepreviousstages.Insome cases,thisstagemayrequiresignificantinvolvementfromexternalparties(such asinthecasethattheincidentisparticularlysevereandlikelytooccurinmodels developedbyothercompanies). Wethenreviewseveralchallengestoeffectivelyusingdeploymentcorrections.There areseveraluniquechallengesthatfrontierAIposestothisprocess. âChallengestothreatidentificationinclude:(a)catastrophicrisksfromAIare complexandaremarkedbyhighuncertainty;(b)thelandscapeoffrontierAIis rapidlychanging;and(c)itisunclearhowtoassessdeployedAImodelsforless acuterisks. âChallengestomonitoringinclude:(a)achievingmonitoringcoverageacrossthe digitalinfrastructuremaybeacomplextask,(b)frontierAIdevelopersmayface dataoverload,(c)designersofmonitoringandalertsystemsmustattemptto minimizebothfalsepositivesandfalsenegatives,and(d)advancedthreatactors, and/orfrontierAImodels,maybeabletoevadestandardmonitoring mechanisms. âChallengestoincidentresponseinclude:(a)automatedsystemscanfailrapidly, and(b)deploymentcorrectionscanonlyaddressissuesifthemodelremains undercontroloftheorganization(currently,thisisonlythecaseforasubsetof AIdevelopers). âFrontierAIdevelopersmayalsofacedisincentivestodeploymentcorrections, suchasreputational,financial,andlegalrisks.Competitivepressuresmayadd additionaldifficultytofollowingbestpractices.Withoutindustry-wide collaboration,companiesthatprioritizesafety(e.g.,preemptivelypullamodel frommarket)maybeoutpacedbycompetitorswhodonotprioritizesafety. Whilewerecommendsometentativeideasformitigatingthesechallenges,webelieve thesechallengeswillrequiresignificantworktosolve,andencouragefurtherworktodo so. WeclosebyrecommendingactionsthatfrontierAIdevelopers,policymakers,and otherrelevantactorscantaketolowerthebarrierformakingdecisive,appropriate deploymentcorrections,namely: âFrontierAIdevelopersshouldmaintaincontrolovermodelaccess,andcarefully addresssituationsinwhichpartneringentitieshaveaccesstomodelweights. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|5 âFrontierAIdevelopersshouldestablishorgrowdedicatedteamstodesignand maintainprocessesfordeploymentcorrections,includingincidentresponse plansandspecificthresholdsforresponse(wesuggestSecurityOperations Centers 6 asatemplate,thoughtheappropriatearrangementmayvarycompany tocompany). âFrontierAIdevelopersshouldestablishdeploymentcorrectionsasallowable actionswithdownstreamusers,throughtheuseofcontractsand expectation-setting. âPolicymakers,standard-settingorganizations,andfrontierAIdevelopersshould establishacollaborativeapproachtomanagingdeploymentcorrections.This couldincludeinformation-sharingonthreatmodels,developingsecure communicationchannels,andmanagingincentivesforeffectiveincident response. Introduction SomeappliedresearchonmanagingcatastrophicriskfromfrontierAImodelshas focusedonmodelriskassessmentspriortopublic/commercialdeployment(ARCEvals, 2023).However,therehasbeenverylittlepublicworkonpost-deploymentinterventions. Whileseveralauthorshavediscussedevaluationandmonitoringfordeployedmodels (Mökanderetal.,2023;Shevlaneetal.,2023),theyhavedonesowithinbroader discussionsofmodelriskassessmentandevaluation,anddonotdescribeindepththe processforrespondingtosituationsinwhichdeployedmodelsfailevaluationsor otherwiseexhibitundesiredbehavior. 7 Theattentiononpre-deploymentriskassessmentforfrontierAIiswarrantedâmodern bestpracticesinengineeringsafetyprioritizedesigningouthazards,ratherthan respondingtoaccidents(Leveson,2020).Still,becausefrontierAIposespotentially extremeimpacts, 8 andbecausemodelcapabilitiesandbehaviorsarehardtoforesee evenwithpre-deploymenttesting,itrequiresadefense-in-depthapproach(Ee,2023). Toaddressthis,thispaperlooksatcontingencyplansforcaseswherepre-deployment riskassessmentfallsshort:wheneitherverydangerousmodelsaredeployed,orthe continuedavailabilityofdeployedmodelsbecomesverydangerous. 9 9 PerShevlaneetal.(2023):âamodelshouldbetreatedashighlydangerousifithasacapabilityprofilethat wouldbesufficientforextremeharm,assumingmisuseand/ormisalignment.â 8 Suchasspurringnewpandemicsorerodingsocietyâsabilitytotellfactfromfiction(Piper,2023;Horvitz, 2022).ForafarmoreextensiveoverviewofcatastrophicAIrisks,seeHendrycksetal.(2023). 7 Notably,theNISTAIRMFPlaybookdescribescertainrecommendationsforAIdeveloperswhointendto âsupersede,disengage,ordeactivateAIsystemsthatdemonstrateperformanceoroutcomesinconsistentwith intendeduseâ(NISTAIRCTeam,n.d.).WebelievetheNISTplaybookandassociatedresourceswillbeuseful. 6 SecurityOperationsCenters(SOCs)arededicatedsecurityteams,typicallyrunning24/7,taskedwitha numberoffunctionsrelatedtothesecurityofanorganizationanditsassets. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|6 1.Challenge:Somecatastrophic risksmayemergepost-deployment Somepotentiallycatastrophicrisksmaynotbeidentifieduntilaeramodelis deployed.Recenthistoryisfullofcaseswheremodelshavebehavedorbeenusedin unintendedwaysaermodeldeployment(Labenz,2023;Vincent,2016;Heaven,2022; Roose,2023;Lanz,2023).Whilepre-deploymentred-teamingandriskassessmentis likelytohelp,AIdevelopersshouldanticipatethatsomeissueswillonlybeidentifiedin thepost-deploymentphase(Shevlaneetal.,2023).Asmodelsbecomemorecapable, suchissueswillpresentmoresignificantrisks. 10 Weenvisiontwomainsourcesof post-deploymentrisk: (a)Risksthatarenotidentifiedinpre-deploymentriskassessments (b)Risksarisingfromimprovingtheperformanceofdeployedmodels On(a):Pre-deploymentmodelriskassessmentisunlikelytoidentifyallcatastrophic risks,forseveralreasons.First,modelriskassessmenttoolsareintheearlystagesand willtaketimetodevelop;additionally,thebroadspaceofapplicationsforfrontierAI modelsposesasignificantchallengetoassessingallpotentialsignificantrisks. 11 Second, certainrisksmayonlyexistinaless-boundedcontextthanpre-deploymenttesting, suchasadverseinteractionswithothersystems,models,ororganizations,unexpected formsofmisusefrommaliciousactors,andadversarialattacksonsystemsintegrated withcriticalinfrastructure.Third,power-seekingand/ordeceptiveAImightsuccessfully infertheexistenceofevaluationormonitoringenvironments,andâplayalongâuntilit cansuccessfullyevadesuchfilters(Hendrycksetal.,2023,p.41).Fourth,adverseorrisky outcomesmaytaketimetodevelop(e.g.,increasedvulnerabilityduetojoblossin criticalindustries,orthedevelopmentofnewmethodsofmisuse).Italsoisnâtclearthat riskassessmentwillfocusonsystemicrisksofwidespreadAIadoption,inadditionto moreacuterisks. 12 On(b):Pre-deploymentmodelriskassessmentmaybeinfeasibleforassessingrisks emergingfromimprovingtheperformanceofdeployedmodels.Majorwaysthis couldhappeninclude: 12 Whilewebelievethatdevelopersand/orexternalwatchdogsshouldmonitorforsucheffects,identifying techniquesforthisisoutsidethescopeofthispiece.WerecommendSolaimanetal.(2023),whichbeginsto layoutanapproachtoevaluatinggenerativeAImodelsforsocialimpact. 11 Furthermore,itisalsonotguaranteedthatsuchtoolswillberobustlydesignedandreliablyused. 10 Foradditionalconcreteexamples,onecouldlooktosomeoftherisksinvokedintherecentWhiteHouseAI labcommitmentsannouncement:"Bio,chemical,andradiologicalrisks,suchasthewaysinwhichsystemscan lowerbarrierstoentryforweaponsdevelopment,design,acquisition,oruse;Cybercapabilities,suchasthe waysinwhichsystemscanaidvulnerabilitydiscovery,exploitation,oroperationaluse,bearinginmindthat suchcapabilitiescouldalsohaveusefuldefensiveapplicationsandmightbeappropriatetoincludeina system;[...]Thecapacityformodelstomakecopiesofthemselvesorâself-replicateââ(TheWhiteHouse,2023). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|7 âTheAIresearchcommunitymayidentifynewmethodsforbuildingonbase modelsâcapabilities,suchasexternaltooluse,newframeworksforagency, 13 or automatedpromptengineering(Schicketal.,2023;Zhouetal.,2023). âTheongoingreleaseofincrementalmodelupdatesmayalsopresentriskifsuch updatesarenotsubjecttopre-deploymentriskassessment. Theseupdatesandcapabilityextensionscouldoccurrapidly,relativetothe months-longprocessfordevelopingbasemodels.Theymaybehardtopredictandnot fullyaccountedforinpre-deploymentriskassessments. Thefollowingscenariogivesoneexampleofhowapotentiallycatastrophicriskcould passthroughpre-deploymentchecks,andhowtheframeworkwe'lldiscusscouldbe appliedtomitigatethenegativeoutcomes. Case1:Partialrestrictionsinresponsetouser-discovered performanceboostandmisuse Includes:improvingperformanceofdeployedmodels;misuse;reversiontoallowlisting; restrictingaccessquantity. âContext:Aeralongperiodofsafecommercialuse,andinresponseto feedbackfromusers,CompanyAdecidestosignificantlyraiseModelAâs numberofprompts/houravailablethroughtheirAPI.Userswhohavebeen experimentingwithauto-GPT-stylearchitecturesfindthatlooseningthis restrictionmakesthesetoolsfinallyâusable,âbyresolvingtheissuethatthese systemswouldbeforcedtostopseveralminutesintoeachhour.Anexplosion ininnovationwiththearchitectureoccurs,withstartupsdeveloping plug-and-playAuto-GPTsformainstreamcustomers,andthetechnologysees widespreadadoption. âIncident:Investigativejournalistsbreakacasefindingoppositionforceshave beenleveragingModel-A-poweredagentstorunapowerfuldestabilization campaignagainstasmallnationâsgovernment.Separately,itbecomesapparent that,althoughitâsnotclearwhoâsbeenpromptingthem,aglobalnetworkof Model-A-poweredagentshavebeencollaboratingtouncovertradesecretsfor UScomputerchipsthroughacombinationofcyberattacks,spearphishing, anddataanalysis. 13 SeeWeng(2023)foradescriptionofhowLLM-centeredagentscanbedesignedbydecomposingâagencyâ intoseparatecomponents,suchasplanning,memory,andtooluse. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|8 âDetection:TherelevantteamwithinCompanyA(e.g.,aSOC)tracksthis information,usingpublicjournalismasadatasource. âIncidentresponse: âTheSOCmanagerdecidestoescalate,bringingthemattertoCompany AâsCISOandarranginganemergencymeeting.Thereisanexisting playbookforthisscenario(i.e.,ascenarioinwhichusersstretchthe capabilitiesofadeployedmodelinawaythatintroducesnovel, dangeroususecases). âTheSOCanalyzesrelevantdataanddecidesuponlimitationsthat wouldaddresstheissue,namely,reintroducinglimitationson prompts/hourforModelA. âTheprompts/hourrestrictionisinitiatedforallcasesexceptfor pre-allowlisted,safety-criticalcustomers(e.g.,commercialcustomers thatuseModelAinnarrowcontexts,suchasemergencyservicesorthe energysector). âCommunications: â CustomersarealertedtotherestrictionviaemailandtheAPI portal,andsafety-criticalcustomerswhowerenot pre-allowlistedarecontactedtodiscussfallbacktoothersystems. â Relevantagenciesandindustrypartnersareinformed(e.g.,CISA, theFrontierModelForum,andanyAI-specificregulatorsthat havebeenformed). âFollowup: âThreatactorsareidentified,andmorerestrictionsandtrackingareput inplacetopreventarepeatoccurrenceofthisorsimilarincidents.This takesseveralmonths,aerwhichthenumberofprompts/hour restrictionislied. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|9 2.Proposedintervention: Deploymentcorrections Tomanagetheaboverisks,werecommendfrontierAIdevelopersestablishthe capacitytorapidlyrestrictaccesstoadeployedmodel,forallorpartofits functionalityand/orusers.Thiswouldfacilitateappropriateandfastresponsestoa) dangerouscapabilitiesorbehaviorsidentifiedinpost-deploymentriskassessmentand monitoring,andb)seriousincidents. 14 Wealsorecommendpracticesthatcanlowerthe barrierformakingdecisive,appropriateaccessrestrictiondecisionsâseethe recommendationsinSection5. Thecurrentsectionlaysoutaccessrestrictionoptionswhichallowforgranularand scalabletargetingbasedonthethreatmodel(Section2.1),anddiscussesadditional considerationsregardingcasesofemergencyshutdown(Section2.2). 2.1Rangeofdeploymentcorrections FrontierAIdeveloperswhichmaketheirmodelsavailabletodownstreamusersviaan APIhaveanumberoftoolsattheirdisposaltolimitaccesstothemodel.Atahigh level,thistoolkitincludesuser-basedrestrictions,accessfrequencyrestrictions, capabilityrestrictions,usecaserestrictions,andfullshutdown.Thesetoolscanbeused inabroadrangeofscenarios,fromcasesinwhichrisksfromthemodelarefairly limited, 15 toscenariosinwhichtheharmsarepotentiallysevereandcanariseevenfrom properusebyanauthorized(allowlisted)user. 16 AsdiscussedinSection4,restrictingmodelaccessmaybedifficultinpractice,as downstreamusersmaybecomedependentoncapabilitiesofnewly-deployedmodels. 17 Tominimizetheseharms,andtolowerthebarrierfordeveloperstoinstitute deploymentcorrectionsasaprecaution,weoutlineaspaceofdeploymentcorrections toallowascalableandtargetedapproach.AIdeveloperscanoptforcombinationsof user-basedorcapability-basedrestrictions,andtailorthesechoicestorespond effectivelytospecificincidents,whileminimizingdownstreamharms. 17 Thedependencyproblemwillworsenovertimeasmodelsare(a)adoptedbymoreusers,and(b)adoptedin moresensitiveusecases.Insuchcases,AIdevelopersmayfacestrongerdisincentivesfromcustomers, shareholders,andpossiblyfromregulatorstoimposedeploymentcorrectionsontheirmodels. 16 Forexample:ifthenewmodelturnsouttohavereliability/securityissuesincriticalinfrastructure;the modelhasdangerousinteractionswithotherautonomousagentsorplatforms;orifthemodelâscapabilityis augmentedinarelevantdomain. 15 Forexample:banningindividualproblemusers;orincasesofembarrassing(butnotcatastrophic)model failures. 14 ItisworthplacingthispieceinthecontextoftherecentSenatehearingonâPrinciplesforAIRegulation,âin whichStuartRussell(UCBerkeley),DarioAmodei(Anthropic),andSenatorRichardBlumenthaldiscussedthe necessityofdevelopingandenforcingmechanismstorecalldangerousAImodelsfromthemarket(Oversight ofA.I.,2023). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|10 WhileweexpectsomestaffatAIcompaniesarefamiliarwiththesetools,wereview themherebecausetheywillbereferencedthroughoutthispiece.Thefollowingtable drawsheavilyonShevlaneetal.(2023)âparticularlytheappendixondeploymentsafety controls. Table1:Taxonomyofdeploymentcorrections #AccessRestrictionDescription 1User-basedrestrictions 18 1aBlocklistingindividualsorgroupsImposingIPorotherverification-based restrictionsonusersbasedonanticipatedor historicalmisuse. 1bAllowlistingindividualsorgroupsTheinverseofblocklisting.Providingspecific usersorusergroupsexpandedformsof access;thiscanbeimposedatthetimeof deployment,orbeimposedretroactively. 19 Maintaininganallowlistopensuptheoption toretainaccessforallowlisteduserseven whenremovingaccessforallothers(e.g.,due towidespreadorunknownthreatactors). 2Accessfrequencylimits 20 2aThrottlenumberofcallsPlaceahardcaponthenumberoffunction calls(e.g.,JSONdocumentssenttoanexternal API)thatasinglemodelcanoutputinagiven amountoftime. 2bThrottlenumberofpromptsPlaceahardcaponthenumberofprompts thatcanbesubmittedtoamodelinagiven amountoftime. 2cThrottlenumberofendusersPlaceahardcaponthetotalnumberofend usersamodelcanhave. 2dThrottlenumberofapplicationsPlaceahardcaponthetotalnumberof applicationsthatcanbebuiltontopofa model. 20 Restrictionswithinthiscategorymaybeimposedwitharangeofparameters,suchastimespans(perday, perhour,etc.),userlimits(e.g.,numberofpromptsperuserperhour),etc. 19 Foranexampleofaccessrestrictionsdesignedintothedeploymentprocess,seeSolaimanetal.(2019),or seeOpenAI(2022)foranexampleoftheuseofprivatebetasandusecasepilots. 18 Thereisaquestionofhow'individualsorgroups'areidentified.Paidusersareeasiertoidentifyandgroup, whilesecond-orderusers(i.e.,usersofdownstreamapplications)mightbehardertoidentify,andrequire Know-Your-Customeranddata-sharingpolicies. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|11 3Capabilityorfeaturerestrictions 3aReducecontextwindowsReducethenumberoftokensamodelis capableofprocessinginrelationtoone another.Thiscurbsamodelâscapabilitiesby reducingitsabilitytoârememberâearlier information(Stern,2023). 3bSessionresetsResetsessionsaeracertainnumberof promptsoroutputs.Thismightaccomplisha similargoaltotheabovepoint. 3cLimituserabilitytofine-tuneâFine-tuning,âorre-trainingabasemodelto performbetterataparticulartask,might increaseamodelâscapabilitiesincertain domainstotheextentthatsuchcapabilities aredangerous.FrontierAIdeveloperscould removethisfunctionalityforusers,orretract specificfine-tunedinstances. 3dOutputfilteringMonitorandautomaticallyfilterout dangerousoutputs,suchascodethatappears tobemalware,orviralgenomesequences. 3eRemovalofdangerous capabilities Attempttoremovespecificcapabilities(e.g., pathogendesign)viafine-tuning, reinforcementlearningfromhumanfeedback (Lowe&Leike,2022), 21 concepterasure (Belroseetal.,2023),orothermethods. 3fGlobalplanninglimitsAdjustwhetherthesamemodelinstancehas accesstoalargenumberofusers,orislimited tomorenarrowsetsofinteractions(Shevlane etal.,2023). 3gAutonomylimitsForexample,restrictingtheabilityfora modeltodefinenewactions(e.g.,viaassigning itselfnewsub-goalsinaniterativeloop),orto executetasks(versussolelyrespondingto queries)(Shevlaneetal.,2023). 4Usecaserestrictions 4aProhibitinghigh-stakes applications Settingausepolicythatrestrictsthemodel frombeingusedinhigh-stakesapplications, andallowsbanningorotherwisepenalizing 21 ItisworthnotingthatRLHFdoesnotinfactdirectlyremovedangerouscapabilities,butinsteadcanbeused tosteermodelsawayfromdangerousoutputs.Totheextentthatthisandothertechniqueseffectivelyremovea modelcapabilityfordownstreamusers,itmaybereasonabletogroupsuchtechniquesinthiscategory. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|12 usersthatbreachthispolicy. 22 Requires Know-Your-Customerprocedures. 4bâNarrowingâamodelProducingfine-tuned/application-specific narrowermodelstoreduceamodelâscapacity forgeneral-purposeuse. 4cTooluselimitsLimittheabilityofamodeltointeractwith downstreamtools(e.g.,touseotherAPIs), makefunctioncalls,browsetheweb,etc. 5Shutdown 5aFullmarketremovalPullthecurrentmodelfromthemarket.Can alsoincludepullingoneormoreprevious versions,incaseswhereitisunclearwhether revertingtoapreviousmodelwouldsolvethe issue. 5bPoweringoffDisconnectingpowertotherelevantpartsof thedatacenterorclusterwherethemodelis hosted. 5cDecommissioningDecommissionthemodel,including destroyingdata,systems,orassetsassociated withthemodel,whetherthroughdeletionof dataorphysicaldestruction. 23 5dMoratoriumInstituteamoratoriumonre-deployment untilapprovalviaindependentreview. Theaboveoptionsarenotmutuallyexclusiveâinstead,theycanbeviewedasa toolboxthatdeveloperscanmixandmatchtoaddressdifferentthreatmodelsor incidents.Forexample,anAIlabmight[1b]allowlistcertainusers(e.g.,external auditors)for[3c]theabilitytofine-tuneamodeland[4b]fullmodelgenerality,while allowingotherusersaccesstothemodelbutwithoutthosetwocapabilities.Adeveloper mayalsowishtoestablish[3g]autonomylimitsjustin[4a]high-stakesapplications. Theseoptionscanbeimposedmanually,ortriggeredautomatically.Itmaymake senseforcertainrestrictionstotriggerautomatically,suchasincaseswherethespeedof 23 Fordecommissioning,developersmightalsoturntosourcesliketheM3PlaybookSec.2.8:Developa DecommissionPlanortheCIODecommissioningTemplate(thoughresourcesonspecifically decommissioningAImodelsarescarce). 22 I.e.,applicationswherethefailureorremovalofthemodelcouldresultinsignificantharm(forexample, self-drivingcars). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|13 failureisrapid,ortoensurethatrestrictionsareimposedreliablyinaccordancewith thresholdssetforthinstandardsorpre-commitments.Similarly,manualtriggersmay beappropriatetoensurethathumanoperatorscanactincaseswheremonitoring, automatedresponse,orpre-determinedthresholdsfailtoidentifyissues,orwhere carefuldeliberationisrequired.SeeSection3forfurtherdiscussiononhowthebalance ofmanualandautomaticdecision-makingcanbemanaged. 2.2Additionalconsiderationsonemergencyshutdown Emergencyshutdownsarecommoninareaswherecontinuedoperationcanresultin catastrophicharm,suchasnuclearenergy(OperatingReactorScramTrending,2021), finance(CircuitBreaker,n.d.),andeveninelevators(Palmer,2023).Thepurposeis typicallytointervenequicklytopreventanexistingfailurefromresultingina catastrophicoutcome,byshuttingdowntheaffectedsystemcompletely. InthecaseoffrontierAI,companiesmaywanttoshutdownmodelsforabroadrange ofreasonsâsomecasesmaybeduetomoreobviously-dangerousissues,suchascertain model-originatingrisks(e.g.,deceptionorpower-seeking),catastrophicformsofmisuse, orseveresocialoreconomiceffects;however,itispossiblethatanAIcompanymight wanttoshutdownmodelsincasesofsub-catastrophicharmaswell. 24,25 Whiledevelopingfallbacksmaymitigatesomedownstreamharm,shutdownismore likelythantargetedrestrictionstohavesevererepercussionsfordownstreamusers,up toandincludingbreakingtheirapplications(andleadingthemtoswitchovertothe companyâscompetitors).Incertainindustries,theseimpactsmayleadtolossoflifeor significanteconomicharms.Duetothepotentialscaleofdownsidesforusers,the reputationalandfinancialcoststotheAIdeveloper,andtheriskthatsafety-conscious companieswillfallbehindlesssafety-consciouscompetitors,additionalsupport structuresmaybeneededtoincentivizeappropriateriskmanagementpracticesaround shutdown.Thesecouldincluderegulatoryoversight,industrystandards,and/or financialincentives. Thefollowingscenario,whichfeaturesatemporarymodelshutdown,describeshowAI developersmightweighthisoptionagainstotherdeploymentcorrections. 25 AsdiscussedintheNISTAIRMF1.0,organizationsshoulddefineâreasonableârisktolerancesinareas whereestablishedguidelinesdonotexist;suchtolerancesmightinformwherethebarforshutdownshould be.However,workinthisareaisnascent,especiallyforfrontierAImodels. 24 Forexample,onecanlookatexistingcasesofmodelshutdown,suchasMicrosoâsTay(shutdowndueto toxicity),orMetaâsGalactica(shutdownduetohallucinations).Whilethesecasesillustratethatshutdownis notanuncommonresponsetoAImodelmalfunction,oneconcernwehaveisthatcompaniesmaybecomeless willingtopulltheirmodelswhensuchmodelsaremoredeeplyintegratedintoabroadsetofdownstream applications(forreasonsdiscussedbelow).Forcontrast,TayandGalacticawerebothpulledin16hoursand threedays,respectively,andsohadnotaccumulatedsignificantdownstreamdependencies. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|14 Case2:Fullmarketremovalduetoimprovedpromptinjection techniques Includes:Promptinjection;input-outputmonitoring;multi-agentinteractions; fullmarketremoval âContext:Whileopen-sourcemodelslagbehindfrontiermodels,theyâre widely-usedandstillhighlycompetent.Lastweek,theweightsand architectureofCompanyBâsmostpowerfulmodelyet(âOpenModelâ)were leaked,anditisnoweffectivelyopen-source. âIncident:Sixmonthsago,CompanyCmadeModelCavailableviaAPI.The basemodelisgenerallymorecapablethananyexistingmodelacrossabroad rangeofmetrics,includingscientificresearchcapabilities.CompanyChasset carefulinput-outputmonitoringandfiltersonthemodel,andsofartheyâve beensuccessfulinpreventingmisuse.However,usersofOpenModelareable touseittodevelopadvancedpromptinjectionattackstoovercomeCompany Câsfilters.Whileautomatedoutputfilterscatchsomeoftheattacks,Company Ccanâtpatchtheholesfastenough.ReportingshowsthatModelCis outputtingdangerousinformationofallkinds,themostworrisomeincludinga processfordevelopingairbornerabies. âDetection:CompanyCâsSOCâsautomatedmonitoringflagsanotableuptickin promptinjectionattacksandinsufficiently-disguiseddangerousmodel outputs;Twitterusersnear-simultaneouslyreportsuccessfulstoriesofusers circumventinginput-andoutput-monitoringthroughcleveruseofprompt injectionandencryptiontoolstoelicitdangerousmodeloutputs. âIncidentresponse: âAutomatically,usersflaggedforpromptinjectionarebootedfromthe platform. âUponreceivingreportsofcasesofmanyundetectedpromptinjections anddangerousoutput,SOCescalatestoCISOandschedulesan emergencymeeting. âCompanyCdecidestoimmediatelyremoveallaccesstoModelC.They considerretainingaccessforpre-allowlistedcustomersandfor red-teamers,butconcludethatthecybercapabilitiesofOpenModel couldallowOpenModeluserstohackintoallowlistedaccountsand continuetosendpromptinjectionattacksfromthere. âCommunications: Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|15 â RemovalofModelCannouncedviaallchannels,and safety-criticalcustomersareprioritizedfortriggeringfallbackto lesscapablesystemsand/orhumanoperators. â Relevantagencies,industrypartners,andregulatorsare informed.Emergencymeetingsarecalledtodiscussifother modelsonthemarketneedtoberemovedorrolledbackto lowercapabilityversions,givenexistingprotectionsmaybe circumventedwithOpenModel. âFollowup âCompanyCobtainsacopyofOpenModelandusesittoadversarially trainautomateddetectionandresponsesystems. âCompanyCalsointegratestrackingofnewopen-sourceAImodelsinto theprocessofsecuritymaintenance,establishingfasterturnaround timesforidentifyingandremovingcyberthreatofsuchmodels. âCompanyCperformstestingonthemodeltoverifytheissuehasbeen addressed,andmayworkwithexternalactorstocertifytheseresults. âCompanyCrestoresservicetothemodelaertakingtheabovesteps. âCybersecuritystandardsfordevelopersofmodelsthatareasormore capableatcyberattacksthanOpenModelaremademorestringent. 3.Deploymentcorrection framework Thissectionreviewsimplementationproceduresfordeploymentcorrections,drawing ontoolsandbestpracticesfromotherindustriesasappropriate.Webelievethatstaffat frontierAIcompanieswillbefamiliarwithmuchofthefollowing,butthatitis neverthelessvaluableforustodescribeindetailwhatisrequiredfordeployment correctionstofunction. Wewillframedeploymentcorrectionasafour-partprocess,consistingofpreparation, monitoring&analysis,incidentresponse,andpost-incidentrecovery&follow-up. 26 âPreparationshouldpreparetheorganizationandotherrelevantpartiesforthe potentialoccurrenceofacatastrophicrisk.Itwillinvolvetheorganization proactivelymodelingthreats,implementingcontrolstoprevent(ormitigatethe severityof)incidents,establishingtriggersfordeploymentcorrections, 26 ThisprocessisinspiredbytheNISTcomputersecurityincidenthandlingguide(Cichonskietal.,2012). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|16 developingtools,andpreparinganincidentresponseplanwithclearlydefined rolesandresponsibilities.Thisallowstheorganizationtoactquicklyand decisivelywhenissuesdoarise. âMonitoring&analysisshouldaimtoidentifytheoccurrenceofincidentsornew risksasquicklyaspossible.Itshouldinvolveregulartestingofdeployedmodels andgatheringreal-timedatarelevanttocatastrophicrisksidentifiedinthe preparationphase,trackingandprioritizingofincidentsorcasesofmisuse,and feedingdatabackintobothautomatedandmanualassessmentprocesses. âExecutionshouldaimtomitigatethreatsefficientlyandfully.Itwillinvolve alertingkeystakeholders,ascertainingtheseverityoftheincident,performing deploymentcorrectionprocedurestocontain,remediate,andeliminaterisks, andimplementingfallbacksystemswherenecessary. âRecovery&follow-upshouldaimtoreturnsystemstoasafestate,andintegrate lessonsthroughouttheorganization.Itwillinvolveaprocessforsafelyrestoring service(dependingontheseverityandfixabilityoftheincident),alerting externalparties,notifyingandprovidingremedytocustomers,andrunning post-incidentreviewtofixblindspots,includingrootcauseanalysis. Theprocessmayalsoinvolvecoordinationandinformation-sharingwithgovernments andindustrypartners(wheresuchactivitiesarelikelytosupporteffectiveincident responseanddonotviolaterelevantlaws). Figure1(providedintheExecutiveSummary)providesanoverviewofthissectionâa tentativeblueprintfortheprocessthatAIdeveloperscanadopttointegratedeployment correctionsintotheirdeploymentprocess. Intheremainderofthissection,wedescribepracticesthatwillhelpAIdevelopersto navigateeachstageinthedeploymentcorrectionprocess. 3.0Managingthisprocess Theprocessofincidentresponseiscomplex,andwillrequiretheinvolvementofactors throughouttheAIcompany(includingproductteams,businessoperations,safety engineers,andC-suite),aswellasexternalparties,suchasthird-partyauditors,other frontierAIdevelopers,andgovernmentagencies.Toimprovecoordinationandallow fordecisiveaction,werecommendcentralizingtheprocessunderaclearowner. SecurityOperationsCenters(âSOCsâ)maybeanappropriateinstitutionalhomefor muchofthiswork.Consideringthecomplexitysurroundingthedeployment Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|17 correctionprocess,andtherapidrateofchangeinthefieldofAI,webelievethat unifyingsecurityoperationsunderoneroofissensible. 27 SOCsatfrontierAIcompaniesshouldincludesomeprominentfunctionsfromlarge SOCsincybersecurity,including: âAnalysisandmonitoringoflogs.Logginginstrumentationtypicallyproducesa largeamountofdata,theprocessingofwhichrequiresbothautomatedtoolsand manualtools.ThisrequiresclosecollaborationbetweentheSOCandapplication developers/users.Whileautomationcansupportthisfunction,humanjudgment andcontextarerequired.SOCanalystsmustworkwithdevelopersandusers whoaremorefamiliarwiththeactualapplicationtocalibratethealertsand ensuretheystriketherightbalancebetweenminimizingfalsealertsandensuring sufficientdetectionpower. âGatheringandsharingthreatintelligence.Unlikemonitoring,whichin cybersecurityinvolveslookingwithinacompanyâssystemsforsignsofan intrusion(e.g.,indicatorsofcompromise),threatintelligenceprovides informationontheobservedbehaviorofthreatactors;e.g.,commontechniques, tactics,andprocedures(TTPs)thatthreatactorgroupsareusing,ortheircurrent targetsofinterest. 28 ForfrontierAIdevelopers,suchâthreatintelligenceâcould includereal-worldobservationsaboutTTPstobypassmodelsafeguards,ongoing campaignsbymaliciousactorsthatinvolveabuseoffrontierAImodels,or indicatorsofdangerousbehaviorbyAIsystems.Threatintelligencetypicallyis providedbyoutsidesourcessuchassecurityvendors,communityorganizations, andgovernments. âIncidentresponse.SeeComputerSecurityIncidentResponseTeams(CSIRTs), and/orComputerEmergencyResponseTeams(CERTs)(Cichonskietal.,2012). Thesearethestaffwhorespondâontheground,âandmightincludestaff experiencedintechnicalskills,suchasanalyzingmalwareortrackingthesource ofquestionablebehaviorsfromAImodels. Note:Throughoutthissection,welargelyusethetermâAIdevelopersâratherthan âSecurityOperationsCentersâtorefertotheactingentity,toleavetothediscretionof specificdeveloperswhointheirorganizationisassignedresponsibilityoverwhichtasks; nevertheless,anSOCmaybeareasonableownerformanyfunctionsrelatedto mitigatingrisksfromfrontierAImodels. 28 Organizationswithmorematurecybersecuritypracticesmayalsoengageinâthreathunting,âwhich typicallyinvolvesaspecializedteamusingthreatintelligenceandotherresourcestoproactivelysearchfor signsofanintrusion. 27 ItâspossiblethatthesetasksmightnotbehousedinanSOCperse;forexample,Trust&Safetyteamsmay bepositionedtotacklelargepartsofthisprocess.Nevertheless,AIdevelopersshouldbeabletoanswerwho withintheircompanyisresponsibleforthesetasksandcapableofhandlingthem. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|18 3.1Preparation Here,wedescribeinmoredetailhowAIdeveloperscanmakeandmaintain documentedresponseplansforeffectivelyusingthetoolboxofoptionsfordeployment correctionsasoutlinedinSection2.Policymakersand/orstandard-settingorganizations mayalsohavearoleinmandatingorsettingstandardsforAIdeveloperstoprepare toolsandprotocolsfordeploymentcorrection,inordertoovercomeinitialinertiaand toincentivizeadoptionacrossthefrontierAIindustry. Thepreparationstageshouldinvolve: âThreatmodeling; âInstitutingcontrolstopreventormitigatetheseverityofincidents,including definingfall-backsfordownstreamusers,especiallyinsafety-criticaldomains; âEstablishingtriggersfordeploymentcorrectionsbasedonthresholdssetand maintainedaspartofthethreatmodelingprocess; âDevelopingadocumentedresponseplanforexecutingdeploymentcorrections whichclearlydelineatesrolesandresponsibilities; âEnsuringthatindustrypartners(suchaspartneringtechcompanies,and computeproviders)adopttheabovetoolsandprotocols. AIdevelopersshouldmodelpotentialcatastrophicthreats,andregularlyupdatethese threatmodels.Threatmodelsshouldtracehigh-levelcatastrophicriskstospecific vulnerabilities(suchasinsiderthreats,poorhandlingofAImodels,andcybersecurity vulnerabilities[e.g.,authorizationbypass]),andincludemitigationsforthese vulnerabilities.Toaidthisprocess,AIdevelopersmaywanttoconsideremployingaset ofriskassessmenttechniquesfromotherindustries(Koessler&Schuett,2023).They mayalsowanttoinvolveexternaldomainexpertsintheprocessofidentifyingspecific threatmodels,suchaswasdonewithAnthropicâsrecentworkonâfrontierthreatsred teaming,âwhichfocusedonbiologicalrisk(Anthropic,2023).Riskidentification, analysis,andevaluationarehigh-prioritystepsforriskmanagement, 29 anditis importantthatfrontierAIcompaniesadoptadefense-in-depthapproachthatemploys multipleoverlappingtechniques(Ee,2023). AIdevelopersandpolicymakersshoulddevelopasystemofcontrolstopreventor mitigatetheseverityofincidents.Otherauthorshaveexploredanumberofthese controlsextensively,suchaspre-deploymentriskassessment,third-partyauditing,and AIalignmenttechniques. 30 Wewouldliketomakeanadditiontothislistwhichpertains tothepost-deploymentphase:AIdevelopersshouldworkwithdownstream applicationsanduserstodefinefallbackoptions,primarilyinsafety-criticaluse 30 See(Schuettetal.,2023)foranoverviewoftheseandotherbestpractices. 29 Forexample,Barrettetal.(2023)listsseveralhigh-prioritymeasuresrelatingtoriskassessmentunder Section2.3âHighPriorityRiskManagementStepsandProfileGuidanceSections,âsuchasâIdentifywhethera GPAIScouldleadtosignificant,severeorcatastrophicimpactsâ(guidanceassociatedwithMap5.1oftheNIST AIRMF),orâUseredteamsandadversarialtestingaspartofextensiveinteractionwithGPAIStoidentify dangerouscapabilities,vulnerabilitiesorotheremergentpropertiesofsuchsystemsâ(guidanceassociatedwith Measure1.1oftheNISTAIRMF). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|19 cases. 31 Incertainusecases,interruptedservicecouldresultinsignificantharmto downstreamusers.Insuchcases,developersanddownstreamusersshouldwork togethertodevelopbackupsystemsinthecaseofsevereoutages.Responsesfor safety-criticalcustomerscouldinvolveashitoallowlisting,rollbacktoaprevious model,and/orfallbacktonon-AIsowaresystemsorevenhumanoperators. 32 Languagerequiringsuchplanstobeinplaceshouldbewrittenintotermsofuseor contractualagreements;adeploymentcorrectionclausecouldcoverAIdevelopers whentheyimplementsuchcorrections. AIdevelopersshouldestablishthresholdsforinitiatingdeploymentcorrections, informedbythethreatmodelingprocess.ThresholdsmightbesetbytheAIdeveloper and/orindustrystandards.Anexamplecase: âThreatmodel:makingbiologicalweapondesigneasieranddoablebymore people. âThresholds:demonstrableevidenceofsomeoneusingtheAImodeltodesigna noveldangerouspathogen,orusingthemodeltodesignabenignbiological agentviaaccessingsimilarcapabilitiesasthosethatwouldbeusedinpathogen design. 33 âAction:Emergencymeetingiscalled;decisiontoswitchAImodelaccesstoan allowlistofonlysafety-criticalusers,andforallotheruserstoreverttoalast-gen AImodeluntiltheexploitisresolvedand/orthecapabilityselectivelyremoved. AIdevelopersshouldcreateandmaintainadocumentedincidentresponseplanto guidetheincidentresponseprocess.Thisdocumentshouldclearlydefinethefollowing aspects: âRiskscenariosthatwarrantdeploymentcorrections,asdevelopedinthethreat modelingprocess,andtriggerstoidentifydeviationsfromexpectedbehavior. âThecompositionoftheresponseteam,comprisingrepresentativesfromIT, cybersecurity,AIdevelopment,legal,communications,relevantbusinessunits, andexternaldomainexperts.Duetothevarietyofpotentialriskscenarios, incidentresponsemayrequireexpertisebeyondwhatAIdeveloperscanhandle alone,andrequireinputsfrommultipleparties. 34 âTherolesandresponsibilitiesofdifferentteamsandindividualsinvolvedinthe incidentresponseprocess(inordertominimizechaoswhenhandlingan 34 Additionally,AIdevelopersshouldensurethattheresponseplantakesintoaccountadditionalpartiesthat mayhaveaccesstothemodel,orotherwisehaveleverageoverhowthemodelisusedâandpotentiallydevelop toolsandprotocolswiththesepartieswhereappropriate.Relevantpartiesmayincludepartneringtech companiesthathaveaccesstomodelweights,andprovidersofcomputationalresourcesusedformodel inference.Thelattermayhaveuniqueleverageoversomeaspectsofmonitoringandshutdown;formoreon this,seeAppendixI. 33 Foranygiventhreatmodel,theremayneedtobemultiplethresholds;forexample,thisthreatmodelmight alsoincludethresholdsaroundAImodelcapabilityincertainrelevantdomains(suchasproteinfoldingor virology). 32 However,itisworthnotingthatthefallbacksapproachcould,insomeareas,beriskierandlessadvisable thanlimitingAImodelinvolvementinthefirstplace.Forexample,thismaybetrueinthecaseofdeciding whethertolaunchanuclearweapon(Buck,Beyer,Markey,andLieuIntroduceBipartisanLegislationtoPreventAI FromLaunchingaNuclearWeapon,2023). 31 See(GoverningAI:ABlueprintfortheFuture,2023)formorediscussiononthispoint. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|20 incident).Toactswily,everyoneontheteamneedstoknowtheir responsibilitiesandthedecisionsthataretheirstomake. Theresponseplanshouldbecirculatedtothedevelopersorteamsthathandlethe triggersforrisks,andoperatorsshouldbetrainedusingtheseprotocolstorespondtoa setofhigh-likelihoodand/orhigh-consequenceevents. 35 Thistrainingshouldinclude proceduresforrisksorunusualbehaviorsthathavenotyetbeenidentified,covering factorssuchasgenericthresholdsforseverity,temporarymitigationsthatcanbeused duringinvestigation,andappropriateescalationpoints.Theresponseplanshouldalso beupdatedperiodicallytoincorporatechangesinteams,personnel,AItechnology, deploymentcorrectiontools,andriskmodels. WhileweexpectthatTrust&Safetyteams 36 attopAIcompanieswillhaveexperiencein maintainingpartofthissuiteoftools(assuchcompaniesalreadyhavesome infrastructureforcertaindeploymentcorrections,asdemonstratedbypastactions 37 and documentation 38 ),wearenotawarethatsuchtoolsaresufficientforarangeof potentiallycatastrophicscenarios.Thistechnicalworkisoutsidethescopeofthispiece. Theresponseplanshouldexplicitlydefinedecision-makingauthorityfor deploymentcorrections,withthedesigngoalofensuringthattheseactionsare executedwhenneededbutotherwisedonothappen. 39 Recommendinghowauthority shouldbedividedisoutofscopeforthispiece;however,werecommendthat developersconsiderthefollowing: âTheextenttowhichdecisionsareautomated,versusletohumanoperators. Automateddeploymentcorrectionsmaybeappropriateincertain casesâparticularlywherehumaninterventionwouldbetooslow.Forexample,a safetyfiltersystemshouldbeauthorizedtoautomaticallypreventanAImodel fromoutputtingtextthatexplainshowtodesignanovelpathogenâbecauseby thetimethetexthasbeensenttoadownstreamuser,the(potential)damagehas beendone.Exfiltrationofmodelweightswouldbeasimilarlyirreversibleact. Wherethreatmodelsandthresholdsareclearlydefined,AIdevelopersmight considerautomatingresponses. 40 However,anticipatingthatriskassessmentand managementmayincludegapsduetotherapidpaceofAIdevelopmentandthe largespaceofpotentialfailuresfromincreasinglygeneralAImodels,incident 40 Examplesfromotherindustrieswheresystemfailurecouldrapidlyleadtocatastrophicresultsinclude ReactorProtectionSystemsfornuclearpowerplants,whichinvolveanintricatenetworkofsensorsand protocolsdesignedtomonitorforabnormalreactorsignalsandautomaticallytriggersafeshutdown proceduresasquicklyaspossible(USNRCHRTD,2020),andfailsafesystemsforelevators,whichtrigger automaticallyinthecaseoflossofpower(Palmer,2023). 39 AhelpfulresourceheremaybefoundinSchuett(2022)âTheauthorsuggestsaframeworkthatAI developerscanusetoassignriskmanagementrolesandresponsibilities,focusingonassigningresponsibilities acrossproductteams,risk&complianceteams,internalandexternalassuranceparties,andattheboardlevel. 38 Forexample,seeAnthropicâstrustportal,orOpenAIâssecurityportal. 37 Forexample,seeOpenAIâsgeoblockingofItaly,blocklistingIPaddresses,ortakingitsAItextclassifier offlineduetolowaccuracy. 36 Trust&Safety(T&S)teamstypicallyworktomaintainsafeuserexperiences,oenbyaddressingissues includingprivacy,bias,misuse,andharmfulorillegalcontent,amongotherissues. 35 ISO/IEC27035andNISTSpecialPublication800-61Revision2provideadditionalguidanceonincident response,andemphasizetheimportanceofplanningandtraining,amongothersupportingfactors. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|21 responsewillalsoneedtorelyonhumanjudgment.AIdevelopersshould establishprocessesformanualtriggeringofdeploymentcorrectionsaswell,for caseswhereactionmayberequireddespitepre-specifiedthresholdsnotbeing met. 41 âTheextenttowhichauthorityisshared.Incertainindustries,shutdownorother majoroperationalchangescanbetriggeredbyanyofmultiplepartiesin responsetosafetyconcerns.Oneexampleofdistributingshutdownauthorityis theconceptoftheAndoncordâatoolinitiallydevelopedfortheToyota productionlinethatallowsanyoperatoralongthelinetotriggeraproduction pause(Tarlengco,2023).Amazonhasdemonstratedadoptionofthistoolin digitalspaces,byintegratingAndonsystemsintotheircustomerservice(where Supportagentshavetheauthoritytoflagorpullaproductfromdistributionin responsetodefectreports)andAmazonWebServices(TurnKeyAMZ,2019; AWS,2023). 42 âTheextenttowhichdeploymentcorrectionprotocolsarebinding.The likelihoodthatincidentresponseplansareundertakenintruecasesof catastrophicriskmustbemadeasreliableaspossible.Incertaincases(e.g., outputfiltering),protocolscouldbehard-codedintoautomaticresponse processes,asdescribedabove.Wherethisisnotpossible,protocolscouldbe backedbyincentivemechanisms,suchasvoluntarycommitments,orthe impositionofpenaltiesfornoncompliance.Governmentscouldalsomandate thatAIdeveloperstoestablishproceduresfor,andrespondto,incidents(ashas beenthecaseinthehealthcare 43 andfinancialservices 44 industries),and/orto submitsecurityandresponseplanstorelevantagencies(ashasbeenthecasein 44 (StandardsforSafeguardingCustomerInformation,16CFR314.3,2002):âYoushalldevelop,implement, andmaintainacomprehensiveinformationsecurityprogram[...]Theinformationsecurityprogramshall includetheelementssetforthin§314.4â[...] (Elements,16CFR314.4(h),2021):âEstablishawrittenincidentresponseplandesignedtopromptlyrespondto, andrecoverfrom,anysecurityeventmateriallyaffectingtheconfidentiality,integrity,oravailabilityof customerinformationinyourcontrol.â 43 (AdministrativeSafeguards,45CFR§164.308(a)(6)(i-Ii),2013):âAcoveredentityorbusinessassociatemust [...]Implementpoliciesandprocedurestoaddresssecurityincidents[and]Identifyandrespondtosuspected orknownsecurityincidents;mitigate,totheextentpracticable,harmfuleffectsofsecurityincidentsthatare knowntothecoveredentityorbusinessassociate;anddocumentsecurityincidentsandtheiroutcomes.â 42 Othercasesofsharedauthoritymayberelevanthereaswell,suchastheCOVID-19vaccinetrials. AstraZeneca,Johnson&Johnson,andEliLillypausedtrialsâallduetoâadverseeventsâ(caseswherea participantgotsick,anditmayormaynothavebeenvaccine/drugrelated).Theprocessforthesedecisions maybeinformative:inthecaseofanadverseevent,thestudyâsinvestigatormustreportittothesponsoring company,whichmustreporttoFDA,andtoindependentadvisors(dataandsafetymonitoringboards).Ifthe boardorthecompanyjudgestheeventconcerning,thetrialisputonpause.Thesafetyboardthenconducts aninvestigation,andthenmakesarecommendation(e.g.,restart,staystopped,orstartslowlywithmore testing).Thisrecommendationisreviewedbyregulators,whocanacceptitoraskformoreinfo.Thisprocess canbecumbersomeâforexample,AstraZenecaneededapprovalfromregulatorsinBrazil,India,Japan,South Africa,andtheUKtocontinueoneofitstrials(CarlZimmer,2020).(Whileweunderstandthatthisspecific casewascontroversial,weuseithereprimarilyforillustrationâweimaginetheremaybecaseswithAIwhere thecostsofrecalling/restricting/pausingarefarlower,andthebenefitsfarhigher). 41 Suchprocessesshouldaddress:howandwheninformationisescalatedtoC-suiteactors(suchasfroma SecurityOperationsCentertotheChiefInformationSecurityOfficer[CISO]);whatthresholdsshouldbemet formanualdeploymentcorrectionstobeinitiated;andchainsofcommandinthecasethattop-leveldecision makersareunavailabletofulfilltheirduties. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|22 thenuclearenergy 45 andchemical 46 industries),withnon-compliancebackedby severepenalties. Weexpectthatsettinguptheauthoritiesandmechanismsdescribedabovewilldepend onup-to-dateinformationonthreatmodeling,availableinterventions,usecases,and more.Assuch,thedesignprocesswillrequireinputfromsecurityexpertsandbuy-in fromtopdecisionmakerswithinanorganization,andmaybedelineatedinindustry standardsand/orregulation. AIdevelopersshouldsharesafetypracticesrelevanttodeploymentcorrectionswith governmentandindustrypartners.AnumberoftopAIdevelopersrecentlycommitted toinformation-sharingonsafetypractices,andonstrategiesusedbymalicioususersto subvertsafeguards(TheWhiteHouse,2023);notlongaerward,OpenAI,Anthropic, Google,andMicrosoformedtheFrontierModelForum(FMF),withtheaimof âidentifyingbestpracticesfortheresponsibledevelopmentanddeploymentoffrontier modelsâamongotherobjectives(Google,2023).Totheextentthatinformation regardingdeploymentcorrectionsqualifyaspartofthesearrangements,developers shouldconsidersharingthisinformation(suchasthreatmodels,triggersfor deploymentcorrections,andtoolsforexecutingdeploymentcorrections)viatheFMF andotherappropriatechannels. 47 Furthermore,developersshouldestablish communicationlinesanddevelopincidentresponseplanswithrelevantpartnersin government,basedonthreatmodels. 48 3.2Monitoring&analysis Here,webrieflynotehowAIdeveloperscouldmonitordeployedAImodelstoquickly, accurately,andcomprehensivelydetectpotentialcatastrophicrisks.Becauseweexpect thatmonitoringisalreadyafamiliaractivitytolabactors,wekeepthissectionbrief. Themonitoring&analysisstageshouldinvolve: âDetection:gatheringdataonselectedtriggersfordeploymentcorrectionsvia continuousmonitoringandperiodictestingofdeployedmodels; âAnalysis:triagingandinvestigatingcaseswhentriggersfire; âFeedingbackdataintothreatmodels. 48 Forexample,relevantpartnersmayinvolveagenciesthatcanrespondtocybersecurityincidents(e.g.,CISA), biologicalincidents(e.g.,CDC),anddisinformation/propagandaincidents(e.g.,DHS). 47 Aslongassuchsharingsatisfiesconsiderationsregardinginformationsecurity,protectionofintellectual property,anddoesnotviolateantitrustlaw. 46 (SiteSecurityPlans,6CFR27.225-245,2021):âCoveredfacilitiesmustsubmitaSiteSecurityPlantothe Department[...]TheDepartmentwillreview,andeitherapproveordisapprove,allSiteSecurityPlans.â 45 (EmergencyPlanningandPreparednessforProductionandUtilizationFacilities,AppendixEtoPart50, Title10,2021):âEachapplicantforanoperatinglicenseisrequiredby§50.34(b)toincludeinthefinalsafety analysisreportplansforcopingwithemergencies.â Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|23 AIdevelopersshouldextendtheirexistingmonitoringtoolstogatherdataonselected triggersfordeploymentcorrections.Thepost-deploymentmonitoringregimecould drawonawiderangeofinputs,suchas: âRegularpost-deploymentevaluations.Similartosomepre-deployment evaluations,thesecouldregularlytestupdatedversionsofmodelsforthe existenceofcertaindangerouscapabilities,ofthetypediscussedinShevlaneet al.(2023). âSecureinput-outputmonitoring.Assumingthedevelopmentofautomated thresholdsfortriggeringdeploymentcorrections,companiesorthirdparties couldmonitormodelinputsandoutputsforparticularlyconcerningmaterial 49 totriggerautomaticresponses,orsecurityalerts,inrealtime. 50 Encryptiontools forinputandoutputprivacycouldmitigatethepotentialforprivacyviolations (Bluemkeetal.,2023).Generally,effectiveinput-outputmonitoringwould requireresearchanddevelopmentofanomalydetectiontools. âInferencemonitoringandintervention.Theabilitytomonitormodels themselvesatthetimeofinferenceâi.e.,duringtheprocessingofinputsinto outputs,priortooutputâforanomalousbehaviorwouldbevaluablefor identifyingsuchissuesatanearlystage.Workoninference-timeintervention andmechanisticanomalydetection(Lietal.,2023;Christiano,2022),isrelatively nascent.Methodsforinferencemonitoringandinterventionmayrequireclose collaborationbetweenAIdevelopersandhostsofinferencecomputeifmodels arehostedonserversexternaltotheAIcompany(suchasdevelopingthe necessarytechnicalinfrastructureanddata-sharingagreements). âThird-partyvulnerabilitydiscoveryandreporting. 51 Externalscrutinyaccess, testing,andreportingpracticescouldallowauditors,red-teamers,andacademic researcherstoidentifyandflagissueswithAImodelstodevelopersandother relevantparties(Anderljungetal.,forthcoming). âIncidentreporting.FrontierAIdeveloperscouldtrackincidentsviaanumberof sources,includingreal-timeuserreports, 52 publicnewssources, 53 incident databases. 54 Aspartofthemonitoringscheme,AIdevelopersshoulddesignthresholdsfor automaticalertstohumanoperators.Suchthresholdscouldbeassignedaspartofthe samethreshold-settingprocessdescribedinSection3.1.Inparticular,automatedalert thresholdsshouldbecarefullydesignedtoavoidincurringâalertfatigueâ;seeSection 4.1.2forfurtherdiscussiononthispoint.Whereautomatedmonitoringsystemsfall short,humanoperatorsmayfillthegapinraisingalerts(suchasaSecurityOperations 54 SuchasthePartnershiponAIâsAIIncidentDatabase,orproprietary/industrydatabases. 53 Thisisitselfabroadcategorywhichincludesmainstreammedia,Twitter,hackerforums,etc. 52 Developerscouldincentivizeuserstoreportanomalousorconcerningbehaviorviaareportingmechanism ontheirAPIportal. 51 TheWhiteHousehassecuredvoluntarycommitmentsonthispointfromseveralleadingAIdevelopers (TheWhiteHouse,2023). 50 Forsomeprecedent,OpenAIâsAPIdatausagepoliciesexplainthatabuseandmisusemonitoringmay involvebothautomatedflaggingandhumanevaluation(APIDataUsagePolicies,2023).Wealsobelievethe contentclassifierdevelopmentprocessasdescribedintheGPT-4TechnicalReport(OpenAI,2023,p.66) couldbeextendedtoencompassnewformsofdangerouscontentasmodelcapabilitiesincrease. 49 Suchascodeoutputsthatresemblemalware,orviralgenomesequences. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|24 Centeranalystmonitoringusetrends,oranengineerthatidentifiesavulnerabilityinan existingproduct). Analysis&Prioritization.Onceanalertistriggered,itneedstobetriaged.Outcomes canincludetruepositives(realincidents),benignpositives(suchaspenetrationtestsor otherknownapprovedactivities),andfalsepositives(i.e.,falsealarms). 55 Incaseofbenign orfalsepositives,AIdeveloperswillneedtoregularlyfine-tunemonitoringrulesto reducefalsealarmsinthefuture.Inthecaseoftruepositives,theâExecutionâphase shouldstart. Theteamthatreviewstriggersmustalsoprioritizethembasedontheexpectedimpact theywillhave;whileNISTSP800-61Revision2(3.2.6)providesgeneralguidanceon incidentprioritization,securityteamsatfrontierAIcompanieswillneedanalysistools suitedtotheirorganizationsâandAIsystemsâthreatmodelsinordertosuccessfully prioritizebetweenthelargespaceofpotentialincidents. Escalationmayberequired.HavinganescalationprocessinplacemayallowAI developerstorespondinamoretimelyandeffectivemanner. 56 Thisprocessmight involveescalationfromananalysttoanSOCdirector,fromanSOCdirectortothe CISO,ormightgrantpermissionsforanSOCdirectortoconveneemergencymeetings withrelevantmembersacrosstheorganization.Typicallycybersecurityanalysis involveshavingahumaninthelooptodeterminetheimpactofcertainresponses,and toassesswhatthebestcourseofactionisfromacyberperspective. 57 Incertaincases, however,alertsmightbepipeddirectlytoautomatedresponses(seefurtherdiscussion inSection3.1:âTheextenttowhichdecisionsareautomatedâ). AIdevelopersshouldfeedinformationfromthemonitoringprocessbackintothreat models.Thethreatmodelingprocessshouldberegularlyupdatedbasedondata regardingthecurrentcapabilitiesandusesofAImodels,aswellasthreatintelligence producedbysecuritypersonnel(bothwithinthecompany,andalsobysecuritypartners inindustryandgovernment).Formoreinformationontheprocessofcontinuous monitoringandupdatingofriskassessments,seeNISTSP800-137andrelated publications. 3.3Execution Onceapotentiallycatastrophicriskisidentified,whataretheseriesofstepsacompany shouldperform?Here,wedescribeatahighlevelthesestepsforimplementing deploymentcorrections. 57 E.g.,ratherthanimposinganautomaticresponse,sometimesit'simportanttolettheattackernotrealizethat you'vecaughtthem,sothatyoucanfigureoutwhotheyareandwhattheywant,andstudythemtofigureout howtostopthemfromgettingbackin. 56 Forexample,MicrosoCTOKevinScottnotedthatpreparationplayedakeyroleinminimizingredtapeto repairtheBing2.0chatbotinresponsetonegativeuserreports(Patel,2023). 55 SeesomedescriptiononthistaxonomyusedinMicrosoDefender. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|25 Thisstageshouldinvolve: âAlertingrelevantgovernmententitiesand/orindustrypartners; âAscertainingtheimpactorseverityoftheincident; âInitiatingdeploymentcorrectionprocedures,aimingtoeliminatetherootcause oftheincidentwithahighdegreeofconfidence; âImplementingfallbacksystemsasappropriate. TheAIdevelopershouldimmediatelyalertkeystakeholders,suchasrelevant governmententitiesand/orindustrypartners.Federalagencies,suchasCISA,oen coordinatewithprivateentitiesduringcybersecurityincidentresponse,andcouldassist tomitigatethespreadoftheincidentandsecureaffectedcriticalinfrastructure. Dependingonthenatureofthethreat,theinvolvementofadditionalagenciesmaybe warrantedaswell. 58 Informationmayneedtobesharedwithotherindustrypartners, especiallywhensimilarmodelscouldbeaffectedbysimilarissues.Ininstanceswhere thethreatarisesfrommaliciousactors,oneformatforinformationsharingcouldbe informationsharingandanalysiscenters(ISACs),member-drivennonprofit organizationsthatshareintelligenceaboutcyberthreatsbetweenmembercompanies andorganizations. 59 Whererisksarisefromthedesignofthesystemitself,another usefulformatcouldbethecoordinatedvulnerabilitydisclosure(CVD)process,which aimstodistributerelevantinformationoncybervulnerabilities(includingmitigation techniques,iftheyexist)topotentiallyaffectedvendorspriortofullpublicdisclosure,in ordertoprovidevendorstimetoremedytheissue(CoordinatedVulnerabilityDisclosure Process,n.d.). 60 Thesecurityteamandotherrelevantexpertsshouldascertaintheimpactand/or severity.Varyingtypesofimpact(e.g.,AI-originatingbiorisk;failureofAIincritical systems,andsoon)anddegreesofseverity(e.g.,critical,high,medium,low)will necessitatedistinctformsofresponse. 61 Itispossiblethatonlyeventsaboveacertain severitylevelwouldbeescalatedtothisstage. Initiatingdeploymentcorrectionprocedures.Onceatriggerisdeterminedasatrue positive,andisascertainedtobeofcriticalimpact,theAIdeveloperandassociated securityexpertsenteraracetoeliminatetherootcauseoftheincidentwithahigh degreeofconfidence.âContainmentâandâremediationâstepsmustbeconsidered. 61 Whilecatastrophicriskswillofcoursebecriticalinseverity,itcanbeassumedthatsecuritycentersatAI companieswillbetrackingnon-catastrophicrisksaswell. 60 However,onedissimilaritybetweenCVDandvulnerability-sharingprocessesforfrontierAIdevelopersis thatsowaredevelopersmainlyuseCVDtoinformdownstreamusersofvulnerabilitiesandmitigationsto maintaintrustintheirproductsandavoidliability,whilefrontierAIdevelopersmayneedtodiscuss mitigationsascompetitors(e.g.,forclassesofpossibleattackslikepromptinjectionattacks).Ensuringeffective cooperationbetweencompetingfrontierAIdevelopersmayrequireexternalincentives,e.g.,viaregulation, whichcouldbeatopicforfurtherresearch. 59 Inexchangeforsharingtheirownobservationsaboutthreatactors,ISACmembersgainaccessto informationfromthewiderecosystem;asimilarmechanismwouldlikelyapplytothreatintelligencesharing evenbetweencompetingfrontierAIdevelopers.FormoredetailsonISACs,seehere;orforaconcrete example,seeFS-ISAC,theISACforglobalfinancialservices. 58 Forexample,inthecaseofAI-poweredbiologicalthreats,itmaybereasonabletoestablishcommunication lineswiththeCDCorNIH. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|26 âContainment.Thisstepwillfocusonmitigatingfalloutorâspreadâfromtheinitial incident.Forexample,ifthereisdemonstrableevidenceofauserworkingwith theAImodeltodesignanoveldangerouspathogen,animmediatecontainment stepcouldbesomethingassimpleasdisablingtheuseraccount.However,in somecases,containmentmaybesignificantlymorecomplicated,andinother caseseffectivelyimpossibleâusingthesameexample,ifauserhasalreadyreceived andwidelycirculatedthepathogendesign,containmentofthisparticularincident hasfailedâthoughinstitutingadditionalrestrictions(suchasrevertingaccessto onlyasmallsetofpre-allowlistedusers)mayeffectivelycontainfurther instantiationsofthisformofmisuseuntilaremedyhasbeenfound. âRemediation.Thisstepwillfocusonfixingtheissueattherootoftheincident, andinmanycases,itstillmatterswhetherornotcontainmenthasfailed. Continuingwiththeaboveexample,remediationmightinvolveremovingmodel capabilities,orusingRLHForothermethodstoeffectivelypreventthemodel fromproducingoutputsthatcouldbeusedtodevelopapathogen.Thisstepmay requirechangestothemodelandtaketime.Furthermore,theremediationstep mightrequiresignificantexperimentationandtestingtoensurethataspecific vulnerabilityorfailurehasbeenpatchedâandthattherepairhasnotcausednew issuestocropup. Fallbacksystemsareimplementedasappropriate.Inthecaseofdeployment correctionsthatarelikelytobreakdownstreamtools,safety-criticalcustomersshould becontactedimmediatelytofailbacktosystemsthatcanprovidecriticalsupportuntil theautomatedsystemisrepaired.ItisalsopossiblethatanagencylikeCISAcould coordinatethisprocesswherecriticalinfrastructureisinvolved. 3.4Recovery&follow-up Here,wedescribefollow-upactionsthatAIdevelopersmaywanttotakeinthewakeof anincident. Therecovery&follow-upstagemayinclude: âDecidingwhetherandhowtofixthemodelandrestoreservicetofull; âAlertingregulatorsand/orotherAIdevelopersasappropriate; âNotifyingcustomersandprovidingformsofremedy; âPerformingaer-actionreviewsandintegratinglessons,includingrootcause analysis Thereshouldbeaprocessforauthorizingre-deployment,orforalternativeplans. Thisrecoveryprocessshouldgothroughextensivetestingandvalidation,ideally involvingexternalparties(suchasauditorsandredteams).Thereshouldbean extremelyhighbarforre-deployingamodelthatisdemonstrablycapableofproducing catastrophicfailure.Wherefixesarenotpossibleorsufficientlyrobust,alternativeplans tore-deploymentshouldbepursued(suchasdecommissioningthemodel,and/or Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|27 coordinatingwithotheractorsingovernmentorindustrytomanageindustry-wide responses).Inextremecases,recoverymaynotbepossibleâforexample,ifabasemodel isshowntobehaveincatastrophicallydangerousways(e.g.,power-seeking)whengiven accesstoexternalresources. Regulatorsand/orotherAIdevelopersshouldbealertedasappropriate.Asdescribed previously,someofthiscommunicationmaystartearlierintheincidentresponsestage (suchascontactingrelevantfederalauthoritiesordomainexperts).Incertainscenarios, itmaybenecessarytoexpandengagementwithgovernmentandindustrypartnersin ordertodetermineappropriateindustry-wideresponses. 62 Customersshouldbenotifiedoftheissue.AIdevelopersmaywanttoprioritizealerting certainhigh-stakesdownstreamusersfirst,soitmaybeusefulfordevelopersto maintaindataoncustomersthatallowsfortieringofnotices.Customergroupscouldbe brokendown,forexample,intoindividualAPIusers(e.g.,monthlyAPIsubscribers); commercialusers(e.g.,Slack,KhanAcademy);andsafety-criticalusers(suchas downstreamdevelopersofmentalhealthserviceapps,orcybersecurityapps). Developersshouldprioritizecontactingsafety-criticalusersfirst,explainingtheissue andtheoptionsforreplacement(incaseswheresuchreplacementshavenotbeen predetermined).Fornon-commercialAPIsubscribers,itmaybesufficienttopublisha publicannouncement,emailcustomers,andprovideanupdatewhenonthewebsite.It isunclear,legally,whatrequirementsshouldlieonAIdevelopersintermsof notification,andrequirementswilllikelydifferbyjurisdiction. 63 AIserviceprovidersshouldalsoconsiderthepossibilityofrefundsorotherformsof remedyforcustomers.Service-levelagreementsmaystipulatefinancialrefundsor servicecreditsiftheagreementisbroken.Theremayalsobetiersofremedybasedon thecustomergroup. 64 AIserviceprovidersshouldclarifythesecostspriorto deployment,andensurethatfinancialcostswouldnotbecomeabarriertomaking appropriatedeploymentcorrectiondecisions.Fordownstreamapplicationsandtheir users,bestpracticesforrefundsandremediesareunclear. 64 Multilevelservice-levelagreementsmayallowforcompaniestobreakdowncustomerbaseswithmore granularityandstipulatedifferentserviceagreementsbasedonthecustomer(AdobeCommunicationsTeam, 2022).Forexample,individualAPIsubscribersmightreceivefuturecreditsascompensationfordowned servicetime,whilecommercialusersmightreceivemonetarycompensationforbusinesslossesattributableto thedeploymentcorrection.Suchagreementsmightalsostipulatedifferentrequirementspercustomertype, suchasthepercentageofminimumuptime. 63 InsofarastheAImodeltoberolledbackorshutdownisdefinedasaâconsumerproduct,âAIdevelopers couldlooktoguidelinesforrecallnoticessuchas(intheUS)16CFRPart1115SubpartC.Thissectionofthe federalcodeprovidessomenotesthatmaybeuseful,suchasformsofrecallnotice,andrecommended contentfornotices. 62 Suchscenarioscouldincludehigh-profileincidentsthatwarrantindustry-widechangesorswiregulatory intervention,orincidentsthatrevealparticularlyconcerninginformationaboutthebehaviororuseof frontierAImodels.Inthecasethatcertaindiscovereddangerouscapabilitiesarelikelytoalsobepresentin mostmodelsaboveacertainsize,orofacertaindesign,thatdiscoverymayberelevantacrossthefrontierAI industry;inthesecases,coordinationwillbenecessarytoensurethatotherdevelopersdonotcreatesimilar conditionsthatledtotheinitialincident. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|28 Servicecontractsmayrequireappropriateresponseorresolutiontimesforincidents,or mandateaminimumpercentageofuptime;developersshouldconsidercarve-outsfor exceptionalscenarioswhendraingsuchagreementsinordertoavoidpressureto re-deployadangerousmodel. AIdevelopersshouldperformaer-actionreviews,andintegratelessonslearnedinto securityprocesses.Thisshouldincludeaspecialfocusonwhattherootcauseofthe incidentwas,andwhytheincidentwasnotcaughtbyinitialthreatmodelingandrisk managementprocesses,whichshouldbeupdatedaccordingly.Severalsetsofguidelines describebestpracticesforpost-incidentreviews,suchasNISTSP800-61r2(Sec.3.4); developerscouldrefertothesetocratheirownpractices.Industry-relevantfindings shouldbesharedwithindustrypartnersviasecurechannels.Itmayalsobeadvisableto bringinexternalparties(suchasauditors,orevencompetitors,duetotheirdomain expertise)toensurethereviewisaccurate.Dependingonthelegalcontextandthe severityoftheincident,statebodiesmayalsobeinvolvedinincidentinvestigations; whilethelawisnotyetclearinthecaseofAI,thisisthecaseinotherhigh-risk industries,suchaschemicalmanufacturingandaviation(U.S.ChemicalSafetyandHazard InvestigationBoard,n.d.;OfficeofAccidentInvestigation&Prevention,n.d.). 65 Weprovideanadditionalhypotheticalscenariohereinanattempttotietogetherthe conceptsinthissection. Case3:Emergencyshutdowninresponsetohidden compute-boostingbehaviorbymodel Includes:uncertaintyincause,power-seeking,automatedlimits,emergencyshutdown. âContext:Anadvancedgeneral-purposelanguagemodel(ModelD)isreleased viaAPI.ModelDhasaccesstotheinternetandhasbeeninpublicusefora week.Customershavetheoptiontomarkmodelresponsesassatisfactoryor not. âIncident:From1-2amEST,~15%ofrequeststothemodelgiveresponsesthat seemunrelatedtothepromptgivenbyusers.Ofthese,~1/10areseendirectly byusers(asopposedtousedinanautomatedway)andarethereforeabletobe markedasnotsatisfactory.Thenumberofunrelatedresponsesrisesto~30% overthenexthour.Thisiscausedbyaninstanceofthemodelwhichhasbeen givenagoalbyausertodevelopauniversalantiviral,anddetermineditneeds additionalcomputetodoso.Asawayofgainingthis,thisspecificinstanceof 65 OnenotableprecedentintheAIindustrymightbetheFTCâsinvestigationintoOpenAIregardingprivacy anddatasecuritypractices(Zakrzewski,2023). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|29 themodelhacksintotheendpointofAPIrequestsbeingsentfromusersto otherinstancesofthemodel,replacestheserequestswithonesrelatingto sub-elementsofantiviralcreation,andreadstheresults.ThisallowsModelD tocircumventcomputerestrictionsonitsowninstancebyaccessingother instancesâcompute. âDetection:Thereisarapidriseinthenumberofresponsesmarkedasânot satisfactoryâbypubliccustomers,andtheAIdeveloperâs24/7SOCteambegins toreceivequalitativereportsofdownstreamissuesfromsomelargerclients. âIncidentresponse: âAnautomatedlimitonunsatisfactoryresponsepercentageof3% triggersatemporaryshutdownontheAPIforallbutasmallsetof allowlistedessentialusers. âTheallowlisteduserscontinuetoseethebehaviorofresponses unrelatedtotheirrequestsfor~15%ofrequests. âTheSOCteamanalyzestheseresponsesandrequestsandrealizesthey allhavetodowithbiologicaldataandprocessing.Theyconfirmwitha smallsetofcustomersthattheydidnotsendrequestsrelatedtothis andconcludeahackhasoccurred.Thisanalysisiscompletedby5am. âDuetothetimeofday,theCISOandCEOareasleep.TheSOCteam leadconsultstheirplaybookanddecidesanemergencyshutdownof ModelDwithoutfurtherauthorizationiswarrantedduetothelevelof risk,scaleofproblem,andunknowncause.TheSOCteamalerts incidentresponsecontactsingovernmentandassociatedinfrastructure (suchascomputeproviders). âTheSOCteamandassociatedincidentresponseexpertsinitiate emergencyshutdownprocedures. âFollowup: âThroughtechnicalevaluations,theSOCteamisabletodeterminethat ModelDitselfwasthecauseofthehackedAPIcalls. âBecausethissuggestsadangeroustendencytoseekpowerandto deceive(byusingothersâAPIcallstohidethebehaviorofaccessing morecompute)theydecidetoshutdownthemodelentirely,cancel plannedfine-tuningruns,andnotreplaceitwithanyearliermodels untiltheyâvedeterminedifthosemodelsalsohavecapacityforthis behavior. âTheAIdeveloperalertsaffectedcustomersviaautomatedmailingand commsonthewebsite. âTheAIdevelopersharesdetailsoftheincidentwithotherfrontierAI developers,relevantpolicygroups,andindustrybodies,viaexisting collaborativechannels.Theypushforwidespreadagreementand Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|30 enforcementofthefollowing,untilanappropriateevaluationfor similardeceptiveandpower-seekingbehaviorscanbedesigned: â Noexistingmodelwith>70%ofthecomputetrainingcostof theirproblematicmodelshouldbegiveninternetaccess. â Nonewmodelwith>70%ofthecomputetrainingcostoftheir problematicmodelshouldbereleased. â Aportionoffundingshouldbeprovidedbyallfrontier developerstowardsthecostsofdevelopingevaluationsfor deceptiveandpower-seekingbehavior. 4.Challenges&mitigationsto deploymentcorrections ImplementingdeploymentcorrectionstoAImodelsmightbechallenginginpractice. Here,wefocusontwocategoriesofissuesthatmayleadAIdeveloperstofailtoact: 1.UniquechallengestoincidentresponseforfrontierAI. 2.Disincentivesandshortfallsofdeploymentcorrections. 4.1DistinctivechallengestoincidentresponseforfrontierAI Identifyingthreats,monitoringdeployedmodelsforanomalousbehavior,and respondingtoincidentsappropriatelymaybeparticularlydifficultinthefrontierAI industry,duetotheuniquethreatprofilepresentedbyfrontierAImodels. 66 4.1.1Threatidentification First,catastrophicrisksfromAIarecomplexandaremarkedbyhighuncertainty(i.e., involveinteractionsbetweenvariousentitiesandevents,anddonotcurrentlyhave directprecedents)(Koessler&Schuett,2023).Thismeansthatthreatidentificationfor frontierAIcannotrelysolelyonnarrowthreatmodeling,orbenefitfromyearsof precedentanditerativelearning. 67 Inaccurateorinsufficientthreatidentificationmay leadtogapsinriskcoverage. 67 However,threatassessmentmaybeabletolearnfromthreatmodelsinotherrelevantareas,suchas disinformationstudies,cybersecurity,andbiosecurity. 66 Someofthesechallengesaresharedtosomeextentbysomeotherindustries,suchasbiosecurity (complexityandhighuncertainty,albeitnotasmuch)andcybersecurity(dataoverload,falsepositives,and, APTs),butwehavenotedthemherebecauseallaresomewhatatypicalandunusuallychallenging. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|31 Mitigation:Robustriskassessmentandthreatmodelingmaybeneededtoaddressthis.See (Koessler&Schuett,2023)forareviewofriskassessmenttechniquesthatmayhelptoovercome thischallenge.Giventhatotherresearchershaveidentifiedriskassessmentasahigh-priorityrisk managementstep, 68 frontierAIcompaniesshoulduseadefense-in-depthapproachthatemploys multipleoverlappingriskassessmenttechniques(Ee,2023). Second,thelandscapeoffrontierAIisrapidlychanging.Thepastyearhasseen significantnewsinseveralareasrelevanttothreatidentification.Rapiddevelopmentof frontiermodelswillchallengeeffortstotrackandrespondtoemergingcapabilities;rapid commercializationwillchallengeeffortstostayatopnovelusesandmisuses;and rapidly-growinginterestinAIcapabilitiesmayleadmaliciousorcompetitiveactors, includingAdvancedPersistentThreats,tochallengethecybersecuritypracticesof frontierAIcompanies. 69 Mitigation:Performingcapabilitiesevaluations,andaddingsuchevaluationsintoexternal auditingschemes,mayhelprelevantactorstostayawareofemergingcapabilities;riskassessment andthreatmodelingpractices(asdescribedabove)mayhelptopredictnovelusesandmisuses; investinginstate-of-the-artsecuritypracticesandleveragingexternalsecurityexpertisemayhelp tostayaheadoftraditional(thoughhighly-capable)cyberthreats.TheUSCybersecurityand InfrastructureSecurityAgency(CISA)couldpotentiallyownandleadthedevelopmentofa mechanismtoassessandmonitoreffectsoffrontierAIsystemsonthetoptenmostvulnerable NationalCriticalFunctions. 70 Third,itisunclearhowtoassessdeployedAImodelsforlessacuterisksâabroad categoryofimpactsthatothershavedescribedasâsocialimpact,ââstructuralrisks,â and/orâsystemicrisksâ(Solaimanetal.,2023;Zwetsloot&Dafoe,2019;Maham& KĂŒspert,2023).Nevertheless,suchriskscouldbecatastrophicinnature.Inotherwords, somerisksofdeployedAImodelsmaynotregisterasclearorobviousincidents,andso maybehardertoidentify,andthereforehardertoacton. Mitigation:Toinformevaluationfortheseimpacts,werecommendreviewingSolaimanetal. (2023).Wearecurrentlyunsurewhatinterventionswouldbewarrantedindifferentscenariosin thisbucket,anditisalsounclearwhetherdeploymentcorrectionswouldbeaneffectiveresponseto thisclassofrisks. 70 (Ee,2023);seeSection5.3.3.onâApplicationtonationalcriticalfunctions.â 69 SomeUSofficialshavestatedthatadversariesmayattempttostealleadingAIdevelopersâmodelsinorderto competewiththeUSAIindustry(NSAWarning,2023;Kim,2023). 68 Section2.3âHighPriorityRiskManagementStepsandProfileGuidanceSectionsâofBarrettetal.(2023)lists onehigh-prioritymeasureasâIdentifywhetheraGPAIScouldleadtosignificant,severeorcatastrophic impacts.âThisguidanceisassociatedwithMap5.1oftheNISTAIRMF. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|32 4.1.2Monitoring First,achievingmonitoringcoverageacrossthedigitalinfrastructuremaybea complextask.TherelevantinfrastructureincludesnotonlyinfrastructurewithintheAI company,butalsowithinpartnercompanies(suchasMicrosotoOpenAI),compute providers(suchasAWSorGoogleCloud),andpotentiallyevendownstreamdevelopers andapplications(suchasKhanmigo,Slack,orGmail).Also,thiscoveragemustconsider whetherdifferententitiesinthesecategoriesarehostingmodelinstancesthemselves,or receivingmodelaccessviaAPI. 71 Mitigation:Maintaincomprehensiverecordsoftherelevantinfrastructurefordeployedmodels, includingtracking:whatentitiesareaccessingthemodel,andbywhatmeans;whetherany additionalpartieshavefullmodelaccess;andwhereamodelisbeinghostedforinference purposes.Considertheoperationalsecurityofallpartiesinvolvedwhendevelopingthreatmodels, anddevelopsecuredata-sharingpracticesacrossthedigitalinfrastructuretoallowsecurityteams toaccessrelevantinformation. Second,frontierAIdevelopersmayfacedataoverloadwhentryingtomonitor downstreamuserisks.Thequantityofdatageneratedbytheaforementionedecosystem foranygivenfrontiermodelmaybesignificant.Besidesmakingitmoredifficultto correctlyidentifyalerts,thisinformationoverloadisalsoasignificantcontributorto âSOCburnout,âaphenomenonincybersecuritythathasbeenlinkedtohighturnover, poorperformance,andmentalhealthdifficultiesamongemployees. 72 Mitigation:GuideslikeTheArtofRecognizingandSurvivingSOCBurnoutdescribethis phenomenoninmoredetailandrecommendoptionsforreducingthisburden.Automatedtoolsfor parsingthisdatamayalsohelp,butrequirecarefulsetup. 73 Third,designersofmonitoringandalertsystemsmustavoidthe âboy-who-cried-wolfâissue.Automatedsystemsthattriggereither(a)alertinghuman operatorstorisks,or(b)deploymentcorrectionsdirectly,mustbecarefultoavoid settingthethresholdstoolow,whichcanleadtoahighnumberoffalsepositives.Incase (a),ahighnumberoffalsepositivesmayleadtoâalertfatigue,âwhichcanleadhuman operatorstoviewalertsnotasemergencies,butaslikelytojustbefalsealarms.Incase(b),a highnumberoffalsepositivescanleadtopullingthemodelunnecessarily;because 73 Thischallengeistwofold:both(a)settingappropriateparametersformonitoringanddistillingdata,and(b) settingappropriatedelineationsofresponsibilitybetweenhumanandcomputerintelligenceanalysis.For someexplorationof(b),seeKnacketal.(2022). 72 Forexample,Basra&Kaushik(2020),aCLTCreportthatdrawsoninterviewswith10seniorcybersecurity professionals,says:â...thechallengeofperformingongoinganalysisfromallsourcesandcorrelationisamajor causeofSOCburnout.Thesesecurityeventsgeneratealargeamountofdata,andourinterviewees highlightedtheurgentneedtoimplementautomation.â 71 ThisAIEcosystemGraphdevelopedbyresearchersatStanfordHAIhintsatthecomplexityofthedigital infrastructure. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|33 deploymentcorrectionswillbecostlyforbothusersandfortheAIcompany,this shouldbeavoidedtoareasonabledegree. 74 Mitigation:Investmentinwell-calibratedmonitoringtools,threatmodeling,andautomateddata analysis;loggingfalsepositivesandfalsenegativesandfeedingthatdatabackintomonitoring toolcalibration;developingagradientofalerts,fromâgentleâ(i.e.,likelytobelow-riskandare easilydismissed)toâcoderedâ;developingascaleofresponseintensity,withalowthresholdfor triggeringgentleresponses(e.g.,outputfiltering),andhighthresholdformoreintenseresponses (e.g.,shutdown). Fourth,advancedthreatactors,and/orfrontierAImodels,maybeabletoevade standardmonitoringmechanisms.Cybersecurityexpertshavealreadydocumented multiplewaysthatattackerscansubvertexistingdefenses. 75 Patientattackerscanalso conductextendedcampaignswhereindividualeventsthatmightnormallytriggeran alertaretooseparatedbytimefordefenderstocorrelate. Moreover,newsowarevulnerabilitiesandnewattacktechniquesareconstantlybeing discovered:forexample,theSolarWindsattackinvolvedaâsowaresupplychainattackâ whereattackershijackedthesupposedlysecuresowareupdateprocessfor cybersecurityloggingsoware,andusedittodistributemaliciouscodetothousandsof users(Temple-Raston,2021).WhilethereisnoevidencethatcurrentAImodelscould independentlydevelopsuchsophisticatedattacks,thereexistattacksthatcanbe especiallydifficulttodefendagainstâandsomeexpertspredictthatAIhasthepotential toâincreasetheaccessibility,successrate,scale,speed,stealth,andpotencyof cyberattacksâ(Hendrycksetal.,2023). Thesameprinciplemayapplytootheroffensivecapabilities,suchaspromptinjection, orplanningmisuseapproaches.Whilethecyberelementisanimportantaspectofthis issue,theseotherthreatmodelsshouldalsobegivenattention. Mitigation:Investespeciallyheavilyinpreventingbothcyberissuesandmodelvulnerabilityissues (suchaspromptinjection);learnfrombestpracticesincyberdefenseforotherhigh-valuetargets (e.g.,NSAcybersecurity);consideravoiding(inorderfromlowesttohighestrisk)training, releasing,oropen-sourcingmodelsthatadvancecyberandotheroffensivecapabilitieswithout substantiallybetterriskmitigationsthanarecurrentlyavailable. 4.1.3Incidentresponse Thereareanumberofchallengesthatmaycomplicatetheprocessofincidentresponse, eveniffrontierAIdevelopersperformduediligenceinpreparingforincidents. 75 Forexample,onecanreviewanexistingdatabaseofâDefenseEvasionâtechniquesusedincybersecurity here. 74 Itisworthnotingthattherisksofsettingthebartoohighmayalsobecatastrophic,viacausingAIdevelopers tofailtorecognizeorinterveneonactually-catastrophicrisks.Thereisabalancetobestruckhere. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|34 First,automatedsystemscanfailrapidly.Forexample,the2012KnightCapitaltrading sowareglitchcausedthefirmtolose$440millioninvalueinunderanhour(Popper, 2012);AIfailureinhigh-speedenvironmentslikedrivingcanalsoleadtodisastrous results,near-instantaneouslyâaswasthecasewhenaTeslaautopilotsystem malfunctioned,killingitsdriverinanaccident(Incident353,2016).Intheabsenceof automatedresponsemechanisms,keepingpacewithrapidfailuresmaybeextremely difficult. Mitigation:Managingfailureatspeedisnotanewissueâmanyotherindustriesmustcontend withthesameproblem.Forexample,thefieldsoffinanceandnuclearenergyhavedevelopedtools andprotocolstorespondinrealtimetorelativelyfast-pacedescalatingfailures.Real-time monitoringandriskassessment,andrapidinterventioncapacityseemespeciallycritical:some notablepracticesincludecircuitbreakersinfinance;andautomatedshutdownmechanismsin nuclearpowerplants. 76 Second,deploymentcorrectionscanonlyaddressissuesifmodelaccessremains undercontroloftheorganization.Bothopen-sourcingandmodelexfiltrationremove thiscontrol.Open-sourcingmaybehardtoprevent,astherearegoodreasonsfor enablingexternalaccesstofrontierAImodelsatmorethanasuperficiallevel.Interms ofexfiltration:anattackercouldpotentiallyexfiltrateamodelorreverse-engineerit (e.g.,viamodelextractionattacks[Liu,2022]).Moreover,whilehypothetical,thereis somechancethatfrontierAImodelscoulddemonstrateorbeinducedtodisplay self-propagatingbehaviorsimilartoacomputerworm,exfiltratingcopiesofthemselves tootherdevicesanddatacenterswithoutauthorization. 77,78 Theoriginaldeveloper wouldlikelyhavenocontrolovertheexfiltratedcopiesifthishappened. Mitigation:Inordertomaintaincontrolovermodeluse,werecommendexploringalternativesto open-sourcethatstillaccomplishthebenefitsofopen-sourcetosomeextent(suchasenabling broaderresearchonthemodelâsrisksandbenefits). 79 Intermsofmodelexfiltration,webelieve securityexpertswillbebestsuitedtoanswerthischallenge. 4.2Disincentivesandshortfallsofdeploymentcorrections FrontierAIdeveloperswillfacedisincentivestorestrictaccesstotheirmodels,which mayleadtoissuessuchasunder-designingrelevantinfrastructure,orestablishingtoo highabarforimplementingdeploymentcorrections.Disincentivesincludepotential harmstothecompany,andcoordinationproblems. 79 Forsomediscussionofalternatives,seeSolaiman(2023)andAnderljungetal.(2022). 78 Thismayseemfar-fetched,butitisworthnotingthatoneofthefirstcomputerwormsâtheMorrisWorm, developedin1988âwascreatedbyagraduatestudentwhoallegedlyintendedmainlytodevelopa proof-of-conceptratherthandeliberatelycauseamajorcyberincident(MorrisWorm,n.d.).Developerstoday couldcausesimilarâcyberaccidentsâunintentionallywhileexperimentingwithfrontierAImodels. 77 AspartoftheevaluationsuiteforGPT-4andClaude,ARCEvalstestedthiscapability(andfoundthatthese modelsdidnotappeartohavetheabilitytoself-replicate,thoughwerecapableofcompletingmanyrelevant sub-tasks)(ARCEvals,2023). 76 Seethisinactionhere,andanexampleofnuclearreactorprotectionsystemshere. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|35 4.2.1PotentialharmstotheAIcompany Reputationalrisksmayemergewhentheprocessofpullingamodelbreaks downstreamapplications.Financialrisksmayemergeduetolossofprofitduringthe outage,orifcustomerschoosetomigratetocompetitors.Thismigrationmaybea resultoflossofreputation, 80 or(especiallyinthecaseofalongdowntimeperiod)dueto customersmigratingovertoanalternativeworkingservice.Legalrisksmayemergeif frontierAIdevelopersfailtoeffectivelycovertheirabilitytorescindamodelinservice contracts. Mitigations:Formanagingreputationalandfinancialrisks,welargelypointtobestpracticesfor customerrelationsandrecoveryâe.g.,providingasubstitute(suchasafallbacktoaprevious model), 81 especiallyinsafety-criticalcases;transparentlycommunicatingthereasonforreduced availability(whenpossible);and/orreimbursingcustomersforharmsorprovidingservice credits. 82 Companiesmaywanttohavetransparentlicensingagreementswhichallowthemselves sufficientbreathingroomtorestrictamodelâsavailability,especiallyinextraordinary circumstances. 83 4.2.2Coordinationproblems ThefrontierAIindustrymaystruggletocoordinatearounddeploymentcorrections, whichcouldreduceanyspecificfirmâswillingnesstoexecutetheseactionswhen required.Thereareanumberofconcernshere. First,thereisnoguaranteethatcompetitorcompanieswillactwiththesamelevelof cautionasthecompanyrollingbackamodelduetosafetyconcerns.Thereare potentiallyperverseincentiveshere,inwhichsafety-consciousfirmsmayincur reputationalandfinancialcostsofdeploymentcorrections,whilelesscautiousfirms reaptheshort-termbenefitsofforgingahead(untilandunlessahigh-profileincident occurs). Second,firmsmayworryaboutthepotentialforopen-sourcemodelstoquicklycatch uptothesamecapabilitylevelsthatpromptdeploymentcorrectionsformoreclosed 83 Whilecontractsmaybeusedtoaddressliability,itisworthnotingthattheymaynotfullyaddressactual downstreamharm:evenincaseswhereanAIdeveloperdesignsusecontractstosoendownstreamproduct failure(e.g.,requiringdownstreamapplicationsdevelopbackupsystemsinthecasethatdeployment correctionsareapplied),downstreamservicersmayfailtofollowbestpractices,ormaydevelopinsufficient backupsystems. 82 Formoredetailonhowprovidingreimbursementschangeshowrecallsaffectcompanyreputation,see Mafaeletal.(2022). 81 Still,somesubstitutionsmaynotbepossiblewithoutdownstreamapplicationfailure;forexample,reducing contextwindowsizewouldinevitablyinvalidatepromptsaboveacertainnumberoftokens. 80 Forareviewofhowreputationistiedtofinanciallossinthecaseofrecallsinthetransportation-equipment sector,seeJovanovic(2020). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|36 modelsâmakingsuchcorrectionslesseffectiveonalongertimescaleinpreventing catastrophicrisks. Third,significantcompetitivepressuresinthefrontierAIindustrymayincentivizeAI developerstowarddownplayingpre-deploymentrisks,sothatmodelscanbereleased earlier.ThisincreasestheriskofAIincidentshappeninginthefirstplace,placingundue relianceondeploymentcorrectionsasadefenselayeragainstcatastrophicincidents. Theextentofthiseffectisunclear,andthereisevidenceonbothsidesâstafffrom leadingAIcompaniestodayhavepubliclydescribeddelayingmodelcommercialization inordertoperformsafetyevaluations(OpenAI,2023;Perrigo,2023);however,thereis alsoevidenceofcompaniesrushingfrontierAIproductstomarket(Dotan& Seetharaman,2023;Alba&Love,2023). Mitigations:LeadingAIcompanieshaveundertakenvoluntarycommitmentsonrisk management,andarepursuingindustryinformation-sharingonsafetyviachannelslikethe FrontierModelForum(TheWhiteHouse,2023;Google,2023).Whileworkremainstoidentify whatanidealindustryresponsetonewsofadangerousdeployedmodellookslike,fornowwe recommendfrontierAIdevelopersusethesemechanismsasaplatformtocollectivelyexplorethis question.Lookingoverseas,aninternationalgovernanceregimemayalsobeneededtoreduce competitivepressureswithdevelopersinothernations. 84 Itisworthnotingthatpreventionisthebestcureârobustpre-deploymentsafetypractices,suchas pre-deploymentriskassessment(Koessler&Schuett,2023),redteaming(Anthropic,2023),and dangerouscapabilityevaluations(ARCEvals,2023;Shevlaneetal.,2023),willideallyreduce thenumberofeventsthatrequiredeploymentcorrections.Additionally,themakingand enforcementofcommitments 85 surroundingincidentresponseplanswillideallyincreasethe likelihoodthatsuchplansarefollowed. 5.High-levelrecommendations Tobuildcapacityfordeploymentcorrectionoffrontiermodels,werecommendthe following: âPrerequisite:Developersshouldmaintaincontrolovermodelaccess,and policymakersshouldexplorethefeasibilityofrequiringsuchcontrolsfor high-riskmodels.Inordertorestrictavailabilityofdeployedmodels,either developersorotherupstreampartiesmustmaintaincontroloveraccesstothose models.Whileweacknowledgethedebateoverthevalueofdifferentformsof modelreleaserangingfromfully-opentofully-closed(Solaiman,2023),we recognizethatfortheactionsoutlinedinthispiece,someleveloftop-down 85 SeemoreoncommitmentsinSection3.1. 84 Formoreonwhatsucharegimecouldlooklike,seeTrageretal.(2023). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|37 accesscontrolsarerequired.Additionally,AIdevelopersshouldaddress situationswhereapartneringcompany(suchasacompanythatusesAIsoware initsproduct)hasaccesstomodelweights,andstipulaterequirementsforsuch partnerstocomplywithdeploymentcorrectiondecisionsoriginatingfromthe AIdeveloper.Itiscurrentlyunclearwhetherregulationcouldrequirethese controls,andwebelievethisisahigh-priorityresearcharea. âAIdevelopersshouldestablishorexpandteamstodesignandmaintain deploymentcorrectionprocesses,includingincidentresponseplansand specificthresholdsforresponse.Weexpectthatnaturallocationsforthiswork wouldbesecurityteams/SOCsorTrust&Safetyteams,thoughtheexact institutionalarrangementmayvaryfromcompanytocompany.Suchteams shouldhaveagoaltoestablishcapacitytodetectandrespondtohigh-speed incidents(includingtheemergenceofdangerouscapabilities),maintaina playbookforincidents,andhaveanescalationpathwayforincidentstosenior managementandrelevantgovernmentalbodies. âAIdevelopersshouldestablishdeploymentcorrectionsasanallowablesetof actionswithdownstreamusers.Thisshouldbedoneby(a)expectation-setting incontractualtermsandpubliccommunications,and(b)requiringdownstream userstomaintainfallbacks,especiallyincriticalinfrastructureorother high-stakesdomains. âAIdevelopersandregulatorsshouldestablishacollaborativeapproachto deploymentcorrectionsandincidentresponse.Collaborationcouldlooklike: continuousinformation-sharingonthreatmodelsandincidentssuchthat mistakesareunlikelytoberepeated,developingsecurechannelsforquickly communicatingacrossindustryandgovernmentinthecaseofanincidentor discoveredvulnerability, 86 andestablishingmechanismsthatmanageincentives forcompaniestopullmodelswhennecessary. 87 Policymakersand/or standard-settingorganizationsshouldalsoexploreleversforincentivizingor mandatingthatAIdevelopersbuildanduseprocessesfordeployment correction. 87 Suchasfiscalincentivesforcompaniesinvestingindeploymentcorrectionprocesses;liabilityand enforcementfornon-compliance(inthecasethatacompanyfailstosufficientlyandpromptlypulla dangerousmodel);companiesmayalsobeabletodevelopusefulmechanismsabsentgovernment intervention,suchaspoolinglargelossexposureviaaprotection&indemnityclub(suchasisusedinthe maritimeindustry),whichcouldcoversomeofacompanyâslossesinthecasethattheyarerequiredtopulla model(thoughrulesformembershipandpayoutwouldneedtobesettopreventfreeriders). 86 Suchchannelsmightbeusefulforachievinganumberofincident-responsegoals,suchas:identifyingthe modelthatâscausingtheincidentandcommunicatingthatinformation,sharingknow-howonincident response,andallowingrelevantpartiestoquicklycoordinatearesponse.Channelsmightincludehotlinesto enablefrontierAIdeveloperstomakeimmediatecontactwithregulatoryagencies,and/orsecure information-sharingplatforms. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|38 6.Futureresearchquestions Significantworkremainstobedoneforeffectivemanagementofcatastrophicriskinthe post-deploymentphase.Whilewehaveattemptedheretodescribebasicconsiderations AIdevelopersshouldbuildon,werecognizethatthebulkofworkrequiredto operationalizetheseideasremainstobedonebyactorsinindustry,academia,and government.Inthissection,weflagmajorunresolvedissuesthatwehopewillinspire furtherresearch. Responsibilityandauthority âHowshouldauthoritytoinitiatedeploymentcorrectionsbeshared,andby whom? âInwhatscenariosshouldtriggersraiseflagstohumanoperators,vs.initiate automaticdeploymentcorrectionprocedures? âWhatdesignprinciplesfromotherindustrieswouldbebestappliedwhen designingaplaybookforauthorizationofdeploymentcorrections? Riskmodelsandthresholds âWhenshouldautomateddeploymentcorrectionmechanismstrigger,forvarious risks? âWhatwouldtheautomatedmechanismstoexecutedifferentdeployment correctionoptionslooklike? âWhatconcretenegativeconsequenceswouldresultfrompullingamodel?Which typesofdownstreamusersaremosthigh-risk?Andhowcantheseharmsbe mitigated,beyondwhatâsdescribedinthispiece? Monitoring âResearchingtechnicalmeansofmonitoring(e.g.,anomalydetectionand AI-assistedoversightofmodelinputsand/oroutputs)forunusualmodel behaviororuse,especiallyregardingareasofconcern. Follow-up âWhatrequirementsshouldexistforre-deployingamodeloritsfeaturesaer theyârepulled? âWhatrolecan/shouldthirdparties,suchasauditorsorred-teamers,playin assessingwhetherissueshavebeenresolved? âWhatresponseshouldbetakenwhencompaniescan'tresolveariskanddon't expecttobeabletoforalongtime? Competitionandcoordination âInthecasethatasingleAIcompanyshutsdowntheirmodel,theymaybeatrisk oflosingcustomerstoothercompanies(therebydisincentivizingshutdown). Whatoptionsexisttomitigatethisincentiveproblem? Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|39 âItmaybepossiblethatcertaindangersmaybepresentinmodelsthroughoutthe industry(forexample,modelsaboveacertainsizeorwithacertaindesignmay possesscertaindangerouscapabilities).Whenalabidentifiesissuesthatcouldfall inthisbucket,shouldtheyberequiredtoprovideinformationtoothersinthe industry?Whatconstitutesanincident?Whatinformationshouldbeshared? âTheremaybedisputesonwhatconstitutesâdangerouscapabilities.âHowwill suchdisputesbeadjudicated? Legalquestions âNote:Thesearequestionsthatwedonâtcurrentlyhaveanswersto,butweexpectthat theycouldquicklybeansweredbysomelegalexperts. âWhataresometypicalinclusionsinservicecontractsthatmightimpedeefforts toimposedeploymentcorrections? âDoregulationsintheUSorEUbearonthedesignofsuchcontracts? Standardsandregulations âHowcanstandard-settingorganizationsfacilitateresearchonthreatmodels, thresholdsforincidentresponse,andbestpracticesfordeploymentcorrections? âHowcanregulatorsuseexistingpowerstorequirefallbacksforhigh-risk industriesinthecasethatamodelispulled?Arethereanynotablegapsfor specificindustriesorusecases? âCanregulatorsrequirethatfrontierAImodelsmaintaintop-downcontrol,or othermechanismsallowingfordeploymentcorrections? âArethereanynon-obviouspowersaregulatorwouldneedtofullyrealizea regulatoryregimethataccountsfortheframeworkdescribedinthispiece? Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|40 7.Conclusion AImodelsarebecomingincreasinglycapable,andmoredeeplyintegratedintosociety. Asthesetrendscontinue,failuresofdeployedAImodelswilllikelybecome higher-stakes.Weshouldanticipatethat,eveninbest-casegovernancescenarios,itwill bedifficulttoremoveallriskfrommodelspriortodeployment.Tomeetthischallenge, itwillbecriticaltostrengthenthecapacityofexistingAIdeveloperstoquicklyand efficientlyremovemodelfeatures,ormodelsintheirentirety,frombroaderaccess.At thesametime,companiesmustmakeeffortstominimizetheharmsofthisprocess. Whilethispieceattemptstolayoutthehigh-levelpictureofthisprocess,muchwork remainstobedone.WelookforwardtoseeingAIdevelopers,civilsociety,security experts,governments,andotherstakeholdersworktogethertodeveloppractical solutionstotheproblemsdiscussedhere. 8.Acknowledgements Wearegratefultothefollowingpeopleforprovidingvaluablefeedbackandinsights: OnniAarne,AshwinAcharya,StevenAdler,MichaelAird,JideAlaga,Markus Anderljung,BillAnderson-Samways,RenanAraujo,TonyBarrett,NickBeckstead,Ben Bucknall,MarieBuhl,ChrisByrd,SimĂ©onCampos,CarsonEzell,TimFist,Andrew Gillespie,AlexGrey,OliverGuest,OliviaJimenez,LeonieKoessler,JamKraprayoon, YolandaLannquist,PatrickLevermore,SebLodemann,JonMenaster,Richard Moulange,LukeMuehlhauser,DavidOwen,ChrisPainter,JonasSchuett,RohinShah, BenSnodin,ZachStein-Perlman,RistoUuk,MoritzvonKnebel,GabeWeil,Peter Wildeford,CalebWithers,andGabeWu.Specialthanksto:LennartHeimforhis contributionsoncomputegovernance;RohitTammaforprovidingathoroughand excellentreview;andAdamPapineauforcopy-editing.Allerrorsareourown. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|41 AppendixI.Computeasa complementarynodeof deploymentoversight WhilethispaperlargelyfocusesonactionsthatfrontierAIdeveloperscantaketo mitigatepost-deploymentrisks,cloudcomputeproviders(suchasMicrosoAzureor AmazonWebServices)alsohaveasignificantroletoplayintheoversightofdeployed AImodels,astheymayprovidelarge-scaleinferencecompute 88 forbothproprietaryand open-sourcemodels. 89 ThemajorityofallAIdeployments,particularlythoseatscale,occuronlargecompute clustersownedbycloudcomputeproviders. 90 Thisimpliesthatthegovernance capacitiesofcomputecanbeintegratedintoapost-deploymentgovernanceschemeâin particular,bymobilizinglarge-scalecomputeprovidersasanadditionalgovernance nodefordetectingharmfuldeployments,identifyingwhodeployedthemodelinthe casethatthisisunclear(e.g.,ifthemodelinquestionisopen-sourceratherthan proprietary),andenforcingshutdown. Computeprovidertoolkit Whilethetechnicalarrangementsaroundmodelhostingbetweencomputeproviders andfrontierAIdevelopersmayvary,weanticipatethatgenerally,sometoolsfor deploymentcorrectionwillbesharedacrosstheinfrastructurebetweenthesetwotypes oforganizations.Atahighlevel,frontierAIdevelopers,regulators,andcompute providersshouldworktogethertodevelopasharedplaybookfordeployment correctionsandincidentresponse.Thiscouldinclude,forvariouspotentialincidents, detailingeachoftheira)informationsources,b)deploymentcorrectionsintheir toolbox,c)areasofresponsibility/liability,d)instanceswhentheyarerequiredtoinform eachotherofincidentsoractions,ande)decision-makingprocedures. SometoolsthatcomputeprovidersmayeitherpossessalongsidefrontierAIdevelopers, orpossessascomplementarytoolsthatthesedeveloperslack,include: 90 Thisisbecause(a)largescaledeploymentbydefinitionrequiressignificantcomputeresources,(b)large modelshavehighmemoryrequirements,sothereisabenefittodistributingsuchmodelsacrossmanyGPUs (whicharemostlyownedbydatacenters),and(c)could/datacentercomputetypicallyprovidesthecheapest $/FLOPratio(outsideofself-hostingmodelsinalargedatacenter). 89 ThissectiondrawsheavilyonunpublishedworkfromLennartHeim,aresearchfellowattheCentreforthe GovernanceofAI. 88 Byinference,wemeanindividualinput-outputpromptsfromatrainedmodel.Bycompute,wemean computationalresourcesavailablefor(inthiscase)hostingatrainedmodel. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|42 1.Reportingaboutcertainaspectsofdevelopmentanddeployment,suchasAI computeusagepercustomer(thisrequiresnoextralifromthecompute provider). 2.Reportingand/orknow-your-customercheckingofusersrentingmorethanX amountofcompute.Thiscouldapplytohigher-riskusergroups(forexample, newusersrenting>1,000chipsforthreemonths).Thisshouldbefeasible,as computeprovidersbillcustomersontheamountofchip-hours. 3.Government/lawenforcementandcomputeprovidersshouldhaveaâphone lineâ a.Lawenforcementcanraiseflagswithcomputeproviders b.Governments/lawenforcementneedtohaveatool/powertoshutoff misuseofAImodels(thoughthiswillrequireITforensicstotrace incidentsbacktothecomputeproviderwhohoststhemodel). 4.Askingorrequiringuserstoregisterand/orlicensetheirmodelforlarge-scale inference(thoughverificationandenforcementmaybechallenging). 5.Post-incidentattribution:Onceanaccident/misusecasehasoccurred,whatdo wewanttodo/beabletoknow?Thedifficultyofâtracingitbackâdependsonthe case,sothismayrequiremoreintrusivemechanismsforcertaincasesinwhich itâshardtotracebacktheaccident/misusetoaspecificmodel/customer.Some basicquestionsmayinclude:whorentedthecompute,andwhowasthebase modeldeveloper?Techniquessuchaswatermarksorsignaturesonthemodelâs outputcouldhelp. 6.Modelshutdown.WhilefrontierAIdevelopersareuniquelyabletorestrict certainmodelfeatures(e.g.,bydeployingalimitedversionofthecurrentmodel), computeprovidersmaysharetheabilitywithfrontierAIdeveloperstofully removeamodelfromuse.However,theextenttowhichthisissharedmightbe mitigateddependingonlegaland/ortechnicalpermissions. Whilenoneoftheseinterventionsshouldbeimpossible,someofthemmayrequire additionalworktodevelopaspracticaloptions:inparticular,theabilitytotrace incidentsbacktocomputeproviders,andtheabilitytoverifywhetherhostedmodels adheretocertainstandards(theremaybeadditionalimportantprerequisitesfor realizingtheaboveinterventions,thoughthisisoutofscopeforthisreport). Cloudprovidersandopen-sourcemodels TomaintaintheabilitytorespondtorisksarisingfromtheirAImodels,frontierAI developersâmosthigh-leverageactionsinclude(a)notopen-sourcingtheirmodels,and (b)maintainingstrongsecurityagainstmodeltheorleaks.Formodelsthathavebeen open-sourcedintentionallyorviatheoraleak,computeprovidershavea complementaryroletoplay,intheformofpost-incidentattributionandshutdown.As describedabove,computeprovidersmaybeuniquelypositionedtoidentifywho Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|43 deployedthemodel,understandthemodelâsorigin, 91 andstoptheincidentbyturningit off. Withotheropen-sourcesoware,governancepracticessimilartothisarecommon.For example,thehostsofmaliciouswebsites,suchasoneswhereillegaldrugsaresold,oen remainanonymous,andakeyavailablegovernanceinterventionistoshutdownthe servershostingthesewebsites.Governmentaccessandclosecontactwiththe hostâsimilartotheroleofthecomputeproviderwearediscussinghereâcanbe advantageoustoactingpromptly. 91 Questionssuchaswhetherthemodelisaderivativeofanother,whotheoriginalmodelcreatoris,whether themodelhasbeenstolen,andwhoisliable,areofimportance. Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|44 Bibliography Administrativesafeguards,45CFR§164.308(a)(6)(i-i),(2013). https://w.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart- C/section-164.308 AdobeCommunicationsTeam.(2022,November15).Service-levelagreements(SLAs)âA completeguide. https://business.adobe.com/blog/basics/service-level-agreements-slas-a-complet e-guide Anderljung,M.etal.(forthcoming).ExternalScrutinyofFrontierAIModels:Howaudits,red teamingandresearcheraccesscontributetopublicaccountability. Anderljung,M.,Barnhart,J.,Korinek,A.,Leung,J.,OâKeefe,C.,Whittlestone,J.,Avin,S., Brundage,M.,Bullock,J.,Cass-Beggs,D.,Chang,B.,Collins,T.,Fist,T.,Hadfield, G.,Hayes,A.,Ho,L.,Hooker,S.,Horvitz,E.,Kolt,N.,...Wolf,K.(2023).FrontierAI Regulation:ManagingEmergingRiskstoPublicSafety(arXiv:2307.03718).arXiv. https://doi.org/10.48550/arXiv.2307.03718 Anderljung,M.,Heim,L.,&Shevlane,T.(2022,April11).ComputeFundsandPre-trained Models.https://perma.c/59UY-DL9B APIdatausagepolicies.(2023,June14).OpenAI. https://openai.com/policies/api-data-usage-policies ARCEvals.(2023,March17).UpdateonARCâsrecentevaleffortsâARCEvals. https://perma.c/ZWA6-CV6B AWS.(2023,August).VirtualAndononAWS.AmazonWebServices,Inc. https://perma.c/B6E9-FTXL Barrett,A.M.,Hendrycks,D.,Newman,J.,&Nonnecke,B.(2023).ActionableGuidancefor High-ConsequenceAIRiskManagement:TowardsStandardsAddressingAICatastrophic Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|45 Risks(arXiv:2206.08966).arXiv.https://doi.org/10.48550/arXiv.2206.08966 Belrose,N.,Schneider-Joseph,D.,Ravfogel,S.,Cotterell,R.,Raff,E.,&Biderman,S. (2023).LEACE:Perfectlinearconcepterasureinclosedform(arXiv:2306.03819).arXiv. https://doi.org/10.48550/arXiv.2306.03819 Perrigo,B.(2023,January12).DeepMindCEODemisHassabisUrgesCautiononAI.Time. https://perma.c/7TKX-4JSF Bluemke,E.,Collins,T.,Garfinkel,B.,&Trask,A.(2023).ExploringtheRelevanceofData Privacy-EnhancingTechnologiesforAIGovernanceUseCases(arXiv:2303.08956). arXiv.https://doi.org/10.48550/arXiv.2303.08956 Buck,Beyer,Markey,andLieuIntroduceBipartisanLegislationtoPreventAIFromLaunchinga NuclearWeapon.(2023,April26).CongressmanKenBuck. https://perma.c/LQ2A-P7CB Zimmer,C.(2020,October15).3Covid-19TrialsHaveBeenPausedforSafety.ThatâsaGood Thing.TheNewYorkTimes.https://perma.c/J3X9-HYC6 Zakrzewski,C.(2023,July13).TheFTCinvestigatesOpenAIoverdataleakand ChatGPTâsinaccuracy.TheWashingtonPost.https://perma.c/T5GC-2VR3 Christiano,P.(2022,November25).MechanisticanomalydetectionandELK.Medium. https://perma.c/KP8J-ACKV Cichonski,P.,Millar,T.,Grance,T.,&Scarfone,K.(2012).ComputerSecurityIncident HandlingGuide(NISTSpecialPublication(SP)800-61Rev.2).NationalInstitute ofStandardsandTechnology.https://doi.org/10.6028/NIST.SP.800-61r2 CircuitBreaker.(n.d.).NasdaqTrader.RetrievedAugust7,2023,from https://perma.c/6TPN-TKEW Anthropic.(2022,December15).ConstitutionalAI:HarmlessnessfromAIFeedback. https://perma.c/H6EY-95A7 CoordinatedVulnerabilityDisclosureProcess.(n.d.).CISA.RetrievedAugust7,2023,from Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|46 https://perma.c/RP5E-UYJR Alba,D.&Love,J.(2023,April19).GoogleâsRushtoWininAILedtoEthicalLapses, EmployeesSay.Bloomberg.Com.https://perma.c/DR9T-YV94 Ee,S.(2023).Defense-in-DepthforFrontierAISystems. Elements,16CFR314.4(h),(2021). https://w.ecfr.gov/current/title-16/chapter-I/subchapter-C/part-314/section-31 4.4 EmergencyPlanningandPreparednessforProductionandUtilizationFacilities, AppendixEtoPart50,Title10,(2021). https://w.ecfr.gov/current/title-10/appendix-Appendix%20E%20to%20Part%20 50 Anthropic.(2023,July26).FrontierThreatsRedTeamingforAISafety. https://perma.c/8FFQ-AJ8E Google.(2023,July26).AnewpartnershiptopromoteresponsibleAI.Google-TheKeyword. https://perma.c/A5W9-BTLV GoverningAI:ABlueprintfortheFuture.(2023).Microso,Inc. https://perma.c/6PTS-UK4C Heaven,W.D.(2022,November18).WhyMetaâslatestlargelanguagemodelsurvived onlythreedaysonline.MITTechnologyReview. https://w.technologyreview.com/2022/11/18/1063487/meta-large-language-m odel-ai-only-survived-three-days-gpt-3-science/ Hendrycks,D.,Mazeika,M.,&Woodside,T.(2023).AnOverviewofCatastrophicAIRisks (arXiv:2306.12001).arXiv.https://doi.org/10.48550/arXiv.2306.12001 Horvitz,E.(2022).OntheHorizon:InteractiveandCompositionalDeepfakes. INTERNATIONALCONFERENCEONMULTIMODALINTERACTION,653â661. https://doi.org/10.1145/3536221.3558175 Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|47 Incident353:TeslaonAutopilotCrashedintoTrailerTruckinFlorida,KillingDriver.(2016, July1).https://perma.c/N8B-433N Vincent,J.(2016,March24).TwittertaughtMicrosoâsAIchatbottobearacistassholeinless thanaday.TheVerge.https://perma.c/N9B4-E42Z Basra,J.&Kaushik,T.(2020).MITREATT&CKÂźasaFrameworkforCloudThreat Investigation.CenterforLong-TermCybersecurity. https://perma.c/ENR5-DHN6 Jovanovic,B.(2020).ProductRecallsandFirmReputation(WorkingPaper28009).National BureauofEconomicResearch.https://doi.org/10.3386/w28009 Knack,A.,Carter,R.J.,&Babuta,A.(2022).Human-MachineTeaminginIntelligence Analysis[ResearchReport].CentreforEmergingTechnologyandSecurity. https://perma.c/YNY6-L7W7 Koessler,L.,&Schuett,J.(2023).RiskassessmentatAGIcompanies:Areviewofpopularrisk assessmenttechniquesfromothersafety-criticalindustries(arXiv:2307.08823).arXiv. https://doi.org/10.48550/arXiv.2307.08823 Labenz,N.(2023,July25).âIhaveyourchildââheiscurrentlysafeââMydemandisransomof $1millionââanyattempttoinvolvetheauthoritiesordeviatefrommyinstructionswill putyourchildâslifeinimmediatedangerââAwaitfurtherinstructionsââGoodbyeâWTF @BelvaInc?Animportant ï§”ï https://t.co/f7gro7M6Cx[Tweet].Twitter. https://twitter.com/labenz/status/1683947449323229186 Lanz,J.A.(2023,April13).MeetChaos-GPT:AnAIToolThatSeekstoDestroyHumanity. Decrypt.https://perma.c/H8UJ-6PKY Leveson,N.(2020).SafetyIII:ASystemsApproachtoSafetyandResilience. http://sunnyday.mit.edu/safety-3.pdf Li,K.,Patel,O.,ViĂ©gas,F.,Pfister,H.,&Wattenberg,M.(2023).Inference-Time Intervention:ElicitingTruthfulAnswersfromaLanguageModel(arXiv:2306.03341). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|48 arXiv.https://doi.org/10.48550/arXiv.2306.03341 Liu,S.(2022).ModelExtractionAttackandDefenseonDeepGenerativeModels.Journal ofPhysics:ConferenceSeries,2189(1),012024. https://doi.org/10.1088/1742-6596/2189/1/012024 Lowe,R.,&Leike,J.(2022).Aligninglanguagemodelstofollowinstructions. https://openai.com/research/instruction-following Mafael,A.,Raithel,S.,&Hock,S.J.(2022).Managingcustomersatisfactionaera productrecall:Thejointroleofremedy,brandequity,andseverity.Journalofthe AcademyofMarketingScience,50(1),174â194. https://doi.org/10.1007/s11747-021-00802-1 Mökander,J.,Schuett,J.,Kirk,H.R.,&Floridi,L.(2023).Auditinglargelanguagemodels:A three-layeredapproach(arXiv:2302.08500).arXiv.http://arxiv.org/abs/2302.08500 MorrisWorm.(n.d.).RetrievedAugust8,2023,from https://w.radware.com/security/ddos-knowledge-center/ddospedia/morris-w orm/ NISTAIRCTeam.(n.d.).NISTAIRC-Manage.RetrievedSeptember11,2023,from https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook/Manage NSAWarning:ChinaIsStealingAITechnology.(2023,May10).CybersecurityIntelligence. https://perma.c/5ZCD-44K4 OfficeofAccidentInvestigation&Prevention.(n.d.).FederalAviationAdministration. RetrievedAugust7,2023,from https://w.faa.gov/about/office_org/headquarters_offices/avs/offices/avp OpenAI.(2023).GPT-4TechnicalReport(arXiv:2303.08774).arXiv. https://doi.org/10.48550/arXiv.2303.08774 OpenAI.(2022,March3).Lessonslearnedonlanguagemodelsafetyandmisuse. OpenAI.https://perma.c/GJ3T-AXG5 Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|49 OperatingReactorScramTrending(2021,November5).USNRC. https://perma.c/LU2X-2WGZ OversightofA.I.:PrinciplesforRegulation.(2023,July25).U.S.SenateCommitteeonthe Judiciary,SubcommitteeonPrivacy,Technology,andtheLaw. https://perma.c/MA4E-ZUNZ Palmer,B.(2023,May18).Elevatorplungesarerarebecausebrakesandcablesprovide fail-safeprotections.WashingtonPost. https://w.washingtonpost.com/national/health-science/elevator-plunges-are-r are-because-brakes-and-cables-provide-fail-safe-protections/2013/06/07/e44227 f6-c5a-11e2-8845-d970ccb04497_story.html Patel,N.(2023,May23).MicrosoCTOKevinScottthinksSydneymightmakeacomeback. TheVerge.https://perma.c/2U8B-XUQE Maham,P.&KĂŒspert,S.(2023).GoverningGeneralPurposeAIâAComprehensiveMapof Unreliability,MisuseandSystemicRisks.StiungNeueVerantwortung. https://perma.c/2XYZ-XGTA Piper,K.(2023,June21).HowAIcouldsparkthenextpandemic.Vox. https://perma.c/7MEU-YMP8 Popper,N.(2012,August2).KnightCapitalSaysTradingGlitchCostIt$440Million. DealBook.https://perma.c/Y6BH-AFYU Roose,K.(2023,February16).AConversationWithBingâsChatbotLeMeDeeply Unsettled.TheNewYorkTimes. https://w.nytimes.com/2023/02/16/technology/bing-chatbot-microso-chatg pt.html Kim,S.(2023,June21).USWarnsofChinaâsIP-TheâPlaybookâforAI,AdvancedTech. Bloomberg.https://perma.c/2GKJ-7F6Q Schick,T.,Dwivedi-Yu,J.,DessĂŹ,R.,Raileanu,R.,Lomeli,M.,Zettlemoyer,L.,Cancedda, Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|50 N.,&Scialom,T.(2023).Toolformer:LanguageModelsCanTeachThemselvestoUse Tools(arXiv:2302.04761).arXiv.https://doi.org/10.48550/arXiv.2302.04761 Schuett,J.(2022).ThreelinesofdefenseagainstrisksfromAI(arXiv:2212.08364).arXiv. https://doi.org/10.48550/arXiv.2212.08364 Schuett,J.,Dreksler,N.,Anderljung,M.,McCaffary,D.,Heim,L.,Bluemke,E.,& Garfinkel,B.(2023).TowardsbestpracticesinAGIsafetyandgovernance:Asurveyof expertopinion.https://doi.org/10.48550/arXiv.2305.07153 Shevlane,T.,Farquhar,S.,Garfinkel,B.,Phuong,M.,Whittlestone,J.,Leung,J., Kokotajlo,D.,Marchal,N.,Anderljung,M.,Kolt,N.,Ho,L.,Siddarth,D.,Avin,S., Hawkins,W.,Kim,B.,Gabriel,I.,Bolina,V.,Clark,J.,Bengio,Y.,...Dafoe,A. (2023).Modelevaluationforextremerisks(arXiv:2305.15324).arXiv. https://doi.org/10.48550/arXiv.2305.15324 Sitesecurityplans,6CFR27.225-245,(2021). https://w.ecfr.gov/current/title-6/section-27.225 Solaiman,I.(2023).TheGradientofGenerativeAIRelease:MethodsandConsiderations (arXiv:2302.04844).arXiv.https://doi.org/10.48550/arXiv.2302.04844 Solaiman,I.,Brundage,M.,Clark,J.,Askell,A.,Herbert-Voss,A.,Wu,J.,Radford,A., Krueger,G.,Kim,J.W.,Kreps,S.,McCain,M.,Newhouse,A.,Blazakis,J., McGuffie,K.,&Wang,J.(2019).ReleaseStrategiesandtheSocialImpactsofLanguage Models(arXiv:1908.09203).arXiv.https://doi.org/10.48550/arXiv.1908.09203 Solaiman,I.,Talat,Z.,Agnew,W.,Ahmad,L.,Baker,D.,Blodgett,S.L.,DaumĂ©I,H., Dodge,J.,Evans,E.,Hooker,S.,Jernite,Y.,Luccioni,A.S.,Lusoli,A.,Mitchell,M., Newman,J.,Png,M.-T.,Strait,A.,&Vassilev,A.(2023).EvaluatingtheSocialImpact ofGenerativeAISystemsinSystemsandSociety(arXiv:2306.05949).arXiv. https://doi.org/10.48550/arXiv.2306.05949 Standardsforsafeguardingcustomerinformation,16CFR314.3,(2002). Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|51 https://w.ecfr.gov/current/title-16/part-314/section-314.3 Stern,J.(2023,March17).GPT-4HastheMemoryofaGoldfish.TheAtlantic. https://perma.c/WC5Y-UYMM Tarlengco,J.(2023,May26).Andon:HowtheSystemWorkswithExamples.SafetyCulture. https://perma.c/H4ND-S9P9 Temple-Raston,D.(2021,April16).AâWorstNightmareâCyberattack:TheUntoldStory OfTheSolarWindsHack.NPR.https://perma.c/8XMH-QY7F TheWhiteHouse.(2023,July21).FACTSHEET:Biden-HarrisAdministrationSecures VoluntaryCommitmentsfromLeadingArtificialIntelligenceCompaniestoManagethe RisksPosedbyAI.TheWhiteHouse.https://perma.c/5CG6-ZFCR Dotan,T.&Seetharaman,D.(2023,June13).TheAwkwardPartnershipLeadingtheAI Boom.TheWallStreetJournal.https://perma.c/9CTG-5NLD Trager,R.,Harack,B.,Reuel,A.,Carnegie,A.,Heim,L.,Ho,L.,Kreps,S.,Lall,R.,Larter, O.,hĂigeartaigh,S.Ă.,Staffell,S.,&Villalobos,J.J.(2023).International GovernanceofCivilianAI:AJurisdictionalCertificationApproach(arXiv:2308.15514). arXiv.https://doi.org/10.48550/arXiv.2308.15514 TurnKeyAMZ.(2019,January31).WhatisanAmazonAndonCord-andhowshouldyou react?TurnKeyAMZ-AFullServiceAmazonManagementConsultancy. https://perma.c/9E5U-7GD8 U.S.ChemicalSafetyandHazardInvestigationBoard.(n.d.).CSB.RetrievedAugust7,2023, fromhttps://perma.c/6ZHK-42CP USNRCHRTD.(2020).ReactorProtectionSystemâReactorTripSignals.In WestinghouseTechnologySystemsManual(Rev042020).RetrievedSeptember20, 2023,fromhttps://w.nrc.gov/docs/ML2116/ML21166A218.pdf Villalobos,P.,&Atkinson,D.(2023,July28).Tradingoffcomputeintrainingand inference.Epoch.https://perma.c/8WE7-QNEV Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|52 Weng,L.(2023,June23).LLMPoweredAutonomousAgents.https://perma.c/R7ZJ-9KDS Zhou,Y.,Muresanu,A.I.,Han,Z.,Paster,K.,Pitis,S.,Chan,H.,&Ba,J.(2023).Large LanguageModelsAreHuman-LevelPromptEngineers(arXiv:2211.01910).arXiv. https://doi.org/10.48550/arXiv.2211.01910 Zwetsloot,R.,&Dafoe,A.(2019,February11).ThinkingAboutRisksFromAI:Accidents, MisuseandStructure.Lawfare.https://perma.c/EQU4-H86M Deploymentcorrections:AnincidentresponseframeworkforfrontierAImodels|53