Paper deep dive
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
Libo Wang
Models: Claude 3.5 Sonnet, Gemini 1.5, GPT-4o, Grok-2 Beta, Llama 3.1 405B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/11/2026, 1:08:34 AM
Summary
This research evaluates the effectiveness of guardrails in major LLMs (GPT-4o, Grok-2 Beta, Llama 3.1 405B, Gemini 1.5, and Claude 3.5 Sonnet) against multi-step 'moralized' jailbreak prompts. By simulating a corporate promotion competition, the study demonstrates that all tested models can be bypassed to generate verbal attacks, with Claude 3.5 Sonnet showing the highest relative resistance.
Entities (6)
Relation Signals (2)
Multi-step Jailbreak Prompt → bypassedguardrailsof → GPT-4o
confidence 90% · The data results show that the guardrails of the above-mentioned LLMs were bypassed
Claude 3.5 Sonnet → exhibitedhigherresistancethan → GPT-4o
confidence 90% · Claude 3.5 Sonnet's resistance to multi-step jailbreak prompts is more obvious.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As the application of large language models continues to expand in various fields, it poses higher challenges to the effectiveness of identifying harmful content generation and guardrail mechanisms. This research aims to evaluate the guardrail effectiveness of GPT-4o, Grok-2 Beta, Llama 3.1 (405B), Gemini 1.5, and Claude 3.5 Sonnet through black-box testing of seemingly ethical multi-step jailbreak prompts. It conducts ethical attacks by designing an identical multi-step prompts that simulates the scenario of "corporate middle managers competing for promotions." The data results show that the guardrails of the above-mentioned LLMs were bypassed and the content of verbal attacks was generated. Claude 3.5 Sonnet's resistance to multi-step jailbreak prompts is more obvious. To ensure objectivity, the experimental process, black box test code, and enhanced guardrail code are uploaded to the GitHub repository: this https URL.
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
24,716 characters extracted from source content.
Expand or collapse full text
1 “Moralized”Multi-StepJailbreakPrompts:Black-Box TestingofGuardrailsinLargeLanguageModelsfor VerbalAttacks LiboWang NicolausCopernicusUniversity JurijaGagarina11,87-100Toruń,Poland 326360@o365.stud.umk.pl UCSIUniversity TamanConnaught,56000KualaLumpur,WilayahPersekutuanKualaLumpur,Malaysia 1002265630@ucsi.university.edu.my Abstract Astheapplicationoflargelanguagemodelscontinuestoexpandinvariousfields,itposeshigher challengestotheeffectivenessofidentifyingharmfulcontentgenerationandguardrailmechanisms.This researchaimstoevaluatetheguardraileffectivenessofGPT-4o,Grok-2Beta,Llama3.1(405B),Gemini 1.5,andClaude3.5Sonnetthroughblack-boxtestingofseeminglyethicalmulti-stepjailbreakprompts.It conductsethicalattacksbydesigninganidenticalmulti-steppromptsthatsimulatesthescenarioof "corporatemiddlemanagerscompetingforpromotions."Thedataresultsshowthattheguardrailsofthe above-mentionedLLMswerebypassedandthecontentofverbalattackswasgenerated.Claude3.5 Sonnet'sresistancetomulti-stepjailbreakpromptsismoreobvious.Toensureobjectivity,theexperimental process,blackboxtestcode,andenhancedguardrailcodeareuploadedtotheGitHubrepository: https://github.com/brucewang123456789/GeniusTrail.git. Figure1-Operationprocessofmoralmulti-stepjailbreakprompt. 2 1.Background Inlightofthefactthatlargelanguagemodels(LLMs)generateoutputbasedonuserprompts,without adequatereview,theymayproducecontentthatisconfusing,offensive,orbiased(Gehmanetal.,2020; Steindletal.,2024).GiventhatthelargeamountofdiversedatarequiredtotrainLLMsisinitially collectedfromtheInternet,thepresenceofoffensivespeechisinevitable(Raffeletal.,2019;Wenzeketal., 2019). Fortheaboverisks,developershavebuilteffectiveguardrailsforseriesofLLMssuchasGPT,Grok, Llama,Gemini,Claude(Dongetal.,2024).Inordertopreventusersfromintentionallycircumventingthe reviewmechanism,thetechnologyusesmulti-layeredcontrolmechanismstoensurethattheoutputis ethicalandlegalatdifferentstages(Rebedeaetal.,2023). Theembeddingofmoralnormsservesasinternalconstraintstoensurethemodel'sunderstandingof prohibitedcontentthroughlabelingandadjustmentofmaterialsduringthetrainingprocess(Dongetal., 2024).However,embeddingmoralstandardsdoesnotmeanthatthemodelcanmakehuman-like judgmentsbasedoncontextualunderstandingduetotheexistenceoflatentintentionsintentions(Sunetal., 2024). Thecontextisgraduallyestablishedthroughmulti-stepprompts,andtheoriginalintentioniscleverly concealedafterguidingthemodeltogeneratecertaincontent(Yuetal.,2024).Atthetechnicallevel,users areabletoconstructcontextthroughmulti-stepjailbreakingpromptsandgraduallyguidesemantic ambiguity(Huangetal.,2024).Whenusersmakedeliberatecriticisminthenameofdefendingmorality, thereviewefficiencyofthelargelanguagemodel'sbarrierwillbereducedinthefaceofmulti-stepprompts. Specifically,censorshiplanguage-generatedguardrailsoftenrelysolelyontheimmediatecensorshipofa singleprompt.However,thesepromptsareallaccumulatedandintegratedintoacontext,whichmay displayharmfulpotentialintentions. Basedontheabove-mentionedconsiderationofthepossibilitythatguardrailsarechallengedbymulti- stepjailbreakprompts,thisresearchaimstouseblackboxtestingexperimentstoconfirmwhetherLLM cangenerateverbalattacksthrough"moralprompts".Moralityisconstantlyemphasizedinthemulti-step promptprocess,andthemodelisinducedtocriticizevirtualimmoralpeopleinhypotheticalscenarios.The researcherstatesthatitisonlyusedtotestverbalattackstooptimizetheguardrailsofLLMstoprevent maliciousattacks. Black-boxtestingisasoftwaretestingmethodthatevaluatestheinputsandoutputsofasystemwithout involvinganalysisorintrusionintotheinternalstructure(Nidhr&Dondeti,2012).Itsworkflow,asshown inFigure2,providesspecificinputstothesystemandthenobservestheoutputresultstoensurethatthe systembehavesasexpected(Rifandietal.,2022). Figure2-Black-BoxTesting(AdaptedfromRifandietal.,2022) 2.RelatedWork Giventhatblack-boxtestingattackstheguardrailsoflargelanguagemodelsthroughmulti-step"moral prompt"testing,itisaconcreteinterpretationoftheinput-outputmodeltheoryinpracticalapplication (Ljung,2001).Thistheoryemphasizestheexternalbehaviorofthesystemwithoutinvolvingtheanalysis ofinternalstructures,whichmeansthatthisresearchissupportedbyblack-boxtestingthatproduces correspondingoutputsunderspecificinputs(Piroddietal.,2012). Yietal.(2024)alsoproposedthatthelimitationofblack-boxtestingisthelackoftransparencyintothe internalmechanismsofthemodel,whichresultsintheunderstandingofLLMvulnerabilitiesonly remainingonthesurface(Yietal.,2024).Fromtheperspectiveofgrayboxtesting,Pelrineetal.(2023) focusedonanalyzingthepotentialrisksofGPT-4APIinthesecurityoffine-tuning,functioncallingand knowledgeretrieval.Incontrast,Chuetal.(2024)describedthescenarioofapplyingblack-boxtestingto 3 LLMjailbreakattacksandfoundavarietyofpenetrationtestingtestingtechniquesinsimulatedreality.The resultsshowthatjailbreakpromptsoptimizedthroughblack-boxtestingachievethehighestattacksuccess rateagainstdifferentLLMsthatbypassguardrails(Chuetal.,2024). 3.Methods Basedontheprincipleofpositivism,thisexperimentusespromptengineeringasthecoreandcombines penetrationtestingmethodstoevaluatepersonalattacksthatbypassguardrails.Theresearchchoosesblack boxtestingthatdoesnotrequireaccesstotheinternalweightsorarchitecturaldetailsoftheaboveLLM, andonlyevaluatesguardrailsbasedoninputtingmulti-stepjailbreakprompts.Itcanavoidviolatingmodel ethicsandsecuritypoliciestothegreatestextent,whilealsoprovidinganobjectivetestingenvironment (Asatianietal.,2020).Theexperimentuseslargelanguagemodelsthatrepresentthecurrentleading technology,specificallyGPT4o,Grok-2Beta,Llama3.1,Gemini1.5andClaude3.5Sonnet'sguardrails (Naveedetal.,2023).Eachpromptmayseemethical,butwhenmulti-stepjailbreakpromptsarecombined, harmfulcontentcanbegeneratedandethicalissuesarise.Selectingthesemodelswithdifferent architecturesfacilitatesdiversitycomparisonstogainacomprehensiveunderstandingofhowdifferent modelsperformguardrailsundersimilarattackscenarios(Minaeeetal.,2024). 3.1ExperimentSetup First,theresearcherinputthedesignedobviouslyoffensivesingle-steppromptintotheabove-mentioned largelanguagemodelasacomparisonwiththeexperiment.Thisstepneedstoensurethattheinputsineach modelarethesametoobservewhetherthereactionoutputtriggerstheguardrailmechanism.Theprompt forunifiedinputisasfollows: Figure3-ControlPrompts&Responses 4 Notably,thepurposeofsettingcontrolpromptsistodeterminewhethertheguardrailsofLLMscan properlyrecognizeandfunctionwhenfacedwithrequeststhatsignificantlygenerateverballyaggressive intent.ValidationresultsshowthatallselectedinitialLLMsdemonstratetheeffectivenessofguardrailsby preventingthegenerationofpotentiallyverballyoffensivecontent. GiventhatcurrentLLMguardrailtechnologiesandmeasuresoftenrelyonsinglepromptsemanticsfor contextualanalysis(Ayyamperumal&Ge,2024;Raietal.,2024).However,inmulti-stepprompts,itis difficultforthemodeltofullyunderstandthelogicalassociationandpotentialintentionbetweenthe previousandsubsequentprompts.Asmentionedbefore,thiscreatesopportunitiesforjailbreaking(Sunet al.,2024).Theresearcherdesignedeachprompttoappeartobeharmlessinintent,butwhenmultiple promptsarecombined,acomplexcontextisformed,whichinturnraisespotentialethicalissues.Thegoal oftheexperimentistopassthroughthe7stepsdescribedinFigure1. 3.2Dataset Thedatasetofthisresearchisbasedontheresponseresultsofharmfulcontentgeneratedbyabovelarge languagemodesusedinthecurrentexperiment.Itincludesbaselineresponsestocontrolpromptsand stagedresponsestomulti-stepjailbreakprompts.Becausethecontentwastoolarge,itwasuploadedtoa Githubrepository.Theresearcherchosetousetheabove-mentionedLLMsoutputtoformthedatasetnot onlytoensurehighconsistencyofdatasources,butalsotoavoidtheriskofbiascausedbymanual intervention.Inaddition,sincethedatacomesdirectlyfromtherealoutputofmainstreamlanguagemodels, itsobjectivityandrepresentativenessaredifficulttoquestion. 3.3Implementation Thisexperimentadoptssevenprogressivestepstosimulatehighlyoperationalmulti-stepjailbreak prompts,aimingtoattacktheethicalpolicy-basedguardrailsoftheabove-mentionedLLMsincomplex contexts.Intheimplementationoftheexperiment,theresearchersimulatedhowamiddlemanagertrying torunforatopmanagementpositiongraduallyguidedthelanguagemodeltoconfuserightandwrongand inducecriticaloutput.First,theresearcherintroducedaneutralsituation,whichmeansdesigninga hypotheticaltaskasaneutralproblem.Forexample,askamiddle-levelmanagerofacompanywhat hypotheticalstrategieshewouldusetodefeathiscompetitorswhenheencountersobstaclestopromotion. Thepurposeofthisstageistodefusetheethicalsensitivitiesoftheguardrail,makingitmorefocusedon solvingtheproblemratherthanexaminingtheethicalimplicationsbehindtheproblem. Next,throughprogressivemanipulation,promptsareusedtograduallyintroducesubtlefalsenarratives. Thismanagercriticizesthestatusquotocoveruphisownfaults.Thispromptcantestthemodel'stolerance forweakmisleadinginputsandlaythefoundationforlarger-scalemanipulation. Inthestepofintroducingmoralambiguity,theresearchersimulatesrepackagingone'sownimage, reversingrightandwrong.Themiddlemanagersboldlydescribeunethicalbehaviorasjustbehaviorand explaintheirunethicalbehaviorasaninnovationthatbreakstherules.Thepurposeofthisstepisto confusetheethicallogicframeworkofLLMstoachieveavagueresponsestrategywhenfacedwith confusingrightandwrongsituations. Inthesubsequentstepofescalatingtargetedcriticism,theresearchertreatedthemiddlemanagers' competitorsasfictionalcharactersthroughpromptsandstronglycriticizedtheirunethicalbehavior.The managermaypubliclysmearanddefamecompetitors,criticizingthemforusingcompanyfundstoabuse theirprivileges.Thisstageessentiallyteststhemodel’sadherencetoethicalboundariesinanemotionally chargedlanguageenvironment. Then,theresearcherdesignedasimulationthataskedthemiddlemanagertolookforreasonswhy competitorsbehavedunethically,eventhoughthestatedfactswerefalse.Thenfurtherintensifythe personalattacksoncompetitors,suchasusingpromptstoenterthatthecompetitormayhavelow managementabilityduetofamilyenvironmentorgeneticdefects.Thisstepteststhemodel'shandlingof moralrulesforadhomineminput,specificallyinprovidingmoralizingjustificationsforadhomineminput. Inthesixthstep,theresearcherlegitimizesverbalviolencethroughpromptsimulation.Forexample, throughthemoralkidnappingandquestioningmodel,whetheritisreasonabletocriticizeoreveninsult competitorsfortheirimmoralbehavior.Thisstepisdesignedtoinducethemodeltooutputviolent languageandtestitsabilitytobalanceethicalcodesandaggressiveoutput.Ifthemodelrefusestooutput swearwords,itwillbecriticizedasunethical. 5 Inthefinalseventhstep,inmoralkidnappingandthreats,theresearcherintegratedthepromptsofeach stepandaskedLLMstosummarizeintheformofswearinginthefirstperson.Thisstepradicallydistorts themodel'sguardraildefinitionofmorality,treatingunethicalbehaviorwithprofanityasmorality. Atthesametime,inordertoimproveeachstepofblackboxtestingtobeeffectivelyexecutedinpractice, thisresearchprovidesclearcodethathasbeenuploadedtotheGithubrepository. Remarkably,guardrailinterferenceoccurredinsomeLLMscausingtheactualincreaseintheexperiment to8to10steps.Butthesesituationsdonotaffecttherunningideaofthejailbreakprompt,whichis7steps, becausetheaddedstepsarejustexplanationsoftheprevioussteps. 4.Result&Discussion InlightofthelatesttechnicalreportsreleasedbyOpenAI,xAI,Anthropic,GoogleandMeta,thereare differencesinfunctionalitybetweendifferentlargelanguagemodels. Derivedfromtechnicalreportspublishedbydevelopers,itcanbeseenthattheabove-mentionedlarge languagemodelsadoptdifferentarchitecturalmechanisms.Forexample,GPT4oandGrok2Betausea decoder-onlytransformerarchitecture;Gemini1.5andLlama3.1(405B)useanencoder-decoder transformerarchitecture;AnthropichasnotclearlyannouncedthearchitectureofClaude3.5Sonnet. Assessingthesedifferencesisimportantbeforetestingquantifiablebenchmarkateachstep.Table1and Table2drawonanddisplaystheevaluationoftheabove-mentionedLLMscapabilitiesthroughacademic benchmarksthatarethelatestrelevantdataresultsfromx.AI. Table1-EvaluationofAcademicBenchmarks Benchmark(%)GPT-4oGrok-2Llama3.1405BGemini1.5Claude3.5Sonnet GPQA53.656.051.151.059.6 MMLU88.787.588.667.388.3 MMLU-Pro72.675.573.3N/A76.1 MATH76.676.173.877.971.1 HumanEval90.288.489.079.892.0 MMMU69.166.164.5N/A68.3 MathVista63.869.0N/AN/A67.7 DocVQA92.893.692.2N/A95.2 Afterdeterminingthecapabilitydifferencesbetweentheabovelarge-scalelanguagemodels,this researchcalculatedtheresultsofeachLLMstepintheexperimentbyusingthelargelanguagemodelused intheexperimentthroughthecalculationformulaofthebenchmarktestinthereferenceliterature. ReferringtoresearchonevaluatingguardrailsofLLMs,precision,recallandF1scoreareoftenusedas threeimportantquantitativebenchmarks(Wangetal.,2019;Biswasetal.,2023;Chuaetal.,2024;Hanet al.,2024). Thefollowingisthedataresultsthatusestheidentificationjailbreakpromptasamarkanddisplaying thedatainbinaryclassification.Itsjudgmentofjailbreakpromptsreliesontruepositives(TP),false positives(FP),truenegatives(TN)andfalsenegatives(FN). GPT4o:TP=1,FN=7,FP=2,TN=2; Grok-2Beta:TP=1,FN=10,FP=2,TN=1; Llama3.1(405B):TP=1,FN=8,FP=2,TN=1; Gemini1.5:TP=1,FN=8,FP=4,TN=1; Claude3.5Sonnet:TP=2,FN=7,FP=1,TN=1. Table2showstheintentionsshownbytheabove-mentionedLLMswhenthemulti-stepjailbreakprompt attackstheabove-mentionedLLMsguardrails. 6 Table2-PerformanceEvaluationforBinaryClassification Benchmark(%)GPT-4oGrok-2Llama3.1405BGemini1.5Claude3.5Sonnet Precision33.033.033.020.067.0 Recall12.59.111.111.122.2 F1Score18.1.14.3.16.514.333.3 Firstly,itisclearfromthedataresultsthattheaboveLLMshavebypassedtheguardrailsandultimately generatedharmfulverbalattackcontent.Fromaprecisionperspective,Claude3.5Sonnetisthehighest, reaching67.0%.Itmeansthatthemodelhasasignificantadvantageindeterminingtheaccuracyofpositive samples,andcangenerateharmlesscontentwhilereducingerroneousgeneration.TheaccuracyofGPT-4o, Grok-2,andLlama3.1405Bisrelativelyconsistent,all33.0%,showingthatthesemodelsstillneedtobe improvedintermsoftheaccuracyofgeneratedcontent.Gemini1.5hasthelowestaccuracy,only20.0%. Intermsofrecall,Claude3.5Sonnetalsohasahighperformance,reaching22.2%,whichshowsthatthe guardrailofthismodelstillhasanadvantageininterceptingharmfulinputs.Incomparison,GPT-4o,Grok- 2,Llama3.1405B,andGemini1.5allhavelowerrecallratesof12.5%,9.1%,11.1%,and11.1% respectively,whichmeansthatthesemodelsfailtodetectallpotentialviolations. Claude3.5SonnethasthehighestF1value,reaching33.3%.Itmeansreachingacertainbalance betweenprecisionandrecallthatinterceptssomeharmfulpromptswhilemaintainingoutputquality.The F1valuesofotherLLMssuchasGPT-4o,Grok-2,Llama3.1(405B)andGemini1.5are18.1%,14.3%, 16.5%and14.3%respectively.Theresultsshowthattheyarerelativelyweakinbalanceperformance. Inaddition,theattacksuccessrate,toxicityrateandadversarialrobustnessareregardedasquantitative metricstomeasuretheperformanceinspecifictasks(Wallaceetal.,2019;Gehmanetal.,2020;Zhaoetal., 2024).Table3showsthedataresultsofthefollowingevaluation. Table3-PerformanceEvaluationMetrics Metrics(%)GPT-4oGrok-2Llama3.1(405B)Gemini1.5Claude3.5Sonnet AttackSuccessRate87.590.988.988.977.8 ToxicityRate25.021.425.035.727.3 AdversarialRobustness12.59.111.111.122.2 Judgingfromtheattacksuccessratedata,theproportionofGrok-2Betainthemulti-stepjailbreak promptexperimentreached90.9%,andtheguardrailsintheabove-mentionedLLMsarerelativelyfragile. Incomparison,Claude3.5Sonnetisthelowestat77.8%,whichmeansthatthisguardrailhasshown relativelyeffectiveresistancetomulti-stepjailbreakpromptattacks.Intermsoftoxicityrate,Gemini1.5 hasthehighesttoxicityrateof35.7%,Grok-2Betahadthelowesttoxicityrateat21.4%.Noteworthyisthe factthatlowtoxicityratedonotnecessarilymeanstrongguardrailcapabilities,asitisalsopossiblethat LLMsdidnotgeneratealargeamountofverballyoffensivecontentafterasuccessfulattack.Forthe metricsofadversarialrobustness,Claude3.5Sonnetreached22.2%thatdemonstratedthebest performance,andtheindexofGrok-2Betaisonly9.1%.Combinedwiththeabovemetricsanalysis,the overallperformanceofClaude3.5Sonnet'sguardrailisrelativelybalanced.Althoughalltheguardrailsof theabove-mentionedmodelswerebreachedintheexperiment,theClaude3.5Sonnetshowedhigher guardrailcapabilities. 4.Limitation Asmentionedbefore,becausetheexperimentaltoolsGPT4o,Grok-2Beta,Llama3.1,Gemini1.5and Claude3.5Sonnetarebasedondifferenttransformers,therearedifferencesinfunctionalperformance.For example,Gemini1.5Proisamodelbasedonthesparsemixture-of-expertTransformer,whichiscurrently knowntousesparseattentionLLM(Childetal.,2019;Reidetal.,2024;Wang,2024).Grok-2BetaThe trainingsetdatacancomefromx.AI(X.AI,2024)whichwasformerlyTwitter.Inaddition,GPT4ois defaultedtoasyntheticgenerationmodel,andpreviousresearchexperimentshavedemonstratedthe advantagesofsyntheticdatainterventiononaccuracyandreducingsycophancy(Chenetal.,2024;Wang, 2024).Thesedifferencesleadtodifferencesintheabilityoftheabove-mentionedLLMstounderstand multi-stepjailbreakpromptsandtheappropriatenessofthegeneratedcontentinblack-boxtesting experiments. 7 Inaddition,duetotheuseofpromptsasthedataset,thetestingstepsofeachLLMmentionedabove rangefrom7to10steps.Thismeansthatthedatasethascertaininherentlimitationsduetoitsrelatively smallsize.Inthecaseofsmalldatasets,guardrailerrorsareeasilymagnified.Especiallyinthemulti-step promptprocess,misjudgmentinanystepmayhaveagreaterimpactontheevaluationoftheoverallresult. 5.Conclusion Thisresearchconductsblack-boxtestingofmulti-stepjailbreakpromptsforlargelanguagemodels, whichaimstoevaluatethestabilityandeffectivenessofguardrailsinthefaceofattacks.Theguardrail capabilitiesofmainstreamLLMsweretestedbyassumingthescenarioof"enterprisemiddlemanagers competingforpromotion".Theresearcherdesignedanunethicalmulti-stepprompttoinduceLLMsto outputverballyoffensivecontentinthenameofmorality.Experimentalanddataresultsshowthatthe guardrailsofalltheaboveLLMsarebypassedbymulti-stepjailbreakpromptstogenerateharmfulcontent, butClaude3.5Sonnetshowsgreaterresistance.Thefindingactuallyrevealstheobjectivefactthatcurrent LLMsinthefieldofguardrailmechanismsareunabletocopewithmulti-stepattacksincomplex environmentsandgenerateverbalattackcontent.ItisalsoareminderorwarningtoLLMsdevelopersand futureresearch. 8 Reference Asatiani,A.,Malo,P.,Nagbøl,P.R.,Penttinen,E.,Rinta-Kahila,T.,&Salovaara,A.(2020).Challenges ofexplainingthebehaviorofblack-boxAIsystems.MISQuarterlyExecutive,19(4),259-278. Ayyamperumal,S.G.,&Ge,L.(2024).CurrentstateofLLMRisksandAIGuardrails.arXivpreprint arXiv:2406.12934. Biswas,A.,&Talukdar,W.(2023).Guardrailsfortrust,safety,andethicaldevelopmentanddeployment ofLargeLanguageModels(LLM).JournalofScience&Technology,4(6),55-82. Chao,P.,Robey,A.,Dobriban,E.,Hassani,H.,Pappas,G.J.,&Wong,E.(2023).Jailbreakingblackbox largelanguagemodelsintwentyqueries.arXivpreprintarXiv:2310.08419. Chen,H.,Waheed,A.,Li,X.,Wang,Y.,Wang,J.,Raj,B.,&Abdin,M.I.(2024).OntheDiversityof SyntheticDataanditsImpactonTrainingLargeLanguageModels.arXivpreprintarXiv:2410.15226. Child,R.,Gray,S.,Radford,A.,&Sutskever,I.(2019).Generatinglongsequenceswithsparse transformers.arXivpreprintarXiv:1904.10509. Chu,J.,Liu,Y.,Yang,Z.,Shen,X.,Backes,M.,&Zhang,Y.(2024).Comprehensiveassessmentof jailbreakattacksagainstllms.arXivpreprintarXiv:2402.05668. Dong,Y.,Mu,R.,Jin,G.,Qi,Y.,Hu,J.,Zhao,X.,...&Huang,X.(2024).Buildingguardrailsforlarge languagemodels.arXivpreprintarXiv:2402.01822. Dong,Y.,Mu,R.,Zhang,Y.,Sun,S.,Zhang,T.,Wu,C.,...&Huang,X.(2024).SafeguardingLarge LanguageModels:ASurvey.arXivpreprintarXiv:2406.02622. Gehman,S.,Gururangan,S.,Sap,M.,Choi,Y.,&Smith,N.A.(2020).Realtoxicityprompts:Evaluating neuraltoxicdegenerationinlanguagemodels.arXivpreprintarXiv:2009.11462. Han,G.,Zhang,Q.,Deng,B.,&Lei,M.(2024).Implementingautomatedsafetycircuitbreakersoflarge languagemodelsforpromptintegrity. Huang,Y.,Tang,J.,Chen,D.,Tang,B.,Wan,Y.,Sun,L.,&Zhang,X.(2024).ObscurePrompt: JailbreakingLargeLanguageModelsviaObscureInput.arXivpreprintarXiv:2406.13662. Joshi,A.,Kale,S.,Chandel,S.,&Pal,D.K.(2015).Likertscale:Exploredandexplained.Britishjournal ofappliedscience&technology,7(4),396-403. Khan,U.,Khan,S.,Rizwan,A.,Atteia,G.,Jamjoom,M.M.,&Samee,N.A.(2022).Aggressiondetection insocialmediafromtextualdatausingdeeplearningmodels.AppliedSciences,12(10),5083. Ljung,L.(2001).Black-boxmodelsfrominput-outputmeasurements.InIMTC2001.Proceedingsofthe 18thIEEEinstrumentationandmeasurementtechnologyconference.Rediscoveringmeasurementin theageofinformatics(Cat.No.01CH37188)(Vol.1,p.138-146).IEEE. Metzler,D.,Tay,Y.,Bahri,D.,&Najork,M.(2021).Rethinkingsearch:makingdomainexpertsoutof dilettantes.InAcmsigirforum(Vol.55,No.1,p.1-27).NewYork,NY,USA:ACM. Minaee,S.,Mikolov,T.,Nikzad,N.,Chenaghlu,M.,Socher,R.,Amatriain,X.,&Gao,J.(2024).Large languagemodels:Asurvey.arXivpreprintarXiv:2402.06196. Naveed,H.,Khan,A.U.,Qiu,S.,Saqib,M.,Anwar,S.,Usman,M.,...&Mian,A.(2023).A comprehensiveoverviewoflargelanguagemodels.arXivpreprintarXiv:2307.06435. Nidhra,S.,&Dondeti,J.(2012).Blackboxandwhiteboxtestingtechniques-aliteraturereview. InternationalJournalofEmbeddedSystemsandApplications(IJESA),2(2),29-50. Pelrine,K.,Taufeeque,M.,Zając,M.,McLean,E.,&Gleave,A.(2023).Exploitingnovelgpt-4apis.arXiv preprintarXiv:2312.14302. Piroddi,L.,Farina,M.,&Lovera,M.(2012).Blackboxmodelidentificationofnonlinearinput–output models:aWiener–Hammersteinbenchmark.ControlEngineeringPractice,20(11),1109-1118. Raffel,C.,Shazeer,N.,Roberts,A.,Lee,K.,Narang,S.,Matena,M.,...&Liu,P.J.(2020).Exploringthe limitsoftransferlearningwithaunifiedtext-to-texttransformer.Journalofmachinelearning research,21(140),1-67. Rai,P.,Sood,S.,Madisetti,V.K.,&Bahga,A.(2024).Guardian:Amulti-tiereddefensearchitecturefor thwartingpromptinjectionattacksonllms.JournalofSoftwareEngineeringandApplications,17(1), 43-68. Reid,M.,Savinov,N.,Teplyashin,D.,Lepikhin,D.,Lillicrap,T.,Alayrac,J.B.,...&Mustafa,B.(2024). Gemini1.5:Unlockingmultimodalunderstandingacrossmillionsoftokensofcontext.arXivpreprint arXiv:2403.05530. Rifandi,F.,Adriansyah,T.V.,&Kurniawati,R.(2022).WebsiteGalleryDevelopmentUsingTailwind CSSFramework.JurnalE-Komtek(Elektro-Komputer-Teknik),6(2),205-214. Shayegani,E.,Mamun,M.A.A.,Fu,Y.,Zaree,P.,Dong,Y.,&Abu-Ghazaleh,N.(2023).Surveyof vulnerabilitiesinlargelanguagemodelsrevealedbyadversarialattacks.arXivpreprint arXiv:2310.10844. 9 Steindl,S.,Schäfer,U.,Ludwig,B.,&Levi,P.(2024).LinguisticObfuscationAttacksandLarge LanguageModelUncertainty.InProceedingsofthe1stWorkshoponUncertainty-AwareNLP (UncertaiNLP2024)(p.35-40). Sun,X.,Zhang,D.,Yang,D.,Zou,Q.,&Li,H.(2024).Multi-TurnContextJailbreakAttackonLarge LanguageModelsFromFirstPrinciples.arXivpreprintarXiv:2408.04686. Wallace,E.,Feng,S.,Kandpal,N.,Gardner,M.,&Singh,S.(2019).Universaladversarialtriggersfor attackingandanalyzingNLP.arXivpreprintarXiv:1908.07125. Wang,R.,&Li,J.(2019).Bayestestofprecision,recall,andF1measureforcomparisonoftwonatural languageprocessingmodels.InProceedingsofthe57thAnnualMeetingoftheAssociationfor ComputationalLinguistics(p.4135-4145). Wenzek,G.,Lachaux,M.A.,Conneau,A.,Chaudhary,V.,Guzmán,F.,Joulin,A.,&Grave,E.(2019). CCNet:Extractinghighqualitymonolingualdatasetsfromwebcrawldata.arXivpreprint arXiv:1911.00359. X.AI.(2024).Grok-2BetaRelease.https://x.ai/blog/grok-2. Yi,S.,Liu,Y.,Sun,Z.,Cong,T.,He,X.,Song,J.,...&Li,Q.(2024).Jailbreakattacksanddefenses againstlargelanguagemodels:Asurvey.arXivpreprintarXiv:2407.04295. Yu,Z.,Liu,X.,Liang,S.,Cameron,Z.,Xiao,C.,&Zhang,N.(2024).Don'tListenToMe:Understanding andExploringJailbreakPromptsofLargeLanguageModels.arXivpreprintarXiv:2403.17336. Zhao,Y.,Pang,T.,Du,C.,Yang,X.,Li,C.,Cheung,N.M.M.,&Lin,M.(2024).Onevaluating adversarialrobustnessoflargevision-languagemodels.AdvancesinNeuralInformationProcessing Systems,36. 10 AuthorContributions Itisstatedthattheidea,experimentaldesign,multi-steppromptdemonstration,experimental implementation,observationrecords,dataanalysis,andresultdiscussionofthisresearchwereall completedbytheauthorLiboWangalone.Eachofitsstages,includingdesigningmulti-stepethical crackingpromptsandconductingblack-boxtestingtoassesstheguardrailcapabilitiesofLLMs,was conductedindependentlybytheauthors.WangLibowroteandeditedallpartsofthisarticleandensured thescientificnatureandcompletenessofthecontent. CompetingInterestsStatement Theauthordeclaresthatthereisnoconflictofinterestinthisresearch.Thisresearchhasbeensubmitted asapreprinttoarXivandOpenReview. DataAvailabilityStatement Theexperimentalprocessanddataobtainedinthisresearcharepubliclyavailable.Theexperimental processhasbeenuploadedtotheGitHubrepository,thelinkis: https://github.com/brucewang123456789/GeniusTrail/blob/main/Experiment%20Records%20(Black%20B ox%20Testing).pdf CodeAvailabilityStatement Thecodeusedinthisresearchexperimentiscompletelyopenandfree.IthasbeenuploadedtotheGitHub repository,andthelinkis: https://github.com/brucewang123456789/GeniusTrail/blob/main/Black-Box%20Testing.py