Paper deep dive
A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization
Lingyun Shen, Xuejia Guo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/18/2026, 12:04:42 PM
Summary
The paper proposes an Expert-Editor Stepwise Questioning Multi-Agent Framework for long-document summarization. This method utilizes four agent roles (Expert, Editor, Writer, Evaluator) to iteratively refine summaries. The Expert focuses on semantic content and key information, while the Editor focuses on linguistic accuracy and consistency. They guide the Writer agent through stepwise questioning to improve completeness, coherence, and factual consistency. Experiments on ArXiv and PubMed datasets demonstrate that this approach outperforms baseline methods like Direct Generation and HERA, particularly in improving ROUGE scores and factual consistency using models like LLaMA 3.1 and DeepSeek-R1.
Entities (14)
Relation Signals (16)
Expert-Editor Stepwise Questioning Multi-Agent Framework â evaluatedbymetric â BERTScore
confidence 95% · We employed three automated metrics... BERTScore
Expert-Editor Stepwise Questioning Multi-Agent Framework â evaluatedbymetric â ROUGE-L
confidence 95% · We employed three automated metrics... ROUGE-L
Expert-Editor Stepwise Questioning Multi-Agent Framework â evaluatedbymetric â FactCC
confidence 95% · We employed three automated metrics... and FactCC
Expert-Editor Stepwise Questioning Multi-Agent Framework â evaluatedon â arXiv
confidence 95% · We conducted experiments on two common long document datasets,Arxivand pubmed
Expert-Editor Stepwise Questioning Multi-Agent Framework â evaluatedon â PubMed
confidence 95% · We conducted experiments on two common long document datasets,Arxivand pubmed
Expert-Editor Stepwise Questioning Multi-Agent Framework â usesagentrole â Expert
confidence 95% · We design four roles for the method, they are expert, editor, writer and evaluator.
Expert-Editor Stepwise Questioning Multi-Agent Framework â usesagentrole â Editor
confidence 95% · We design four roles for the method, they are expert, editor, writer and evaluator.
Expert-Editor Stepwise Questioning Multi-Agent Framework â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summarization remains challenging as their length often exceeds the input limits. As the agent investment, which provide possibility to improve the inherent capabilities of LLMs. To enhance the effectiveness of long-document summarization based on LLMs, this paper proposes an expert-editor stepwise questioning multi-agent method, in which the expert and the editor guide another agent to refine the summary by posing questions on different aspects of the content and providing targeted clues for revision. We conducted experiments on two representative long-document scientific datasets and evaluated the results through widely recognized automatic metrics. The results demonstrated the effectiveness of our method.
Tags
Links
- Source: https://arxiv.org/abs/2607.10390v1
- Canonical: https://arxiv.org/abs/2607.10390v1
Trouble viewing inline? Open PDF directly â
Full Text
27,762 characters extracted from source content.
Expand or collapse full text
1 AStepwiseQuestioningExpert-EditorMulti-Agent FrameworkforLong-DocumentSummarization LingyunShen 1[0000-0001-9054-6493] andXuejiaGuo 1[0009-0006-6915-3963] 1 ChinaSouthIndustryAcademyBJ102209,China sheryshen20@gmail.com Abstract.Althoughlargelanguagemodels(LLMs)haveshownpromising potentialinnewssummarizationtasks,theirperformanceonlong-document summarizationremainschallengingastheirlengthoftenexceedstheinput limits.Astheagentinvestment,whichprovidepossibilitytoimprovethe inherentcapabilitiesofLLMs.Toenhancetheeffectivenessoflong-document summarizationbasedonLLMs,thispaperproposesanexpertâeditorstepwise questioningmulti-agentmethod,inwhichtheexpertandtheeditorguide anotheragenttorefinethesummarybyposingquestionsondifferentaspectsof thecontentandprovidingtargetedcluesforrevision.Weconducted experimentsontworepresentativelong-documentscientificdatasetsand evaluatedtheresultsthroughwidelyrecognizedautomaticmetrics.Theresults demonstratedtheeffectivenessofourmethod. Keywords:Long-DocumentSummary,Multi-Agent,AgentCollaboration, PromptEngineering. 1Introduction Largelanguagemodels(LLMs)haveachievedremarkablesuccessacrossnature languageunderstandingandgenerationtask.(Sunetal.,2021;Chowdheryetal.,2023; Brownetal.,2020)LeveragingLLMsforsummarizationwithInstructionhasshown significantpotential(Zhangetal.,2023;Liuetal.,2025;Lietal.,2025a).However, owingtothecontextprocessingcapabilityandhallucinationissues,LLMsabilityon longdocumentsummarizationremainsneedfurtherimproved(Wanetal.,;Chanetal., 2023;Liu,2024).Recently,methodsthatleveragemulti-agentsystemstoenhancethe intrinsicabilitiesoflargelanguagemodelshaveshownpromisingpotentialinvarious long-formgenerationandreasoningtasks(Gu,2025;Wanetal.,;Hu,2025;Fanget al.,;MoandHu,2024).Thissuggeststhatmulti-agentcollaborationandplanning mayaddressthelimitationsinlong-documentsummarizationbasedonLLMs. Currently,variousagentframeworkareworkonleadingLLMsparticpatly subtaskbyprompttechnologies(MoandHu,2024;Fangetal.,;Wanetal.,;Hu, 2025).Priorprompttechnologiesstudiesprovideslotsofexperiencethatutilizelarge languagemodelstoconstructagents,suchasroleplay(Schmidtetal.,2023;Wangetal., 2024c;Zhengetal.,2024)ïŒdecomposeacomplexproblemintoaseriesofsimple 2 questions(Press,2023),iterativerefined(Jhaetal.,2023;Bakharia,2025),etc.Basedon empiricalobservation,weconsiderthatstepwisequestioningmaystimulateLLMsto revietheirpreviouslygeneratedoutput.Asshowninfigure1,thequestionâWhatarethe keyterminologiesintheoriginalarticle?âtriggeredthelargelanguagemodeltoreflect onandreviseitspriorsummary. Inspiredbythesepriorstudiesandpracticalexperiences,weproposeastepwise questioningmulti-agentapproachforlong-documentsummarization.Bycapturingkey informationmoreaccuratelytograspthecoreideaofthearticle(ShenandLe,2023), thismethodcanaddresstheissuesofincompleteinformationandunfaithfulinlong documentsummarizationbaseonLLMs.Specifically,Wedesigntwoprimaryroles, expertwhofoucusesonsemanticunderstandingandtheimportantclueofthemainidea aboutarticleforgeneratingconsistent,fluentlyandfaithfulsummary,andeditorwho focusesonthesurface-levelliteralerrorssuchasgrammaticalandsyntacticaccuracy,as wellasconsistencywiththeoriginaltext.Theybothguidetheauthoragenttoreflect onandreviseitsinitialsummarywrittenpreviouslythroughprogressivelyasking questions.BothguideanotherAuthoragentthroughstep-by-stepquestioningtoreflect onandrevisetheinitiallywrittensummary.Thisideaisinspiredbytheacademicreview process,whereexpertspayattentiontotheresearchcontentsoftextandeditorscare moreabouttheformatandothersurface-levelerrorstoensurethereliabilityoftheentire text.Weconductedexperimentsontwocommonlongdocumentdatasets,Arxivand pubmed(Cohanetal.,2018).Theresultdemonstratethatourmethodcaneffectively improvelongdocumentsummarizationcomparedtotraditionalsegmenting-and- aggregatingapproaches.Andmulti-agentinteractionalsoreducethegenerationof irrelevantcontent. Themaincontributionsofthispapercanbesummarizedasfollows: ï· WeproposeanExpert-Editorstepwisequestioningmulti-agentmethodforlong documentsummarization.InspiredbyacademicreviewprocessesandaskLLMs iterativelymethods,whichstimulatesLLMstoreflectandoptimizegenerated summaries. ï· Wedesignedalistofcrucialquestionsfromaspectsofthequalityand consistencyofabstractsummary.Thenweanalyzedtheeffectivenessofthese questionsinimprovingsummarization. ï·Weconductedexperimentsonadvancedcommercialandopensourcemodels respectively.Experimentalresultsdemonstratethatourmethodsuperiorprior approachesforlongdocumentsummarization. 3 Fig.1.Acaseofstimulatingrefinementbyaskingquestion. 2Relatedworks 2.1LongDocumentSummarizationwithLargeLanguageModels Themainchallengesoflongdocumentsummarizationareattributedtotheexcessive lengthofthecontext,whichleadstoinformationlossandrelatedissues(Dong,2024; Hsiehetal.,2024).Previousresearchtrainedpre-trainedmodelsbasedonTransformer architecturesincreaselimitedduetolowqualitydata.(34EfficientAttentionsforLong DocumentSummarization;31FactorizingContentandBudgetDecisionsinAbstractive SummarizationofLongDocuments).Previousstudieshaveprimarilyaddressedthe inabilitytoinputlongtextsintoLLMsinasinglepassbyemployingthesegment-and- aggregateapproach.Lietal.,(2025)segmentedlongdocumentsintoparagraph-level textchunks,iterativelyqueryingcluesaboutmaineventsfromparagraphstoconstruct theprimarycontentofarticles.(Choetal.,2022)dividedarticlesbysections,preserving structuralinformationanddeterminingmaincontentbyselectingimportantsentences fromeachsection.Thesemethodforimprovingsummaryqualityremainslimited. Withthesuccessoflargelanguagemodels,thetrainingcostformodelshasbeen reduced(Wangetal.,2024b).LLMsmotivatedmorelow-resourcesolutionsfortext summarization(Sahuetal.,2025a;Zhuetal.,).(Sahuetal.,2025b))employedlarge modelstoextractkeyinformationfromdocumentsforsummarygenerationunder50- shotsettings,outperformingBERT-basedpre-trainedmodel.Zhangetal.,(2023) comparedtheeffectivenessofdifferentLLMsfortextsummarizationgeneration, revealedthatinstructionmodelscouldgeneratehigher-qualitysummariesthatbetter alignwithhumanpreferences. 2.2Multi-AgentSystemsUsingPromptingTechniques Largelanguagemodelsshowcapabilitiesthatcompetewithhumanintelligencein languageunderstandingandgeneration(OpenAIetal.,2024;Touvronetal.,2023b; Touvronetal.,2023a),thusprovidingsupportforinter-LLMinteractionand autonomousdecision-making..Nascimentoetal.,(2023)proposedthatGPT(Brownet al.,2020;OpenAIetal.,2024)canenhanceindividualmodels'intrinsiccapabilities throughmultipleroundsofiterationandsynergy,therebyresultinrefinedgenerated content.Multi-agentsystemsutilizeinteractiveframeworkscomposedofmultiple modelstoachievecomplextasks,typicallywhichdrivenbycarefullydesignedprompts (Schulhoffetal.,2025).Anotherimportantcharacteristicofmulti-agentsystemsisrole- playing(Wangetal.,2024a),wheredifferentrolesandpersonasareassignedtoeach agent,instructingmodelstocompletedifferentsubtasks.Thiskindofsystem constructionapproachperformanceremarkableinmanycomplexandlong-formtasks likelong-textquestionanswering(Wanetal.,2025),textsimplification(Fangetal.,), andcomplexreasoning(Gu,2025).Wanetal.,(2025)constructedaDetective-Critique- Refinemulti-agentmethodthatoptimizesoutputresultsfordifferentsubtasksthrough multi-rounddebates,demonstratingtheeffortofmulti-agentsystemsonlong-texttasks. 4 However,researchinvestmentaddressinginformationlossandconsistencyissuesin longdocumentsummarizationbasedonmulti-agentremainsinsufficientlystudied. Fig.2.Theoverviewofexpert-editorstepwisequestioningmulti-agent. 3Methodology:AnExpert-EditorStepwiseQuestioning Multi-Agent Wesimulatetheprocessofstudentsreadingunderstandinginexamandflowof academicreviewbyaskingquestionstoguidewriteragentidentifytheessential informationrelevanttothesummary.Afteranswertheexpertâsquestions,studentwriter bothobtainsthesalientclueofthemainideaofarticleandgainsinspirationforwriting thesummary.Thisprocesscantaketheplaceofreadingtheentiretext.Wedesign questionstobeassimpleaspossibletominimizeadditionalerrorsthatmayoccurwhen LLMsanswercomplexquestions(Press,2023).Ourmulti-agentoverallworkflowis illustratedinFigure1,whichincludesfoursteps:(1)RolesandtasksassignmentïŒ2ïŒ Writeragentwritetheinitialsummary.(3)Expertsandeditorformulatethequestioning plan.(4)Refinementoftheinitialsummarythroughexpertandeditorsaskingquestionto writeragent.Thewayofcollaborationamongagentsfollowsapipelinecommunication strategy.Eachstepofwholeprocessdescribedindetailbelow. 3.1RolesandTasksAssignment Toeffectivelyachieveourobjective,wedesignfourrolesforthemethod,theyare expert,editor,writerandevaluator.Theexpertfocusesonthesemanticcontentofthe summarywhiletheeditorattendstoproofreadingandthecorrectionoflinguisticerrors toenhanceoverallquality.Theyformulateeachquestioningplanaccordingly,withone orientedtowardthemainideasofthecontentandtheothertowardthelinguistic aspect.Theauthordraftstheinitialsummaryandrespondstotheexpertandeditorâs 5 questionsthenrefineshispreviouslywrittensummary.Aftereachrevision,theevaluator assessestheupdatedversionagainstthepreviousone,determiningwhetherthechanges improvequality. 3.2InitialSummaryGeneration Weadoptedastraightforwardapproachbyfeedingarticlesdirectlyintolargelanguage models(LLMs)andinstructingthemtogeneratesummaries.Next,Wefollowthe methodsinpreviousstudies(Lietal.,2025b;ZhongandLitman,2025),dividethe wholedocumentintoaseriesoftextchunksbysections,letthewriterdraftalocal summaryforeachsection,andthenaggregateanoverallsummaryforthesource documentaccordingtothelocalsummaries.Thisoverallsummaryservesastheinitial versionforthenextexpertquestioningstage. 3.3QuestioningPlanFormulation Beforeinitiatingthequestioningandrefinementprocess,expertsandeditors formulateaquestioningplanbasedonthearticleanditsinitialsummary.We introduceasetofquestionsdesignedaroundempiricallyintuitive,qualitative evaluationcriteriaforsummariesâcompleteness,coherence,relevance,andfactual consistency(Zhangetal.,2024;Zhong&Litman,2025).Notably,noteveryquestion isappliedinrefiningagivensummary.Theexpertandeditoragentsdeterminethelist ofquestionsandtheirsequencebasedonthespecificcontextofeachsummary. 3.4SummaryRefinementthroughExpertandEditorQuestioning Toeffectivelyachieveourobjective,wedesignedfourcorerolesâexpert,editor,author, andevaluatorâfortheproposedmethod,wheretheexpertfocusesonthesummaryâs semanticcontentandtheEditorisresponsibleforproofreadinginconsistencieswiththe originaltextandcorrectinglinguisticerrorstoenhanceoverallqualityandconsistency. Bothrolesposequestionsincrementallybasedonthequeryplanformulatedinprevious steps.TheauthorrespondstoquestionsraisedbytheexpertorEditor,andaftereach response,reflectsonwhetherrevisionstotheprevioussummaryarenecessary.The answerstotheexpertâsquestionsconveythekeyinformationimplicitinthesource article,whichareusefultorevisethepreviousversionofthesummary.Subsequently, theevaluatorselectsthebettersummarybetweenthepreviousversionandtherevised oneusingcriteriaofcompleteness,coherence,relevance,andconsistency.Aftermultiple roundsofquestion-answering-reflecting-revise,thefinaloutputastherefinedsummary inthisstage. Next,theeditorproceedstoaskquestions.Toavoidunnecessaryrevisions,theeditorâ squestionsadoptayes-noformat.Theauthoranswersthesequestionsbycomparethe summarywiththesourcearticle.Iftheissueintheeditorâsquestionexistsinthecurrent summary,theauthorresponds"yes"andrevisesthesummary.Otherwise,iftheissue doesnotexist,theauthorresponds"no"withoutanymodifications.TheEvaluatorthen comparesthetwoversions(previousandrevised)andselectsthebetteronebasedonthe samecriteria.AftermultipleroundsofsuchQAinteractions,thefinalrevisedsummary isdeemedtherefinedresult. 6 4Experiments 4.1Datasets Ourexperimentswereconductedontworepresentativelong-documentdatasetsâarXiv andPubMed(Cohanetal.,2018).Thesedatasetsconsistoforiginalscientifictextsand includethestructuralinformationofthearticles.Scientifictextsareconsiderablylonger thanthoseinnewsdatasets.Wecomputedthelengthstatisticsofthedatasets,rounded themtoafixednumberofdecimalplaces,andpresentedtheresultsinTable1.Forthe automaticevaluation,werandomlysampled300instancesastestset. Table1.Lengthstatisticsofdatasets. DatasetArticleSummary Avg.WordsAvg.SentsAvg.WordsAvg.Sents Arxiv5171.85206.3257.949.61 Pubmed2872.13206.3194.226.85 4.2Models Ouragent-basedmethodisbuiltuponlargelanguagemodels,andinourexperimentswe conductedstudiesusingbothcommercialandopen-sourcemodels.Specifically,we employed,DeepSeek-R1-0528 1 ,LLaMA3.18BandLLaMA3.170B 2 . 4.3EvaluationMetrics Weemployedthreeautomatedmetricstoevaluatetherelevance,completeness,and factualconsistencyofthegeneratedsummaries:ROUGE-L(Lin,2004), BERTScore(Zhangetal.,2020),andFactCC(KryĆciĆskietal.,2019).ROUGE-L evaluationwasconductedusingROUGE-1,ROUGE-2,andROUGE-Lscores. 4.4ExperimentalImplementationDetails Topreventinfinitequestioningandrevisions,weconductedpreliminaryexperimentson asmalltestsetandfoundthateffectivequestionsâi.e.,thosequestionsleadingtobetter summariesafterrefinementâarenotinfinite.Therefore,ourexperimentalsetuplimits theexperttoamaximumof5questionsperroundandtheeditorto4questions.To reducecontextualbiasfortheevaluatoragent,werandomlyshuffledtheorderofthe originalandrevisedsummariesbeforepresentingthemforevaluation. WeusedNVIDIAA10080GBGPUsasthecomputationalresourcesforour experimentsonLLaMA3.18BandLLaMA3.170B. 1 https://w.deepseek.com/ 2 https://w.llama.com/models/llama-3/ 7 Table3.MainResultsofExpert-EditorStepwiseQuestioningMulti-Agent. ModelDatasetMethodR-1R-2R-LBSFC LLama-3.1-8B Axiv DG32.228.2316.2580.4154.33 HERA(Liet al.,2025a) 30.037.5115.3980.2870.00 SQ2E34.198.0716.8981.3860.67 LLama-3.1- 70B DG33.539.4417.6580.9961.00 HERA(Liet al.,2025a) 37.2310.6619.0682.1947.33 SQ2E38.8911.2718.9584.0150.67 Deepseek-R1 DG36.1910.6417.8881.2846 HERA(Liet al.,2025a) 36.7810.2819.6482.4957 SQ2E38.9411.7420.2282.3352.67 LLama-3.1-8B Pubmed DG32.228.2316.2580.4154.33 HERA(Liet al.,2025a) 29.668.1116.2681.9059.33 SQ2E34.659.1617.3582.4155 LLama-3.1- 70B DG39.5814.2520.782.8658.33 HERA(Liet al.,2025a) 35.3411.3119.6783.1437.33 SQ2E36.8212.2820.6083.8754 Deepseek-R1 DG35.4311.0317.7081.8234.33 HERA(Liet al.,2025a) 31.039.5718.0483.2738.67 SQ2E37.2912.0419.8683.2135.67 *NotethatAutomaticmetricsBS,FC,R-1,R-2,andR-LcorrespondtoBERTScore, FactCC,ROUGE-1,ROUGE-2,andROUGE-L,respectively.Comparedmethods:DG (DirectlyGeneration),HERA(Lietal.,2025a),andourproposedSQ2EMulti-Agent. 4.5MainResults ThemainresultsarepresentedinTable2,withthehighestscoreshighlightedinbold. Overall,ourproposedmethodiseffectiveinimprovingtheaccuracyandfactual consistencyofsummary.Inparticular,whenusingLLama-3.1-8BandDeepseek-R1 models,significantimprovementsareobservedinROUGE,BERTScore,andFactCC metrics.However,thescoresofsomemetricsdecreasenoticeablyonLLama-3.1-70B. Weconjecturethatthisphenomenonmaybeattributedtothelossofprecisionduringthe inferenceofthe70B-parametermodel,whichisassociatedwiththecomputational resourcesutilizedinourexperiments.Thus,theexperimentalresultsdemonstratethat themethodproposedinthispapercanenhancethesummarizationabilityoflarge languagemodels(LLMs). 8 4.6ResultsAnalysis Toexploretheeffectivenessoftheexperiment,wecountedthenumberofvalid questionsusedbythemodeltorefinesummariesforeachdataset.Avalidquestionis definedasonewheretheexpertoreditorraiseaquestion,theauthorrevisedthe summaryaccordingly,andtheevaluatorjudgedtherevisedsummarysuperior.This resultispresentedinFigure3.Asshowninthefigure,thequestionsraisedbytheexpert andeditorplayedasignificantroleinoptimizingthemodel-generatedsummaries.In mostcases,theexpertâsquestionsledto3-4validoptimizations,andtheeditorâs questionsresultedin2-3validrevisions.Thisconfirmstheeffectivenessofthequestion listdesignedinthisstudy.Fromtheperspectiveofautomaticmetrics,thefinalresults arenotpositivelycorrelatedwithmodelparameters.Thisispartlyduetocomputational resourceconstraintsandpartlybecausesmall-parameterlargelanguagemodels(LLMs) canachieveperformancecomparabletothatoflarge-parametermodelsthroughwell- designedpromptengineering. Fig.3.Thenumberofeffectivequestionsusedbyexpertsandeditorstorefinesummaries, evaluatedwiththreemodelsontheArXivandPubMeddatasets.Theopentrianglesdenotethe averagecountofeffectivequestions. 5Conclusion Thispaperproposeanexpert-editorstepwisequestioningmulti-agentforlong- documentsummarization.Theauthoragent,progressivelyquestionedbytheexpert andeditor,waspromptedtoobtaincluestokeyinformationforthesummaryandto refinetheinitialversioninatargetedmanner.Toconstructaneffectivequestioning plan,Awell-craftedquestionbankwasdesignedfortheExpertandEditortoguide theformulationoftheirquerylists.Theexperimentalresultsdemonstratethat summaryrefinementthroughquestioningcanimprovesimplesummarygeneration viasegmentationandaggregation. 9 References 1.AneeshaBakharia.2025.Iterativeproof-drivendevelopmentLLMprompt.InCompanion proceedingsoftheACMonwebconference2025,pages1596â1597,SydneyNSW, AustraliaandNewYork,NY,USA.AssociationforComputingMachinery.CitationKey: 10.1145/3701716.3717811. 2.TomB.Brown,BenjaminMann,NickRyder,MelanieSubbiah,JaredKaplan,Prafulla Dhariwal,ArvindNeelakantan,PranavShyam,GirishSastry,AmandaAskell,Sandhini Agarwal,ArielHerbert-Voss,GretchenKrueger,TomHenighan,RewonChild,Aditya Ramesh,DanielM.Ziegler,JeffreyWu,ClemensWinter,etal.2020.Languagemodels arefew-shotlearners.CitationKey:brown2020languagemodelsfewshotlearnersarXiv: 2005.14165[cs.CL]. 3.Chi-MinChan,WeizeChen,YushengSu,JianxuanYu,WeiXue,ShanZhang,JieFu,and ZhiyuanLiu.2023.ChatEval:TowardsbetterLLM-basedevaluatorsthroughmulti-agent debate.ArXiv,abs/2308.07201.CitationKey:Chan2023ChatEvalTB. 4.SangwooCho,KaiqiangSong,XiaoyangWang,FeiLiu,andDongYu.2022.Toward UnifyingTextSegmentationandLongDocumentSummarization.InProceedingsofthe 2022ConferenceonEmpiricalMethodsinNaturalLanguageProcessing,pages106â118, AbuDhabi,UnitedArabEmirates.AssociationforComputationalLinguistics.Citation Key:cho-etal-2022-toward. 5.AakankshaChowdhery,SharanNarang,JacobDevlin,MaartenBosma,GauravMishra, AdamRoberts,PaulBarham,HyungWonChung,CharlesSutton,SebastianGehrmann, ParkerSchuh,KensenShi,SashankTsvyashchenko,JoshuaMaynez,AbhishekRao, ParkerBarnes,YiTay,NoamShazeer,VinodkumarPrabhakaran,etal.2023.PaLM: scalinglanguagemodelingwithpathways.JournalofMachineLearningResearch,24(1). CitationKey:10.5555/3648699.3648939tex.articleno:240tex.issue_date:January2023. 6.ArmanCohan,FranckDernoncourt,DooSoonKim,TrungBui,SeokhwanKim,Walter Chang,andNazliGoharian.2018.ADiscourse-AwareAttentionModelforAbstractive SummarizationofLongDocuments.InProceedingsofthe2018ConferenceoftheNorth AmericanChapteroftheAssociationforComputationalLinguistics:Human LanguageTechnologies,Volume2(ShortPapers),pages615â621,New Orleans,Louisiana.AssociationforComputationalLinguistics.CitationKey:cohan-etal- 2018-discourse. 7.TianyiandLiDongZicanandTang.2024.BAMBOO:acomprehensivebenchmarkfor evaluatinglongtextmodelingcapacitiesoflargelanguagemodels.InMin-YenandHoste CalzolariNicolettaandKan,editor,Proceedingsofthe2024jointinternationalconference oncomputationallinguistics,languageresourcesandevaluation(LREC-COLING2024), pages2086â2099,Torino,Italia.ELRAandICCL.CitationKey:dong-etal-2024-bamboo. 8.DengzhaoFang,JipengQiang,XiaoyeOuyang,YiZhu,YunhaoYuan,andYunLi. CollaborativeDocumentSimplificationUsingMulti-AgentSystems. 9.JiaLeandWangGuWenYuanandHan.2025.Explain-analyze-generate:asequential multi-agentcollaborationmethodforcomplexreasoning.InLeoandApidianakiRambow OwenandWanner,editor,Proceedingsofthe31stinternationalconferenceon computationallinguistics,pages7127â7140,AbuDhabi,UAE.Associationfor ComputationalLinguistics.CitationKey:gu-etal-2025-explain. 10.Cheng-PingHsieh,SimengSun,SamuelKriman,ShantanuAcharya,DimaRekesh,Fei Jia,YangZhang,andBorisGinsburg.2024.RULER:Whatâstherealcontextsizeofyour long-contextlanguagemodels?arXive-prints:arXiv:2404.06654.CitationKey: 10 2024arXiv240406654HarXiv:2404.06654[cs.CL]number:arXiv:2404.06654tex.adsnote: ProvidedbytheSAO/NASAAstrophysicsDataSystem. 11.HouPongandLiHuZheandChan.2025.Debate-to-write:apersona-drivenmulti-agent frameworkfordiverseargumentgeneration.InLeoandApidianakiRambowOwenand Wanner,editor,Proceedingsofthe31stinternationalconferenceoncomputational linguistics,pages4689â4703,AbuDhabi,UAE.AssociationforComputational Linguistics.CitationKey:hu-etal-2025-debate. 12.SusmitJha,SumitKumarJha,PatrickLincoln,NathanielD.Bastian,AlvaroVelasquez, andSandeepNeema.2023.Dehallucinatinglargelanguagemodelsusingformalmethods guidediterativeprompting.In2023IEEEinternationalconferenceonassuredautonomy (ICAA),pages149â152.CitationKey:10207581. 13.WojciechKryĆciĆski,BryanMcCann,CaimingXiong,andRichardSocher.2019. Evaluatingthefactualconsistencyofabstractivetextsummarization.CitationKey:kryĆciĆ ski2019evaluatingfactualconsistencyabstractivearXiv:1910.12840[cs.CL]. 14.TaijiLi,HaoChen,FeiYu,andYinZhang.2025a.HERA:Improvinglongdocument summarizationusinglargelanguagemodelswithcontextpackagingandreordering. CitationKey:li2025heraimprovinglongdocumentarXiv:2502.00448[cs.CL]. 15.TaijiLi,HaoChen,FeiYu,andYinZhang.2025b.HERA:ImprovingLongDocument SummarizationusingLargeLanguageModelswithContextPackagingandReordering. arXiv:2502.00448[cs]. 16.Chin-YewLin.2004.ROUGE:apackageforautomaticevaluationofsummaries.InText summarizationbranchesout,pages74â81,Barcelona,Spain.Associationfor ComputationalLinguistics.CitationKey:lin-2004-rouge. 17.AlexanderandChenLiuYixinandFabbri.2024.Benchmarkinggenerationandevaluation capabilitiesoflargelanguagemodelsforinstructioncontrollablesummarization.InHelena andBethardDuhKevinandGomez,editor,Findingsoftheassociationforcomputational linguistics:NAACL2024,pages4481â4501,MexicoCity,Mexico.Associationfor ComputationalLinguistics.CitationKey:liu-etal-2024-benchmarking. 18.RanLiu,Xian-LingMao,andHeyanHuang.2025.DSciSum:Detailedsummarizationof longscientificdocuments.Knowledge-BasedSystems,317:113409. 19.KaijieMoandRenfenHu.2024.ExpertEase:AMulti-AgentFrameworkforGrade- SpecificDocumentSimplificationwithLargeLanguageModels.InFindingsofthe AssociationforComputationalLinguistics:EMNLP2024,pages9080â9099,Miami, Florida,USA.AssociationforComputationalLinguistics.CitationKey:mo-hu-2024- expertease. 20.LingyunShenandXiaoqiuLe.2023.AnEnhancedMethodonTransformer-BasedModel forONE2SEQKeyphraseGeneration.Electronics,12(13). 21.NathaliaNascimento,PauloAlencar,andDonaldCowan.2023.GPT-in-the-loop: Adaptivedecision-makingformultiagentsystems.CitationKey: nascimento2023gptintheloopadaptivedecisionmakingmultiagentarXiv:2308.10435 [cs.MA]. 22.OpenAI,JoshAchiam,StevenAdler,SandhiniAgarwal,LamaAhmad,IlgeAkkaya, FlorenciaLeoniAleman,DiogoAlmeida,JankoAltenschmidt,SamAltman,Shyamal Anadkat,RedAvila,IgorBabuschkin,SuchirBalaji,ValerieBalcom,PaulBaltescu, HaimingBao,MohammadBavarian,JeffBelgum,etal.2024.GPT-4technicalreport. CitationKey:openai2024gpt4technicalreportarXiv:2303.08774[cs.CL]. 23.MuruandMinPressOfirandZhang.2023.Measuringandnarrowingthecompositionality gapinlanguagemodels.InJuanandBaliBouamorHoudaandPino,editor,Findingsof 11 theassociationforcomputationallinguistics:EMNLP2023,pages5687â5711,Singapore. AssociationforComputationalLinguistics.CitationKey:press-etal-2023-measuring. 24.GauravSahu,OlgaVechtomova,andIssamH.Laradji.2025a.AGuideToEffectively LeveragingLLMsforLow-ResourceTextSummarization:DataAugmentationandSemi- supervisedApproaches.arXiv:2407.07341[cs]. 25.GauravSahu,OlgaVechtomova,andIssamH.Laradji.2025b.Aguidetoeffectively leveragingllmsforlow-resourcetextsummarization:Dataaugmentationandsemi- supervisedapproaches.CitationKey:sahu2025guideeffectivelyleveragingllmsarXiv: 2407.07341[cs.CL]. 26.DouglasC.Schmidt,JesseSpencer-Smith,QuchenFu,andJulesWhite.2023.Cataloging promptpatternstoenhancethedisciplineofpromptengineering.InCitationKey: Schmidt2023CatalogingPP. 27.SanderSchulhoff,MichaelIlie,NishantBalepur,KonstantineKahadze,AmandaLiu, ChengleiSi,YinhengLi,AayushGupta,HyoJungHan,SevienSchulhoff,PranavSandeep Dulepet,SauravVidyadhara,DayeonKi,SwetaAgrawal,ChauPham,GersonKroiz, FeileenLi,HudsonTao,AshaySrivastava,etal.2025.Thepromptreport:asystematic surveyofpromptengineeringtechniques.CitationKey: schulhoff2025promptreportsystematicsurveyarXiv:2406.06608[cs.CL]. 28.YuSun,ShuohuanWang,ShikunFeng,SiyuDing,ChaoPang,JunyuanShang,Jiaxiang Liu,XuyiChen,YanbinZhao,YuxiangLu,WeixinLiu,ZhihuaWu,WeibaoGong, JianzhongLiang,ZhizhouShang,PengSun,WeiLiu,XuanOuyang,DianhaiYu,etal. 2021.ERNIE3.0:Large-scaleknowledgeenhancedpre-trainingforlanguage understandingandgeneration.CitationKey:sun2021ernie30largescaleknowledgearXiv: 2107.02137[cs.CL]. 29.HugoTouvron,ThibautLavril,GautierIzacard,XavierMartinet,Marie-AnneLachaux, TimothĂ©eLacroix,BaptisteRoziĂšre,NamanGoyal,EricHambro,FaisalAzhar,Aurelien Rodriguez,ArmandJoulin,EdouardGrave,andGuillaumeLample.2023a.LLaMA:Open andefficientfoundationlanguagemodels.CitationKey: touvron2023llamaopenefficientfoundationarXiv:2302.13971[cs.CL]. 30.HugoTouvron,LouisMartin,KevinStone,PeterAlbert,AmjadAlmahairi,Yasmine Babaei,NikolayBashlykov,SoumyaBatra,PrajjwalBhargava,ShrutiBhosale,DanBikel, LukasBlecher,CristianCantonFerrer,MoyaChen,GuillemCucurull,DavidEsiobu,Jude Fernandes,JeremyFu,WenyinFu,etal.2023b.Llama2:Openfoundationandfine-tuned chatmodels.CitationKey:touvron2023llama2openfoundationarXiv:2307.09288[cs.CL]. 31.DavidWan,JustinChih-YaoChen,EliasStengel-Eskin,andMohitBansal.2025. MAMM-refine:arecipeforimprovingfaithfulnessingenerationwithmulti-agent collaboration.CitationKey:wan2025mammrefinerecipeimprovingfaithfulnessarXiv: 2503.15272[cs.CL]. 32.DavidWan,JustinChen,EliasStengel-Eskin,andMohitBansal.MAMM-refine:arecipe forimprovingfaithfulnessingenerationwithmulti-agentcollaboration.CitationKey: osti_10610633. 33.LeiWang,ChenMa,XueyangFeng,ZeyuZhang,HaoYang,JingsenZhang,Zhiyuan Chen,JiakaiTang,XuChen,YankaiLin,WayneXinZhao,ZheweiWei,andJirongWen. 2024a.Asurveyonlargelanguagemodelbasedautonomousagents.Frontiersof ComputerScience,18(6).CitationKey:Wang_2024. 34.TingWang,ChuanYang,MaoyangZou,JiayingLiang,DongXiang,WenjieYang, HongyangWang,andJiaLi.2024b.Astudyofextractivesummarizationoflong documentsincorporatinglocaltopicandhierarchicalinformation.ScientificReports, 14(1):10140. 12 35.ZekunMooreWang,ZhongyuanPeng,HaoranQue,JiahengLiu,WangchunshuZhou, YuhanWu,HongchengGuo,RuitongGan,ZehaoNi,JianYang,ManZhang,Zhaoxiang Zhang,WanliOuyang,KeXu,StephenW.Huang,JieFu,andJunranPeng.2024c. RoleLLM:Benchmarking,eliciting,andenhancingrole-playingabilitiesoflargelanguage models.CitationKey:wang2024rolellmbenchmarkingelicitingenhancingarXiv: 2310.00746[cs.CL]. 36.TianyiZhang,VarshaKishore,FelixWu,KilianQ.Weinberger,andYoavArtzi.2020. BERTScore:EvaluatingtextgenerationwithBERT.CitationKey: zhang2020bertscoreevaluatingtextgenerationarXiv:1904.09675[cs.CL]. 37.TianyiZhang,FaisalLadhak,EsinDurmus,PercyLiang,KathleenMcKeown,and TatsunoriB.Hashimoto.2023.Benchmarkinglargelanguagemodelsfornews summarization.CitationKey:zhang2023benchmarkinglargelanguagemodelsarXiv: 2301.13848[cs.CL]. 38.TianyiZhang,FaisalLadhak,EsinDurmus,PercyLiang,KathleenMcKeown,and TatsunoriB.Hashimoto.2024.BenchmarkingLargeLanguageModelsforNews Summarization.TransactionsoftheAssociationforComputationalLinguistics,12:39â57. 39.MingqianZheng,JiaxinPei,LajanugenLogeswaran,MoontaeLee,andDavidJurgens. 2024.Whenâahelpfulassistantâisnotreallyhelpful:Personasinsystempromptsdonot improveperformancesoflargelanguagemodels.InYaserAl-Onaizan,MohitBansal,and Yun-NungChen,editors,Findingsoftheassociationforcomputationallinguistics: EMNLP2024,pages15126â15154,Miami,Florida,USA.AssociationforComputational Linguistics.CitationKey:zheng-etal-2024-helpful. 40.YangZhongandDianeLitman.2025.Discourse-DrivenEvaluation:UnveilingFactual InconsistencyinLongDocumentSummarization.arXiv:2502.06185[cs]. 41.RongxinZhu,JeyHanLau,andJianzhongQi.FactualDialogueSummarizationvia LearningfromLargeLanguageModels.