Paper deep dive
First Demonstration of Multi-Agent LLM System for Million-Scale Optical Link Management in Global Production AIDCs
Jingyi Su, Yihao Zhang, Dianxuan Fu, Leiyan Fei, Juan Wang, Mengfan Dai, Qing Liu, Xiong Wu, Yufeng Jiang, Cheng Chen, Bowen Zhang, Peilong Wang, Xi Chen, Zonglong He, Hongchen Yu, Zhicheng Ye, Weisheng Hu, Qunbi Zhuge
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/25/2026, 6:19:18 AM
Summary
The paper introduces OptiMIND, the first production-deployed LLM-powered multi-agent system for autonomous fault management in large-scale AI Data Centers (AIDCs). It manages millions of optical links using a Plan-and-Act architecture with Planner, Evaluator, Executor, and Diagnoser agents. The system utilizes Supervised Fine-Tuning (SFT) on seven months of operational data and a self-evolving memory mechanism. In field validation over ten weeks, OptiMIND achieved a 97.7% F1 score for fault prediction and reduced fault incidents by over 60%, outperforming state-of-the-art LLMs and traditional ML models.
Entities (14)
Relation Signals (13)
OptiMIND → achieves → 97.7% F1 score
confidence 98% · demonstrate a 97.7% F1 score in failure prediction
OptiMIND → comprisesagents → Diagnoser
confidence 95% · the Diagnoser for failure mode classification, fault pattern correlation, and root cause diagnosis.
OptiMIND → comprisesagents → Executor
confidence 95% · the Executor for link failure localization and remediation suggestion
OptiMIND → comprisesagents → Reflector
confidence 95% · All structured outputs are forwarded to the Reflector
OptiMIND → comprisesagents → Planner
confidence 95% · Specifically, the Planner acts as the central orchestrator.
OptiMIND → comprisesagents → Evaluator
confidence 95% · the Evaluator for fault risk assessment and prediction accuracy monitoring
OptiMIND → deployedat → Baidu
confidence 95% · Field validations across Baidu’s production AIDC networks, spanning ten weeks and millions of optical links...
OptiMIND → reduces → fault incidents
confidence 95% · reducing the number of fault events by over 60%
→ →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present the first LLM-powered multi-agent system for autonomous fault management across millions of optical links in production AIDCs. Refined via SFT and continuous memory evolution, it achieves 97.7% F1 and over 60% fault-incident reduction, outperforming SOTA LLMs on a ten-week field data evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2608.23145v1
- Canonical: https://arxiv.org/abs/2608.23145v1
Trouble viewing inline? Open PDF directly →
Full Text
13,072 characters extracted from source content.
Expand or collapse full text
FirstDemonstrationofMulti-AgentLLMSystemforMillion- ScaleOpticalLinkManagementinGlobalProductionAIDCs JingyiSu (1) ,YihaoZhang (1) ,DianxuanFu (1) ,LeiyanFei (1) ,JuanWang (2) ,MengfanDai (2) ,QingLiu (2) , XiongWu (2) ,YufengJiang (2) ,ChengChen (2) ,BowenZhang (2)* ,PeilongWang (2) ,XiChen (3) , ZonglongHe (3) ,HongchenYu (3) ,ZhichengYe (3) ,WeishengHu (1) ,andQunbiZhuge (1)* (1) StateKeyLaboratoryofPhotonicsandCommunications,SchoolofInformationScienceandElectronic Engineering,ShanghaiJiaoTongUniversity,Shanghai,200240,China,*qunbi.zhuge@sjtu.edu.cn (2) SystemsDepartment,Baidu,Beijing,China,*zhangbowen@baidu.com (3) OpticalResearchDepartment,HuaweiTechnologies,Dongguan,China. AbstractWepresentthefirstLLM-poweredmulti-agentsystemforautonomousfaultmanagement acrossmillionsofopticallinksinproductionAIDCs.RefinedviaSFTandcontinuousmemoryevolution, itachieves97.7%F1andover60%fault-incidentreduction,outperformingSOTALLMsonaten-week fielddataevaluation.©2026TheAuthor(s) Introduction Therapidadvancementsinlargelanguage models(LLMs)havedrivenAIdatacenters (AIDCs)tobecomethecoreofglobalcomputing infrastructure,wheretherequirementsfor networkstabilityandreliabilityfarexceedthose oftraditionalDCs[1].AsGPUclustersscale towardthemillion-GPUlevel,evenmillisecond- ormicrosecond-levelopticalsignaldegradations ortransientlinkflapsinhigh-speedoptical interconnectscanleadtohoursofinference latencyortraininginterruptions[2],resultingin massiveeconomiclossesandinflatedoperating expenses.Consequently,establishingarobust failuremanagementsystemhasbecomecritical toensuringthestableandefficientoperationfor AIDCnetworks[3]. Recently,data-driventechniqueshavebeen leveragedtofacilitatetroubleshootingandauto- matecertainworkflowsindatacenternetworks (DCNs)[4].However,theseapproachesstill requireengineerstomanuallyinspectmultiple metricsandlogswhenportalarmsaretriggered, therebylimitingbothflexibilityandefficiency. Concurrently,LLMsexhibitsuperiorproficiency inhandlingcomplextasksoverlarge-scaledata [5-6].IntegratingLLMagentstosafeguardoper- ationalAItraininginfrastructurewithinproduc- tionenvironmentspresentsapromisingavenue thatremainsunexploredincurrentliterature. Inthiswork,wepresentthefirstproduction networkdemonstrationofOptiMIND(Optical Multi-agentIntelligentNetworkDiagnoser),an LLM-drivenframeworkforautonomousAIDC networkfailuremanagement.OptiMINDorches- tratesacollaborativemulti-agentarchitecture foundedonadomain-adaptedLLMaugmented withapersistentmemorysystem.Tooptimize diagnosticreasoning,weestablishaself- evolvingmechanismutilizingsevenmonthsof operationaldataforsupervisedfine-tuning(SFT) andcontinuousmemoryenrichmentthrough operationallygroundedfeedback.Fieldvalid- ationsacrossBaidu’sproductionAIDCnetworks, spanningtenweeksandmillionsofopticallinks, demonstratea97.7%F1scoreinfailure prediction,whilereducingthenumberoffault eventsbyover60%,establishingareliable safeguardforlarge-scaleAIinfrastructure. LinkFailureIncidentProcessing Fig.1(a)illustratestheglobalnetworkinfras- Fig.1:OverviewoftheproductionAIDCnetworkarchitecture.(a)GlobalDCInetwork.(b)Overviewof DCNopticaltransceiverdeployment.(c)Thelinkfailureincidentprocessor. tructureinourdemonstrationthatsupportsa widerangeofAIservices,includingdistributed computingandLLMtraining/inference.As showninFig.1(b),theopticalinterconnectplat- formcomprisesseveralmilliontransceivers, covering100Gand400Gratesoverbothsingle- modeandmulti-modefibers.Per-linkperfor- mancemetricsfrombothlocal-endandremote- endtransceiversarecollectedovera14-day observationwindow,includingDDM,pre-FEC BER,SNR,andFDR.Alongsidefaulttickets recordingremediationactionsandreturnmer- chandiseauthorization(RMA)inspectionreports, thisheterogeneousoperationaldatasetcons- titutesthedatafoundationofOptiMIND. Weadoptalayeredarchitecturethatdecou- plesdeterministicanomalyscreeningfromLLM- basedreasoning.Specifically,asdepictedinFig. 1(c),alinkfailureincidentprocessornarrowsthe searchspacefromfleet-scalelinkstoasmallset ofhigh-riskcandidatesthroughthreestages:① Rule-basedDetection:rawmetricsarefilteredto identifysamplesexceedingpredefinedthresh- oldsorexhibitinganomalousdistributions;② MachineLearning(ML)-basedFaultPrediction: buildinguponourpriorFuture-GuidedLearning model[7],weperformlink-levelprioritization throughspecializedpredictionmodels;and③ DescriptionAggregationforAgents:identified anomaliesarepackagedintostructuredfault caseswithcontextualmetadata,uponwhichthe multi-agentsystemperformscausalreasoning andgeneratesactionablediagnosticconclusions. TheOptiMINDFramework TheMulti-AgentSystem OptiMINDemploysamulti-agentarchitecture, asdepictedinFig.2.ThesystemadoptsaPlan- and-Actparadigm[8]inwhicheachagentis configuredwitharole-specificprofileandacc- essesdomaincontextviaretrieval-augmented generation(RAG)fromapersistentmemory systemincorporatingOperationsandMaint- enance(O&M)knowledge,standardoperating procedures(SOPs),andoperationalfault records.Specifically,thePlanneractsasthe centralorchestrator.Itreceivesrisklinkcases fromtheupstreamdataprocessor,decomposes themintodependentsub-tasks,anddispatches eachtoappropriatespecialistagents.Three categoriesofsub-tasksaredefinedtocoverthe fullfaultmanagementlifecycle:theEvaluatorfor faultriskassessmentandpredictionaccuracy monitoring,theExecutorforlinkfailureloca- lizationandremediationsuggestion,andthe Diagnoserforfailuremodeclassification,fault patterncorrelation,androotcausediagnosis. Eachagentisharnessedwithadedicatedsetof domainskillsaccessiblethroughmodelcontext protocol(MCP)[9],andcriticaloperationsare subjecttohuman-in-the-loopreview.Allstruc- turedoutputsareforwardedtotheReflector, whichperformsmulti-dimensionalvalidation againstproductionevidenceandsynthesizes verifiedresultsintoconsolidateddiagnostic reports.Resolvedcasesarepersistedwithinthe memorysystem,progressivelyenrichingthe episodicknowledgebasetoserveasthe foundationoftheself-evolvingmechanism. TheSelf-EvolvingMechanism Tocontinuouslyenhancethediagnosticcapa- bilityofOptiMIND,wedesignaself-evolvingme- chanism,asshowninFig.2(pink).Thismech- anismcombinesdomain-adaptedSFTforcold- startinitializationwithacontinuouslyenriched memorysystem[10]forsubsequentevolution. Duringthecold-startphase,weconstructahigh- qualitySFTdatasetof~1,000productionfault cases,eachannotatedwithchain-of-thought (CoT)rationales[11].Thisdatasetisusedto fine-tuneaPolicyLLMthatissharedacrossall agentsthroughrole-specificprompts,enablingit toacquirefoundationaldiagnosticreasoning overhistoricalfaultpatterns.Toenablefurther evolutionfromliveinferenceoutcomes,we proposeamemorysystemthatautomatically Fig.2:OverviewoftheproposedOptiMINDframework:themulti-agentsystem(blue)andtheself-evolvingmechanism(pink). reviewsandpersistsresolvedcases,where eachoutcomeisvalidatedagainstproduction- verifiedmetricscoveringpredictionaccuracy, remediationcorrectness,anddiagnosticcomp- leteness.Guidedbythisoperationallygrounded feedback,theReflectorautonomouslyupdates agentskillsacrossthefullopticallinkmanage- mentlifecycle,triggersiterationofexistingfault predictionalgorithms,andincrementallyenri- chestheretrievalcorpusforsubsequentagent reasoning.Thisclosed-loopprocessestablishes OptiMINDasarobust,outcome-drivenself- evolvingsystem. Results Wecollectsevenmonthsofmulti-source telemetrydatafromBaidu'sproductionAIDC infrastructureovermillionsofopticallinks.Each identifiedfaultcaseinthisdatasetispairedwith completeremediationticketsandrootcause annotations.ThedatasetisusedforSFTon Qwen3-8B[12]andinitialknowledgebasecons- tructionforthememorysystem.Fig.3(a)illus- tratestheoverallperformanceofOptiMIND comparedwithstate-of-the-art(SOTA)general- purposeLLMbaselinesacrossfivecoreoptical linkfaultmanagementtasks.OptiMINDachie- vesthehighestaccuracyoneverytask,witha particularlynotablegaininrootcausediagnosis, surpassingDeepSeekV3.2by47.6%.Tovalid- atethecontributionofeachcorecomponent,we conductanablationstudy.AsshowninFig.3(b), theincorporationofRAGwithastaticknow- ledgebaseenhancesthemodel'scapabilityto capturecontextualsemanticassociationsand integrateexternaldomainknowledge.Introdu- cingdomain-adaptedparadigmsviaSFTequips themodelwithfoundationalreasoningcapabi- lities,yetperformanceremainsconstrained.The proposedmemoryevolutionmechanismfurther enhancesadaptabilitytounfamiliarfaultpatterns byaccumulatingverifiedoperationalexperience andrecognizedfailuremodes,therebyyielding significantimprovementsacrossalltasks. Tovalidatereal-timeinferenceperformance, weevaluateOptiMINDontenweeksoflive networkdata.Fig.3(c)comparesitsfaultpre- dictionaccuracyagainstthepreviouslydeployed MLalgorithmsacrossdifferentrisklevels. OptiMINDexhibitssignificantlyhigherprediction andlocalizationaccuracy,particularlyforhigh- riskopticallinkfailuresthathaveextensive impact.Furthermore,Fig.3(d)illustratestheper- formanceofOptiMINDthroughouttheten-week period.Itdemonstratesacontinuousself- evolvingtrajectory,withitspredictionF1-score steadilyreaching97.7%anditsfailureincident reductionrate(FIRR)progressivelyascendingto over60%—coveringalmostallportfailures exceptforfanfailures,powersupplyfailures, anddevicecrasherrors. Conclusions Wepresentthefirstproductionnetworkdemon- strationofanLLM-poweredmulti-agentsystem forautonomousmanagementofmillionsofopti- callinksinAIDCs.Thesystemisrefinedthrough SFTonaseven-monthoperationaldatasetand continuouslyenrichesitsdiagnosticmemory throughgroundedfeedback.Overaten-week periodoffieldvalidationacrossglobalDCNs, thesystemachievesastable97.7%prediction F1-score,providingarobustfoundationforthe reliableoperationoflarge-scaleAIinfrastructure. Fig.3:PerformanceofOptiMIND.(a)AccuracycomparisonwithSOTALLMbaselinesacrosscorediagnostictasks. (b)Ablationstudyonkeycomponents.(c)Faultpredictionaccuracyunderdifferentrisklevels.(d)Self-evolvingmechanism. Acknowledgements This work was supported by Shanghai Pilot Pro- gram for Basic Research-Shanghai Jiao Tong University (21TQ1400213) and National Natural Science Foundation of China (62175145). References [1] C. Xie, A. Wang, R. Lu, Q. Chen, P. Wang, Y. Bao, and C. Wang, “Optical interconnects for AI computing appli- cations,” in Optical Fiber Communication Conference (OFC) 2025, San Francisco, CA, USA, 2025, paper Tu3G.5, DOI: 10.1364/OFC.2025.Tu3G.5 [2] K. Qian, Y. Xi, J. Cao, J. Gao, Y. Xu, Y. Guan, B. Fu, X. Shi, F. Zhu, R. Miao, C. Wang, P. Wang, P. Zhang, X. Zeng, E. Ruan, Z. Yao, E. Zhai, and D. Cai, “Alibaba HPN: A data center network for large language model training,” in Proceedings of the ACM SIGCOMM 2024 Conference, Sydney, NSW, Australia, 2024, p. 691– 706, DOI: 10.1145/3651890.3672265 [3] K. Abdelli, “Prompt Once, Manage All: A Unified LLM Framework for Multi-Task Optical Link Management,” Journal of Lightwave Technology, vol. 44, no. 8, p. 2869–2879, 2026, DOI: 10.1109/JLT.2026.3660177 [4] F. Musumeci and M. Tornatore,“Failure management in optical networks with ML: A tutorial on applications, challenges, and pitfalls,” Journal of Optical Communica- tions and Networking, vol. 17, no. 8, p. C144–C155, 2025, DOI: 10.1364/JOCN.551910 [5] Z. Wang, S. Lin, G. Yan, S. Ghorbani, M. Yu, J. Zhou, N. Hu, L. Baruah, S. Peters, S. Kamath, J. Yang, and Y. Zhang, “Intent-driven network management with multi- agent LLMs: The Confucius framework,” in Proceedings of the ACM SIGCOMM 2025 Conference, Coimbra, Por- tugal, 2025, p. 347–362, DOI: 10.1145/3718958.3750537 [6] Y. Zhang, Q. Qiu, X. Liu, X. Yu, D. Fu, X. Liu, Z. Wang, H. Lin, Y. Chen, L. Yi, W. Hu, and Q. Zhuge, “AI agent for autonomous optical networks: architectures, technol- ogies, and prospects,” Journal of Optical Communica- tions and Networking, vol. 18, no. 2, p. A159–A178, 2026. DOI: 10.1364/JOCN.576017 [7] J. Su, D. Fu, Q. Qiu, X. Ding, J. Wang, C. Chen, B. Zhang, P. Wang, X. Chen, Z. He, H. Yu, Z. Feng, W. Hu, and Q. Zhuge, “Analysis and future-guided prediction for optical transceiver failures in AI data center networks,” in Optical Fiber Communication Conference (OFC) 2026, Los Angeles, CA, USA, 2026, paper W4H.2, DOI: 10.1364/OFC.2026.W4H.2 [8] L. E. Erdogan, N. Lee, S. Kim, S. Moon, H. Furuta, G. Anumanchipalli, K. Keutzer, and A. Gholami, “Plan-and- Act: Improving planning of agents for long-horizon tasks,” in Proceedings of the 42nd International Conference on Machine Learning (ICML), Vancouver, Canada, 2025, p. 15419–15462, DOI: 10.48550/arXiv.2503.09572 [9] Anthropic, “Model Context Protocol Specification,” https://modelcontextprotocol.io/specification, accessed on 12 April 2026. [10] W. Pan, S. Liu, X. Zhou, S. Zhang, W. Shi, M. Xu, and X. Jia, “M⋆: Every task deserves its own memory har- ness,” arXiv preprint, arXiv:2604.11811, 2026, DOI: 10.48550/arXiv.2604.11811 [11] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, vol. 35, New Orleans, LA, USA, 2022, p. 24824–24837, DOI: 10.48550/arXiv.2201.11903 [12] Alibaba Cloud, “Qwen3-8B,” https://qwen.ai/ blog?id=qwen3, accessed on 12 April 2026.