Paper deep dive
TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought
Jianmin Li, Ying Chang, Su-Kit Tang, Yujia Liu, Yanwen Wang, Shuyuan Lin, Binkai Ou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 9:48:59 AM
Summary
The paper introduces TCM-DiffRAG, a Retrieval-Augmented Generation (RAG) framework designed for Traditional Chinese Medicine (TCM) syndrome differentiation. It addresses the limitations of standard RAG in handling complex, individualized TCM reasoning by integrating Knowledge Graphs (KG) with Chain of Thought (CoT) reasoning. The method constructs a dual-level knowledge graph: a general TCM knowledge graph derived from textbooks and a personalized knowledge graph derived from clinical cases and CoT decomposition. Experiments show significant performance improvements over native LLMs and other RAG baselines on TCM-specific datasets.
Entities (10)
Relation Signals (9)
TCM-DiffRAG → uses → Knowledge Graph
confidence 95% · TCM-DiffRAG, an innovative RAG framework that integrates knowledge graphs (KG) with chains of thought (CoT).
TCM-DiffRAG → uses → Chain-of-Thought
confidence 95% · TCM-DiffRAG, an innovative RAG framework that integrates knowledge graphs (KG) with chains of thought (CoT).
TCM-DiffRAG → evaluatedon → Jingfang-SD
confidence 90% · TCM-DiffRAG was evaluated on three distinctive TCM test datasets... Jingfang-SD
TCM-DiffRAG → evaluatedon → TCM-MCQ
confidence 90% · TCM-DiffRAG was evaluated on three distinctive TCM test datasets... TCM-MCQ
TCM-DiffRAG → evaluatedon → TCM-SD
confidence 90% · TCM-DiffRAG was evaluated on three distinctive TCM test datasets... TCM-SD
Qwen2.5-7B-Instruct → finetunedfor → TCM-DiffRAG
confidence 90% · We selected the Qwen2.5-7B-instruct model for fine-tuning... obtain a specialized model LLM cot
TCM-DiffRAG → outperforms → native LLMs
confidence 90% · TCM-DiffRAG achieved significant performance improvements over native LLMs.
TCM-DiffRAG → outperforms → SFT LLMs
confidence 90% · TCM-DiffRAG outperformed directly supervised fine-tuned (SFT) LLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Retrieval augmented generation (RAG) technology can empower large language models (LLMs) to generate more accurate, professional, and timely responses without fine tuning. However, due to the complex reasoning processes and substantial individual differences involved in traditional Chinese medicine (TCM) clinical diagnosis and treatment, traditional RAG methods often exhibit poor performance in this domain. Objective: To address the limitations of conventional RAG approaches in TCM applications, this study aims to develop an improved RAG framework tailored to the characteristics of TCM reasoning. Methods: We developed TCM-DiffRAG, an innovative RAG framework that integrates knowledge graphs (KG) with chains of thought (CoT). TCM-DiffRAG was evaluated on three distinctive TCM test datasets. Results: The experimental results demonstrated that TCM-DiffRAG achieved significant performance improvements over native LLMs. For example, the qwen-plus model achieved scores of 0.927, 0.361, and 0.038, which were significantly enhanced to 0.952, 0.788, and 0.356 with TCM-DiffRAG. The improvements were even more pronounced for non-Chinese LLMs. Additionally, TCM-DiffRAG outperformed directly supervised fine-tuned (SFT) LLMs and other benchmark RAG methods. Conclusions: TCM-DiffRAG shows that integrating structured TCM knowledge graphs with Chain of Thought based reasoning substantially improves performance in individualized diagnostic tasks. The joint use of universal and personalized knowledge graphs enables effective alignment between general knowledge and clinical reasoning. These results highlight the potential of reasoning-aware RAG frameworks for advancing LLM applications in traditional Chinese medicine.
Tags
Links
- Source: https://arxiv.org/abs/2602.22828v1
- Canonical: https://arxiv.org/abs/2602.22828v1
Trouble viewing inline? Open PDF directly →
Full Text
49,915 characters extracted from source content.
Expand or collapse full text
TCM-DiffRAG:PersonalizedSyndromeDifferentiationReasoning MethodforTraditionalChineseMedicinebasedonKnowledgeGraph andChainofThought JianminLi(1),YingChang(3,6),SU-KITTANG(1),YujiaLiu(3,6),YanwenWang(4),Shuyuan Lin(3,6)#,BinkaiOu(2,5)# #Correspondingauthor 1.FacultyofAppliedSciences,MacaoPolytechnicUniversity,R.deLuísGonzagaGomes,999078, Macao,China. 2.GuangdongInstituteofIntelligenceScienceandTechnology,Hengqin,Zhuhai,519031, Guangdong,China. 3.SchoolofBasicMedicalSciences,ZhejiangChineseMedicalUniversity,548BinwenRoad, BinjiangDistrict,Hangzhou310053,China. 4.HangzhouGanzhicaoTechnologyCo.,Ltd,10thFloor,BlockB,No.11jugongRoad,Binjang District,Hangzhou310053,China. 5.BoardWareInformationSystemLimited 6.ZhejiangChineseMedicalUniversity–GANCAODOCTORInstituteofArtificialIntelligence forChineseMedicine Abstract Background:Retrieval-augmentedgeneration(RAG)technologycanempowerlargelanguagemodels (LLMs)togeneratemoreaccurate,professional,andtimelyresponseswithoutfine-tuning.However,dueto thecomplexreasoningprocessesandsubstantialindividualdifferencesinvolvedintraditionalChinese medicine(TCM)clinicaldiagnosisandtreatment,traditionalRAGmethodsoftenexhibitpoorperformance inthisdomain. Objective:ToaddressthelimitationsofconventionalRAGapproachesinTCMapplications,thisstudy aimstodevelopanimprovedRAGframeworktailoredtothecharacteristicsofTCMreasoning. Methods:WedevelopedTCM-DiffRAG,aninnovativeRAGframeworkthatintegratesknowledgegraphs (KG)withchainsofthought(CoT).TCM-DiffRAGwasevaluatedonthreedistinctiveTCMtestdatasets. Results:TheexperimentalresultsdemonstratedthatTCM-DiffRAGachievedsignificantperformance improvementsovernativeLLMs.Forexample,theqwen-plusmodelachievedscoresof0.927,0.361,and 0.038,whichweresignificantlyenhancedto0.952,0.788,and0.356withTCM-DiffRAG.The improvementswereevenmorepronouncedfornon-ChineseLLMs.Additionally,TCM-DiffRAG outperformeddirectlysupervisedfine-tuned(SFT)LLMsandotherbenchmarkRAGmethods. Conclusions:TCM-DiffRAGshowsthatintegratingstructuredTCMknowledgegraphswithChain-of- Thought–basedreasoningsubstantiallyimprovesperformanceinindividualizeddiagnostictasks.Thejoint useofuniversalandpersonalizedknowledgegraphsenableseffectivealignmentbetweengeneral knowledgeandclinicalreasoning.Theseresultshighlightthepotentialofreasoning-awareRAG frameworksforadvancingLLMapplicationsintraditionalChinesemedicine. Keywords:TraditionalChineseMedicine;retrieval-augmentedgeneration;knowledgegraph;Multi- StepReasoning;Chain-of-Thought 1.Introduction Sincetheendof2022,largelanguagemodels(LLMs)havemaderemarkableprogress,andmedical researchershavesubsequentlyusedLLMsinvariousclinicalspecialties.Thesepreliminarystudies[1-4] haveshownthatLLMs,withtheirpowerfulsemanticunderstandingcapabilities,exhibitsignificant advantagesovertraditionaldeeplearningmodelsinmedicalnaturallanguageprocessingtasks.However, whilegeneral-purposeLLMshavesomeeffectivenessinthemedicalfield,theystillfallfarshortofexpert levels[5],andsometimestheymayproduceincorrectormisleadingcontentdueto"hallucinations,"which mayberelatedtothequalityoftrainingdata,modelarchitecture,trainingprocesses[6]. Althoughbothfine-tuning[7-14]andRAG[15-25]havebeenwidelypracticedinthemedicalfield,they stillfacemanychallengesinthefieldofTCM.Unlikemodernmedicine,thefocusofTCMdiagnosisisnot ondiseases,butonsyndromes.Asyndromeisamethodofclassifyingpathologicalsymptomsandsignsto determinethebody'sfundamentaldisorders.TCMobtainsapatient'ssyndromethroughsyndrome differentiationandtreatment.TCMsyndromedifferentiationtheoryincludestheEightPrinciples,Zang-Fu organs,meridians,Qiandblood,ortheSanjiaotheory[26].Itanalyzesthecharacteristicsofsyndromes anddiseasesbasedondifferentparameterssuchasthefunctionofZang-Fuorgans,thecombinationofYin andYangmeridians,andthecirculationofQiandblood[27].Althoughtheprinciplesofsyndrome differentiationandtreatmentarethesame,TCMcanbedividedintodifferentschools,suchastheClassical Formulaschool,theEarthschool,andtheWarmDiseaseschool,duringthespecificdiagnosisand treatmentprocess,withcertaindifferencesamongthem[28].Thissituationisreferredtoas"treating differentdiseaseswiththesamemethod"and"treatingthesamediseasewithdifferentmethods." PreviousTCMlargelanguagemodels[14,29]havedonesomeworkincontinuouspre-training,supervised fine-tuning,andreinforcementlearning.However,duetothelackofdiagnosticandtreatmentdatafrom differentTCMschools,itisdifficulttoreflectthedifferencesinclinicalpractice.Atthesametime,the RAGapproachalsofacesnumerouschallengesinclinicalscenarios.TCMclinicalproblemsinvolvealarge amountofpotentialreasoning.Forproblemswithhighcomplexityandstronglogic,embeddingmodelscan onlymatchtextwithsurfacesimilarityandcannotidentifypotentiallogicalstructures.Inaddition,the knowledgebaseinRAGoftencomesfromtextbooks,whichisrelativelytheoreticalandhasacertaingap fromrealclinicalpractice.LLMsfinditdifficulttoeffectivelyderivecoherentanswersfromfragmentedor marginallyrelevantknowledgesnippets[30]. ToaddressthechallengesofRAGinthecontextofTCMdiagnosisandtreatment,recentstudieshave proposedmoreadvancedRAGmethods.Forinstance,KnowledgeGraphRAGtransformstheknowledge baseintographstructure[31,32],andReasoning-basedRAGinitiatesmulti-stepretrievaltoensure iterativeoptimizationandenhancedreasoning[33].However,neitheroftheseRAGmethodsadequately reflectsclinicalthinking.Thecoreofclinicalthinkingliesinthedialecticalunityofcomprehensiveness andfocus:itrequiresnotonlyahigh-sensitivityanalysisofmedicalhistoryandsymptomsbutalsotargeted, efficientprobabilityassessmentanddecision-makingbasedonhypotheses[34,35]. TheKnowledgeGraphRAGmethodcanmanageknowledgehierarchically,connectingdifferentnodesand graphicalpathstoenhancethelogicofrecalledtexts.However,itdemandshigh-qualitygraphsandcannot achievedeepiterativereasoning.Reasoning-basedRAGcancontinuouslydecomposesub-questions, conductingin-depthreasoningretrieval,butthequalityofheuristicsub-queriesisdifficulttocontrol, makingitchallengingtoachievepersonalizedsub-questiondecompositionfordifferentTCMschools. Toaddresstheseissues,weproposetheTCM-DiffRAGmethod,withthefollowingmaincontributions: 1.Weproposeamethodforconstructingadual-levelknowledgegraphspecificallyfortextbooks,further improvingthequalityoftheRAGknowledgebase. 2.Bycombininggeneralknowledgegraphswithactualcases,weconstructapersonalizedknowledge graph,compensatingforthelackofmedicalreasoningdataandpersonalizeddiagnosisand treatmentthinkingintheknowledgebase. 3.WetrainaTCMthinkingchainmodel,whichsolvestheproblemsofTCMschooladaptationandthe unityofretrievalcomprehensivenessandfocusinRAGpractice. 4.WeconstructasetofevaluationdatasetstoverifytheperformanceoftheTCMthinkingchainmodel anddifferentknowledgebases. 2.Methods 2.1.RelatedWork 2.1.1.Retrieval-AugmentedGeneration RAGtechnologyrequiresadditionalmodeltrainingandcandynamicallyintegrateexternalknowledge basestooptimizemodeloutput,makingitaresearchhotspotinmedicalLLMapplications[36-38].The initialRAG,alsoknownasNaiveRAG,canbedividedintothreecomponents:knowledgebase,retriever, andlargelanguagemodel.Fig.1ashowstheworkflowofNaiveRAG,wheretheuserinputsaquestion, andtheretriever(suchastheBM25algorithmorembeddingmodel)retrievesdocumentsfromastatic dataset.Then,theretrieveddocumentsserveascontexttoenhancethegenerationcapabilitiesofthelarge languagemodel.NaiveRAGhassomevariants,suchasHydeRAG[39,40].Earlystudiesfocusedon improvingdifferentRAGcomponentstoenhancetheeffectivenessofRAG[41,42].Researchhasshown thatthepreprocessingschemeoftheknowledgebase,thematchingdegreebetweentheknowledgebase andthequestion,andthecapabilityofthelargelanguagemodelarecloselyrelatedtotheeffectivenessof RAG[17,43].Wehavenoticedthatinsometestsets,theperformanceofstrongermodelsactually decreasesafteraddingRAG[44,45],whichalsooccursinourresearch.Thismaybebecausethelarge modelitselfhascertainmedicalcommonsense,andlow-qualitydocumentrecallcannegativelyimpactthe capabilitiesofthelargemodel. 2.1.2.KnowledgeGraphandRAG Knowledgegraphs,asstructuredknowledgebasesforrepresentingandlinkingentitiesandtheir relationshipsintherealworld,canplayrolessuchasevidence-basedsupport,linkprediction,andflexible multi-hopqueriesinmedicalscenarios[46].BeforetheriseofLLMs,therewererelatedstudiesonthe combinationofknowledgegraphsandmedicalquestion-answeringsystems[47-51].AftertheriseofLLMs, knowledgegraphs,duetotheirstructuredsemanticassociations,provideanewparadigmfordeep reasoningofmedicaldataandareoftencombinedwithRAGasaknowledgebasecomponent.Asshownin Fig.1b,comparedtoNaiveRAG,KnowledgeGraphRAGcancapturetherelationshipsbetweenentitiesin thequestionandrecalldocumentswithtightersemanticlogicfromtheknowledgebase.Inaddition, knowledgegraphscanperformhierarchicalknowledgemanagement,improvingthegranularityofretrieval [31,33,52,53].DespitethepracticaleffectivenessofKnowledgeGraphRAG,itstillfacessignificant challengesinconstructinghigh-qualityknowledgegraphs[54],balancingthesizeoftheretrievalsubgraph withcomputationaloverhead[55],andthelackofmulti-stepreasoning[56]. 2.1.3.ReasoningRAG NaiveRAGcannotdecomposecomplexquestionsstepbystepandsearchforrelevantinformation.Some studieshaveproposedreasoning-basedRAG.Self-BioRAG[57]hastheabilitytoreflect,judgewhether thequestionneedsfurtherretrieval,whethertheretrievedparagraphsarerelevanttothequestion,and whetherthegeneratedcontentisreasonable.i-MedRAG[43]allowsLLMstoiterativelygenerate subsequentqueries,graduallydeepeningtheexplorationofinformationbasedonpreviousretrievalresults, formingan"informationsearchhistory,"whichisultimatelyusedtogeneratethefinalanswer.HiRMed[32] alsoadoptsatreestructure,performingmedicalreasoningateachtreenodeduringRAG.However,mostof thesestudiesrelyonmanuallydesignedpromptsorheuristicmethods,whichnotonlyconsumehuman resourcesbutalsolackscalabilityformorecomplexproblems[58-60].Withtheemergenceofpowerful largereasoningmodels(LRMs)suchaso1andDeepSeek-R1[61],RAGhasbeguntoentertheeraof intelligentagentsanddeepretrieval.Recentstudies[43,62-64]haveintegratedtheconceptsofintelligent agentsandreinforcementlearningintoreasoning-basedretrieval.Althoughreasoning-basedRAGcan performin-depthretrievalandreasoning,itstilllackstheclinicalthinkingofarealdoctor:theabilityto systematicallyintegratecluestobuildaholisticpictureandtheclosed-loopreasoningabilitytoselectively focusonkeycluesandverifyhypotheses. Basedonthis,weproposetheTCM-DiffRAGmethod,whichcombinestheadvantagesofKnowledge GraphRAGandreasoning-basedRAG.Bytrainingathinkingchainmodelandconstructingapersonalized knowledgegraphasaknowledgebase,weaddresstheadaptationofTCMschools,andtheunityof comprehensivenessandfocusinretrievalinpractice. Fig.1.ComparisonofseveralcommonRAGmethodswiththeTCM-DiffRAGmethod. (Fig.1a.NaiveRAG,whichgeneratesanswersafteronlyoneroundofretrieval.Commonapproaches includegeneratinganswersafterretrievalorretrievingaftergeneratingapreliminaryanswer. Fig.1b.RAGmethodsusingdifferentdatastructuresasknowledgebases.Commonapproachesinclude thosebasedonknowledgegraphsandthosebasedonthinkingtrees,whichcanimprovethegranularityof retrieval. Fig.1c.Reasoning-basedRAG.Commonapproachesincludeheuristic-basedmethodsandreinforcement learning-basedmethods.Theformerreliesonthecapabilitiesoflargemodelstograduallydecompose complexproblems,whilethelatterincorporatestheconceptsofintelligentagentsandreinforcement learningtotrainaRAGsystemthatcanautonomouslyoptimizeretrievalstrategiesandgenerationquality. Fig.1d.TCM-DiffRAG.CombiningthecharacteristicsofKnowledgeGraphRAGandreasoning-based RAG,itusesathinkingchainmodeltodecomposecomplexproblemsintotriples,whicharethenretrieved andmatchedwiththeknowledgegraph.) 2.2.Methodology 2.2.1.ConstructionofaGeneralTraditionalChineseMedicineKnowledgeGraph ThequalityofRAGiscloselyrelatedtothequalityoftheknowledgebase[17,44].However,most previousstudiesonmedicalknowledgegraphsandRAG[31,52,53]donotdelvedeeplyintohowto constructaknowledgegraph.Therefore,weproposea"macro-micro"knowledgegraphconstruction methodspecificallyformedicaltextbooks.AsshowninFig.2,wecollected580classicChinesemedicine textbooks,famousmedicalcases,andotherbooks,andusedpreviousdocumentparsingresearch[65-68] forpreprocessing.Atthemacrolevelofthebooks,weutilizedadocumentlayoutmodeltoidentifythe elementsofeachPDFpageandextractedthetitlesandcorrespondingparagraphtexts,constructinga knowledgegraphsimilartoatreediagram.Thenodesconsistofthetitlesofthebooks,andthenode relationshipsareautomaticallygeneratedthroughtheparent-childstructureofthetitles.Althoughthis structuresacrificesthetraditionalontologicalconceptofknowledgegraphs,losingsomerigor,itendows therapeuticknowledgewithnaturalsemanticindexingcapabilities.Atthemicrolevelofmedicalentities, weextractedentitiesandrelationshipsfromparagraphtextsusingalargelanguagemodel.Themacrotitle nodesserveasstructuralhubs,establishingbidirectionalmappingswithmicroentities. LetthedocumentsetDbecomposedofachapterhierarchystructureℋandacontentcollectionP: D=ℋ,P Where ℋ isatree-likechapterhierarchy(e.g., TraditionalChineseMedicineInternalMedicine→Chapter4:LungSystemDiseases→Section2:Cough ), and P isthetextcontentcorrespondingtoeachchapter.Thisstructureisconstructedusingadocument parsingmodel.ByapplyingaLLMtoD,weextractasetofmedicallogictriplets G book : G book =LLM extract D Eachtriplet t k ∈G book satisfiest k = e sub ,r,e obj ,where e sub isthesubjectentity,ristherelation,and e obj is theobjectentity.Thereexistsamany-to-manymappingbetweendocumentsandtriplets, ℳ 1 t k →d 1 ,d 2 ,...,d k , ℳ 1 d k →t 1 ,t 2 ,...,t k ,where d 1 ,d 2 ,...,d k ∈D . Fig.2.SchematicDiagramoftheGeneralKnowledgeGraphConstruction. 2.2.2.EnhancementandTransferoftheGeneralTraditionalChineseMedicineKnowledgeGraphto aPersonalizedKnowledgeGraph AlthoughwehaveconstructedageneralTCMknowledgegraphthroughthestudyofTCMclassics,there aresignificantdifferencesinclinicaldiagnosisandtreatmentamongdifferentTCMschoolsand practitioners.ThegeneralTCMknowledgegraphandexistingRAGsolutionscannoteffectivelycapture andreflectthesestylisticdifferencesinclinicalpractice.Integratingdiverseclinicalthinkingpatternsinto theretrievalandreasoningmechanismsofRAGremainsacorechallenge.Additionally,theexternal knowledgebaserequiredforcomplexclinicalquestionsisnotexplicitlyavailableandneedstobeenhanced basedonactualcasestoimprovethematchingdegreeoftheRAGknowledgebase[69-74].Drawing inspirationfromAgentHospital[75]andMedReason[76],weenhanceandtransferthegeneralTCM knowledgegraphtoobtainapersonalizedknowledgegraphbyanalyzingthediagnosisandtreatmentcases ofdoctorsfromdifferentschools.ThespecificmethodisshowninFig.3: Fig.3.TheConstructionProcessofthePersonalizedKnowledgeGraph. 2.2.2.1.ChainofThoughtDecomposition InputthegivenquestionandanswerintotheQwen2.5-72B-instructmodeltogenerateamulti-hop reasoningchainanddecomposeitintostructuredtriplets. Let Q bethesetofquestionsfromdifferentschoolsofthought,and A gold bethesetoftheirstandard answers,derivedfromthetrainingsetsofTCM-MCQ,TCM-SD,andJingfang-SDdatasets.Forq i ∈Qand a i ∈A gold ,generatethegeneralthinkingchaintriplets: G query i =LLM decomp q i ,a i Where G query i isthesetoftripletsdecomposedfromtheithquestion. 2.2.2.2.TripletMatchingandTraceability Thegeneratedtripletsarealignedwiththeentitiesandmappedtotherelationshipsinthegeneral knowledgegraphtolocatetheoriginaltextbasisintheTCMclassics.Firstly,byusingvectorsimilarity, retrievetherelevanttripletsfromthegeneralknowledgegraph G book ,andobtaintherecalledset G reca i : G recall i =argtop-k t j ∈ G book simφt j ,φG query i Where φ⋅ istheembeddingoperation,andweuseAlibabaCloud'stext-embedding-v3. sim⋅ isthe cosinesimilarity.Then,wemap ℳ 1 torecallthedocumentsintheTCMclassicscorrespondingtothe triplets: D tuple i = ⋃ t k ∈G recall i ℳ 1 t k 2.2.2.3.QuestionsandDocument Meanwhile,foragivenquestion q i ,findthekmostrelevanttextsnippetsfromthedocumentcollectionD. D snippets i = argtop-k d j ∈ D sim φd j ,φq i 2.2.2.4.GeneratingReasoningThinkingProcess UsethealignedtripletsandrelatedclassictextsascontexttodrivetheQwen2.5-72B-instructmodelto generatethecompletereasoningprocessfromquestiontoanswer. C i =LLM gen (q i ,a i ,D snippets i ∪D tuple i ) 2.2.2.5.GeneratingPersonalizedKnowledgeGraphs Extractnewentitiesandrelationshipsfromthereasoningtext,integratethemwiththeoriginalgeneral knowledgegraph,andformaknowledgegraphcontainingpersonalizedclinicaldiagnosisandtreatment features.Similarto G book , C and G personal haveamany-to-manyrelationshipwithamappingrelationship ℳ 2 betweenthem,where ℳ 2 t k →c 1 ,c 2 ,...,c k and ℳ 2 c k →t 1 ,t 2 ,...,t k ,with c 1 ,c 2 ,...,c k ∈C and t 1 ,t 2 ,...,t k ∈ G personal . G personal i =LLM extract C i Theadvantageofconstructingapersonalizedknowledgegraphliesinthedecompositionofthereasoning chainthroughtheanalysisofdoctors'actualdiagnosisandtreatmentcases(Q&A),explicitlycapturingtheir personalizedreasoninglogic(suchasschoolpreferences,emphasisonsyndromedifferentiation),and avoidingthehomogenizationdefectsofthegeneralknowledgegraph.Atthesametime,constrainedbythe generalknowledgegraph,whileintroducingpersonalizedknowledge,itstrictlyadherestothecore authorityofthetraditionalChinesemedicinetheoreticalsystem.MorespecificexamplescanbeseeninFig. 4. Fig.4.SpecificExamplesoftheConstructionofaPersonalizedKnowledgeGraph. 2.2.3.ConstructionoftheTCM-DiffRAGRAGArchitecture Aftertheconstructionofthepersonalizedknowledgegraph,weproposetheTCM-DiffRAGRetrieval- AugmentedGenerationarchitecture,whosecoreinnovationliesin:decomposingclinicalquestionsinto multi-hoptriplepathsequencesthroughthechain-of-thoughtreasoningmodel,andperformingsemantic alignmentandevidencegenerationbasedonthepersonalizedknowledgegraph.Thespecificstepsareas follows: 2.2.3.1Chain-of-ThoughtModelTraining Inthepreviousstep,foreachquestionandanswer,wegeneratedadataset C containingquestion-answer pairswithreasoningprocesses,aswellasthetripletsG style obtainedbydecomposingthereasoningprocess. Thesetwopartsofthedatacontainasubstantialamountofreasoningcontent,whichweusetoconstructa superviseddataset: D SFT =q i ,C i ⏟ Question-answerpair ∪q i ,G style i ⏟ Chainofthoughttriplet Byfine-tuningthelargemodelparameterswithdomainsupervision,weobtainaspecializedmodel LLM cot thatpossessestheabilitytoreasonabouttraditionalChinesemedicinediagnosisandtreatment: ℒ SFT θ=− q i ,y i ∈D SFT logP θ y i ∣q i WeselectedtheQwen2.5-7B-instructmodelforfine-tuning.Thespecifictrainingequipmentandparameter settingsareasfollows:Weused8A80080GGPUsfortraining,conductedfullparameterfine-tuningbased ontheLLaMAFactoryframework,withabatchsizeof2perGPU,alearningrateof1e-4,awarm-upratio of0.1,andacutoff_lensetto2046.WeusedDeepSpeedZeRO-2toacceleratetraining. 2.2.3.2Multi-hopRetrievalandKnowledgeEnhancement Chain-of-ThoughtDecomposition:Fortheinputquestion q i ,use LLM cot toparseitintoamulti-hop reasoningpath. LLM cot q i =s 1 ,r 1 ,o 1 ,...,s k ,r k ,o k =T query i PersonalizedKnowledgeRecall:Thetriplets T query i frommulti-hopreasoningarematchedwiththe personalizedknowledgegraph G style forsemanticsimilarity,recallingtherelevanttriplets T recall i . T recall i =argtop-k t j ∈G style sim φt j ,φT query i RetrievalofProvisions:ℳ 2 realizesthemappingfromtripletstoreasoningtextprovisionsC,recallingthe relevanttextsnippets C tuple i . C tuple i =⋃ t k ∈T recall i ℳ 2 t k 2.2.3.3TraceableDiagnosisandTreatmentDecisionMaking Forcomplexmedicalquestionsinputbyusers,thelargemodelgeneratesenhancedresponsesbasedonthe recalledpersonalizedknowledgegraphanditsassociatedclassictexts.Thankstothedeepgraphtraversal capabilityofthegraph(supportingmulti-hopreasoning)andtheimplicitassociativityandscalability,the recalledcontentensuresboththebreadthofstructuredknowledgeandthedepthoftraceablereasoning, therebyintegratingthedualadvantagesofbothknowledgegraph-basedRAGandreasoning-basedRAG. y i = LLM gen q i ,T recall i ⏟ Reasoningpath ,C tuple i ⏟ Textsource 3.Results 3.1.DatasetandKnowledgeGraphConstruction Wedividethecorpusdatasetintofourtypes:TCMbooks,TCM-MCQ,TCM-SD,andJingfang-SD,as showninTable1. TheTCMBooksCorpusconstitutesthefoundationalGeneralTCMKnowledgeGraph.Thiscorpus includes580TCMworks,whicharedividedintomacroandmicroknowledgegraphsthroughspecific methods.Thetextsnippetsinthecorpus(totaling433,950)correspondtodifferenthierarchicaltitlesinthe books,formingthenodesofthemacroknowledgegraph.Eachtextsnippet(withanaveragelengthof about330tokens)representsthespecifictextcontentunderacertainhierarchicaltitle.Weuselarge languagemodelstoextracttripleknowledgefromeachtextsnippet,withanaverageof8triplesextracted persnippet. Theremainingthreecorpora(TCM-MCQ,TCM-SD,Jingfang-SD)representdifferentdifficultylevelsfor Retrieval-AugmentedGeneration(RAG)taskevaluationbenchmarks: a.TCM-MCQCorpus:FocusesontestingthemasteryofgeneralTCMknowledge.Thisdatasetis derivedfromTCMmedicalexaminationquestionbooks,andthetaskrequiresselectingtheonly correctanswerfromfiveoptions.Mostoftheanswerinformationcanbedirectlyretrievedfrom theTCMbookscorpus,makingthisdatasetrepresentthelowesttaskdifficulty. b.TCM-SDCorpus:Thisopen-sourcedatasetoriginatesfromtherealmedicalrecordsofXuzhou HospitalofTCM[26].Thetaskistodeterminetheonlycorrectanswerfrom148candidate syndromes.ComparedtoTCM-MCQ,theRAGdifficultyofthisdatasetissignificantly increased,asanswersaretypicallynotdirectlyobtainablefromthebookscorpusandrequire reasoningbasedontherecalledbooksnippets. c.Jingfang-SDCorpus:ThisprivatedatasetcomesfromtheoutpatientcasesoftheSecondAffiliated HospitalandtheThirdAffiliatedHospitalofZhejiangChineseMedicalUniversity.Thetask requiresselectingtheonlycorrectanswerfrom42candidatesyndromes.Itssyndrome differentiationapproachdiffersfromTCM-SD(basedongeneralprinciplesliketheEight Principles,Zang-Fuorgans,meridians,Qi,blood,bodyfluids,orSanjiao)andisrootedinthe TCMClassicalFormulaschool,withdistinctschoolcharacteristics.Duetotherelative scarcityofliteraturerecordingsuchcharacteristicsyndromedifferentiationexperiences,onlya smallnumberofrelevanttextsnippetsareavailableforrecallinthebookscorpus,posingthe greatestchallengetotheRAGsystem. Theabovethreeevaluationbenchmarks(TCM-MCQ,TCM-SD,Jingfang-SD)aredividedintotrainingsets andtestsets.Thetrainingsetsareusedtoconstructthepersonalizedknowledgegraph,whilethetestsets areusedtoevaluatetheeffectivenessoftheRAGsystem. Table1CompositionoftheCorpusDatabase. CorpusSnippetsAverageTokensAverageTriples TCMBooks4339503308 TCM-MCQ_TrainingSet2166010312 TCM-MCQ_TestSet600117/ TCM-SD_TrainingSet4308540916 TCM-SD_TestSet5486416/ Jingfang-SD_TrainingSet2004919416 CorpusSnippetsAverageTokensAverageTriples Jingfang-SD_TestSet5012194/ 3.2.EvaluationoftheGeneralTraditionalChineseMedicineKnowledgeGraphEffectiveness Thequalityofthecorpus,preprocessing,andthemethodofgraphconstructionhaveasignificantimpacton RAGperformance.GiventhatthegeneralTraditionalChineseMedicine(TCM)knowledgegraphisthe foundationalpremiseforconstructingthepersonalizedknowledgegraph,itiscrucialtoevaluatethe advancementofthisgeneralknowledgegraphconstructionmethodthroughexperiments.Forthispurpose, wechoseRAGAS[77]astheevaluationframework.Thisisatooldesignedtoautomatetheevaluationof RAGsystemeffectiveness.AsshowninTable2,thisstudyusesOpenAI'sgpt-3.5-turbo-16kasthebase largelanguagemodel(LLM),thetestsetisselectedfromtheTCM-MCQtestset,AlibabaCloud'stext- embedding-v3isusedastheembeddingmodel,andthenumberofrecalleddocumentsissettok=20.A systematicevaluationwasconductedondifferentcorpusprocessingmethods.Theexperimentcompared thefollowingfourmethods: a.WithoutRAG:RelyingsolelyontheLLM'sownknowledgetogeneratepredictedanswers(without retrievalenhancement). b.FixedCharacterSegmentation:TheTCMbookscorpustextissegmentedintofixedlengths (approximately330tokens),resultingin435,316documentsegments.Theinputquestionsand documentsegmentsarerecalledbasedonsemanticsimilarityafterembeddingprocessing. c.MacroKnowledgeGraphSegmentation:Segmentationisperformedbasedontheoriginaltitle hierarchyofthebooks,generating433,950documentsegments.Thismethodtypicallyproduces segmentsthatincludetitlesandtheircorrespondingtextcontent,offeringbettersemantic completeness.Documentretrievaliscalculatedbasedoncosinesimilarity: D snippets i = argtop-k d j ∈ D sim φd j ,φq i ,where φ representstheembeddingfunction. d.MicroKnowledgeGraphSegmentation: 1.First,performsemanticmatchingretrievalofquestionsinthetripletsetofthegeneralTCM knowledgegraph G book :G recall i = argtop-k t j ∈ G book sim φt j ,φq i . 2.Subsequently,retrievetheoriginaltextsnippetscorrespondingtothematchedtripletsbased onapredefinedtriplet-textsnippetmapping ℳ 1 : D tuple i =⋃ t k ∈G recall i ℳ 1 t k . e.Macro-MicroKnowledgeGraphIntegratedRetrieval:Combinethemacro-levelsnippetretrieval D snippets i frommethod(3)withthemicro-leveltriplet-associatedsnippetretrievalD tuple i from method(4),andtaketheirunionasthefinalretrieveddocumentset: D final i = D snippets i ∪D tuple i . Table2RAGASEvaluationoftheGeneralTraditionalChineseMedicineKnowledgeGraphontheTCM- MCQTestSet. Accuracy Answer Similarity Context Precision Context Recall Context EntityRecall withoutRAG0.4030.786\\\ FixedCharacter Segmentation 0.5400.8560.6210.8290.173 Accuracy Answer Similarity Context Precision Context Recall Context EntityRecall MacroKnowledgeGraph Segmentation 0.6400.8630.8080.8480.188 MicroKnowledgeGraph Segmentation 0.6270.8710.7820.8360.192 Macro-MicroKnowledge GraphIntegratedRetrieval 0.6870.8850.8460.8870.244 Theexperimentalresultsshowthatthemacro-microknowledgegraphintegrationmethodleadsinall evaluationindicators,verifyingthesignificantadvancementoftheknowledgegraphmethodwe constructed. Currently,fixedcharactersegmentationisstillthemethodusedbymostresearch.Toensurethe comparabilityofdocumentlengthwithsubsequentknowledgegraphsegmentationmethods,thisstudy uniformlysetsthesegmentationatapproximately330characters.Althoughthismethodissimpleto implement,iteasilyleadstosemanticfragmentationandentitydisconnection.Atypicalproblemisthata largenumberoftablesintextbookscarrykeyinformation,andrelatedquestionsinthetestsetrequire completetablestoobtainanswers.Fixedcharactersegmentationoftentruncatestables,destroyingtheir structuralintegrityandsemanticcoherence. Incontrast,macroknowledgegraphsegmentationreliesontheoriginalhierarchicalstructureofbooks, effectivelyensuringtheintegrityofsemanticunits.Thismethodperformsbetterwhendealingwith questionsthatrequirecross-documentcomparison(e.g.,whatisthepreferredtreatmentplanforlowback paincausedbydampnessandheat?).However,itsretrievaldepthhaslimitations,anditseffectivenessmay bereducedwhenthequestioninvolvesknowledgepointsburiedintextdetails(e.g.,intheGuizhi Decoctionanditsmodifiedformulas,howdoestheproportionofGuizhiandShaoyaochangeaccordingto clinicalconditions?)orwhenthehierarchicalstructurecauseskeyentityinformationtobeineffective. Theadvantageofmicroknowledgegraphsegmentationisitsabilitytorelativelyaccuratelymatchentity relationships.However,itshouldbenotedthatthecurrentmethoddirectlymatchesthesimilaritybetween questionsandtriplesusingvectorsimilarity,lackinganexplicitstepforextractingkeyentitiesinthe question,whichresultsinitsoverallperformancebeingslightlyinferiortothatofthemacroknowledge graphsegmentation,showingsomeimprovementonlyintheContextEntityRecallindicator. Themacro-microknowledgegraphintegrationmethodcombinestheadvantagesofboth:themacrolevel retainsthelogicalframeworkandsemanticintegrityofsyndromedifferentiationandtreatment,whilethe microtriplespreciselylockthecoreknowledgeentities.Thissynergisticeffectenablesittoachievethe bestcomprehensiveperformanceamongthefourRAGmethods. 3.3.EvaluationoftheThinkingChainModel'sEffectiveness UsingthedatasetsCand G style ,whichcontainthethinkingprocessbehindquestions,wefine-tunedthe qwen-2.5-7B-instructmodeltoobtainLLM-cot-7B.ToevaluateLLM-cot-7B'sabilitytodecompose questions,weselecteddeepseek-r1asthereferencemodeltoassessthequalityofthequestion-related triplesgeneratedbyqwen-2.5-7B-instructandLLM-cot-7B.WeusedtheLikertScaleastheevaluation metric,scoringthetriplesgeneratedbythetwomodelsonascalefrom0to5.Asshowninthefigure,the qualityofthetriplesgeneratedbythefine-tunedLLM-cot-7Bissuperiortothatofqwen-2.5-7B-instruct(P <0.05). Inadditiontohavingbetterquestiondecompositioncapabilities,LLM-cot-7Balsohastheabilitytodirectly answerquestions.TheresultsontheTCM-SDtestsetshowthattheeffectivenessofLLM-cot-7B(0.74)is significantlybetterthanthatofpreviousstudies(0.52).Thisconfirmsthatfine-tuningLLMswithdataon thethinkingprocess,inadditiontothefinalanswerdata,canachievebetterresults.Furtherablation experimentsindicatethatTCM-DiffRAG,usingLLM-cot-7Basthethinkingchainmodelincombination withapersonalizedknowledgegraph,outperformstheuseofLLM-cot-7Baloneinallthreetestsets. Fig.5.ComparisonoftheQualityofTripletDecompositionbyDifferentThinkingChainModels. 3.4.EvaluationofTCM-DiffRAG'sAblationExperiment AsseeninFigures6,7,and8,whenrelyingsolelyonthemodels'owncapabilities,qwen-plusand deepseek-r1,whichareprimarilytrainedonChineselanguagedatasets,significantlyoutperformgpt-4o- miniandgemini-2.5-flash-previewontheTCM-MCQandTCM-SDtasks.However,ontheJingfangjing- SDtestset,allfourLLMsperformpoorly,indicatingthatnoneofthemhavelearnedthistypeof personalizedsyndromedifferentiationthinkingduringtraining.Basedontheperformanceofthelarge modelsalone,thethreetestsetsrepresentthreedifferentlevelsofdifficulty.TheperformanceoftheLLMs issignificantlyimprovedwhentheTCM-DiffRAGmethod,basedonStylized-KGandLLM-cot-7B,is added. IntheTCM-MCQtestset(Fig.6),bothgpt-4o-miniandgemini-2.5-flash-previewachievesignificant improvementswithdifferentRAGmethods,whileqwen-plusanddeepseek-r1onlyseeimprovements whenusingtheTCM-DiffRAGmethodwithLLM-cot-7Bandapersonalizedknowledgegraph.Thismay bebecauseqwen-plusanddeepseek-r1alreadypossessexcellentgeneralTCMknowledgecapabilities,and ordinaryRAGmethodsintroducenoisyrecalls,leadingtonegativeeffects.OnlybyaddingLLM-cot-7Bas thethinkingchainmodelandusingamorespecializedpersonalizedknowledgegraphastheknowledge basecanthesemodelsachievecertainimprovements. ThedifficultyoftheTCM-SDdatasetishigher(Fig.7),andordinaryRAGmethodsarenoteffectivein improvingtheperformanceofLLMs.SignificantimprovementsareobservedonlyafterapplyingtheTCM- DiffRAGmethod.WhenusingLLM-cot-7Basthethinkingchaingenerationmodel,thereisanotable performanceimprovementcomparedtotheoriginalmodel.WebelievethisisbecauseTCM-SD,asa clinicalpracticetestset,requiresahighlevelofreasoningabilityfromthemodel.LLM-cot-7Bcan decomposeinputqueriesintofiner-grainedtriplesandreasonthroughthem.Theseinterconnectedtriples formaknowledgegraphstructurewithaclinicalthinkingchain.Thismethodbalancesthebreadthof informationretrievalofRAGwiththedeepreasoningcapabilitiesofthethinkingchain. Jingfang-SDisthemostchallengingamongthethreetestsets(Fig.8).WithoutusingRAG,thefour generationmodelscanonlyachieveanaccuracyof0.03-0.07,andthereisnosignificantimprovementeven withordinaryRAGmethods.However,whenusingLLM-cot-7Bincombinationwithapersonalized knowledgegraph,theaccuracyisincreasedto0.35-0.38.Thepossiblereasonsforthis,inadditiontothe LLMs'lackofreasoningabilityinclassicalformuladiagnosisandtreatment,maybethegeneralknowledge graph'slackofclassicalformuladiagnosisandtreatmentdata.Moresignificantimprovementscanbe achievedafterusingapersonalizedknowledgegraph. Fig.6.PerformanceofDifferentRAGMethodsontheTCM-MCQTestSet. Fig.7.PerformanceofDifferentRAGMethodsontheTCM-SDTestSet. Fig.8.PerformanceofDifferentRAGMethodsontheJingfang-SDTestSet. 4.Discussion ThisstudyproposesaRAGmethodnamedTCM-DiffRAG,whosecoreinnovationliesintheintroduction ofastructuredknowledgebaseconstructionworkflowandatrainingmethodbasedonChain-of-Thought (CoT). Intermsofknowledgebaseconstruction,wefullyutilizethechapterhierarchystructureoftraditional Chinesemedicine(TCM)bookstosegmentdocuments,therebyconstructingauniversalTCMknowledge graph.Comparedtotraditionalfixed-lengthcharactersegmentationmethods,thissemanticunit-based preprocessingapproachhassignificantadvantages. Tobridgethegapbetweengeneralknowledgeandspecificclinicalpractice,consideringthedistinct individualizedcharacteristicsofTCMdiagnosticthinking,wefurtherbuildapersonalizedknowledge graphanditsaccompanyingCoTmodelbasedontheaforementioneduniversalknowledgegraph.ThisCoT modelcanautomaticallydecomposeinputquestionsintomulti-hopquerytriplesequencesthatconformto aspecificstyleandretrievesubgraphsaccordingly.TCM-DiffRAGcombinesthebroadretrievalrange advantageoftraditionalknowledgegraphRAGswiththedeepreasoningadvantageofreasoning-based RAGs. AblationtestresultsdemonstratethatTCM-DiffRAG,combinedwithapersonalizedknowledgegraphand CoTmodel,achievesoptimalperformanceonthreebenchmarkdatasets. Ourresearchconclusionsaregeneralizableandcanbeextendedtoothernon-medicalfields: a.TCM-MCQ(SimpleDomainKnowledgeQuestionAnswering):Largelanguagemodels(LLMs) alreadyperformwellontheirowninsuchtasks,andgeneralRAGmethodsofferlimitedimprovement andmayevenhaveanegativeeffect. b.TCM-SD(DomainGeneralInferenceQuestionAnswering):WhileLLMshavemasteredgeneral industryknowledge,theirabilitytohandlecomplexreasoningproblemsremainsinsufficient.The introductionofaCoTmodelcansignificantlyenhancetheperformanceofLLMs. c.Jingfang-SD(Domain-SpecificInferenceQuestionAnswering):LLMslacktrainingonsuch specificprivatedataandperformpoorly.Thisisatypicalapplicationscenarioformostenterprises, whichrequirereasoningfrominternalbusinessdata.Inthisscenario,TCM-DiffRAG,combinedwitha CoTmodelandpersonalizedknowledgegraph,canbringsignificantperformanceimprovements. However,thisstudyalsohascertainlimitations,suchas:1.Theevaluationdimensionsoftheknowledge graphneedfurtherenrichmentandimprovement;2.TheperformanceofTCM-DiffRAGundersmall sampledataconditionsrequiresfurtherresearch;3.Thedifferentialevaluationofdifferentsizesand basesofLLMsasCoTmodelsremainstobeexplored.Thesewillbeimportantdirectionsforfuture research. Dataavailability Thedatasetsusedand/oranalyzedduringthecurrentstudyareavailablefromthecorrespondingauthor uponreasonablerequest. Reference 1.WilhelmTI,RoosJ,KaczmarczykR.Largelanguagemodelsfortherapyrecommendationsacross3 clinicalspecialties:comparativestudy.JMedInternetRes2023;25:e49324.[doi:10.2196/49324] 2.BuschF,HoffmannL,RuegerC,vanDijkEH,KaderR,Ortiz-PradoE,etal.Currentapplicationsand challengesinlargelanguagemodelsforpatientcare:asystematicreview.CommunMed2025;5(1):26. [doi:10.1038/s43856-024-00717-2] 3.ZhangK,MengX,YanX,JiJ,LiuJ,XuH,etal.Revolutionizinghealthcare:thetransformative impactoflargelanguagemodelsinmedicine.JMedInternetRes2025;27:e59069. [doi:10.2196/59069] 4.DengL,WangT,Yangzhang,ZhaiZ,TaoW,LiJ,etal.Evaluationoflargelanguagemodelsinbreast cancerclinicalscenarios:acomparativeanalysisbasedonchatgpt-3.5,chatgpt-4.0,andclaude2.IntJ Surg2024;110(4):1941-1950.[doi:10.1097/JS9.0000000000001066] 5.HagerP,JungmannF,HollandR,BhagatK,HubrechtI,KnauerM,etal.Evaluationandmitigationof thelimitationsoflargelanguagemodelsinclinicaldecision-making2024;30(9):2613-2622. [doi:10.1038/s41591-024-03097-1] 6.JiZ,LeeN,FrieskeR,YuT,SuD,XuY,etal.Surveyofhallucinationinnaturallanguagegeneration. AcmComputSurv2023;55(12):248.[doi:10.1145/3571730] 7.ChenY,WangZ,XingX,ZhengH,XuZ,FangK,etal.Bianque:balancingthequestioningand suggestionabilityofhealthllmswithmulti-turnhealthconversationspolishedbychatgpt.ArxivE- Prints2023:2310-15896.[doi:10.48550/arXiv.2310.15896] 8.LiY,LiZ,ZhangK,DanR,JiangS,ZhangY.Chatdoctor:amedicalchatmodelfine-tunedonalarge languagemodelmeta-ai(llama)usingmedicaldomainknowledge.Cureus2023;15(6):e40895. [doi:10.7759/cureus.40895] 9.SinghalK,AziziS,TuT,MahdaviSS,WeiJ,ChungHW,etal.Largelanguagemodelsencodeclinical knowledge.Nature2023;620(7972):172-180.[doi:10.1038/s41586-023-06291-2] 10.SinghalK,TuT,GottweisJ,SayresR,WulczynE,AminM,etal.Towardexpert-levelmedical questionansweringwithlargelanguagemodels.NatMed2025;31(3):943-950.[doi:10.1038/s41591- 024-03423-7] 11.WangH,ZhaoS,QiangZ,LiZ,LiuC,XiN,etal.Knowledge-tuninglargelanguagemodelswith structuredmedicalknowledgebasesfortrustworthyresponsegenerationinchinese.AcmTrans KnowlDiscovData2025;19(2):53.[doi:10.1145/3686807] 12.YangS,ZhaoH,ZhuS,ZhouG,XuH,JiaY,etal.Zhongjing:enhancingthechinesemedical capabilitiesoflargelanguagemodelthroughexpertfeedbackandreal-worldmulti-turndialogue. ProceedingsoftheThirty-EighthAAAIConferenceonArtificialIntelligenceandThirty-Sixth ConferenceonInnovativeApplicationsofArtificialIntelligenceandFourteenthSymposiumon EducationalAdvancesinArtificialIntelligence:AAAIPress;2024.p.2159. 13.YeQ,LiuJ,ChongD,ZhouP,HuaY,LiuF,etal.Qilin-med:multi-stageknowledgeinjection advancedmedicallargelanguagemodel.ArxivE-Prints2023:2310-9089. [doi:10.48550/arXiv.2310.09089] 14.JiaY,JiX,WangX,ZhangH,MengZ,ZhangJ,etal.Qibo:alargelanguagemodelfortraditional chinesemedicine.ExpertSystAppl2025;284:127672. [doi:https://doi.org/10.1016/j.eswa.2025.127672] 15.XuR,HongY,ZhangF,XuH.Evaluationoftheintegrationofretrieval-augmentedgenerationinlarge languagemodelforbreastcancernursingcareresponses.SciRep2024;14(1):30794. [doi:10.1038/s41598-024-81052-3] 16.GeJ,SunS,OwensJ,GalvezV,GologorskayaO,LaiJC,etal.Developmentofaliverdisease-specific largelanguagemodelchatinterfaceusingretrievalaugmentedgeneration.Hepatology 2024(80(5)):1158-1168.[doi:10.1101/2023.11.10.23298364] 17.KresevicS,GiuffrèM,AjcevicM,AccardoA,CrocèLS,ShungDL.Optimizationofhepatological clinicalguidelinesinterpretationbylargelanguagemodels:aretrievalaugmentedgeneration-based framework.NpjDigitMed2024;7(1):102.[doi:10.1038/s41746-024-01091-y] 18.LiuS,McCoyAB,WrightA.Improvinglargelanguagemodelapplicationsinbiomedicinewith retrieval-augmentedgeneration:asystematicreview,meta-analysis,andclinicaldevelopment guidelines.JournaloftheAmericanMedicalInformaticsAssociationjournaloftheAmericanMedical InformaticsAssociation2025;32(4):605-615.[doi:10.1093/jamia/ocaf008] 19.MalikS,KharelH,DahiyaDS,AliH,BlaneyH,SinghA,etal.Assessingchatgpt4withandwithout retrieval-augmentedgenerationinanticoagulationmanagementforgastrointestinalprocedures.Ann Gastroenterol2024;37(5):514-526.[doi:10.20524/aog.2024.0907] 20.MiaoJ,ThongprayoonC,SuppadungsukS,GarciaVO,CheungpasitpornW.Integratingretrieval- augmentedgenerationwithlargelanguagemodelsinnephrology:advancingpracticalapplications. Medicina(Kaunas)2024;60(3).[doi:10.3390/medicina60030445] 21.RauS,RauA,NattenmüllerJ,FinkA,BambergF,ReisertM,etal.Aretrieval-augmentedchatbot basedongpt-4providesappropriatedifferentialdiagnosisingastrointestinalradiology:aproofof conceptstudy.EurRadiolExp2024;8(1):60.[doi:10.1186/s41747-024-00457-x] 22.WangD,LiangJ,YeJ,LiJ,LiJ,ZhangQ,etal.Enhancementoftheperformanceoflargelanguage modelsindiabeteseducationthroughretrieval-augmentedgeneration:comparativestudy.JMed InternetRes2024;26:e58041.[doi:10.2196/58041] 23.ZakkaC,ShadR,ChaurasiaA,DalalAR,KimJL,MoorM,etal.Almanac-retrieval-augmented languagemodelsforclinicalmedicine.NejmAi2024;1(2).[doi:10.1056/aioa2300068] 24.ZelinC,ChungWK,JeanneM,ZhangG,WengC.Rarediseasediagnosisusingknowledgeguided retrievalaugmentationforchatgpt.JBiomedInform2024;157:104702. [doi:10.1016/j.jbi.2024.104702] 25.WooJJ,YangAJ,OlsenRJ,HasanSS,NawabiDH,NwachukwuBU,etal.Customlargelanguage modelsimproveaccuracy:comparingretrievalaugmentedgenerationandartificialintelligenceagents tononcustommodelsforevidence-basedmedicine.Arthroscopy2025;41(3):565-573. [doi:10.1016/j.arthro.2024.10.042] 26.RenM,HuangH,ZhouY,CaoQ,BuY,GaoY.Tcm-sd:abenchmarkforprobingsyndrome differentiationvianaturallanguageprocessing.ChineseComputationalLinguistics:21stChina NationalConference,CCL2022,Nanchang,China,October14–16,2022,Proceedings;Nanchang, China:Springer-Verlag;2022.p.247-263. 27.CaoH,BourchierS,LiuJ.Doessyndromedifferentiationmatter?Ameta-analysisofrandomized controlledtrialsincochranereviewsofacupuncture.MedAcupunct2012;24(2):68-76. [doi:10.1089/acu.2011.0846] 28.ChenK,XieY,LiuY.Profilesoftraditionalchinesemedicineschools.ChinJIntegrMed 2012;18(7):534-538.[doi:10.1007/s11655-012-1147-2] 29.ZhangH,ChenJ,JiangF,YuF,ChenZ,ChenG,etal.Huatuogpt,towardstaminglanguagemodelto beadoctor.FindingsoftheAssociationforComputationalLinguistics:EMNLP2023;1990; Singapore:AssociationforComputationalLinguistics;2023.p.10859-10885. 30.ZhaoS,YangY,WangZ,HeZ,QiuLK,QiuL.Retrievalaugmentedgeneration(rag)andbeyond:a comprehensivesurveyonhowtomakeyourllmsuseexternaldatamorewisely.ArxivE-Prints 2024:2409-14924.[doi:10.48550/arXiv.2409.14924] 31.WuJ,ZhuJ,QiY,ChenJ,XuM,MenolascinaF,etal.Medicalgraphrag:evidence-basedmedical largelanguagemodelviagraphretrieval-augmentedgeneration.Proceedingsofthe63rdAnnual MeetingoftheAssociationforComputationalLinguistics;Vienna,Austria:Associationfor ComputationalLinguistics;2025.p.28443-28467. 32.YangY,HuangC.Tree-basedrag-agentrecommendationsystem:acasestudyinmedicaltestdata. ArxivE-Prints2025:2501-2727.[doi:10.48550/arXiv.2501.02727] 33.SinghA,EhteshamA,KumarS,TalaeiKhoeiT.Agenticretrieval-augmentedgeneration:asurveyon agenticrag.ArxivE-Prints2025:2501-9136.[doi:10.48550/arXiv.2501.09136] 34.Joplin-GonzalesP,RoundsL.Theessentialelementsoftheclinicalreasoningprocess.NurseEduc 2022;47(6):E145-E149.[doi:10.1097/NNE.0000000000001202] 35.YazdaniS,HoseiniAM.Fivedecadesofresearchandtheorizationonclinicalreasoning:acritical review.AdvMedEducPract2019;10:703-716.[doi:10.2147/AMEP.S213492] 36.LewisP,PerezE,PiktusA,PetroniF,KarpukhinV,GoyalN,etal.Retrieval-augmentedgenerationfor knowledge-intensivenlptasks.Proceedingsofthe34thInternationalConferenceonNeural InformationProcessingSystems;Vancouver,BC,Canada:CurranAssociatesInc.;2020.p.793. 37.GuuK,LeeK,TungZ,PasupatP,ChangM.Realm:retrieval-augmentedlanguagemodelpre-training. Proceedingsofthe37thInternationalConferenceonMachineLearning:JMLR.org;2020.p.368. 38.JiangX,FangY,QiuR,ZhangH,XuY,ChenH,etal.Tc–rag:turing–completerag’scasestudyon medicalllmsystems.Proceedingsofthe63rdAnnualMeetingoftheAssociationforComputational Linguistics;Vienna,Austria:AssociationforComputationalLinguistics;2025.p.11400-11426. 39.GaoL,MaX,LinJ,CallanJ.Precisezero-shotdenseretrievalwithoutrelevancelabels.Proceedingsof the61stAnnualMeetingoftheAssociationforComputationalLinguistics(Volume1:LongPapers); Toronto,Canada:AssociationforComputationalLinguistics;2023.p.1762-1777. 40.JostmannM,WinkelmannH.Evaluationofhypotheticaldocumentandqueryembeddingsfor informationretrievalenhancementsinthecontextofdiverseuserqueries.Wirtschaftsinformatik2024 Proceedings;2024.p.115. 41.KeYH,JinL,ElangovanK,AbdullahHR,LiuN,SiaATH,etal.Retrievalaugmentedgenerationfor 10largelanguagemodelsanditsgeneralizabilityinassessingmedicalfitness.NpjDigitMed 2025;8(1):187.[doi:10.1038/s41746-025-01519-z] 42.YangQ,ZuoH,SuR,SuH,ZengT,ZhouH,etal.Dualretrievingandrankingmedicallargelanguage modelwithretrievalaugmentedgeneration.SciRep2025;15(1):18062.[doi:10.1038/s41598-025- 00724-w] 43.XiongG,JinQ,WangX,ZhangM,LuZ,ZhangA.Improvingretrieval-augmentedgenerationin medicinewithiterativefollow-upquestions.ArxivE-Prints2024:2408-2727. [doi:10.48550/arXiv.2408.00727] 44.XiongG,JinQ,LuZ,ZhangA.Benchmarkingretrieval-augmentedgenerationformedicine.Findings oftheAssociationforComputationalLinguistics:ACL2024;Bangkok,Thailand:Associationfor ComputationalLinguistics;2024.p.6233-6251. 45.VishwanathK,AlyakinA,AlberDA,LeeJV,KondziolkaD,OermannEK.Medicallargelanguage modelsareeasilydistracted.ArxivE-Prints2025:1201-2504.[doi:10.48550/arXiv.2504.01201] 46.EbeidIA.Medgraph:asemanticbiomedicalinformationretrievalframeworkusingknowledgegraph embeddingforpubmed.FrontBigData2022;5:965619.[doi:10.3389/fdata.2022.965619] 47.LiZQ,FuZX,LiWJ,FanH,LiSN,WangXM,etal.Predictionofdiabeticmacularedemausing knowledgegraph.Diagnostics(Basel)2023;13(11).[doi:10.3390/diagnostics13111858] 48.WangL,XieH,HanW,YangX,ShiL,DongJ,etal.Constructionofaknowledgegraphfordiabetes complicationsfromexpert-reviewedclinicalevidences.ComputAssistSurg(Abingdon) 2020;25(1):29-35.[doi:10.1080/24699322.2020.1850866] 49.ZhaoX,WangY,LiP,XuJ,SunY,QiuM,etal.Theconstructionofatcmknowledgegraphand applicationofpotentialknowledgediscoveryindiabetickidneydiseasebyintegratingdiagnosisand treatmentguidelinesandreal-worldclinicaldata.FrontPharmacol2023;14:1147677. [doi:10.3389/fphar.2023.1147677] 50.ZhouG,EH,KuangZ,TanL,XieX,LiJ,etal.Clinicaldecisionsupportsystemforhypertension medicationbasedonknowledgegraph.ComputMethodsProgramsBiomed2022;227:107220. [doi:10.1016/j.cmpb.2022.107220] 51.LyuK,TianY,ShangY,ZhouT,YangZ,LiuQ,etal.Causalknowledgegraphconstructionand evaluationforclinicaldecisionsupportofdiabeticnephropathy.JBiomedInform2023;139:104298. [doi:10.1016/j.jbi.2023.104298] 52.ZhaoX,LiuS,YangS,MiaoC.Medrag:enhancingretrieval-augmentedgenerationwithknowledge graph-elicitedreasoningforhealthcarecopilot.ProceedingsoftheACMonWebConference2025; SydneyNSW,Australia:AssociationforComputingMachinery;2025.p.4442-4457. 53.ChenQ,NiL.Tcmmlkg-rag:traditionalchinesemedicineintelligentdiagnosisbasedonmulti-layer knowledgegraphretrieval-augmentedgeneration.20244thInternationalConferenceonElectronic InformationEngineeringandComputerCommunication(EIECC);2024.p.958-962. 54.BuiT,TranO,NguyenP,HoB,NguyenL,BuiT,etal.Cross-dataknowledgegraphconstructionfor llm-enablededucationalquestion-answeringsystem:acasestudyathcmut.Proceedingsofthe1st ACMWorkshoponAI-PoweredQ&ASystemsforMultimedia;Phuket,Thailand:Associationfor ComputingMachinery;2024.p.36-43. 55.LiM,MiaoS,LiP.Simpleiseffective:therolesofgraphsandlargelanguagemodelsinknowledge- graph-basedretrieval-augmentedgeneration.ArxivE-Prints2024:2410-20724. [doi:10.48550/arXiv.2410.20724] 56.MatsumotoN,MoranJ,ChoiH,HernandezME,VenkatesanM,WangP,etal.Kragen:aknowledge graph-enhancedragframeworkforbiomedicalproblemsolvingusinglargelanguagemodels. Bioinformatics2024;40(6).[doi:10.1093/bioinformatics/btae353] 57.JeongM,SohnJ,SungM,KangJ.Improvingmedicalreasoningthroughretrievalandself-reflection withretrieval-augmentedlargelanguagemodels.Bioinformaticsbioinformatics 2024;40(Supplement_1):i119-i129.[doi:10.1093/bioinformatics/btae238] 58.ShaoZ,GongY,ShenY,HuangM,DuanN,ChenW.Enhancingretrieval-augmentedlargelanguage modelswithiterativeretrieval-generationsynergy.FindingsoftheAssociationforComputational Linguistics:EMNLP2023;Singapore:AssociationforComputationalLinguistics;2023.p.9248- 9274. 59.PressO,ZhangM,MinS,SchmidtL,SmithN,LewisM.Measuringandnarrowingthe compositionalitygapinlanguagemodels.FindingsoftheAssociationforComputationalLinguistics: EMNLP2023;Singapore:AssociationforComputationalLinguistics;2023.p.5687-5711. 60.TrivediH,BalasubramanianN,KhotT,SabharwalA.Interleavingretrievalwithchain-of-thought reasoningforknowledge-intensivemulti-stepquestions.Proceedingsofthe61stAnnualMeetingof theAssociationforComputationalLinguistics;Toronto,Canada:AssociationforComputational Linguistics;2023.p.10014-10037. 61.DeepSeek-AI,GuoD,YangD,ZhangH,SongJ,ZhangR,etal.Deepseek-r1:incentivizingreasoning capabilityinllmsviareinforcementlearning.ArxivE-Prints2025:2501-12948. [doi:10.48550/arXiv.2501.12948] 62.ChenM,LiT,SunH,ZhouY,ZhuC,WangH,etal.Research:learningtoreasonwithsearchforllms viareinforcementlearning.ArxivE-Prints2025:2503-19470.[doi:10.48550/arXiv.2503.19470] 63.SongH,JiangJ,MinY,ChenJ,ChenZ,ZhaoWX,etal.R1-searcher:incentivizingthesearch capabilityinllmsviareinforcementlearning.ArxivE-Prints2025:2503-5592. [doi:10.48550/arXiv.2503.05592] 64.GuanX,ZengJ,MengF,XinC,LuY,LinH,etal.Deeprag:thinkingtoretrievestepbystepforlarge languagemodels.ArxivE-Prints2025:1142-2502.[doi:10.48550/arXiv.2502.01142] 65.ShenZ,ZhangR,DellM,LeeBCG,CarlsonJ,LiW.Layoutparser:aunifiedtoolkitfordeeplearning baseddocumentimageanalysis.DocumentAnalysisandRecognition–ICDAR2021:16th InternationalConference,Lausanne,Switzerland,September5 – 10,2021,Proceedings,PartI; Lausanne,Switzerland:Springer-Verlag;2021.p.131-146. 66.LuoC,ShenY,ZhuZ,ZhengQ,YuZ,YaoC.Layoutllm:layoutinstructiontuningwithlarge languagemodelsfordocumentunderstanding.ArxivE-Prints2024:2404-5225. [doi:10.48550/arXiv.2404.05225] 67.WangB,XuC,ZhaoX,OuyangL,WuF,ZhaoZ,etal.Mineru:anopen-sourcesolutionforprecise documentcontentextraction.ArxivE-Prints2024:2409-18839.[doi:10.48550/arXiv.2409.18839] 68.BlecherL,CucurullG,ScialomT,StojnicR.Nougat:neuralopticalunderstandingforacademic documents.ArxivE-Prints2023:2308-13418.[doi:10.48550/arXiv.2308.13418] 69.ZhangT,MadaanA,GaoL,ZhengS,MishraS,YangY,etal.In-contextprinciplelearningfrom mistakes.ArxivE-Prints2024:2402-5403.[doi:10.48550/arXiv.2402.05403] 70.PangC,CaoY,DingQ,LuoP.Guidelinelearningforin-contextinformationextraction.Proceedings ofthe2023ConferenceonEmpiricalMethodsinNaturalLanguageProcessing;Singapore: AssociationforComputationalLinguistics;2023.p.15372-15389. 71.StammerW,FriedrichF,SteinmannD,BrackM,ShindoH,KerstingK.Learningbyself-explaining. ArxivE-Prints2023:2309-8395.[doi:10.48550/arXiv.2309.08395] 72.YangL,YuZ,ZhangT,CaoS,XuM,ZhangW,etal.Bufferofthoughts:thought-augmented reasoningwithlargelanguagemodels.Proceedingsofthe38thInternationalConferenceonNeural InformationProcessingSystems;Vancouver,BC,Canada:CurranAssociatesInc.;2025.p.3607. 73.ZelikmanE,WuY,MuJ,GoodmanND.Star:self-taughtreasonerbootstrappingreasoningwith reasoning.Proceedingsofthe36thInternationalConferenceonNeuralInformationProcessing Systems;NewOrleans,LA,USA:CurranAssociatesInc.;2022.p.1126. 74.SunH,JiangY,WangB,HouY,ZhangY,XieP,etal.Retrievedin-contextprinciplesfromprevious mistakes.Proceedingsofthe2024ConferenceonEmpiricalMethodsinNaturalLanguageProcessing; Miami,Florida,USA:AssociationforComputationalLinguistics;2024.p.8155-8169. 75.LiJ,WangS,ZhangM,LiW,LaiY,KangX,etal.Agenthospital:asimulacrumofhospitalwith evolvablemedicalagents.ArxivPreprintArxiv:2405.029572024.[doi:arXiv:2405.02957] 76.WuJ,DengW,LiX,LiuS,MiT,PengY,etal.Medreason:elicitingfactualmedicalreasoningsteps inllmsviaknowledgegraphs.ArxivE-Prints2025:2504-2993.[doi:10.48550/arXiv.2504.00993] 77.EsS,JamesJ,EspinosaAnkeL,SchockaertS.Ragas:automatedevaluationofretrievalaugmented generation.Proceedingsofthe18thConferenceoftheEuropeanChapteroftheAssociationfor ComputationalLinguistics:SystemDemonstrations;St.Julians,Malta:AssociationforComputational Linguistics;2024.p.150-158.