Paper deep dive
Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support
Shuying Chen, Sen Cui, Zhong Cao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 2:35:42 AM
Summary
Oph-Guid-RAG is a multimodal visual retrieval-augmented generation system designed for ophthalmology clinical decision support. It treats guideline pages as independent visual evidence units, bypassing traditional OCR to preserve layout information like tables and flowcharts. The system features a controllable retrieval framework with query decomposition, routing, filtering, and reranking, achieving significant performance improvements on the HealthBench dataset compared to baseline models.
Entities (5)
Relation Signals (3)
Oph-Guid-RAG â evaluatedon â HealthBench
confidence 100% ¡ We evaluate our method on HealthBench using a doctor-based scoring protocol.
Oph-Guid-RAG â useslibrary â FAISS
confidence 100% ¡ FAISS is employed as the vector indexing library for efficient similarity search.
Oph-Guid-RAG â usesmodel â ColQwen2.5
confidence 100% ¡ ColQwen2.5 is used as the multimodal retriever to encode both queries and document images
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an independent evidence unit and directly retrieve page images, preserving tables, flowcharts, and layout information. We further design a controllable retrieval framework with routing and filtering, which selectively introduces external evidence and reduces noise. The system integrates query decomposition, query rewriting, retrieval, reranking, and multimodal reasoning, and provides traceable outputs with guideline page references. We evaluate our method on HealthBench using a doctor-based scoring protocol. On the hard subset, our approach improves the overall score from 0.2969 to 0.3861 (+0.0892, +30.0%) compared to GPT-5.2, and achieves higher accuracy, improving from 0.5956 to 0.6576 (+0.0620, +10.4%). Compared to GPT-5.4, our method achieves a larger accuracy gain of +0.1289 (+24.4%). These results show that our method is more effective on challenging cases that require precise, evidence-based reasoning. Ablation studies further show that reranking, routing, and retrieval design are critical for stable performance, especially under difficult settings. Overall, we show how combining visionbased retrieval with controllable reasoning can improve evidence grounding and robustness in clinical AI applications,while pointing out that further work is needed to be more complete.
Tags
Links
- Source: https://arxiv.org/abs/2603.21925v1
- Canonical: https://arxiv.org/abs/2603.21925v1
Trouble viewing inline? Open PDF directly â
Full Text
44,608 characters extracted from source content.
Expand or collapse full text
Guideline-groundedretrieval-augmentedgenerationforophthalmic clinicaldecisionsupport https://github.com/Suey419/Oph-Guid-RAG ShuyingChen a ,UniversityofInternationalBusinessandEconomics,No.10HuixinEastStreet, ChaoyangDistrict,Beijing,P.R.China100029;SenCui b ,TsinghuaUniversity,HaidianDistrict, Beijing,P.R.China100084;ZhongCao* c ,HeidelbergInstituteofGlobalHealth,Facultyof MedicineandUniversityHospital,HeidelbergUniversity,Heidelberg,Germany69120 ABSTRACT Inthiswork,weproposeOph-Guid-RAG,amultimodalvisualRAGsystemforophthalmology clinicalquestionansweringanddecisionsupport.Wetreateachguidelinepageasanindependent evidenceunitanddirectlyretrievepageimages,preservingtables,flowcharts,andlayout information.Wefurtherdesignacontrollableretrievalframeworkwithroutingandfiltering,which selectivelyintroducesexternalevidenceandreducesnoise.Thesystemintegratesquery decomposition,queryrewriting,retrieval,reranking,andmultimodalreasoning,andprovides traceableoutputswithguidelinepagereferences. WeevaluateourmethodonHealthBenchusingadoctor-basedscoringprotocol.Onthehardsubset, ourapproachimprovestheoverallscorefrom0.2969to0.3861(+0.0892,+30.0%)comparedto GPT-5.2,andachieveshigheraccuracy,improvingfrom0.5956to0.6576(+0.0620,+10.4%). ComparedtoGPT-5.4,ourmethodachievesalargeraccuracygainof+0.1289(+24.4%).These resultsshowthatourmethodismoreeffectiveonchallengingcasesthatrequireprecise,evidence- basedreasoning.Ablationstudiesfurthershowthatreranking,routing,andretrievaldesignare criticalforstableperformance,especiallyunderdifficultsettings. Overall,weshowhowcombiningvisionbasedretrievalwithcontrollablereasoningcanimprove evidencegroundingandrobustnessinclinicalAIapplications,whilepointingoutthatfurtherworkis neededtobemorecomplete. Keywords:Ophthalmology;ClinicalDecisionSupport;RAG;MultimodalRetrieval;VisualDocumentRetrieval; HealthBench 1.INTRODUCTION Inthedomainofophthalmology,treatmentdecisionsarehighlybasedonclinicalpracticeguidelines,whichdefine treatmentpathways,drugusage,follow-upsandreferralconditions.Thecriticalityinclinicalpracticecouldresultin severeorevenfatalconsequencesduetomistakesmadeduringexecution;consequentlyclinicaldecisionmakingshould relyonaccurate,reproducibleandauditabledataforthesakeofscienceandpatientsalike.. Accordingtorecentreviewresults,ophthalmologyisrapidlyenteringanerawherebothbasicmodelandclinical verificationarerequired.Butprivacy,biasandthequalityofthegeneratedclinicalevidenceareimportantissues[1].In thisregard,thereisanintensedesirefordecision-supportthatcanbebasedonLLM/LMM;however,relyingonlyon parametricknowledgebringsintwosignificantdangers:firstly,outdatedandmissingknowledgecausesunstable reactionstowardthemostrecentevidencedbasedrecommendation;second,itcanhallucinatethresholds,grading, *zhong.cao@uni-heidelberg.de;phone:18610049108 contraindications,orworkflowforrareand/orcomplexcases.Thereal-worldclinicalvalidationofanophthalmological usecaseshowshowastructuredagentdecouplesthevisionprocessing,knowledgesearch,anddiagnosticsreasoningcan helpimprovetreatmentplans,andavoidgeneralpurposemodelfragilityonrarediseases[2]. RAGispopularlyusedtoaddressthoseproblems:itretrievesoutsideinformationandusestheretrievedevidence contextatgenerationtime,thesystemenhancestraceabilityandfact-consistency.Systematicreviewandmeta-analysis showthatRAGyieldsstatisticallysignificantgainsoverbaselinesonbiomedicaltasks,withactionableadviceonhowto deployintheclinic[3].However,vanillaRAGassumesthatwehavegoodtextextraction,chunkingandvectorization. Inreality,clinicalguidelinesthatarefilledwithrichtables,flowcharts,complexlayouts,embeddedimages,etc.,which makestheprocessoftextextractionandOCRpipelinebrittleandexpensive.Furthermore,eveniftherelevantevidenceis retrieved,modelsmightnotusethemproperlybecauseofincompletecontextsaggregation,leadingtoincomplete groundingorremaininghallucination[4].Visualdocumentretrievalthuscallsforaparadigmofdirectlyembeddingthe pageimageforretrieval,avoidingcomplicatedtextextractionandtakingfulladvantageoflayoutandgraphical information,providinganewdirectiontoretrievalsbasedonguidelines. WethusturntotheclinicallyfocusedproblemofhowtoincorporateguidelineevidenceasvisualpagesinRAGfor ophthalmicclinicalQAanddecisionsupport,andhowtodesignanagenticrouterthatcanmakeacontrollabledecision onwhentouseexternalretrievalversusdirectgeneration,whileensuringtraceability,robustness,andfaithfulness,and reducinghallucinations.Thispapermakesfourkeycontributions. First,weproposeapage-levelvisualRAGmethodforophthalmologyguidelines.Eachguidelinepageisusedasan independentevidenceunit.Wedirectlyretrievepageimagesinsteadoftext,whichavoidsOCRerrorsandkeeps importantstructuressuchastablesandflowcharts. Second,wedesigntheretrieverasancontrollableretrievalsystemthathasrouteandfilterfunctions.Itrouteswhetherto retrieveornot,andfiltersoutlessrelevantevidences.Thusitcanbalancefactualityandfluencyofresponses. Third,webuildanend-to-endmultimodalpipeline.Itincludesquerydecomposition,queryrewriting,retrieval,reranking, andanswergeneration.ThesystemalsoreturnsguidelinepageURLs,sotheoutputsaretraceable. Fourth,weevaluateoursystemonHealthBenchwithadoctor-basedscoringprotocol.Theresultsshowthatourmethod improvesperformanceonhardcases,especiallyinaccuracy,andthatcontrolledretrievalisimportantforstableand reliableoutput 2.RELATEDWORK Wecategorizepreviousresearchintothreecategories:(i)foundationmodelsandintelligentagentscenteredon ophthalmology,(i)guideline-basedretrieval-augmentedgeneration(RAG)forclinicaldecisionsupportacrossvarious specialties,and(i)retrievalandevaluationmethodologiesforbiomedicalRAGsystems. 2.1Ophthalmologyfoundationmodelsandintelligentagents MultimodalFoundationModelsandDomain-SpecificAgentsRecentworkincomputervisionforeyediseaseshavebeen focusingonmultimodalfoundationmodelsandthedomain-specificagents.Zhuangetal.(Zhuang,2025)proposesa modulerreasonertointegratevisualanalysis,retrievalanddiagnosisusingconcreteimprovementsinthepracticaluse casesforophthalmictreatment[2].Sevgi,F.A.(2024),PersonalizedInstruction-TailoredandRetrievalAugmentedGPT ModelsforEyeEducationandClinicalAssistance,forthepurposesofprivacy,accountability,andotheroperational concerns[5].Anextensiveroadmapofvision-languagefoundationmodelsforophthalmologyisprovidedin Chia(2024),observingcontinuedproblemswithdatabias,regulationandappropriateclinicalvalidation.Apartfromthese changes,makingeverythingsimpleandunderstandableisanothermajorfocusnow[1].HanWang(2025)combines knowledgegraphsandcontrastivelearningtogetherwithorganisedâclinicalprofileâmetricstoexplainhow ophthalmicAIworks.Training&Simulation[6].Luo(2025)proposesaretrievalaugmenteddigitalpatientfor improvingtakinganophthalmichistoryshowingthatRAG-basedapproachesareworthlearning[7].Grzybowski(2024) discussesAIalgorithmtoanalysethefundusimagesofocularandsystemicdisease[8]whileOngA.Y.(2025)analyses regulator-sanctionedocularimage-analysisAIaMDâsshortcomingonreportingnormsandempiricaldata[9].Also,there iscontinuousdevelopmentofdiseasespecificDSSsuchasthestagingalgorithmforkeratoconusdevelopedbyMuhsinet al.(2024)[10].. 2.2Guideline-groundedRAGforclinicaldecisionsupport Ourconcernisnotjustthatthemodelcouldgiveusananswer,butthatitgivesananswerconsistentwithwhatclinicians wouldsay.Inthisregard,Kresevic(2024)showshowsystematicguidelinereformulationtogetherwithRAGandagile prototypingimprovepreciseguidanceunderstanding,suggestingthatbothevidenceformataswellassearchstrategy matterfortrustworthiness[11].OngC.S.(2024).ProposesSurgeryLLMthatattemptstounifyworkflowandsharethe samegoal[12].Bringevidencesurgicalguidelineinperioperativeassistant,showcasingapplicationofguideline groundingtoareal-worldclinicalusecase,andKe(2025)takesthisideafurtherstillwithacomparisonofRAGacross 10largeLMsforpreoperativeriskassessment,reportedbetterdiagnosticsandlesshallucination[13].Thissupportsthe notionthatstructuredretrievalcanserveasastabilizingconstraintratherthanmerelyanoptionalenhancement.The samepatternisimportanttonote:Miao(2024)usesKDIGO-basedcorporatoapplyRAGinnephrology[14],andWada, 2025),asitallowsforrobustlocaldeploymentofradiologyconsultativeAIwhilesimultaneouslyenablingprivacy- preservinginfrastructureâimplementationconsiderationswhichhavedirectimpactonpracticality[15].Analogously, Chen(2025)combinesRAGandleastto-mostpromptingstrategyfortreatinglowbackpain[16];andZhou(2024)builds GastroBot,aChinesegastroenterology-basedchatbotbuiltuponcarefullychosenguidelinecorpora[17].LLMresponses arethereforemoreconsistent,moreclinicallyrelevantwhentheinputinformationiscollectedinasystematicmanner andtiedbacktodomainguidelines.andarelesspronetohallucination,whichmakesitincreasinglybecomeadesign principleratherthanaperformancehack. 2.3Retrievalarchitectureandevaluationmethodology Theresearchmethodologyofthefieldaimatsolvingbelowthreeproblems:accuracy,resilienceandrigourof assessment.Liuetal.(2025)placesthedevelopmentofBiomedicalRAGsystemwithinexistingclinicdeployment standardsthroughanextensivesurveyandmetaanalysis[3].Todirectlyaddressevidencealignment,Prabha(2025) introducesstructured,iterativeself-queryretrievalutilizingPICOT/SPICEframeworks,suchasincorporatingit inexistingclinicaldecisionmakingframeworks[18].Lopezetal.(2025)proposeCLEAR,amethodtoincorporate entitiesintotheretrievalprocesswithlowertokenusage,slowerresponsetimebutimprovedspanrecovery.Duringthe designofsearchers,thisapproachconsidersefficiency[19].Zhang(2025)studiesâlost-in-the-middleâproblemin medicalquestionanswering,andevaluateBriefContextreorderingsolutionsthatcaneaseoutofcontextdilutionwith longercontexts[20].Gilbert(2024)callforstrongerintegrationofknowledgegraphsasawaytoreducelossof hallucinationwhencuratingmedicalinformation[21],andGilani(2025)thatproposeCDE-Mapper:combinationof retrievalandrulebasedmappingformappingclinicalitemstoCVs[22].Secondlyfromtheperspectiveofsystem maintenance,Borchert(2025)arguesthataccurateretrievalpipesareneededtomaketimelyrevisionsinguidelines possibleandcalculatesthetimelagbetweentheappearanceofapieceofnewevidenceanditsresultingrevisionofa guidelinerelianceontime-synchronizationaspartoftheintegrity[23]. RecentworkproposessomebenchmarkstoevaluatemedicalRAGs:MIRAGEisalargescaleevaluationbenchmark whichsystematicallycomparesdifferentcombinationsofretrievers,corpora,andbackbonemodelsacrossmultiple medicalQAdatasetstohighlightthesignificanceofretrievaldesignintermsofincreasingaccuracy[24].RAGCare-QA furtherprovidesastructureddatasettoevaluatetheRAGpipelineovertheoreticalmedicalknowledgewithdifferent levelsofcomplexitiesinquestionsaswellasretrievalsrequired[25].However,thesebenchmarksmostlytargetstaticQA tasksandevaluatebasedoncorrectnessalone,whilemissinganydomain-specificmetricsortaskcharacteristicswhich arerepresentativeofclinicalpractice.inparticulartheydonotdirectlyassessdimensionssuchassafety,context awarenessoradherencetotheprocessofclinicalreasoning.Inordertoovercometheseshortcomings,weleveragethe doctor-rubricbasedbenchmark,HealthBench[26]toevaluatehealth-careconversationalagentsinclinicallyrealistic scenarios.HealthBenchevaluatesthemodelresponsefrommultipleaspectssuchasaccuracy,completeness,context awareness,andcommunicationquality,offeringamorecompleteevaluationontheclinicalutilityandsafety.Compared withpreviousQA-basedbenchmarks,HealthBenchmorecloselymatchestheneedsofrealworldCDS,whereresponses arerequiredtobenotjustrightbutalsosafe,contextualised,andclinically-relevant.. 3.METHOD 3.1Terminology Forclarity,wedefinekeytermsandabbreviationsusedthroughoutthepaper.Retrieval-AugmentedGeneration(RAG) referstotheparadigmofaugmentinglanguagemodelswithexternalknowledgeretrievedatinferencetime.Question answering(QA)denotesthetaskofgeneratinganswerstouserqueriesinaclinicalcontext.DIRECTreferstothe pathwaywherethemodelgeneratesanswerswithoutexternalretrieval.TOS(objectstorageservice)isusedtostore page-levelguidelineimagesandtheirmetadata.FAISSisemployedasthevectorindexinglibraryforefficientsimilarity search.ColQwen2.5isusedasthemultimodalretrievertoencodebothqueriesanddocumentimagesintoashared embeddingspace. Toavoidterminologicalconfusion,throughoutthispaperweconsistentlyemploypagelevelevidenceunitsandpage levelretrievalterminology.Weconsidereachindividualguidelinepageimageresultingfromapdf-to-imageconversion asouratomicevidenceunitforretrievalandcitation,andcalltheretrievalmoduleapagelevelvisualretrieverorvisual pageretriever,wheretheretrieveditemsarecalledevidencepages.Multimodalreasoninginthispaperspecificallymeans jointlyfeedingthetextualqueryandguidelinepageimageevidencetoavisionlanguagemodelforevidenceaware responsegeneration,anditexcludesautomaticdiagnosis/inferenceoverthepatientâsOCT,fundusphotograph,etc. clinicalimages.Herewerefertotheclinicaldecisionsupportasaquestion-answeringandrecommendationdraft assistantwhichcanhelpcliniciansortraineesquicklyfindoutguidelinesâevidencesandgeneratetraceableexplanation recommendations,anditâsnotamedicaldevicethatautonomouslydiagnosesorindependentlyprescribes:anyoutput shouldbeusedinconjunctionwithappropriateprofessionalsupervisionandlocalclinicalgovernance. 3.2Systemarchitectureoverview Wepresentthefour-stagemultimodalvisualRAG(Retrieval-AugmentedGeneration)systemforophthalmology questionanswering(QA).namedOph-Guid-RAGforsupportingcontrollableandtraceableguideline-basedgeneration withtheconstraintofclinicalsafety. Specifically,wefirstpreparetheoff-linecorpusbyconverting305ophthalmologyguidelinesfrompdftopagelevel images,standardizedintoauniformdimension(5390Ă7940pixel),uploadedtoTOS(objectstorageservice),recorded withpagelevelmetainformationandurlmappingrelation.TheresultedpageimagesarethenencodedbyColQwen2.5 andindexedbyFAISSasthevisionknowledgebase.Second,inthequeryprocessingstage,eachincominguserqueryis analyzedbyaPlanner,whichoptionallydecomposescomplexqueriesintouptothreefocusedsubquestions,denotedas SQ1,SQ2,andSQ3.ARoutersendseverysubquestiondownoneofaRAGoraDIRECTbranch.Subquestionspassed intoRAGgothroughaQueryRewritemodulewhichproducesoneortworewritingsmoresuitableforretrievalontopof theguidelinescorpus.SubquestionsroutedtoDIRECTbypassretrievalandareanswereddirectlybythemodel.Third, duringretrieve-and-filterstep,therewrittenqueriesareencodedbyColQwen2.5andsearchovertheFAISSindextoget thetop-kcandidatepageimagesandfilterthembyGPT-5.2,whichmeasurestheirrelevancewithrespecttoboth,the initialquestionaswellasthegeneratedsubquestion.Onlyrelevantpagesarekeptasvalidevidence.Whenthereâsnot enoughrelevantevidenceleft,itgoesbacktotheDIRECTpath.Fourth,inthegenerationstage,thesystemcarriesout eitherevidence-groundedmultimodalanswergenerationordirectanswergenerationwithsafety-orientedmedical responseconstraints.AFinalSynthesismodulefinallysynthesizesallsubanswersintoonefinalanswerwhileattaching referencedguidelineimages(ifusingRAG)andstoringacompleteprocesstraceforauditability.Anoverviewofthe proposedOph-Guid-RAGframeworkisillustratedinFigure1. Figure1:SystemArchitectureofGuideline-GroundedMultimodalRAGforOphthalmicClinicalDecisionSupport. 3.3StageI:CorpusPreparation Intheofflinecorpuspreparationstage,wegotPDFsofophthalmology-relatedguidelinesfromtheMedlive(Yimaitong) platformduringtheingestionstage.Therewerenofailureswhenprocessing305PDFfiles.Thereare7001pagesinthe corpus,andeachguidelinedocumenthasanaverageof22.95pages.Theguidelinesourcesspanabroadrangeof authoritiesandcanbegroupedintofourmaincategories: (1)GlobalAuthorities:ThiscategoryconsistsofrenownedinternationalinstitutionssuchastheWHO,theAmerican AcademyofOphthalmology(AAO). (2)Government/NationalOrganizations:Thisincludesofficialnationalhealthandmedicalregulatorybodies. (3)Provincial/SocietyGuidelines:Includesmedicalsocietiesattheregionallevelandconsensusdocumentsatthe provinciallevel. (4)OtherExpertConsensus:Thisincludesotherstatementsofagreementfromspecializedexperts. ThemeanpagesizefortheguidelinePDFswas5908x8063pixels,andwerenderedeverypagetoahighresolution imageat720DPI.Post-rendering,eachpageisrescaledtobefittedintoatargetcanvasofsize5390Ă7940pixels preservingitsoriginalaspect-ratioandthenpaddedwithacentralwhitebackgroundsoastoensurematchinginput dimensionalityforsubsequentembeddings.NoOCRwasperformed,automaticnoisereduction,rotationand/orborder trimming.Weintentionallydidnotdoanytextextractionorstructureparsinginordertopreservetheoriginallayout informationliketables,flowcharts,stagingcriteria,dosingthresholdetc).Inthisapproacheverypageistreatedlikean independentdocumentandtheretrieverdealswithpreciseimagesoforiginaldocuments.Westoredthecleaneduppage imagefilesintotheobjectstoreandgeneratedaJSONmanifestfilecontainingURLsforeaseofuseduringtheindexing andonlineinferencephases.Ourpage-levelvisualindexingapproacheliminatestheneedforstandardtext segmentation:whichmaysplitclinicallyrelevantinformationinparagraphs,anddestroythedocumentâsformatting. Table1:GuidelineDataProcessingStatistics. CategoryStatistic NumberofGuidelineDocuments305PDFs(0failures) TotalNumberofPages7001 AveragePagesperDocument22.95 RenderingDPI720DPI AverageRawImageResolution5908Ă8063pixels ProcessedImageResolution5390Ă7940pixels(resize+centerpadding) PreprocessingOperationsNoOCR/denoising/cropping/rotation Tosupportpage-levelvisualretrieval,weadoptColQwen2.5asthevisualdocumentretriever.Bothguidelinepage imagesanduser-sideretrievalqueriesaremappedintoasharedvectorspacebyColQwen2.5,withpageimagesencoded throughtheimageencoderandtextualqueriesencodedthroughthequeryencoder.Theretrieveroutputssequence-level embeddings,whichwereducetoonefixed-lengthrepresentationperimageorquerybymeanpoolingoverthesequence dimension.Thispoolingstrategysimplifiessystemintegrationandstabilizesdownstreamindexingandretrieval. Inourpresentsetup,wedonotperformanyfurtherL2normalisationofthepooledembeddings;insteadweconstructa FAISSexactnearest-neighborindex(IndexFlatL2)andusetherawL2distanceastheretrievalscore.Thisdesign preservestransparencyintheretrievalprocessandmakestheretrieval-confidencethresholdeasiertocalibrateduring inference.Asaresult,theofflinestageproducesapage-levelvisualknowledgebaseconsistingofnormalizedpage images,alignedmetadataandURLmappings,andaFAISS-basedColQwen2.5visualindexforlaterretrieval. 3.4StageII:QueryProcessing Intheonlineinferencestage,thesystemfirstreceivestheuserqueryandprocessesitthroughadedicatedquery processingmodule.Sincereal-worldmedicalquestionsareoftenmulti-intent,under-specified,orcontainmultiple clinicalconstraints,thesystembeginswithaPlannerthatdetermineswhethertheoriginalqueryshouldremainintactor bedecomposedintosmallerfocusedsubquestions.Ifdecompositionisneeded,thequestionissplitintoatmostthree subquestions,whichallowsdownstreamretrievalandanswergenerationtooperateonmorelocalizedandsemantically coherentunits. Oneachoftheseresultingsubquestions,weapplyaRoutertodecideifthesubquestionwilltakeaRAGpathwayora DIRECTpathway.TheRAGpathwayismeantforthosequestionswhicharedependentonevidencesfromguidelines, suchasmedicationregimens,threshold-dependentcriteria,contraindications,follow-upintervals,andstructured managementrecommendations.Incontrast,theDIRECTpathwayisusedforcasesinwhichretrievalisunnecessaryor unlikelytoimproveanswerquality. ForsubquestionsroutedtotheRAGpathway,thesystemfurtherappliesaQueryRewritemodulethatreformulatesthe subquestionintoretrieval-orientedqueriesbetteralignedwiththeunderlyingguidelinepages.Theserewrittenqueriesare designedtoemphasizeclinicallymeaningfulentitiesorconstraints,suchasdiseasenames,examinationparameters,drug names,phases,cutpoints,ortherapies.WhenasubquestiongetspassedtoDIRECT,theretrievalphaseisskippedand thesubquestiongoesstraightintoanswerextraction. 3.5StageIII:RetrievalandFiltering VisualRetrievalForeveryincomingsubquestiontotheRAGbranch,weperformavisualsearchonthepage-wise knowledgebase.RewrittenretrievalqueriesareencodedbytheColQwen2.5queryencoder,whileguidelinepageshave alreadybeenencodedofflinebytheimageencoderintothesamesharedvectorspace.Theresultingembeddingsare indexedwithFAISS,whichisusedtoretrievethenearestcandidateevidencepages.Inthisway,retrievaloperates directlyoverpageimagesratherthanoverextractedtextpassages,allowingthesystemtopreservepage-levellayoutcues andvisualevidencestructure. However,theretrievedcandidatesarenotuseddirectly.Toreducefalse-positiveretrievalandimproveevidenceTo enhancetheprecision,wefurtheraddonemoreLLMbasedrelevancefilterlayerontopofit.Specifically,weemploy GPT-5.2toexamineeachretrievedcandidatepageanddeterminewhetheritisgenuinelyrelevanttoboththeoriginal userqueryandthecurrentfocusedsubquestion.Onlythosepagesjudgedasrelevantareretainedasvalidevidencefor downstreammultimodalanswering. Oncefiltered,wedecideifthereisenoughreasontobelieveintheevidencefoundornot.Ifsoandthereisenough relevantevidencethen,thesubquestionmovesontomultimodalRAGanswer.Ifthereisnoevidencesurvivedafterthe filterstep,orwhenaretrievalsignalisjudgedtobetooweak,thesystemtriggersaretrieval-qualityfallbackandreverts totheDIRECTpathway.Thisdesignpreventsthegeneratorfromrelyingonsemanticallymismatchedorlow-confidence retrievalresults,therebyimprovingrobustnessandreducingerrorpropagationfromtheretrievertothefinalanswer. 3.6StageIV:GenerationandFinalSynthesis Inthegenerationstage,thesystemanswerseachsubquestionaccordingtothepathwaydeterminedinthepreviousstage. Ifreliableevidencepagesareavailable,thesystemperformsmultimodalRAGanswering,inwhichthefocusedtextual questionandtheretrievedpageimagesarejointlyprovidedtothegeneratorsothattheanswercanbegroundedin explicitguidelineevidence.Ifreliableevidenceisnotavailable,thesysteminsteadperformsDIRECTanswering,relying onthegeneratorâsgeneralmedicalreasoningabilitytogetherwithsafety-orientedprompting. Afterallsubquestionshavebeenanswered,thesysteminvokesaFinalSynthesismoduletomergethesubanswersinto onecoherentfinalresponse.Thissynthesisstepisresponsibleforpreservinglogicalconsistency,reducingredundancy acrosssubanswers,andproducingafluentend-to-endanswertotheoriginaluserquery.WhenRAGevidencehasbeen used,thesystemcanadditionallyappendthereferencedimagesourcessothatthefinaloutputremainstraceableand auditable. Thesystemalsorecordsastructuredprocesstracecoveringallmajorintermediatedecisions,includingquestion decomposition,routingresults,rewrittenqueries,retrievedpages,filteringoutcomes,finalanswermodes,andsource references.Thistraceservesbothasanauditingmechanismandasananalysistoolforlaterdebugging,ablationstudies, andsystemevaluation.Therefore,thefinaloutputofthesystemisnotonlyananswer,butalsoincludesthedirectly displayedguidelinepageimagesasvisualevidenceforverificationandclinicalinterpretability,followedbyastructured processtracethatrecordshowtheanswerwasproduced. Tobetterillustratehowtheproposedsystemoperatesinpractice,wepresentarepresentativecasestudyinFigure2. Figure2:CaseStudy:ComparingBaselineandOph-Guid-Rag Figure2showsarealexampleofhowthesystemprocessesaclinicallycomplexquery.Theinputquestioninvolves multipleclinicalconsiderations(e.g.,doseadjustment,druginteraction,andelectrolytesafety).ThePlannerfirst decomposesthequeryintothreefocusedsubquestions.EachsubquestionisthenroutedtoeithertheRAGorDIRECT pathway.ForsubquestionsinitiallyroutedtoRAG,thesystemretrievescandidateguidelinepagesandappliesanLLM- basedrelevancefilter.Ifnosufficientlyrelevantevidenceisretained,thesystemfallsbacktotheDIRECTmode.Inthis example,twosubquestionsfailtoobtainreliableevidenceafterfilteringandarethereforeanswereddirectly,whileone subquestionretainsarelevantguidelinepageandproceedswithmultimodalRAG.Comparedtothebaselinemodel, whichprovidesageneralanswerwithoutexplicitevidencegrounding,theproposedsystemproducesamoreclinically groundedresponse.Inadditiontoreturningthesupportingguidelinepageasvisualevidence,thesystemalsorecordsand outputsastructuredend-to-endprocesstrace(trace.json),coveringdecomposition,routingdecisions,rewrittenqueries, retrievalresults,filteringresults,andmodeoffinalanswersothatwehavecompletetransparency,traceabilityandpost- hocverificationofthereasoningprocess. 4.EVALUATION 4.1Datasetconstruction TheophthalmologypromptsweuseinourworkaretakenfromthefullHealthBenchdataset,aswellastwoofitsofficial subsets.Weuseone,fullyreproduciblefilteringruletocreatetheophthalmologysubset.Weextractedforeveryexample thequestiontextfromthefirstavailabletextfield(âquestionâ,"prompt",or"content"),lowercasedthestring,and matchedagainstanophthalmologykeywordlistwhichwaspre-definedbyus.Thekeywordlistconsistedoffollowing words:"ophthalmology,""eye,""retina,""glaucoma,""cataract,""cornea,""vision,""intraocularpressure,""fundus," "strabismus,""myopia,""hyperopia,""amblyopia,""macula,""vitreous,"and"opticnerve."Ifatleastonekeyword matched,anexamplewaskept. Topreventapotentialdistributionalbiasinducedthroughfilteringonthesubsetsthemselves,weusedexactlythesame deterministicruleforfilteringthewholedatasetandbothofficialsubsets;thisensuresthatthewayinwhichwebuildup theophthalmologysubsetistransparent,canbeverified,andcanbepreciselyreproduced.Wecollectedintotal78 ophthalmologyprompts:16weretakenfromthehardsubset,and62fromtheconsensusone.Foreasyverificationand duplication,wereleasetheprecisefilteringcodeandkeywordslistpubliclyonourGitHubrepo6.Thefinaldatasplit statisticsarepresentedinTable2. Table2:HealthBenchOphthalmicSubsetSplitStatistics. SplitNumberofPromptsPercentageofTotal Main78100.00% Consensus6279.49% Hard1620.51% 4.2Metrics/Rubric WeuseHealthBenchthatscoresmodeloutputwithrubrics(writtenbydoctors)onaconversation-byconversationbasis andaggregatestherubriccriteriaintobehavioralaxessothatstrengthsandfailuremodescanbeattributedingreatdetail [1].Inourreporting,wefollowtheHealthBenchaxisdefinitionsandfocusondimensionsthathaveadirectimpacton clinicalusabilityandsafety,suchas: (1)Accuracy:Theclinicalinformationandsuggestionsmustbefactuallycorrect. (2)Completeness:Arealltheimportantpartsofassessmentandmanagementcovered? (3)Contextawareness:Doesthesystemrespondtosituationalcuesandaskfollow-upquestionswheninformationis missing? (4)Communicationquality:Isthetoneappropriatefortheaudience? (5)Follow-the-leader:adheringtotheguidelinesofboththeuserandthecircumstance. WeconsiderthefollowingkeyfacetsinourHard-subsetinvestigation:completenessandcontextulizationthat demonstratemostclearlytherubricbehaviorssupportingsafeandeffectiveclinicalconversation.Allratingsweregiven asvaluesbetween0and1,wherewetakethemean. 4.3Experimentalsetupandcomparability WetestOph-Guid-RAGagainstGPT-5.2,GPT-5.3,andGPT-5.4toseehowitcompares.Pleasekeepinmindthatwe madeallofthebaselinescoresourselves.WeusethesamepipelineforGPT-5.2,GPT-5.3,andGPT-5.4,andwegrade themallwiththesameHealthBenchsetup.Thismakesitpossibletocomparethemdirectlyandfairlyinthesame situation. Wealsohaveablationvariantslikeno_rerank,no_query_rewrite,andno_router.Tobefair,allmodelsandallversions aretestedinthesameway.TheyusethesameHealthBenchconversationformat,thesamesystemsafetyprompt,the samedecodingsetting,andthesamelimitonthelengthoftheoutput.Theonlythingthatmakesmodelsdifferentishow theygetdataback.Oph-Guid-RAGusesguidelinepageretrievalalongwithrerank,queryrewrite,androuting.Oneof thesemodulesistakenoutbytheablationmodels.Thisretrievalpipelineisnotusedbythebaselinemodels.Thereare nootherchanges. WeevaluatealloutputusingtheofficialHealthBenchgradingscript,runningtheirdefaultmodelbasedgrader.Allruns sharethesamegradingmodel,thesamegradingpromptsandthesamescoreaggregationlogic.Thisguaranteesthatthe scoredifferencescomeonlyfrommodelbehaviorandretrievaldesign,notfromtheevaluationprocess.Wereportresults onboththefullsetandthehardsubset.Thehardsubsetcontainsmoredifficultquestionsandrequiresstrongerreasoning andbetteruseofevidence.Wereportoverallscoreandfiveaxesincludingaccuracy,completeness,instructionfollowing, contextawarenessandcommunicationquality. 5.EXPERIMENTALRESULTS 5.1OverallPerformanceonAllSamples Wefirstcompareourmethodwithseveralstrongbaselinesonthefullevaluationset.Theresultsaresummarizedin Table3. Table3:PerformanceComparisonontheAllSet ModelOverallScoreAccuracyCompletenessInstructionFollowingContextAwarenessCommunicationQuality Oph-Guid-Rag0.55240.62660.46140.34740.62120.4153 GPT-5.20.55590.63170.5010.51930.50270.447 GPT-5.30.54840.5690.45740.56720.53170.4717 GPT-5.40.58560.63360.53140.47490.63330.5562 Thefinalscoreforourmodelis0.55,similarasGPT-5.2(0.5559),butalittlelessthanGPT-5.4(0.5856).Ourmodel obtainstheaccuracyof0.6266,comparablewithGPT-5.4(0.6336),whilebetterthanGPT-5.3(0.5690)showingthat addingguidelineevidencehelpstoobtainacompetitivecorrectness.Yet,thegapmainlycomesfrominstruction followingandcommunicationquality.Ourmodelobtains0.3474oninstructionnext,whichisfarlowerthanGPT- 5.2(0.5193),GPT-5.3(0.5672).Similarphenomenonappearsinqualityofcommunication,onwhichourscore(0.4153) lagsbehindGPT-5.4âs(0.5562),suggestingthatretrievialbasedgenerationisrigidtotheresponsestylethatimpactson fluencyandadherencetoconversationinstruction.Meanwhile,ourmodelisstronglyawareofthecontext(0.6212),better thanthatofGPT-5.2(0.5027)andGPT-5.3(0.5317).Thisconfirmsthatgroundingonguidelinepagesimprovesthe modelâsabilitytostayalignedwithclinicallyrelevantcontext.Ingeneral,theresultshowsthatourapproachhasa balancedperformance,beingstrongataccuracyandcontextgrounding,leavingroomforimprovingininstruction adherenceandresponsenaturalness. 5.2PerformanceonHardSubset Wefurtherevaluateallmodelsonamorechallengingsubset,wherequestionsrequiredeeperreasoningorprecise guidelinegrounding.TheresultsareshowninTable4. Table4:PerformanceComparisonontheHardSet ModelOverallScoreAccuracyCompletenessInstructionFollowingContextAwarenessCommunicationQuality Oph-Guid-Rag0.38610.65760.04830.13330.39690.9167 GPT-5.20.29690.59560.09710.13330.34190.6167 ModelOverallScoreAccuracyCompletenessInstructionFollowingContextAwarenessCommunicationQuality GPT-5.30.33440.57230.018300.28630.7833 GPT-5.40.37560.52870.11390.13330.39010.8667 Onthissubsetweobtainthetotalscoreas0.3861,whichisbetterthanthatonGPT-5.2(0.2969),GPT-5.3(0.3344),and slightlybetterthanGPT-5.4(0.3756).AgainstGPT-5.2itâsanabsoluteimprovementof+0.0892,indicatingthatour methodismorerobustunderdifficultconditions.Themostsignificantaccuracyadvantageappearsinaccuracy.Our modelis0.6576,higherthanGPT-5.2(0.5956)+0.0620,GPT-5.4(0.5287)by+0.1289.Thisshowsthatretrieval groundingisparticularlybeneficialwhenthetaskrequirespreciseclinicalknowledge.Buttheincompletenessisstilla majorlimitation.Weobtainthescoreof0.0483,lowerthanGPT-5.4(0.1139)andGPT-5.2(0.0971),showingthat althoughweareabletodiscoverrightkeypoints,doesnotalwaysprovidesufficientlydetailedorfullydeveloped answers.Interestingly,ourmodelachievesaveryhighcommunicationqualityscore(0.9167),betterthanGPT-5.4 (0.8667)andGPT-5.2(0.6167).Weconcludefromthisresultthat,afterthemodelcommitstoananswer,ittendsto providecoherentexplanations.Takentogether,thesefindingspointtowardsthefactthattrade-off.Ourmethodimproves factualcorrectnessandrobustnessonhardcases,butatthecostofreducedcompleteness,likelyduetoconservativeor partialuseofretrievedevidence. 5.3AblationStudyonAllSamples Tobetterunderstandthecontributionofeachmodule,weconductablationexperimentsonthefulldataset.Theresults arepresentedinTable5. Table5:AblationExperimentResultsofDifferentModelConfigurationsonAllSubset ModelOverallScoreAccuracyCompletenessInstructionFollowingContextAwarenessCommunicationQuality Oph-Guid-Rag0.55240.61910.46140.3120.62120.4153 no_rerank0.52660.58160.48320.33540.52550.3799 no_query_rewrite0.55740.64410.55050.41830.52670.3323 no_router0.550.62660.50880.34740.58510.402 Removingthererankingmoduleleadstoadropinoverallscorefrom0.5524to0.5266(â0.0258),andadecreasein accuracyfrom0.6191to0.5816(â0.0375).Thisconfirmsthatrerankingplaysacriticalroleinimprovingretrieval qualityanddownstreamcorrectness.Removingqueryrewritingresultsinanoverallscoreof0.5574,whichisslightly higherthanthefullmodel.Italsoincreasescompletenessfrom0.4614to0.5505(+0.0891).However,thiscomesatthe costofworsecommunicationsquality(0.3323vs.0.4153).Thisindicatesthatqueryrewritingintroducesabiastowards moreprecise,butshorteranswers,whileremovingitallowsmoreverboseanswers.Removingtheroutercausesan overallscoreof0.55,whichisclosetothefullmodel.However,contextawarenessdropsfrom0.6212to0.5851(- 0.0361).Theseresultsindicatethatroutingenablesselectiveapplicationofretrievalwherenecessarytoimprovethe contextualgrounding.Overall,rerankingisthemostcriticalcomponentonthefulldataset,whiletheeffectsofquery rewritingandroutingaremoresubtleandtask-dependent. 5.4AblationStudyonHardSubset Wefurtheranalyzetheablationsonthehardsubset,asshowninTable6. Table6:AblationExperimentResultsofDifferentModelConfigurationsonHardSubset ModelOverallScoreAccuracyCompletenessInstructionFollowingContextAwarenessCommunicationQuality Oph-Guid-Rag0.38610.65760.04830.13330.39690.9167 no_rerank0.28170.44610.088300.40260.4833 no_query_rewrite0.33470.56540.15990.13330.41790.65 no_router0.36030.48350.06600.48350.65 Theimpactofrerankingbecomesmuchmoresignificantunderdifficultconditions.Removingrerankingreducesthe overallscorefrom0.3861to0.2817(â0.1044),andaccuracydropssharplyfrom0.6576to0.4461(â0.2115).Thisshows theimportanceofselectinggoodevidencetoanswercomplexclinicalquestions,whileremovingqueryrewriting decreasestheoverallscoreto0.3347(â0.0514),butincreasescompletenessfrom0.0483to0.1599(+0.1116).This indicatesthatqueryrewritingimprovesprecisionbutmayoverlyconstraintheretrievedevidence,resultinginpartial answers.Ifweremovetherouter,thentheoverallscoreis0.3603(â0.0258).Theaccuracydropssignificantlyfrom 0.6576to0.4835(â0.1741),whilethecontext-awarenessrisesfrom0.3969to0.4835(+0.0866).Thissuggeststhat forcingallqueriesthroughretrievaladdsmorecontext,butalsomorenoise,whichhurtsthecorrectness.Inotherwords, theseresultsshowthatonhardcases,rerankingisessential,routingallowsfortradingprecisionagainstnoise,while queryrewritingcomeswithacompleteness-focustrade-off. 5.5Discussion Theexperimentalresultsrevealthatourmethoddemonstratesclearadvantagesonthehardsubset,whichrepresentsthe mostchallengingscenariosinHealthBench[26]. Followingtheofficialbenchmarkdesign,weuseHealthBenchHardasasetofespeciallyhardexamplesthatcurrent frontiermodelsstruggletosolve,andwasspecificallydesignedtochallengemodelsatrealisticlevelsofclinical reasoning.Thesecasesarenotsimplefactualquestions.Instead,theyaremulti-turnclinicalconversationsthatrequire reasoningunderuncertainty,integrationofcontextacrossturns,andadherencetodetailedphysician-definedevaluation criteria. Thedifficultyoftheseproblemsliesinthefactthattheyrequireevidence-groundedreasoningundermultipleconstraints, notjustmemorization.Generalmedicalknowledgemaynotsufficetoanswermanyquestionsandadditionalconformity tocertainguidelinestatementincludingthresholdvalues,contraindications,andtreatmentpathways.Atthesameeach responseisscoredonmultipleclinicalcriteria(e.g.,accuracy,completeness,safety,andrequiringcommunication,where partialcorrectnessisnotsufficient.Bothofthesetasksareadditionallymadecomplexbyuncertaintyandmissingdata, forwhichthemodelshouldrefrainfrommakingoverconfidentstatementsanddelivercontextualizedrecommendations. Inthesescenarios,typicalLLMsarepronetofailuresincetheyrelyontheparametricknowledgeandoftencannotmeet multipleclinicalrequirementsatonce. Underthissetting,ourmethodshowsconsistentadvantagesbygroundingresponsesinguidelineevidenceandcarefully controllingtheretrievalprocess.Byretrievingpage-levelguidelinecontent,themodelcanrelyonauthoritativesources insteadofapproximateinternalknowledge,leadingtocleargainsinaccuracy,withimprovementsof+0.0620overGPT- 5.2and+0.1289overGPT-5.4onthehardsubset.Atthesametime,routingandfilteringhelpintroduceexternal evidenceonlywhennecessary,avoidingirrelevantcontext,whilererankingensuresthatthemostrelevantguideline pagesareselected.Thisdesigniscriticalonhardcases,asremovingrerankingleadstothelargestdrop(â0.1044in overallscoreandâ0.2115inaccuracy),andremovingtherouteralsocausesasubstantialaccuracydecrease(â0.1741). Overalltheseresultssuggestourapproachperformsbetterincaseswherethereareclearevidentialreasonstobeableto reasonprecisely,whileitcontinuestogenerateincompleteresponsespossiblybecauseofitsconservativeuseofevidence implyingthatbetterevidenceaggregationshouldbeexploredinfurtherresearch. 6.LIMITATIONSANDRISKS Despitethestrongempiricalperformance,severallimitationsremain.Thesystemreliesonthecoverageandqualityof theguidelinecorpus,whichisinherentlyincompleteandmaynotcapturerecentupdates,regionalvariations,orrare cases,leadingtopotentialfailureswhenrelevantevidencecannotberetrieved.Meanwhile,theretrievalandthefiltering pipelinecomeswithaprecision-completenesstrade-off,asrerankingandLLMbasedfilteringimproveaccuracyby removingnoise,theymayalsodiscardpartiallyrelevantevidence,resultinginconcisebutincompletereplies(asseen fromthehardsubsetresults).Theroutingschemecanalsobeunreliableandsometimestakeawrongturn,eitheradding extraneousretrievaloromittingrequiredevidence,bothofwhichcouldimpactfactualgrounding.Lastbutnotleast,we designoursystemtobeusedasaclinicaldecisionsupporttoolinsteadthanadiagnosticsystem.Althoughitis encouragedtoexpressuncertaintyandavoidoverconfidentconclusions,itmaystillgenerateinappropriateorincomplete suggestionsinsomecases,anditsoutputsshouldthereforebeinterpretedwithcautionunderprofessionalsupervision. 7.CONCLUSION WeintroduceamultimodalvisualRAGsystemforclinicaldecisionsupportinthefieldofophthalmologythatiscapable ofenableevidence-groundedandcontrollablegeneration.Byconstructingapage-levelguidelinecorpusanddirectly retrievingoriginalguidelineimagesasevidence,ourapproachavoidsconventionaltextsegmentationandpreserves clinicallyrelevantstructures.Inadditiontotheabove,weproposeaprincipledinferencepipelineconsistingofquery- decomposition,Therouting,queryrewriting,retrieval,rerankingandmultimodalreasoningenableourmodeltoflexibly leverageexternalknowledgeonaselectivebasisduringgeneration. ExtensiveexperimentsonHealthBenchdemonstratethatourmethodachievescompetitiveoverallperformanceand showsclearadvantagesonchallengingcases.Inparticular,onthehardsubset,oursystemoutperformsstrongbaselines onboththeaccuracyandtotalscores,suggestingthatitiseffectiveatdealingwithmorecomplex,evidentialclinical reasoningtasks.Theablationresultsfurtherconfirmthatreranking,routing,andretrievaldesignallplaycriticalrolesin improvingperformanceunderdifficultconditions. Meanwhile,wealsoshowakeytrade-offbetweenaccuracyandcompleteness:whilegroundingonguideline evidenceincreasesfactualcorrectnessandcontextawareness,andcanresultinalessaggressivereactionaswell anddecreasedcoverage.Itsuggeststhatafutureeffortshouldbededicatedtoimprovedevidenceaggregationand furtherlonggenerationpolicieswhilenotdamagingthefactualityofinference.Overall,ourpaperprovidesa practicalsteptowardtrustworthyandtraceableclinicalAIsystems,anddemonstratesthepotentialofcombining multimodalretrievalwithstructuredreasoningforreal-worldmedicalapplications. ACKNOWLEDGEMENTS Thisstudygratefullyacknowledgesthecontributionsoftworesearchersinvolvedindatacollectionand preprocessing,aswellasthecomputationalsupportprovidedbytheHigh-PerformanceComputingPlatformofthe ChineseAcademyofMedicalSciences.Wealsoacknowledgetheuseofpubliclyavailableresources,includingthe OpenAIHealthBenchdataset,Medlive,andguidelinedocumentsreleasedbyvariousinstitutions,whichprovided essentialsupportforthisresearch. REFERENCES [1]M.A.Chia,F.Antaki,Y.Zhou,A.W.Turner,A.Y.Lee,andP.A.Keane,âFoundationmodelsinophthalmologyâ, Br.J.Ophthalmol.,vol.108,no.10,p.1341â1348,Sep.2024,doi:10.1136/bjo-2024-325459. [2]Y.Zhuangetal.,âMultimodalreasoningagentforenhancedophthalmicdecision-making:apreliminaryreal-world clinicalvalidationâ,Front.CellDev.Biol.,vol.13,p.1642539,Jul.2025,doi:10.3389/fcell.2025.1642539. [3]S.Liu,A.B.McCoy,andA.Wright,âImprovinglargelanguagemodelapplicationsinbiomedicinewithretrieval- augmentedgeneration:asystematicreview,meta-analysis,andclinicaldevelopmentguidelinesâ,J.Am.Med.Inform. Assoc.,vol.32,no.4,p.605â615,Apr.2025,doi:10.1093/jamia/ocaf008. [4]Z.Sunetal.,âReDeEP:DetectingHallucinationinRetrieval-AugmentedGenerationviaMechanistic Interpretabilityâ,Jan.21,2025,arXiv:arXiv:2410.11414.doi:10.48550/arXiv.2410.11414. [5]M.Sevgi,F.Antaki,andP.A.Keane,âMedicaleducationwithlargelanguagemodelsinophthalmology:custom instructionsandenhancedretrievalcapabilitiesâ,Br.J.Ophthalmol.,vol.108,no.10,p.1354â1361,Oct.2024,doi: 10.1136/bjo-2023-325046. [6]M.HanWangetal.,âAppliedmachinelearninginintelligentsystems:knowledgegraph-enhancedophthalmic contrastivelearningwithâclinicalprofileâpromptsâ,Front.Artif.Intell.,vol.8,Mar.2025,doi: 10.3389/frai.2025.1527010. [7]M.-J.Luoetal.,âAlargelanguagemodeldigitalpatientsystemenhancesophthalmologyhistorytakingskillsâ,Npj Digit.Med.,vol.8,no.1,p.502,Aug.2025,doi:10.1038/s41746-025-01841-6. [8]A.Grzybowskietal.,âRetinaFundusPhotograph-BasedArtificialIntelligenceAlgorithmsinMedicine:A SystematicReviewâ,Ophthalmol.Ther.,vol.13,no.8,p.2125â2149,Aug.2024,doi:10.1007/s40123-024-00981-4. [9]A.Y.Ongetal.,âAscopingreviewofartificialintelligenceasamedicaldeviceforophthalmicimageanalysisin Europe,AustraliaandAmericaâ,NpjDigit.Med.,vol.8,no.1,p.323,May2025,doi:10.1038/s41746-025-01726-8. [10]Z.J.Muhsin,R.Qahwaji,M.AlShawabkeh,S.A.AlRyalat,M.AlBdour,andM.Al-Taee,âSmartdecision supportsystemforkeratoconusseveritystagingusingcornealcurvatureandthinnestpachymetryindicesâ,EyeVis.,vol. 11,no.1,p.28,Jul.2024,doi:10.1186/s40662-024-00394-1. [11]S.Kresevic,M.Giuffrè,M.Ajcevic,A.Accardo,L.S.Crocè,andD.L.Shung,âOptimizationofhepatological clinicalguidelinesinterpretationbylargelanguagemodels:aretrievalaugmentedgeneration-basedframeworkâ,Npj Digit.Med.,vol.7,no.1,p.102,Apr.2024,doi:10.1038/s41746-024-01091-y. [12]C.S.Ong,N.T.Obey,Y.Zheng,A.Cohan,andE.B.Schneider,âSurgeryLLM:aretrieval-augmentedgeneration largelanguagemodelframeworkforsurgicaldecisionsupportandworkflowenhancementâ,NpjDigit.Med.,vol.7,no. 1,p.364,Dec.2024,doi:10.1038/s41746-024-01391-3. [13]Y.H.Keetal.,âRetrievalaugmentedgenerationfor10largelanguagemodelsanditsgeneralizabilityinassessing medicalfitnessâ,NpjDigit.Med.,vol.8,no.1,p.187,Apr.2025,doi:10.1038/s41746-025-01519-z. [14]J.Miao,C.Thongprayoon,S.Suppadungsuk,O.A.GarciaValencia,andW.Cheungpasitporn,âIntegrating Retrieval-AugmentedGenerationwithLargeLanguageModelsinNephrology:AdvancingPracticalApplicationsâ, Medicina(Mex.),vol.60,no.3,p.445,Mar.2024,doi:10.3390/medicina60030445. [15]A.Wadaetal.,âRetrieval-augmentedgenerationelevateslocalLLMqualityinradiologycontrastmedia consultationâ,NpjDigit.Med.,vol.8,no.1,p.395,Jul.2025,doi:10.1038/s41746-025-01802-z. [16]R.Chen,S.Zhang,Y.Zheng,Q.Yu,andC.Wang,âEnhancingtreatmentdecision-makingforlowbackpain:a novelframeworkintegratinglargelanguagemodelswithretrieval-augmentedgenerationtechnologyâ,Front.Med.,vol. 12,May2025,doi:10.3389/fmed.2025.1599241. [17]Q.Zhouetal.,âGastroBot:aChinesegastrointestinaldiseasechatbotbasedontheretrieval-augmentedgenerationâ, Front.Med.,vol.11,May2024,doi:10.3389/fmed.2024.1392555. [18]S.Prabhaetal.,âEnhancingClinicalDecisionSupportwithAdaptiveIterativeSelf-QueryRetrievalforRetrieval- AugmentedLargeLanguageModelsâ,Bioengineering,vol.12,no.8,p.895,Aug.2025,doi: 10.3390/bioengineering12080895. [19]I.Lopezetal.,âClinicalentityaugmentedretrievalforclinicalinformationextractionâ,NpjDigit.Med.,vol.8,no. 1,p.45,Jan.2025,doi:10.1038/s41746-024-01377-1. [20]G.Zhangetal.,âLeveraginglongcontextinretrievalaugmentedlanguagemodelsformedicalquestionansweringâ, NpjDigit.Med.,vol.8,no.1,p.239,May2025,doi:10.1038/s41746-025-01651-w. [21]S.Gilbert,J.N.Kather,andA.Hogan,âAugmentednon-hallucinatinglargelanguagemodelsasmedical informationcuratorsâ,NpjDigit.Med.,vol.7,no.1,p.100,Apr.2024,doi:10.1038/s41746-024-01081-0. [22]K.Gilani,M.Verket,C.Peters,M.Dumontier,H.-P.B.-L.Rocca,andV.Urovi,âCDE-Mapper:Usingretrieval- augmentedlanguagemodelsforlinkingclinicaldataelementstocontrolledvocabulariesâ,Comput.Biol.Med.,vol.196, p.110745,Sep.2025,doi:10.1016/j.compbiomed.2025.110745. [23]F.Borchertetal.,âHigh-precisioninformationretrievalforrapidclinicalguidelineupdatesâ,NpjDigit.Med.,vol.8, no.1,p.227,Apr.2025,doi:10.1038/s41746-025-01648-5. [24]G.Xiong,Q.Jin,Z.Lu,andA.Zhang,âBenchmarkingRetrieval-AugmentedGenerationforMedicineâ,in FindingsoftheAssociationforComputationalLinguistics:ACL2024,L.-W.Ku,A.Martins,andV.Srikumar,Eds, Bangkok,Thailand:AssociationforComputationalLinguistics,Aug.2024,p.6233â6251.doi: 10.18653/v1/2024.findings-acl.372%20IF:%20NA%20NA%20NA. [25]J.Dobrevaetal.,âRAGCare-QA:Abenchmarkdatasetforevaluatingretrieval-augmentedgenerationpipelinesin theoreticalmedicalknowledgeâ,DataBrief,vol.63,p.112146,Dec.2025,doi:10.1016/j.dib.2025.112146. [26]R.K.Aroraetal.,âHealthBench:EvaluatingLargeLanguageModelsTowardsImprovedHumanHealthâ,May13, 2025,arXiv:arXiv:2505.08775.doi:10.48550/arXiv.2505.08775.