Paper deep dive
Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis
Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/2/2026, 1:23:38 PM
Summary
The paper introduces Diagnostic Evidence Network (DENet), a multi-task framework for bearing fault diagnosis that extends standard classification outputs with physically verifiable evidence, including predicted characteristic frequencies and temporal impulse localization. This approach addresses the lack of trustworthiness in AI-based diagnosis by providing inference-time validation signals independent of softmax confidence scores. Additionally, the work employs a QLoRA-adapted Large Language Model (LLM) constrained to translate diagnostic evidence into reports, significantly reducing hallucination rates compared to generative approaches.
Entities (8)
Relation Signals (6)
BiMS-Mamba â backboneof â Diagnostic Evidence Network
confidence 95% · a bidirectional multi-scale Mamba encoder (BiMS-Mamba) is adopted as the default backbone
Diagnostic Evidence Network â produces â Characteristic Frequency
confidence 95% · DENet extends the output to a structured evidence record: the classification, a predicted characteristic frequency...
Diagnostic Evidence Network â improves â Trustworthiness
confidence 94% · Trustworthy deployment of AI-based diagnosis... hinges on validation... DENet addresses both problems from the output side.
QLoRA-adapted LLM â reduces â Hallucination
confidence 93% · reducing unsupported-claim rates from 10-12% to 2% and eliminating fabricated quantities.
QLoRA â appliedto â Large Language Model
confidence 92% · a QLoRA-adapted language model is constrained to translate, but never generate, diagnostic content
Characteristic Frequency â validates â Softmax Confidence
confidence 90% · the deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal... remains discriminative in the high-confidence regime where confidence-derived detectors are blind.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in two ways. Their standard output, a class label with a softmax confidence score, is an internal statistic of the classifier, offering nothing checkable against independent physical knowledge; and the growing use of generative language models in maintenance reporting adds a second risk: hallucinated content entering reports on which decisions rest. Taking bearing fault diagnosis as the testbed, this work addresses both problems from the output side. The proposed Diagnostic Evidence Network (DENet) is an encoder-agnostic multi-task framework extending the output to a structured evidence record: the classification, a predicted characteristic frequency comparable against the theoretical value determined by bearing geometry and shaft speed, and a temporal localization of transient impulses inspectable on the raw waveform. Across four encoders and three public datasets, this evidence incurs no statistically significant accuracy cost, with a frequency error of about 6 Hz on 1,024-point segments where spectral estimation is structurally inapplicable. Centrally, the deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal: it detects misclassifications with AUROC values of 0.970 and 0.871, and remains discriminative in the high-confidence regime where confidence-derived detectors are blind. Finally, a QLoRA-adapted language model is constrained to translate, but never generate, diagnostic content, reducing unsupported-claim rates from 10-12% to 2% and eliminating fabricated quantities.
Tags
Links
- Source: https://arxiv.org/abs/2607.22797v1
- Canonical: https://arxiv.org/abs/2607.22797v1
Trouble viewing inline? Open PDF directly â
Full Text
82,492 characters extracted from source content.
Expand or collapse full text
1 PhysicallyVerifiableEvidenceandLLM-BasedReportingfor BearingFaultDiagnosis Author YuntongChen a ,JianyuLiu a ,GuobinZhao b ,ZiangWang a ,ChaoChen c ,JuHuang d ,XitianTian a ,Lijiang Huang a* aSchoolofMechanicalEngineering,NorthwesternPolytechnicalUniversity,Xiâan,710072,China bDepartmentofChemistry,NationalUniversityofSingapore,3ScienceDrive3,Singapore117543, Singapore cSchoolofMechanicalEngineeringandAutomation,BeihangUniversity,Beijing,100191,China dChemicalEngineering&AppliedChemistry,UniversityofToronto,Toronto,ON,M5S3E5,Canada Abstract TrustworthydeploymentofAI-baseddiagnosisinsafety-criticalmechanicalsystemshingesonvalidation:whethera predictioncanbecheckedagainstphysicalrealitybeforeitisactedupon.Currentintelligentfaultdiagnosersfailthisstandardin twoways.Theirstandardoutputâaclasslabelwithasoftmaxconfidencescoreâisaninternalstatisticoftheclassifier, offeringnothingcheckableagainstindependentphysicalknowledge;andthegrowinguseofgenerativelanguagemodelsin maintenancereportingaddsasecondrisk,hallucinatedcontententeringreportsonwhichdecisionsrest.Takingbearingfault diagnosisasthetestbed,thisworkaddressesbothproblemsfromtheoutputside.TheproposedDiagnosticEvidenceNetwork (DENet)isanencoder-agnosticmulti-taskframeworkextendingtheoutputtoastructuredevidencerecord:theclassification,a predictedcharacteristicfrequencycomparableagainstthetheoreticalvaluedeterminedbybearinggeometryandshaftspeed,and atemporallocalizationoftransientimpulsesinspectableontherawwaveform.Acrossfourencodersandthreepublicdatasets, thisevidenceincursnostatisticallysignificantaccuracycost,withafrequencyerrorofabout6Hzon1,024-pointsegments wherespectralestimationisstructurallyinapplicable.Centrally,thedeviationbetweenpredictedandtheoreticalfrequency constitutesalabel-free,inference-timevalidationsignal:itdetectsmisclassificationswithAUROCof0.970and0.871,and remainsdiscriminativeinthehigh-confidenceregimewhereconfidence-deriveddetectorsareblind.Finally,aQLoRA-adapted languagemodelisconstrainedtotranslateâbutnevergenerateâdiagnosticcontent,reducingunsupported-claimratesfrom 10â12%to2%andeliminatingfabricatedquantities. Keywords Verificationandvalidation;Largelanguagemodel;GenerativeAI;Trustworthymachinelearning;Bearingfaultdiagnosis; Characteristicfrequency 2 1.Introduction Rotatingmachineryunderpinssafety-criticaloperationsinaerospace,energy,andmanufacturing,and bearingfaultdiagnosisfromvibrationsignalshasthereforebeenstudiedintensivelyfordecades[1,2]. Drivenbydeeplearning,classificationaccuracyhasclimbedbeyond99%onwidelyusedbenchmarks [3,4],leavinglittleheadroomonthisaxis.Industrialadoption,however,hasnotkeptpace,andrecent workattributesthisgaptotheformofthemodeloutputratherthanitscorrectness[5,6].Thestandard outputisaclasslabelwithasoftmaxconfidencescore.Theconfidencescoreisaninternalstatisticofthe classifierâitreferencesnothingoutsidethemodelâandthusoffersnomeansofindependent verificationbeforeactingonadiagnosis.Fromaverificationandvalidation(V&V)perspective, validationaskstowhatdegreeapredictionisconsistentwithphysicalreality[7];Measuredagainstthis standard,currentpipelinesfailattwopoints:eachdeployment-timepredictioncarriesnophysically checkablequantity,andthegenerativelanguagemodelsincreasinglyemployedtocommunicatediagnoses addstatementsthatcannotbetracedtoaverifiablesource.Thepracticalbottleneckhasshiftedfrom accuracytotheverifiabilityofoutputsâboththequantitiesandthelanguage. Existingresearchaddressesthisbottleneckfromfivedirections.Thefirstcontinuestheevolutionof encoderarchitectures,fromconvolutionalnetworks[8]throughTransformers[9]toMamba-based diagnosers[3].Thesemethodssteadilyimproveaccuracyandefficiency,buttheoutputinterfaceâa labelandaconfidencescoreâisinheritedunchanged.Thesecondappliespost-hocattributionsuchas Grad-CAMandSHAP[10,11]totraineddiagnosers[12],revealingwhichinputregionsdriveadecision. Yetanattributionmapcarriesnotheoreticalreferencevalueagainstwhichitcouldbechecked,soits physicalplausibilityrestsonsubjectivevisualinspection. Thethirddirectioninjectsphysicalknowledgeintotraining,throughinterpretablekernels[13],wavelet- constrainedobjectives[14],orphysics-guidedprobabilisticlabels[15].Theinjectedquantitiesregularize therepresentationbutareconsumedinsidethenetwork;theuserstillreceivesabareclassification.The fourthdirectionacceptsthelabelâconfidenceinterfacebutseekstomaketheconfidenceitselftrustworthy, throughBayesianandevidentialuncertaintyquantification[5,6].Thesemethodsdemonstrablyimprove calibration;yethoweverwellcalibrated,theresultingestimateremainsafunctionofthemodel'sown outputdistributionâwhenaclassifierisconfidentlywrong,distribution-derivedscoresareblindby construction.Thefifthemployslargelanguagemodels,asdiagnosersoperatingonmeasurementdata[16] orasquestion-answeringfront-ends[17].Thesesystemsimproveaccessibility,butcouplingthediagnosis toagenerativeprocessmakesindividualstatementsdifficulttoattributetoaverifiablesource,and hallucinatedcontentcanentertheveryreportsonwhichmaintenancedecisionsrest.Acrossallfive 3 directions,thequantityreachingtheuserisnotdesignedtobecomparedagainstanindependentlyknown physicalreferenceâthediagnosiscanbeexplained,regularized,calibrated,orverbalized,butnot checked. Thisworkapproachestheproblemfromtheoutputside.WeproposetheDiagnosticEvidenceNetwork (DENet),amulti-taskframeworkthatextendstheoutputfromalabelâconfidencepairtoastructured evidencerecordwiththreeindependentlycheckablequantities:thefaultclassification,apredicted characteristicfrequencycomparableagainstthetheoreticalvaluedeterminedbybearinggeometryand shaftspeed,andatemporallocalizationoftransientimpulsesinspectableagainsttherawwaveform.The predictedfrequencycarriesinformationtheconfidencescorestructurallylacks:itsdeviationfromthe theoreticalvalueactsasaper-predictionvalidationmetricintheV&Vsenseof[7],exposing misclassificationsnearlyaswellasthecalibratedconfidenceitselfandremaininginformativeinthehigh- confidenceregimewhereconfidence-basedchecksareblindâatnoextrainferencecostandwithout ground-truthlabels.DENetisencoder-agnostic;abidirectionalmulti-scaleMambaencoder(BiMS- Mamba)isadoptedasthedefaultbackboneforitslinearcomplexityanditsfittothetwotemporalscales ofbearingfaultsignatures.AQLoRA-adaptedlanguagemodelthentranslatestheevidenceintonatural- languagereportsunderastrictdivisionoflabor:theLLMrendersdecisionsmadebyDENetand contributesnoneofitsown,soeverystatementistraceabletoaspecificevidencefield. Themaincontributionsofthisworkaresummarizedasfollows. (1)DENet,anencoder-agnosticevidenceframeworkthatextendsthediagnosticoutputfromalabelâ confidencepairtoper-samplephysicalevidenceâcharacteristicfrequencypredictionandtransient impulselocalizationâatnostatisticallysignificantaccuracycost.ThefrequencyheadattainsanMAEof about6Hzon1,024-pointsegments,aregimewherespectralestimationisstructurallyinapplicable. (2)Anoperationalreliabilitysignalindependentofsoftmaxconfidence:thedeviationbetweenpredicted andtheoreticalfrequencydetectsmisclassificationswithAUROC0.970/0.871ontwopublicdatasets(one mixed-speed)despitereferencingnoclassifierstatistics,andremainsdiscriminativewithinthehigh- confidencesubsetwhereconfidence-baseddetectorsareblindâcomputableatinferencetimewithout ground-truthlabels. (3)AframeworkfortheresponsibleuseofgenerativeAIinmaintenancereporting,inwhichaQLoRA- adaptedLLMactsstrictlyasatranslatorofDENet'sdecisions,witheverystatementtraceabletoaspecific evidencefield.Detailedprompt-levelprohibitionsstillleave10â12%ofreportscontainingunsupported claims,whereaslightweightfine-tuningreducesthisto2%andeliminatesfabricatedquantities. 4 Theremainderofthispaperisorganizedasfollows.Section2reviewsrelatedwork.Section3detailsthe BiMS-Mambaencoder,theDENetframework,andthereportgenerationpipeline.Section4presents classificationbenchmarks,ablationandencoder-agnosticstudies,diagnosticevidenceanalysis,andreport qualityevaluationonthreepublicdatasets.Section5concludesthepaper. 2.Relatedwork Thissectionreviewsthefourlinesofresearchintroducedaboveâencoderarchitectures(Section 2.1),interpretabilityandphysicalknowledgeintegration(Section2.2),anduncertaintyquantification togetherwithLLM-baseddiagnosis(Section2.3)âexaminingeachagainstasinglecriterion:whether thequantitydeliveredtotheusercanbecheckedagainstareferenceindependentofthemodelitself. Section2.4summarizesthegaps. 2.1Deeplearningmethodsformachineryfaultdiagnosis Deeplearninghasprogressivelyreplacedhand-craftedfeatureengineeringinvibration-basedfault diagnosis.Convolutionalarchitecturesoperatingdirectlyonrawsignalsestablishedstronganti-noise performanceanddomainadaptability[8],andattention-baseddesignsfurtherimprovedfeaturefusion acrosstimeandfrequencydomains,exemplifiedbycross-attentionTransformers[9].Morerecently, selectivestatespacemodels(SSMs)haveenteredthefield:VibrMamba[3]introducedalightweight bidirectionalMambabackboneforone-dimensionalvibrationsignals,MD-BiMamba[18]combined modaldecompositionwithbidirectionalfeaturefusionforaero-engineinter-shaftbearings,andWCamba [19]pursuedlightweightdeploymentforaero-enginediagnosis.TheseSSM-baseddiagnosersinheritthe linearO(T)complexityofMambawhilematchingorexceedingTransformeraccuracy,andtheyrepresent thecurrentfrontierofencoderdesign. Twoobservationssummarizethisline.First,classificationaccuracyhaseffectivelysaturatedonthe standardbenchmarksâmultiplerecentmethodsexceed99.5%ontheCWRUdataset[3,4,9]âleaving littleheadroomontheaccuracyaxis.Second,andmoreimportantlyforthiswork,theevolutionfrom CNNstoSSMshasbeenanevolutionoftheencoderonly:theoutputinterfacehasremainedaclasslabel accompaniedbyasoftmaxconfidencescore,aninternalstatisticthatgivesthemaintenanceengineer nothingtocheckagainstindependentknowledgeofthemachine. 5 2.2Interpretabilityandphysicalknowledgeintegration Effortstoopentheblackboxfallintotwogroups.Thefirstappliespost-hocattributiontotrained diagnosers.Grad-CAM[10]andSHAP[11],originallydevelopedforvisionandtabularmodels,have beenadaptedtovibrationdiagnosisthroughmultilayergradientattribution[12],Fourier-domainanalyses ofwhat1D-CNNslearn[20],andstudiesrelatingSHAPvaluestofaultcharacteristicfrequencies[21]. Thesemethodsrevealwhichinputregionsorfrequencybandsdriveadecision,andinfavorablecasesthe highlightedbandscoincidewithknownfaultfrequencies.Thestructurallimitation,however,isthatan attributionmapcarriesnotheoreticalreferencevalueofitsown:thereisnogroundtruthforwhatthemap shouldlooklike,soitsphysicalplausibilityrestsonsubjectivevisualinspection.Attributionanswers wherethemodellooked,butnotwhetherwhatitsawwascorrect. Thesecondgroupbuildsphysicalknowledgeintothemodelitself.WaveletKernelNet[13]parameterizes thefirstconvolutionallayerasinterpretablewaveletkernels;variationalattentionhasbeenusedto constructintrinsicallyinterpretableTransformers[22];wavelet-constrainedobjectives[14],physics- informeddenoisingmodules[23],andphysics-guidedprobabilisticlabelstructures[15]injectfault mechanismknowledgeintotraining.Thesedesignsdemonstrablyregularizetherepresentationandcan improverobustness.Yettheinjectedquantitiesâkernelcenterfrequencies,envelopeconstraints, probabilisticpriorsâareconsumedinsidethenetworkduringtraining;atinferencetimetheuserstill receivesabareclassification.Thephysicalknowledgeshapesthemodelbutisnotreturnedtotheuserin checkableform. Arelatedlineemploysmulti-tasklearning,attachingauxiliarytaskssuchasoperating-condition recognitionorseverityidentificationtoasharedencoder[24,25].Thesedesignsconfirmthatphysically meaningfulauxiliaryobjectivescancoexistwithclassification.Theirauxiliaryoutputs,however,serveas trainingscaffolding:theyaredesignedtoimprovetheprimarytaskandarediscardedatinference.The frameworkproposedinthisworkinvertsthisroleâtheauxiliaryoutputs(characteristicfrequency, impulselocalization)arethemselvesthedeliverable,retainedatinferenceasper-sampleevidence. 2.3Towardstrustworthydiagnosis:uncertaintyquantificationandLLM-based approaches Afourthlineacceptsthelabel-plus-confidenceinterfacebutseekstomaketheconfidencetrustworthy. Bayesianandevidentialformulationsquantifypredictiveuncertaintyandcalibrateconfidence[5,6,26], andfromaV&Vperspectivesuchcalibrationisaprerequisitefordeploymentinsafety-criticalsettings [7].Inthebroadermachinelearningliterature,misclassificationdetectionhasbeenformalizedaround 6 scoresderivedfromtheclassifier'soutputdistributionâmaximumsoftmaxprobability[27],energy scores[28]âandaroundselectiveclassificationframeworksthattradecoverageforrisk[29].These approachesareeffectiveand,inthecaseofMSP,remarkablystrongbaselines.Theirsharedstructural property,however,isthateveryscoreisafunctionoftheclassifier'sownstatistics:whenamodelis confidentlywrong,confidence-deriveddetectorsareblindbyconstruction.BayesianUQadditionally requiresdedicatedprobabilisticarchitecturesormultipleforwardpasses.Thefrequency-deviationsignal proposedinthisworkiscomplementaryratherthansubstitutive:itisazero-costby-productofa deterministicnetwork,referencesanexternalphysicalconstantratherthantheoutputdistribution,andâ asshowninSection4.5.1âremainsdiscriminativepreciselyinthehigh-confidenceregimewhere distribution-derivedscoressaturate. Thefifthlineintroduceslargelanguagemodels.LLMshavebeenemployedasdiagnosersthatingest encodedmeasurementdata[16],withsystematicfine-tuningstudiesoncomplexsystems[30], multimodalreasoningframeworksforexplainablebearingdiagnosis[31],andvision-languagequestion- answeringfront-ends[17];furthervariantsuseLLMsfordataaugmentation[32]orjointdiagnosisand prognosisinasinglegenerativemodel[33].Recentsurveyschartarapidexpansionofthisdirection [34,35].Thesesystemsmarkedlyimproveaccessibilityfornon-expertusers.Theunresolveddifficultyis faithfulness:whenthediagnosisiscoupledtoagenerativeprocess,individualstatementsintheoutput cannotbeattributedtoaverifiablesource,andhallucinatedcontentcanentertheveryreportsonwhich maintenancedecisionsrest.Thisconcernisnothypotheticalâintheadjacentdomainofradiology,a systematicreviewofLLM-generatedreportsfoundpersistenthallucinationsandmissingfindings,with fine-tunedmodelsoutperformingpromptedonesonfaithfulnessandexpertalignment[36].Withinfault diagnosis,quantitativeevaluationofreportfaithfulnessâunsupported-claimrates,fabricatedquantities âremainslargelyabsent,andthedivisionoflaborbetweentheneuraldiagnoserandthelanguagemodel israrelymadeexplicit. 2.4Summaryofresearchgaps Thereviewedlinesconvergeonacommonmissingelement.Encoderevolutionhassaturated accuracywithoutchangingtheoutputinterface(Section2.1).Attributionexplainsdecisionswithouta referencetocheckagainst,andphysics-informedtrainingconsumesphysicalknowledgeinternally withoutreturningittotheuser(Section2.2).Uncertaintyquantificationcalibratestheconfidencescore butcannotescapeitsself-referentialcharacter,andLLM-basedsystemsimprovereadabilityatthecostof attributability(Section2.3).Acrossallofthem,noexistingmethoddelivers,foreachindividual prediction,aquantitythatcanbecomparedagainstanindependentlyknownphysicalreference.Table1 7 summarizesrepresentativemethodsagainstthiscriterion.Thisworkfillsthegapfromtheoutputside: DENetextendsthediagnosticoutputtoastructuredevidencerecordwhosefieldsarecheckableagainst bearinggeometry,shaftspeed,andtherawwaveform,andafine-tunedLLMtranslatesâbutnever generatesâdiagnosticcontent. Table1.Comparisonofrepresentativeintelligentfaultdiagnosismethodswithrespecttooutput verifiability. 8 MethodTypeCls Freq. evidence Temporal loc. Reliability signal NL report LLMrole WDCNN[8]CNNââ Softmax (internal) ââ TwinsTransformer [9] Transformerââ Softmax (internal) ââ VibrMamba[3]Mambaââ Softmax (internal) ââ MD-BiMamba [18] Mambaââ Softmax (internal) ââ WCamba[19]Mambaââ Softmax (internal) ââ MLIFN[37] Multi-source fusion ââ Softmax (internal) ââ MultilayerGrad- CAM[12] Post-hoc XAI ââ Saliency map(no reference value) Softmax (internal) ââ WaveletKernelNet [13] Interpretable kernel â Kernel parameters (training- internal) â Softmax (internal) ââ PIprobabilistic network[15] Physics- informed â Consumed intraining â Posterior (internal) ââ EvidentialGRU [26] UQââ Evidential uncertainty (internal) ââ FD-LLM[16]LLMââ Token likelihood (internal) âDiagnostician FaultGPT[17]VLMââââDiagnostician FR-LLM[33] LLM(multi- task) ââââDiagnostician Ours Mamba+ LLM â â(vs. theoretical valuefrom bearing geometry) â (inspectable onraw waveform) External: physical deviation |Îf| âTranslator Notes:"Reliabilitysignal"denotesthequantityavailableatinferencetimeforjudgingwhetheran individualpredictionshouldbetrusted;internalsignalsarefunctionsofthemodel'sownoutput 9 distribution,whereastheproposed|Îf|referencesatheoreticalconstantdeterminedbybearinggeometry andshaftspeed,externaltothemodel."Freq.evidence"and"Temporalloc."indicatewhetherfrequency- levelandtime-domainevidencearedeliveredtotheuseratinference,ratherthanmerelyusedduring training. 3.Methodology Thissectionpresentstheproposeddiagnosticevidenceframework.AsillustratedinFig.1,theframework consistsofthreemodules:(1)theBiMS-Mambaencoderforfeatureextraction,(2)theDiagnostic EvidenceNetwork(DENet)forphysics-constrainedmulti-tasklearning,and(3)adomain-adaptedLLM fornaturallanguagereportgeneration.Theencoderservesasthedefaultbackboneandisreplaceable (Section4.4);theevidence-generatingcomponentsarethecontributionofthiswork. 3.1OverallFramework Givenarawvibrationsignal x â â 1ĂT oflength T ,theproposedframeworkproducesthreeoutputs:afault classificationlabel y cls ,apredictedcharacteristicfrequency y freq ,andatemporalimpactprobabilitymap y imp â [0,1] T .Theseoutputsarefurtherassembledintoastructureddiagnosticevidencerecordand translatedintoanaturallanguagediagnosticreportbyafine-tunedLLM. Theoverallpipelineisformulatedas: h global ,h seq =f enc (x;Ξ enc ) y cls ,y freq ,y imp =f DENet (h global ,h seq ;Ξ DENet ) r=f LLM (Evidence(y cls ,y freq ,y imp );Ξ LLM ) where f enc denotestheBiMS-Mambaencoder, f DENet denotestheDiagnosticEvidenceNetwork,and f LLM denotesthereportgenerationmodule.h global â â d istheglobalfeaturerepresentationusedfor classificationandfrequencyregression, h seq â â TĂd isthesequence-levelrepresentationusedforimpact localization,andristhegeneratednaturallanguagereport. TheencoderandDENetaretrainedjointlyinanend-to-endmannerwithamulti-taskloss.TheLLMis fine-tunedseparatelyonstructuredevidence-reportpairsgeneratedfromthetrainedmodel. 10 Fig.1.Overallarchitectureoftheproposeddiagnosticevidenceframework:(a)theBiMS-Mamba encoder,(b)theDiagnosticEvidenceNetwork(DENet)producingtheclassification(primarytask,bold border),thepredictedcharacteristicfrequencyversusitstheoreticalvalue(|Îf|),andthelocalized 11 transientimpulses,assembledintoastructuredevidencerecord,and(c)aQLoRA-adaptedLLMthat translatestherecordintoareportâalldiagnosticdecisionsaremadein(b).Illustratedwiththe Paderbornthree-classtask. 3.2BiMS-MambaEncoder TheBiMS-Mambaencoderisdesignedtocapturebothlocaltransientpatterns(e.g.,impactimpulses)and globaltemporaldependencies(e.g.,periodicfaultsignatures)invibrationsignals.Itincorporatesthree keydesignchoices:selectivestatespacemodeling,bidirectionalscanning,andmulti-scaleinput processing. 3.2.1SelectiveStateSpaceModel ThecorebuildingblockoftheencoderistheMambablock[GuandDao,2024],basedontheSelective StateSpaceModel(S6).Theinputisfirstexpandedthroughalinearprojectionanda1-Dconvolution withSiLUactivation,yieldingz.UnlikeconventionalSSMswithfixedparameters,theselective mechanismcomputesthestate-spaceparametersasfunctionsoftheinput:. Givenaninputsequence x â â TĂD ,theMambablockfirstprojectstheinputtoanexpandeddimension throughalinearlayeranda1-Dconvolution: z=Ï(Conv1d(Linear(x))) where Ï denotestheSiLUactivationfunction.Theexpandedrepresentationisthenprocessedbythe selectiveSSM,whichcomputesinput-dependentparameters: Bt=LinearB(zt), Ct=LinearC(zt), Î t =softplus(LinearÎ(zt)) Thecontinuous-timestatespaceequationisdiscretizedusingthezero-orderhold(ZOH)rule: A =exp(Î t A), B t=(exp(Î t A)âI)(Î t A) â1 Î t Bt Thehiddenstateisupdatedrecurrently: ht=A htâ1+B tzt, yt=Cth t where A â â NĂN isinitializedasaHiPPOmatrixand N isthestatedimension.Thisformulationachieves linearcomputationalcomplexity O(T) withrespecttosequencelength,comparedtothequadratic O(T 2 ) complexityofTransformer-basedattentionmechanisms. 3.2.2BidirectionalScanning 12 Asingleforwardscancapturescausaldependenciesbutmissesfuturecontextthatmaybeinformativefor faultpatternrecognition.FollowingthebidirectionaldesigninVisionMamba[Zhuetal.,2024],the BiMS-MambaencoderemploystwoindependentMambabranches:aforwardbranchthatscansthesignal from t=1 to t=T ,andabackwardbranchthatscansthereversedsequencefrom t=T to t=1 . Eachbranchconsistsof$L$stackedlayerswithresidualconnectionsandlayernormalization: hfwd (l) =hfwd (lâ1) +Dropout(MambaBlock(LayerNorm(h fwd (lâ1) ))) hbwd (l) =hbwd (lâ1) +Dropout(MambaBlock(LayerNorm(h bwd (lâ1) ))) Thebackwardbranchoutputisreversedtoalignwiththeoriginaltimeaxis,andthetwobranchesare concatenated: hbi=[h fwd (L) |flip(h bwd (L) )]â â TĂ2D where | denotesconcatenationalongthefeaturedimensionand D isthemodeldimension. 3.2.3Multi-ScaleInputProcessing Bearingfaultsignaturesmanifestatdifferenttemporalscales:high-frequencyimpulsesoccuratthescale ofindividualsamplingpoints,whiletherepetitionpatternoftheseimpulsesreflectsthecharacteristic frequencyatacoarserscale.Tocapturebothscalessimultaneously,theencoderprocessestheinputsignal atmultipleresolutions. Givenasetofscalefactors S= s 1 , s 2 ,..., s K (inourimplementation, S=1,2,4 ),theinputsignalis downsampledbyeachfactor: x (k) =Downsample(x,s k )ââ 1ĂâT/s k â Eachdownsampledsignalisprojectedtothemodeldimensionviaashared1-Dconvolutionand processedbythebidirectionalMambaencoderwithsharedweights: h (k) =f biâmamba (Conv1d(x (k) )) â â âT/ s k âĂ2D Forscales $s k >1 ,theoutputisinterpolatedbacktotheoriginallength$T$usinglinearinterpolationto ensuredimensionalconsistency: h (k) =Interpolate( h (k) ,T)â â TĂ2D Themulti-scalefeaturesareconcatenatedalongthefeaturedimensionandfusedviaachannelattention mechanism: 13 h cat =[h (1) |h (2) |âŻ|h (K) ] â â TĂ2KD w=Ï(FC2(ReLU(FC1(GAP(h cat ))))) hseq=w â hcat whereGAPdenotesglobalaveragepooling,FC1andFC2arefullyconnectedlayerswithareduction ratior=4,andÏisthesigmoidfunction.Theglobalfeaturerepresentationisobtainedbytemporalaverage pooling: hglobal= 1 T t=1 T h seq,t â â 2KD 3.3DiagnosticEvidenceNetwork(DENet) Tofurnisheachdiagnosiswithindependentlycheckableevidence,theDiagnosticEvidenceNetwork (DENet)augmentstheclassificationheadwithtwophysics-constrainedauxiliarytasks:characteristic frequencyregressionandimpacteventlocalization.Beyondproducingauditableevidence,theseauxiliary objectivesshapethesharedencodertowardphysicallymeaningfulrepresentations(Section4.5.3). 3.3.1Task1:FaultClassification(PrimaryTask) Theclassificationheadmapstheglobalfeature h global toclasslogitsthroughatwo-layerMLP: y cls =Linear 2 (Dropout(ReLU(Linear 1 (h global ))))ââ C where$C$isthenumberoffaultclasses.Theclassificationlossisthestandardcross-entropy: â cls =â c=1 C y c logp c , p c =softmax(y cls ) c 3.3.2Task2:CharacteristicFrequencyRegression(AuxiliaryTask) Eachbearingfaulttypehasatheoreticallyderivablecharacteristicfrequencydeterminedbythebearing geometryandrotationalspeed.Forexample,theballpassfrequencyoftheouterring(BPFO)is: f BPFO = n 2 f r 1â d D cosα where n isthenumberofrollingelements, f r istheshaftrotationfrequency, d istheballdiameter, D isthe pitchdiameter,andαisthecontactangle. Thefrequencyregressionheadpredictsthenormalizedcharacteristicfrequencyfrom h global : 14 y freq =Linear 2 (ReLU(Linear 1 (h global )))ââ Adual-targetlabelingschemeisadopted:forfaultsamples,theregressiontargetisthetheoretical characteristicfrequencyofthefaulttype;fornormalsamples,thetargetistheshaftrotationfrequencyf_r, whichisthedominantphysicallymeaningfulfrequencyinahealthybearingsignal.Anchoringnormal samplestof_r,ratherthanexcludingthemfromtheregressiontask,ensuresthatthefrequencyhead receivesawell-definedphysicaltargetforeverysampleandpreventsitsoutputonnormalsamplesfrom driftingarbitrarily.Bothpredictedandtargetfrequenciesarenormalizedbyadataset-specificmaximum frequencyconstantf_max(200HzforbothPaderbornandJNU). Becausethenormalizedtargetsspanroughlyanorderofmagnitudeâfromtheshaftfrequencyofnormal samples(0.05â0.125afternormalization)tothefaultcharacteristicfrequencies(0.3â0.62)âaplain meansquarederrorwouldunder-weightthegradientcontributionofsmalltargets.Thefrequencylossis thereforedefinedastherelativesquarederroroverallsamplesinthebatch: â freq = 1 B i ( y freq (i) ây freq (i) y freq (i) ) 2 whichplacestargetsofdifferentmagnitudesonanequalgradientfooting. Toensurethatthefrequencyheadinfersthefrequencyfromthesignalcontentratherthanmemorizinga class-to-frequencylookup,atime-stretchaugmentationisappliedduringtraining:withprobability0.5,a trainingsegmentistemporallyrescaledbyafactorrdrawnuniformlyfrom[0.8,1.25],anditsfrequency targetisscaledaccordinglytor·y_freq.Underthisaugmentation,samplesofthesameclasscarrya continuumoffrequencytargets,sothelabelcannolongerbepredictedfromtheclassidentityaloneâ thenetworkisforcedtoextractthefrequencyfromthetemporalstructureoftheinput.Toavoid boundary-paddingartifacts,segmentsarestoredwithalengthofâ1.25Ă1,024â=1,280points,andthe rescaledwindowisalwayscroppedfromrealsignalbeforebeinginterpolatedtothefixedinputlengthof 1,024points;per-samplez-scorenormalizationisappliedafterthisfinalcropping.Validationandtest dataareneveraugmented.Thesignal-groundednatureofthefrequencyheadisverifiedposthocbytwo independentchecks:onthemixed-speedJNUdataset,thelearnedheadoutperformsthebestclass-wise constantpredictorbyafactorofâ1.5(Section4.5.1,Table9),andfeedingthefrequencyheadarandom vectorinplaceoftheencoderfeaturedestroysitspredictionaccuracywhileleavingclassificationlargely intact(Section4.3,Table7). Thistaskprovidestwoformsofdiagnosticevidence:(1)thepredictedfrequencycanbecomparedagainst thetheoreticalvaluetoassessphysicalplausibility,and(2)theaugmentation-enforcedsignalgrounding shapestheencodertowardfrequency-sensitiverepresentationsratherthansuperficialstatisticalfeatures. 15 3.3.3Task3:ImpactEventLocalization(AuxiliaryTask) Bearingfaultsproduceperiodicimpulseresponsesinvibrationsignals.Localizingtheseimpactevents providestemporalevidenceforthediagnosis.Theimpactlocalizationheadoperatesonthesequence-level featuresandoutputsaprobabilityofimpactoccurrenceateachtimestep: y imp,t =Ï(Linear 2 (ReLU(Linear 1 (h seq,t )))) â [0,1] Thegroundtruthimpulselabelsaregeneratedthroughanautomatedsignalprocessingpipeline,as illustratedinFig.2:(1)Hilberttransformtoextracttheanalyticsignalenvelope,(2)low-passsmoothing oftheenvelopeusinga4th-orderButterworthfilterwithanormalizedcutofffrequencyof0.1,(3)peak detectiononthesmoothedenvelopewithaheightthresholdof1.5Ăthemeanvalueandaminimuminter- peakdistanceof20samples,and(4)binarylabelassignmentwithinawindowof±5samplesaroundeach detectedpeak.Thispipelineidentifiestime-domainregionsofelevatedtransientactivity,whichmayarise fromfault-inducedimpulses,structuralresonances,orbenignmechanicalexcitation.Importantly,the labelsencodethespatialdistributionoftransientenergyratherthanfault-specificperiodicity;even healthybearingsexhibittransientactivityfromnormalmechanicalsources. 16 Fig.2.Automatedimpulselabelgenerationpipelineillustratedonarepresentativeouter-racefaultsample: (a)z-scorenormalizedvibrationsignal;(b)Hilbertenvelopeandlow-passsmoothedenvelopewith detectionthreshold;(c)peakdetection;(d)binarylabelassignmentwithin±5-samplewindows. Duetothesevereclassimbalanceinimpactlabels(positivesamplestypicallyaccountforonly~8%of timesteps),weemployadynamicallyweightedbinarycross-entropyloss: â imp =â 1 T t=1 T w t + y t log y t +w t â (1ây t )log(1ây t ) wherethepositiveclassweightw t + =(1â p )/ p andthenegativeclassweightw t â =1.0,with p beingthemean positiveratiowithineachbatch.Thisdynamicweightingcounteractsthesevereclassimbalancewithout requiringmanualthresholdtuning. 3.3.4Multi-TaskTrainingObjective Theoveralltraininglossisaweightedcombinationofthethreetasklosses: 17 â=â cls +λ freq â freq +λ imp â imp where λ freq and λ imp arehyperparametersthatcontroltherelativeimportanceoftheauxiliarytasks.Inour experiments,wesetλ freq =0.05andλ imp =0.02,chosenonthePaderbornvalidationsetsuchthatthe weightedauxiliarylossesremainroughlyoneorderofmagnitudebelowtheclassificationlossthroughout training,ensuringtheprimarytaskdominatesgradientupdates. 3.3.5StructuredEvidenceGeneration Aftertraining,theDENetproducesastructureddiagnosticevidencerecordforeachinputsignal.This recordincludesthefollowingfields:(1)thepredictedfaultclassandclassifierconfidence,(2)the predictedcharacteristicfrequency(inHz)anditsdeviationfromthetheoreticalvalue,enablingdirect physicalverification,(3)thenumberofdistincttransientimpulseregionsdetectedinthesignalsegment (n_impact_events),thefractionoftimestepsflaggedasimpulsive(impact_coverage),andthemean impulse-activityprobabilityacrossthesegment(impact_intensity),whichtogethercharacterizewhereand howmuchtransientactivitythemodelhaslocalized.Theseevidencefieldsenableamaintenanceengineer toauditthemodel'sreasoningâverifying,forinstance,thatthepredictedfrequencymatchesthe expectedBPFOfortheinstalledbearingâratherthanrelyingsolelyonaclassificationconfidencescore. 3.4LLM-BasedReportGeneration WhilethestructuredevidencefromDENetisinterpretabletodomainexpertsfamiliarwithvibration analysisterminology,itremainsinaccessibletogeneralmaintenancepersonnelwhomustmaketimely decisionsbasedondiagnosticoutputs.Tobridgethisgap,weemployadomain-adaptedlargelanguage model(LLM)totranslatethestructuredevidenceintonaturallanguagediagnosticreportsthatareboth technicallyaccurateandoperationallyactionable. 3.4.1TaskFormulation Weformulatereportgenerationasaconditionaltextgenerationtask.Givenastructuredevidencerecorde producedbyDENetâcomprisingthepredictedfaultclass,confidencescore,predictedcharacteristic frequency,frequencymatchscore,andimpulselocalizationstatisticsâtheLLMgeneratesadiagnostic reportr: r â =argmax P(r|e; Ξ LLM ) Thegeneratedreportmustsatisfythreerequirements:(1)factualconsistencyâallnumericalvaluesmust exactlymatchtheinputevidence,withnohallucinatedquantities;(2)diagnosticreasoningâthereport shouldarticulatethephysicallogicconnectingtheevidencetothediagnosis(e.g.,explainingthatthe 18 predictedfrequencycloselymatchesthetheoreticalBPFO,therebysupportingtheouterracefault hypothesis);and(3)actionablerecommendationsâthereportshouldconcludewithqualitative maintenancesuggestionscommensuratewiththediagnosticconfidence. ItisimportanttoemphasizethattheLLMservesstrictlyasatranslator,notadiagnostician.Alldiagnostic decisionsâclassification,frequencyprediction,andimpactlocalizationâaremadebytheBiMS- MambaencoderandDENet.TheLLMreceivesthesedecisionsasstructuredinputandrenderstheminto readableprose.Thisdesignensuresthatthediagnosticaccuracyisdecoupledfromthelanguagemodel's capabilitiesandthateveryclaiminthegeneratedreportistraceabletoaspecificevidencefield. 3.4.2DomainAdaptationviaQLoRA WeadoptQwen2.5-7B-InstructasthebaseLLMandfine-tuneitusingQLoRA(QuantizedLow-Rank Adaptation)[Dettmersetal.,2024]forparameter-efficientdomainadaptation.QLoRAquantizesthepre- trainedmodelweightsto4-bitNF4precisionandtrainslow-rankadaptermatricesinjectedintothe attentionlayers,reducingGPUmemoryrequirementstoapproximately8GBandenablingfine-tuningon asingleconsumer-gradeGPU. Trainingdataconstruction.Thefine-tuningdatasetconsistsof498evidence-reportpairs,constructed throughatwo-stagepipeline.First,werunthetrainedBiMS-Mamba+DENetmodelonthePaderborn datasettogeneratestructuredevidencerecordsforeachsample.Wethenapplystratifiedsamplingto selectapproximately500sampleswithbalancedclassdistribution(Normal,OR,IR),coveringdiverse confidencelevelsandincludingasmallproportion(~10%)ofmisclassifiedsamplestoteachtheLLMto expressappropriateuncertainty.Second,eachstructuredevidencerecordistranslatedintoareference diagnosticreportusingacommercialLLMAPI(ClaudeOpus4.6(Anthropic))withacarefullydesigned systempromptthatenforcesthethreerequirementsdefinedinSection3.4.1.Theprompttemplate specifiesthereportstructure(DiagnosisResultâFrequencyAnalysisâImpactAnalysisâConclusion &Recommendation),mandatestheuseofexactnumericalvaluesfromtheevidence,andrequires professionalvibrationanalysisterminology. Fine-tuningconfiguration.LoRAadaptersareappliedtothequery,key,value,andoutputprojection matrices( Wq , Wk , Wv , Wo )inallattentionlayers,withrank r=16 ,scalingfactor α=32 ,anddropoutrate 0.05.Thisresultsin10.1Mtrainableparameters(0.13%ofthetotal7.6B).Themodelistrainedfor3 epochswithalearningrateof 2Ă10 â4 ,cosinelearningrateschedule,effectivebatchsizeof16(micro- batchsize2withgradientaccumulationover8steps),andpagedAdamWoptimizerwith8-bitstates.The trainingconvergesinapproximately6minutesonasingleNVIDIARTX4090GPU. 19 4.Experimentsandresults 4.1ExperimentalSetup 4.1.1Datasets Threepublicbearingdatasetsareemployedtovalidatetheproposedframework.Theirkeyspecifications andexperimentalsettingsaresummarizedinTable2. Table2.Summaryofthedatasetsandexperimentalsettings. ItemCWRUPaderbornJNU Bearingtype SKF6205-2RS (driveend) 6203N205/NU205 Samplingrate12kHz64kHz50kHz Operating condition 0HP,1,797rpm N15_M07_F10(1,500rpm, 0.7N·m,1,000N) Mixed(600/800/1,000rpm) Damagetype Artificial(EDM, singlepoint) Artificial(EDM,engraving, drilling) Artificial(wire-cut;0.3Ă0.25 mmforOR/IR,0.5Ă0.15m forball) No.ofclasses 10(Normal+ IR/OR/BallĂ3 sizes) 3(Normal/OR/IR)4(Normal/IR/OR/Ball) Bearingsusedâ K001âK006/ KA01,03,05,07,09/ KI01,03,05,07 N205(Normal,OR,Ball)/ NU205(IR) Segmentlength/ stride 1,024/5121,024/5121,024/512 NormalizationPer-sampleZ-scorePer-sampleZ-scorePer-sampleZ-score Totalsamples 11,832(7,099/ 2,366/2,367) 149,972(59,885/50,182/ 39,905) 17,577(8,793/2,928/2,928/ 2,928) Train/Val/Test6:2:2(stratified) 89,983/29,994/29,995 (stratified) 10,546/3,515/3,516(6:2:2, stratified) CWRUdataset.TheCaseWesternReserveUniversity(CWRU)bearingdatasetiscollectedfromdrive- endbearingsat12kHz.Followingthecommonprotocol,a10-classtaskisconstructedunderthe0HP loadcondition,comprisingonehealthyconditionandninefaultconditions(innerrace,outerrace,andball faults,eachwiththreedamagediametersof0.007,0.014,and0.021inches). 20 Paderborndataset.ThePaderbornUniversity(PU)bearingdatasetcontainsvibrationsignalsfromtype 6203ballbearings,witheachrecordinglasting4s.A3-classtaskisconstructedunderasingleoperating conditionusingthebearingslistedinTable2,allwithartificiallyinduceddamagesproducedbyelectric dischargemachining,electricengraving,anddrilling.Theaccelerationsignalmeasuredatthebearing housingisused. JNUdataset.TheJiangnanUniversity(JNU)bearingdatasetiscollectedfromcylindricalrollerbearings (typesN205andNU205)atasamplingrateof50kHz.A4-classtaskisconstructedcomprisingone healthyconditionandthreefaultconditions(innerrace,outerrace,andball),withartificialdamage introducedbywire-cutelectricaldischargemachining.Datafromthreeshaftspeeds(600,800,and1,000 rpm)aremixedfortrainingandevaluation,makingthistheonlymulti-speedbenchmarkamongthethree datasets. SinceCWRUiscollectedatasinglespeedanditsthreedamagediameterssharethesametheoretical characteristicfrequencyforeachfaultlocation,frequencyregressiondegeneratesintoanapproximate class-to-constantlookup;frequencyevidenceisthereforeevaluatedonlyonthePaderbornandJNU datasets. 4.1.2ImplementationDetails TheproposedframeworkisimplementedinPyTorchandtrainedonasingleNVIDIARTX4090GPU. ThedetailedhyperparameterconfigurationislistedinTable3.Toaccountforthestochasticityoftraining, allreportedresultsofBiMS-Mambaareaveragedoverfiveindependentrunswithdifferentrandomseeds andreportedasmean±standarddeviation.TheLLMfine-tuningconfigurationfollowsSection3.4.2. Table3.Hyperparameterconfigurationoftheproposedframework. ModuleParameterValue BiMS-MambaencoderModeldimensionD64 StatedimensionN16 Localconvolutionwidth4 Expansionfactor2 NumberoflayersL4 ScalefactorsS[1,2,4](sharedweights) Dropoutrate0.1 Totalparametersâ411k DENetLossweightλ_freq0.05 Lossweightλ_imp0.02 21 ModuleParameterValue Time-stretchranger[0.8,1.25] Augmentationprobability0.5 Storagewindow/outputlength1,280/1,024 ImpactlabelgenerationButterworthfilterorder4 Normalizedcutofffrequency0.1 Peakdetectionthreshold1.5Ămean(envelope) Minimumpeakdistance20samples Labelwindowhalf-widthw5samples TrainingBatchsize64 Optimizer/learningrateAdam/1Ă10â»Âł(cosineschedule) Weightdecay1Ă10â»âŽ Maxepochs/patience100/15 Independentruns5 4.1.3EvaluationMetrics Classificationperformanceisevaluatedbyaccuracyandmacro-averagedF1-score.Thefrequency regressiontaskisevaluatedbythemeanabsoluteerror(MAE)inHzbetweenpredictedandtheoretical characteristicfrequencies.ReportgenerationqualityisevaluatedbythesixmetricsdefinedinSection 4.6.1. 4.2ClassificationPerformanceComparison Althoughtheprimarycontributionofthisworkistheconstructionofanauditablediagnosticevidence chainratherthanrawaccuracyimprovement,itisfirstnecessarytoverifythattheproposedframework doesnotsacrificeclassificationperformance.Tables4â6reportthecomparisonresultsontheCWRU, Paderborn,andJNUdatasets,respectively. CWRU(Table4).Forfairnessandreproducibility,theresultsofcomparisonmethodsarequoteddirectly fromtheoriginalpublications.BiMS-Mambaachieves99.81±0.12%accuracy,rankingfirstamongall quotedmethodsincludingtherecentMamba-basedVibrMamba(99.77%).Nevertheless,giventhat severalrecentmethodsexceed99.5%onthisdataset,theremainingdifferencesliewithinstatisticalnoise, andCWRUprimarilyservesasasanitycheck. 22 Paderborn(Table5).Table5includestwoquotedmulti-sourcemethods(DecisionfusionandMLIFN) asreferencepoints,followedbysixself-implementedbaselinestrainedunderidenticalsettingstoensurea strictlycontrolledcomparison.Amongtheself-implementedmethods,BiMS-Mamba(98.68±0.14%) outperformsthestrongestbaseline(ResNet,98.20±0.23%)by0.48percentagepointsandthevanilla Mambaencoder(98.11±0.16%)by0.57percentagepoints,confirmingthatthebidirectionalmulti-scale designprovidesameaningfuladvantageovertheunidirectionalsingle-scaleSSM.Thetwoquotedmulti- sourcemethodsachievehigheraccuracy(98.96%and99.82%)butexploitcomplementarytorque measurementsandfour-times-longerinputsegments;theyareincludedforcontextratherthandirect comparison. JNU(Table6).TheJNUdatasetintroducestwoadditionalchallengescomparedwiththeprevious benchmarks:(1)fourconditionclasses,includingballfaults,and(2)mixed-speedtrainingacrossthree shaftspeeds.Thelatterbroadenstheintra-classvariabilityoffault-relatedfrequencycomponentsand makesfrequencyestimationconsiderablymorechallenging.Allsixmethodsareself-implementedand evaluatedunderidenticalsettings.Overfiveruns,BiMS-Mambaachievesthehighestmeanaccuracyof 99.47±0.11%,outperformingthestrongestbaseline,ResNetat99.27±0.07%,by0.20percentagepoints. TheTransformerencoderobtainsthelowestbaselineaccuracyat98.94±0.20%,followedbyVanilla Mambaat99.10±0.20%,indicatingthatattentionmechanismsandbasicstate-spacemodelingdonot automaticallyprovideanadvantageonthismixed-speedtaskwithoutbidirectionalmulti-scale augmentation.Therepresentativerunusedfortheconfusion-matrixanalysisachieves99.43%accuracy, with20misclassifiedsamplesoutof3,516testsamples,whichisconsistentwiththefive-runaverage. Moreimportantly,ashighlightedinthelasttwocolumnsofallthreetables,noneofthecomparison methodsprovidesfrequency-levelortemporal-leveldiagnosticevidence:theiroutputsterminateataclass labelandaconfidencescore.Theproposedframeworkattainscompetitiveorsuperioraccuracywhile additionallyproducingaphysicallyverifiableevidencechain,whichisexaminedindetailinSections 4.4â4.6. Table4.ComparisonresultsontheCWRUdataset(10-classclassification).Resultsofcomparison methodsarequotedfrom[ref:VibrMamba,Yietal.,Measurement2025]. MethodAccuracy(%)Freq.Imp. 1D-CNN99.32â ResNet99.42â BiLSTM98.86â Transformer99.57â 23 MethodAccuracy(%)Freq.Imp. TRA-ACGAN99.70â MTF-MARN99.50â SE-ResNet99.10â 1DMamba98.04â VibrMamba99.77â BiMS-Mamba(ours)99.81±0.12â Table5.ComparisonresultsonthePaderborndataset.Resultsofcomparisonmethodsarequotedfrom [ref:MLIFN,D.Wangetal.,Adv.Eng.Inform.66(2025)103405]. MethodInputAccuracy(%)Freq.Imp. DecisionfusionVib.+torque98.96±0.29â MLIFNVib.+torque99.82±0.07â 1D-CNNVibration97.35±0.55â ResNetVibration98.20±0.23â TransformerVibration97.63±0.42â BiLSTMVibration98.09±0.08â VanillaMambaVibration98.11±0.16â BiMS-Mamba(ours)Vibration98.68±0.14â Notes:Thetwoquotedmethods(Decisionfusion,MLIFN)fusevibrationandtorquesignalswith4,096- pointinputsegments;theremainingmethodsuseasinglevibrationchannelwith1,024-pointsegments andarethereforenotstrictlycomparabletothequotedresults.Allself-implementedbaselinessharethe sametrainingprotocolasBiMS-Mambaanddifferonlyinencoderarchitecture. Table6.ComparisonresultsontheJNUdataset(4-classclassification,mixed-speedtraining).All methodsareself-implementedunderidenticalsettings(5runs). MethodAccuracy(%)Freq.Imp. 1D-CNN99.11±0.18â ResNet99.27±0.07â 24 MethodAccuracy(%)Freq.Imp. Transformer98.94±0.20â BiLSTM99.22±0.10â VanillaMamba99.10±0.20â BiMS-Mamba(ours)99.47±0.11â Confusionmatrices.Tocharacterizetheper-classbehaviorofthefinalmodel,Fig.3presentsthe confusionmatricesobtainedfromonerepresentativerunonthethreetestsets.Theaccuraciesderived fromthesematricesthereforecorrespondtothatspecificrun,whereasTables4â6reportthemeanand standarddeviationoverfiveindependentruns. OntheCWRUdataset,therepresentativerunmisclassifiesonly5outof2,367testsamples, correspondingtoanaccuracyof99.79%.Threeoftheseerrorsoccuramongball-faultclasseswith differentdamagediameters:oneB021sampleisclassifiedasB007,whiletwoB021samplesare classifiedasB014.TheremainingtwoerrorsconsistofoneIR014sampleclassifiedasIR007andone B014sampleclassifiedasOR014.Thus,mostresidualerrorsarisefromdistinguishingdamageseverities withinthesamefaultcategoryratherthanfromconfusionbetweenfundamentallydifferentbearing conditions. OnthePaderborndataset,therepresentativerunachievesanoverallaccuracyof98.63%.Thedominant confusionoccursbetweenNormalandIR:193NormalsamplesareclassifiedasIR,while178IRsamples areclassifiedasNormal,forminganearlysymmetricconfusionpair.Incontrast,ORachieves99.8%per- classaccuracy,withonly19of10,037samplesmisclassified.Thispatternisconsistentwiththephysical observationthatmildinner-racedamagemayproducetransientpatternsresemblinghealthybearing signatures,whereasouter-racefaultsgenerallygeneratemorepronouncedandspatiallylocalizedBPFO- relatedimpulses.Notably,the182faultsamplespredictedasNormal(178IRand4OR)constitute misseddetections,thecostliesterrortypeinmaintenancepractice;extendingthefrequency-deviation checktoNormalpredictionsisanaturaldirectionforscreeningsuchcases(seeSection5). OntheJNUdataset,therepresentativerunmisclassifies20outof3,516testsamples,correspondingto anoverallaccuracyof99.43%.ThedominantconfusionoccursbetweenIRandBallfaults:sixBall samplesareclassifiedasIR,whilefourIRsamplesareclassifiedasBall.Thisisphysicallyplausible becausebothfaulttypescanproducestrongimpulsiveresponseswithoverlappingspectralcomponents undervaryingoperatingconditions.TheremainingerrorsincludethreeIRsamplesclassifiedasNormal, twoORsamplesclassifiedasBall,twoBallsamplesclassifiedasNormal,andoneBallsampleclassified 25 asOR.Normalsamplesarenearlyerror-free,withonlyonesampleclassifiedasIRandoneclassifiedas Ball. Fig.3.ConfusionmatricesontheCWRU(left,10-class),Paderborn(righttop,3-class),andJNU(right bottom,4-class)testsets.Diagonalcellsreportrow-normalizedclassificationrateswithsamplecounts; colorintensityreflectssamplecountonalogarithmicscale. 4.3AblationStudy Boththeablationstudyandtheencoder-agnosticvalidation(Section4.4)areconductedonthePaderborn datasetâthelargestofthethreebenchmarks(150ksamples)âwhererun-to-runvariationissmallest andcomponent-leveldifferencesof0.1â0.3percentagepointsremainstatisticallyresolvable;thecross- datasetvalidityofthefullframeworkisseparatelyestablishedinTables4â6.Toquantifythecontribution ofeachcomponentofDENet,fivevariantsaretrainedunderidenticalencoderarchitectureandtraining configuration,differingonlyintheauxiliarytasklosses:(1)thefullmodel;(2)w/oFreqHead(λfreq=0); (3)w/oImpactHead(λimp=0);(4)w/oDENet(bothauxiliarylossesremoved,i.e.,aconventional classification-onlymodel);and(5)FreqHeadw/RandomInput,inwhichthefrequencyheadreceivesa randomvectorinsteadoftheencoderfeature,toverifythatfrequencypredictionreliesonlearnedphysical representationsratherthanatrivialclass-to-frequencymapping.Eachvariantistrainedfivetimeswith independentrandomseeds,andresultsarereportedasmean±standarddeviation. Table7.AblationstudyonthePaderborndataset(5runs,mean±std). 26 VariantParamsAccuracy(%)F1(%)Freq.Imp. FullBiMS-Mamba411k98.68 ± 0.1498.62 ± 0.15â w/oFreqHead386k98.54 ± 0.1898.47 ± 0.19ââ w/oImpactHead386k98.43 ± 0.1798.36 ± 0.18ââ w/oDENet(clsonly)361k98.52±0.2098.44±0.21â FreqHeadw/RandomInput411k98.35 ± 0.1998.27 ± 0.20ââ ThreeobservationscanbedrawnfromTable7.First,thefullmodelachievesthehighestmeanaccuracy amongallvariants,whiletherandom-inputvariantperformstheworst,withagapof0.33percentage points.Giventhatalldifferencesliewithinroughlyonestandarddeviation,wedonotclaimanaccuracy benefitfromtheauxiliarytasks;rather,theresultestablishesthatthephysics-constrainedobjectives imposenoaccuracypenaltywhileshapingphysicallymeaningfulrepresentations(Sections4.5.1and 4.5.3).Second,removingtheimpactheadcausesalargerdegradation,from98.68%to98.43%(â0.25%), thanremovingthefrequencyhead,from98.68%to98.54%(â0.14%).Bothsingle-headvariantsremain withinthesameband,withthedifferencesagaincomparabletotherun-to-runvariation.Third,the random-inputvariantstillachievesrelativelyhighclassificationaccuracy(98.35%),butlosestheability topredictthecharacteristicfrequency,asexpectedgiventhatitsfrequencyheadreceivesnosignal information.Togetherwiththeconstant-predictorcomparisononthemixed-speedJNUdataset(Table9), thisestablishesthatthefrequencyheadofthefullmodelgenuinelyextractsfrequency-relevant informationfromthevibrationsignalratherthanmemorizingalookuptablefromclasslabelsto frequencies.Overall,allvariantsliewithinanarrowaccuracybandof98.35â98.68%,demonstratingthat thediagnosticevidencecapabilitiesareobtainedatnomeasurableaccuracycostâconsistentwiththe encoder-agnosticfindinginTable8. 4.4Encoder-AgnosticValidation AcentralclaimofthisworkisthattheDiagnosticEvidenceNetwork(DENet)isageneral-purposeoutput framework,notamoduletiedtoaspecificencoderarchitecture.Tovalidatethisclaim,wereplacethe BiMS-Mambaencoderwiththreerepresentativearchitecturesâa1D-CNN(305kparams),a Transformerencoder(408kparams),andaBiLSTM(612kparams)âandtraineachwithandwithout theDENetauxiliarytasksunderidenticalsettingsonthePaderborndataset(5seeds).Theclassification- only(cls)variantusesthesameencoderwithasingleclassificationhead,whilethe+DENetvariantadds thefrequencyandimpactheads.TheresultsarereportedinTable8. 27 Table8.Encoder-agnosticvalidationonthePaderborndataset(5runs).Î=(+DENet)â(clsonly). EncoderParams(cls/+DENet)clsAccuracy(%)+DENetAccuracy(%)Î 1D-CNN305k/338k97.35±0.5597.21±0.44â0.14 Transformer408k/425k97.63±0.4297.46±0.30â0.17 BiLSTM612k/645k98.09±0.0898.14±0.05+0.05 BiMS-Mamba361k/411k98.52±0.2098.68±0.14+0.16 TwoconclusionscanbedrawnfromTable8.First,theabsoluteaccuracydifferences|Î|areonly0.05â 0.17percentagepoints,whicharecomparabletoorsmallerthantherun-to-runstandarddeviationsofthe correspondingencoders.Apairedcomparisonacrossall20runs(4encodersĂ5seeds)confirmsno statisticallysignificantaccuracydegradationatthep=0.05level,establishingthattheDENetauxiliary tasksimposenomeasurablecostonclassificationperformanceregardlessoftheencoderbackbone. Second,theparameteroverheadintroducedbyDENetismodest,rangingfrom17kfortheTransformerto 50kforBiMS-Mamba,correspondingtoapproximately4.1%â13.9%oftheclassification-onlyencoder parameters. Theseresultsvalidatetheencoder-agnosticdesignofDENet:thediagnosticevidenceoutputsâ characteristicfrequencypredictionandimpulselocalizationâcanbegraftedontodifferentsequence encoderswithoutsacrificingtheprimaryclassificationtask.ThechoiceofBiMS-Mambaasthedefault encoderinthisworkreflectsitssuperiorabsoluteaccuracy,achievingthehighestclassificationaccuracy bothintheclassification-onlysetting(98.52%)andwithDENet(98.68%),aswellasitsstructural suitabilityforvibrationsignals,ratherthanacouplingbetweenDENetandaspecificencoder. 4.5DiagnosticEvidenceAnalysis ThissectionexaminesthequalityofthestructuredevidenceproducedbyDENet,whichconstitutesthe coreinterpretabilityclaimofthiswork. 4.5.1CharacteristicFrequencyPrediction Fig.4presentsthepredictedcharacteristicfrequenciesforallfaultsamplesinthePaderborntestset, groupedbyclassificationoutcome,withtheoreticalvaluesindicatedbydashedlines(BPFO=76.32Hz forouterracefaults,BPFI=123.68Hzforinnerracefaults).Overallfault-classsamples,thefrequency MAEis1.26HzforOR(n=10,037)and12.35HzforIR(n=7,981),foranoverallMAEof6.17Hz.As Fig.4shows,predictionsforcorrectlyclassifiedsamplesclustertightlyaroundthetheoreticalvalues, whereasmisclassifiedsamplesexhibitdrasticfrequencydeviations.Notably,themisclassifiedIRsamples 28 inFig.4(b)concentrateneartheshaft-frequencyanchorratherthanscatteringrandomly,indicatingthat thefrequencyheadtracksthephysicalanchorassociatedwiththeerroneouslypredictedclassâwhichis preciselywhythedeviationfromthetheoreticalvalueofthetruefaultrenderssucherrorsvisible.This contrastforeshadowsthereliabilitystratificationanalyzedlaterinthissection.Thisagreementprovides frequency-domainevidenceforeachdiagnosis:amaintenanceengineercandirectlyverifywhetherthe predictedfrequencymatchestheexpectedcharacteristicfrequencyoftheinstalledbearing,thereby auditingthephysicalplausibilityofthemodel'sdecisioninsteadofblindlytrustingaconfidencescore. Fig.4.PredictedcharacteristicfrequenciesonthePaderborntestset,groupedbyclassificationoutcome: (a)outerracefault,(b)innerracefault.Violinwidthencodessampledensityforcorrectlyclassified samples;misclassifiedsamplesareshownindividuallyasjitteredpoints.Dashedlinesindicatetheoretical values(BPFO=76.32Hz,BPFI=123.68Hz).ThethinlowertailofcorrectlyclassifiedIRsamples correspondstosegmentswithinthetroughofshaft-frequencyamplitudemodulation(cf.Table9); misclassifiedIRsamplesconcentrateneartheshaft-frequencyanchorofthepredictedclass.Resultsare fromarepresentativerun(seed1). JNUfrequencyprediction.Table9reportstheper-classfrequencyMAEontheJNUtestset,which constitutesthemostchallengingfrequencyregressionbenchmarkamongthethreedatasetsduetomixed- speedtraining(600/800/1,000rpm).Despitethesubstantiallybroadenedintra-classfrequencyvariance, DENetachievesanoverallMAEof6.10Hzâcomparabletothesingle-speedPaderbornresultâwith per-classvaluesof10.21Hz(IR),4.62Hz(OR),and3.49Hz(Ball).TheBallfaultclassattainsthe lowesterror,whiletheIRclassexhibitsthehighesterror,consistentwiththephysicalobservationthat innerracefaultimpulsesareamplitude-modulatedbytheshaftrotationandarethereforemoredifficultto characterizespectrallyinshortsignalsegments.ThegapbetweenthemeanandmedianerrorsinTable9 29 revealsthattheerrorstructuresofthetwodatasetsdifferinaninstructiveway.OnPaderborn,theIRerror isdominatedbyaminorityofsegments:halfoftheIRpredictionsdeviatebylessthan4.21Hz,andthe elevatedMAEstemsfromsegmentsfallingwithinthetroughoftheshaft-frequencyamplitude modulation,wherelittlecharacteristicrhythmispresentinthe16mswindow.OnJNU,bycontrast, mixed-speedtrainingbroadenstheIRerrormoreuniformly(medianabsoluteerror8.63Hz),reflecting thecompoundeddifficultyofinferringboththefaultsignatureandtheoperatingspeedfromasingleshort segment. Table9.CharacteristicfrequencypredictiononthePaderbornandJNUtestsets(faultclassesonly;MAE /medianabsoluteerror). DatasetCondition IRMAE/ Med.(Hz) ORMAE /Med. (Hz) BallMAE /Med. (Hz) Overall MAE/Med. (Hz) Const. baseline MAE(Hz) Paderborn Singlespeed(1,500 rpm) 12.35/ 4.21 1.26/0.78â6.17/1.37â JNU Mixedspeed (600/800/1,000rpm) 10.21/ 8.63 4.62/3.883.49/2.736.10/4.339.37 Notes:Medianabsoluteerror(Med.)isreportedalongsideMAEbecausetheerrordistributionsareheavy- tailed:themeanisdominatedbyaminorityoflarge-errorsegments(seetext). Undermixed-speedtraining,alabel-to-frequencylookupisfundamentallylimitedbythespreadofper- speedtheoreticalfrequencies:thebestclass-wiseconstantpredictorattainsanMAEof9.37Hz,roughly 1.5timesthatofDENet(6.10Hz),confirmingthatthefrequencyheadextractsspeed-dependent informationfromthesignalitselfratherthanmemorizingaclass-to-frequencymapping. Comparisonwithspectralresolutionlimits.TocontextualizetheaboveMAEvalues,Table10compares theDENetfrequencypredictionaccuracyagainstthefundamentalfrequencyresolutionofconventional FFT-basedspectralanalysisappliedtothesame1,024-pointinputsegments.TheFFTfrequency resolutionisdeterminedbyÎf=f_s/N,wheref_sisthesamplingrateandNisthesegmentlength. Table10.DENetfrequencyMAEinthesub-period-segmentregime,withtheFFTbinwidthofthesame segmentlengthasascalereference. Dataset f_s (kHz) Segmentduration (ms) Periodspersegment (BPFO) FFTÎf (Hz/bin) DENetMAE (Hz) Paderborn6416.01.2262.50 6.17 JNU5020.5â0.8â1.448.836.10 30 Notes:FFTÎf=f_s/Nisreportedasanintuitivescalereference;MAEandbinwidtharenotdirectly commensurablequantities.Periodspersegment=segmentduratio n/theoreticalBPFOperiod(13.10msforPaderborn;14.7â24.4msforJNU,spanningthethreeshaft speeds). OnthePaderborndataset,each1,024-pointsegmentspansonly16msâapproximately1.2repetition periodsoftheBPFOimpulsetrainandfewerthantwoperiodsoftheBPFI.Atthissegmentlength, periodicity-basedspectralestimationisnotmerelyimprecisebutstructurallyinapplicable:the62.5Hzbin widthplacesthetwofaultcharacteristicfrequencies(BPFO=76.32Hz,BPFI=123.68Hz)inadjacent bins,andenvelope-spectrummethodsrequiremultipleimpulserepetitionsthatthesegmentdoesnot contain.DENetneverthelessrecoversthecharacteristicfrequencywithanMAEof6.17Hz,indicating thattheencoderexploitsinformationotherthanspectralperiodicityâplausiblythemorphologyof individualimpulseresponses(resonancedecayratesandmodulationpatterns),whichencodesthe excitationcontextwithinasinglerepetitionperiod.OntheJNUmixed-speeddatasetthesamesub-period regimeholds,andtheMAEof6.10Hzconfirmsthatthislearnedinferencegeneralizesacrossoperating conditions. Thiscapabilityiscomplementaryto,notareplacementfor,classicalenvelopespectrumanalysis,which achievessub-hertzaccuracyonsufficientlylongrecordings(e.g.,4-secondsegmentsat64kHzyieldÎf= 0.25Hz).Rather,itsvalueliesinprovidinganembeddedphysicalconsistencycheckthatisproduced simultaneouslywiththeclassificationdecisionon1,024-pointsegmentsâasignallengthatwhich traditionalspectralmethodslacktheresolutiontoidentifycharacteristicfrequencies.Thisco-produced frequencyestimateenablesautomatedcross-validationoftheclassificationoutputwithoutrequiringa separatesignalprocessingpipeline.Thisaccuracylevelisalsocommensuratewithitsintendedroleasa consistencycheck:theoverallMAEremainsbelow15%oftheminimumspacingbetweenadjacentfault characteristicfrequencies(â47HzbetweenBPFOandBPFIonPaderborn),sothepredictedfrequency reliablydiscriminatesbetweencompetingfaulthypothesesatthesamplelevel. Frequencydeviationasanindependentreliabilitysignal.Beyondservingasphysicallyverifiableevidence, thepredictedcharacteristicfrequencyenablesanoperationalreliabilitycheck.Wedefinethefrequency deviation|Îf|astheabsolutedifferencebetweenthepredictedfrequencyandthetheoreticalcharacteristic frequencyofthepredictedclass;sinceitreferencesthepredictedratherthanthetruelabel,|Îf|is computableatinferencetimewithoutanyadditionalmeasurementorgroundtruth. Table11(a)benchmarks|Îf|againsttwostandardconfidence-basedmisclassificationdetectors,maximum softmaxprobability(MSP)andtheenergyscore.Although|Îf|isapurelyphysicalsignalthatmakesno 31 referencetotheclassifier'soutputdistribution,itdetectsmisclassificationsnearlyaswellasthecalibrated confidenceitself(AUROC0.970/0.871onPaderborn/JNU,vs.0.983/0.969forMSP):underthesignal- groundedfrequencyhead,erroneousdiagnosesproducefrequencyestimatesthatdeviatedrasticallyfrom thetheoreticalvalueofthepredictedclass,renderingthemvisibleinthephysicaldomain.Combiningthe twosignalsyieldsaconsistentbutmodestfurthergain(AUROC0.984/0.971),withamorepronounced improvementintheoperatingpointonJNU(FPR@95%TPR0.086,vs.0.149forMSPalone)âmost errorsarevisibletobothsignalssimultaneously,sotheheadroomforcombinationissmall.The distinctivevalueof|Îf|liesinsteadwhereconfidenceisstructurallyblind,asquantifiedbelow. Table11(b)makesthestratificationexplicit.OnPaderborn,79%offault-classpredictionsfallwithin5 Hzofthetheoreticalvalue,wheretheerrorrateisonly0.03%,supportingautomatedacceptanceof frequency-consistentdiagnoses.Theerrorrateclimbsmonotonicallywith|Îf|(0.03%â0.52%â1.67% â15.31%),witherrorsconcentratingoverwhelminglyinthetail:205ofthe230misclassificationsfallin the>20Hzbin.TheintermediatebinsarepopulatedmainlybycorrectlyclassifiedIRsegmentsbelonging totheheavytailcharacterizedinTable9,whichdilutesbutdoesnotbreakthegradient;thecorresponding stratificationonJNUspans0.10%â10.77%.Critically,thestratificationsurviveswithinthehigh- confidencesubset:amongsampleswithsoftmaxconfidenceexceeding0.99,diagnoseswith|Îf|>10Hz aremisclassifiedroughly19Ă(Paderborn)and10Ă(JNU)moreoftenthanfrequency-consistentones, with7and2suchhigh-confidenceerrorsflaggedsolelybythefrequencycheckâpreciselythefailures thatnoconfidence-basedmechanismcandetect. Weemphasizethatnocausalclaimismade:alarge|Îf|mayreflectanintrinsicallyambiguoussignal degradingbothheadssimultaneously.Thepracticalvalueisunaffectedâregardlessofmechanism,|Îf| isazero-costby-productofthediagnosisthatstratifiessamplesintorisktiersdifferingbymorethanan orderofmagnitude,enablingasimpleescalationrule(e.g.,|Îf|>10Hzâexpertreview)evenwhenthe classifieritselfreportshighconfidence. Table11.MisclassificationdetectionviafrequencydeviationonthePaderbornandJNUtestsets. (a)Detectionperformanceofunreliabilityscores DatasetScoreAUROCAUPRFPR@95%TPR PaderbornMSP0.9830.380.061 Energy0.9710.270.089 FreqDev0.9700.330.121 MSP+FreqDev0.9840.390.069 JNUMSP0.9690.230.149 32 DatasetScoreAUROCAUPRFPR@95%TPR Energy0.9680.220.173 FreqDev0.8710.330.426 MSP+FreqDev0.9710.320.086 (b)Errorratestratifiedby|Îf| Dataset |Îf|bin (Hz) nErrors Errorrate (%) High-conf.(n/ Errors) High-conf.rate (%) Paderborn0â514,31550.0315,960/70.04 5â101,72090.52 10â20658111.67905/70.77 >201,33920515.31 Total18,0322301.28 JNU0â51,00810.101,382/10.07 5â1041830.72 10â2026341.52278/20.72 >2065710.77 Total1,754150.86 Notes:FreqDev|Îf|=|predictedfrequencyâtheoreticalfrequencyofthepredictedclass|,computableat inferencetimewithoutground-truthlabels.MSP=1âmaximumsoftmaxprobability;Energy=âlogÎŁ exp(logits).Thecombinedscoreisalogisticregressionoverrank-normalizedMSPandFreqDev,fittedon thevalidationsetonly.Theanalysiscoversfault-classpredictions(18,032/29,995samplesforPaderborn; 1,754/3,516forJNU);Normalpredictionsareexcludedastheirregressionanchoristheshaftfrequency (Section3.3.2).ForJNU,theoreticalfrequenciesfollowthemeasuredshaftspeed.Thehigh-confidence subset(confidence>0.99)isstratifiedintotwobinsowingtolimitederrors;withonly15JNU misclassifications,intermediate-binratesfluctuate,andthetailconcentration(10.77%vs.0.10%)isthe robustfinding.Bearinggeometry:6203(Paderborn),N205/NU205(JNU,Section4.1.1). 4.5.2TransientImpulseLocalization Fig.5visualizestheimpulselocalizationresultsforrepresentativeOR,IR,andNormalsamples.Time- domainregionswherethepredictedimpulseprobabilityexceeds0.7arehighlightedontherawvibration waveform.Thehighlightedregionsalignwithvisuallyidentifiabletransientbursts,demonstratingthatthe modelhaslearnedtolocalizetime-domainintervalsofelevatedtransientenergy.Theannotatedstatistics 33 (numberofdiscreteimpulseeventsandtemporalcoverage)confirmthatthemodelfaithfullycapturesthe spatialdistributionoftransientactivityencodedbythelabelgenerationpipeline(Fig.2). Itshouldbenotedthattheimpulselocalizationprovidestemporalevidenceâshowingwheretransient activityoccursâratherthanfault-specificperiodicityinformation.AsshownintheNormalpanelofFig. 5,healthybearingsalsoexhibitmeasurableimpulseactivityarisingfrombenignmechanicalexcitation (valveaction,flowturbulence,structuralresonances).Thediagnosticdiscriminationbetweenfaultand healthyconditionsrestsontheclassificationheadandthefrequencyregressionhead,notontheimpulse statisticsalone.Thevalueofimpulselocalizationliesinenablinganengineertoinspectthespecific signalregionsthatthemodelconsideredinformative,ratherthantreatingtheentirediagnosisasan opaqueprediction. Fig.5.Impulselocalizationresultsforrepresentativesamples:(a)outerracefault,(b)innerracefault,(c) healthybearing.Shadedregionsindicatemodel-predictedimpulseeventsoverlaidontherawvibration waveform. 4.5.3FeatureSpaceVisualization Fig.6comparesthet-SNEprojectionsofthelearnedglobalfeaturesbetweentheclassification-only model(w/oDENet)andtheproposedfullmodel.Withoutthephysics-constrainedauxiliarytasks,the 34 featureclustersexhibitnoticeableinter-classoverlap,particularlybetweenNormalandIRsamples,and eachclassfragmentsintolooselyconnectedsub-clusterswithdiffuseboundaries.Incontrast,thefull modelproducesmarkedlytighterandbetter-separatedclusters:sub-clusterswithineachclassreflect systematicintra-classvariationâattributabletodifferencesindamagegenerationmethods(EDM, engraving,drilling)andmanufacturingtolerancesacrosstestbearingsâwhileinter-classboundaries remainclearlyseparated,indicatingthatthelearnedrepresentationorganizesphysicallymeaningful variationhierarchicallyratherthancollapsingit.Thisvisualevidencecorroboratestheablationfinding thattheDENetauxiliarytasksreshapethelearnedfeaturestowardphysicallystructuredrepresentationsat nocosttoclassificationaccuracy. Fig.6.t-SNEvisualizationoflearnedglobalfeaturesonthePaderborntestset.(a)Classification-only model(w/oDENet);(b)FullBiMS-MambawithDENet.Eachpointrepresentsa1,024-pointtestsegment; colorsindicatefaultclasses. 4.6LLM-BasedReportGenerationEvaluation 4.6.1QuantitativeComparison Toassessthevalueofdomain-specificfine-tuning,wecomparetheQLoRA-adaptedmodelagainstthree baselinesbuiltonthesamebasemodel(Qwen2.5-7B-Instruct),orderedbyincreasingsupervision:(1)ZS- shortâthesameminimal28-tokensystempromptusedbythefine-tunedmodel,specifyingthetaskbut noformattingrulesorcontentconstraints;(2)ZS-fullâthecompleteeight-rulesystempromptusedto generatethetrainingdata(â1,100tokens),thestrongestprompt-engineeringbaseline;and(3)Few-shot âZS-fullplusonetraining-setexemplarreport(â1,500tokens).Allsettingssharethesameheld-out 35 evaluationsetâthelast10%oftheevidenceâreportpairs(n=50),neverseenduringtrainingâand generationparameters(temperature0.7,max600tokens,onegenerationpersample). Reportsareassessedbysixmetricsintwogroups(Table12):afaithfulnessgroupmeasuringwhether generatedcontentstayswithintheboundsoftheinputevidence,andanadaptationgaingroupmeasuring whatfine-tuningaddsbeyondinstructionfollowing;metricdefinitionsandannotationprotocolaregiven inthetablenotes.Thethreeannotation-basedmetricsareratedblindtomodelidentity,withthetwo faithfulnesscountsindependentlyverifiedbyasecondrater. Table12.Reportgenerationqualityacrossfoursettings.ZS=zero-shot;FT=QLoRAfine-tuned.All ratescomputedovern=50held-outsamples;report-levelcounting(areportwithmultipleviolations countsonce). GroupMetricZS-shortZS-fullFew-shotFT FaithfulnessFormatcompliance(%)0.0100.0100.0100.0 Numericalconsistency(%)100.0100.0100.0100.0 Unsupportedclaimrate(%)1812102 Fabricatedquantityrate(%)10640 AdaptationgainRecommendationcalibration(%)26547696 Styleinternalization(ROUGE-L,%)16.030.641.561.4 TheresultsinTable12revealamonotonicsupervisionladderacrossthefoursettings,withthetwometric groupstellingcomplementarystories. Withinthefaithfulnessgroup,formatcomplianceandnumericalconsistencyaresaturatedoncedetailed instructionsareprovided:ZS-shortproducesnostructurallycompliantreports,confirmingthatthereport formatisentirelylearnedbehavior,whereasallothersettingsreach100%onbothmetricsânumerical grounding,inparticular,isguaranteedbythestructuredevidenceinputratherthanbyanyadaptation.The twostricterfaithfulnessmetrics,however,exposearesidualgapthatinstructionscannotclose.The unsupportedclaimratedecreasesmonotonicallywithsupervisionstrength(18%â12%â10%)yet remainsat10%(5/50)evenforthestrongestprompt-basedsetting,despitethesystempromptexplicitly prohibitingsuchcontent;fine-tuningreducesitto2%(1/50)andeliminatesfabricatedquantitiesentirely (0/50vs.4â10%forprompt-basedsettings).Thiscontrastindicatesthatbehavioralconstraintsarerented bypromptcontextbutinternalizedbyfine-tuningâadistinctionwithdirectsafetyimplications,asthe 36 residualunsupportedclaimsinprompt-basedreportsincludefabricatedremaining-useful-lifeestimates, preciselythetypeofcontentthatmustneverreachmaintenancedecisions. Theadaptation-gaingroupshowsthesameladderatlargeramplitude.Recommendationcalibration improvesfrom26%to54%withdetailedinstructionsandto76%withoneexemplar,buteventhefew- shotsettingleavesnearlyaquarterofreportsrecommendingactionsmisalignedwiththeevidence strengthâmostcommonly,hedgedverificationrequestsfordiagnoseswithconfidenceabove0.99and frequencymatchabove0.95.Fine-tuningclosesthisgapto96%(48/50),indicatingthatcalibrated phrasingâanimplicitprofessionalconventionthatresistsexplicitcodificationâisacquiredfrom examplesratherthanfollowedfromrules.ROUGE-Lexhibitsthesametrend(16.0â30.6â41.5â 61.4)andisreportedasasupplementaryindicatorofstylisticinternalizationratherthanfactualquality. Fromadeploymentperspective,thefine-tunedmodelrequiresonlya28-tokensystempromptcompared toâ1,100tokensforZS-fullandâ1,500tokensforFew-shotâa97%reductioninper-inferenceprompt overhead.Combinedwiththesix-minuteadaptationcostonasingleconsumer-gradeGPU,and consideringthatnoprompt-basedsettingmatchesthefine-tunedmodeloneitherstrict-faithfulnessor calibrationmetrics,QLoRAadaptationisbothpracticallynegligibleincostandfunctionallynecessaryfor reliabletranslation. 4.6.2CaseStudy Fig.7presentsanend-to-endcasestudyofthecompletepipelineforanouterracefaultsample, illustratinghowthesamestructuredevidencerecordisrenderedâfaithfullyorotherwiseâunderthree supervisionsettings. TheDENetevidencerecord(Fig.7b)reportsanORclassificationwithconfidence1.0,apredicted characteristicfrequencyof74.63HzagainstthetheoreticalBPFOof76.32Hz(matchscore0.9779, absolutedeviation1.69Hz),andonelocalizedimpulseeventwithanimpactcoverageof0.0791.Bythe reliabilitycriteriaestablishedinSection4.5.1,thisconstitutesastrong-evidencediagnosis:theclassifieris maximallyconfident,andthefrequencydeviationfallswellwithinthe5Hzbandwheretheempirical errorrateislowest. Thethreegeneratedreports(Fig.7c)exhibitthefailurehierarchyquantifiedinTable12.TheZS-short report,generatedwithouttask-specificguidance,producesafree-formnarrativethatomitstherequired four-sectionstructureentirelyâastructuralfailureconsistentwithits0%formatcompliance.TheZS- fullreportfollowstheprescribedstructureandreferencesthenumericalvaluescorrectly,butcontains unsupportedclaims:itassertsthatthefrequencyevidenceis"consistentwithresonancebehavior"and attributesthelocalizedimpulsestoa"defect-inducedimpactsource"âphysicalinterpretationsthat 37 appearinnofieldoftheevidencerecordand,inthelattercase,oversteptheinterpretiveboundaryofthe impulselabelsestablishedinSection4.5.2.Notably,thiscontentappearsdespitetheâ1,100-tokensystem promptexplicitlyprohibitingit,exemplifyingatthelevelofasinglereporttheresidualunsupported- claimrate(10â12%)thatinstructionsalonecannoteliminate.Itsconcludingrecommendationalsohedges towardfurtherverificationdespitethemaximalevidencestrengthâthecalibrationfailuremodethat dominatestheprompt-basedsettingsinTable12. Thefine-tunedreportsatisfiesallthreerequirementsdefinedinSection3.4.1.Itfollowsthefour-section structureandreferenceseverynumericalvaluefaithfully(factualconsistency).Itarticulatesthereasoning chainconnectingevidencetodiagnosisâthepredictedfrequencyof74.63Hzmatchesthetheoretical BPFOof76.32Hzwithin1.69Hz,supportingtheouterracefaulthypothesisâanddescribesthe localizedimpulseactivityastemporalevidenceoftransientenergywithoutoverclaimingfault-specific periodicity,consistentwiththeinterpretiveboundaryestablishedinSection4.5.2(diagnosticreasoning). Finally,itconcludeswithmaintenancerecommendationscommensuratewiththestrongevidenceâ recommendingscheduledcorrectiveactionratherthanhedgedre-verificationâaninstanceofthe calibrationbehaviormeasuredinTable12(actionablerecommendations).Everyclaiminthereportis traceabletoaspecificfieldoftheevidencerecord,andnocontentexceedswhattheevidencesupports, fulfillingthedesignprinciplethattheLLMservesasatranslatorratherthanadiagnostician. 38 Fig.7.End-to-endcasestudyforanouterracefaultsample(Paderborn,Sample4920):(a)rawvibration waveformwithpredictedimpulseregions(p>0.7)overlaid;(b)thestructuredevidencerecordproduced 39 byDENetanditsmappingtoreportclaims;(c)reportsgeneratedfromthesameevidencerecordunder threesupervisionsettingsâZS-short,ZS-full,andQLoRAfine-tunedâwithper-reportcompliance summarizedbeloweachpanel.Statementsunsupportedbytheevidencerecordaremarkedinreditalics. 4.6.3ExpertEvaluation WhiletheautomaticmetricsinTable12assessformatandlexicalfidelity,theydonotmeasurewhether thegeneratedreportsareactuallyusefultopractitioners.Tothisend,ablindexpertevaluationis conducted.Threeevaluatorswithbackgroundsinvibrationanalysisandrotatingmachinerymaintenance independentlyrate50reportpairsrandomlysampledfromtheheld-outevaluationset.Foreachevidence record,thezero-shot(ZS-full)andfine-tunedreportsarepresentedsidebysideinrandomizedorder withoutmodelidentity,andeachreportisratedonafive-pointLikertscalealongthreedimensions:(1) technicalaccuracyâwhetherthestatedreasoningisconsistentwiththeevidenceandwithvibration analysispractice;(2)readabilityâwhetherthereportisclearandwell-organizedformaintenance personnel;and(3)actionabilityâwhetherthemaintenancerecommendationsarespecificand commensuratewiththeassessedseverity.Table13summarizestheresults. Table13.Expertevaluationresults(mean±std,5-pointLikertscale). DimensionZS-fullFine-tuned Technicalaccuracy 3.54±0.164.50±0.07 Readability 3.64±0.164.51±0.04 Actionability 3.83±0.074.48±0.04 AsshowninTable13,thefine-tunedmodeloutperformsthezero-shotbaselineacrossallthree dimensions.Thelargestgapappearsintechnicalaccuracy(4.50vs.3.54),whereevaluatorsnotedthatthe fine-tunedreportsconsistentlyemployprecisevibrationanalysisterminologyandprovidephysically groundedinterpretationsâforexample,correctlyattributingthepredictedfrequencyofNormalsamples tonon-faultspectralcomponentsratherthanleavingthemunexplained.Theactionabilitygap(4.48vs. 3.83)reflectsarecurringpatterninZS-fullreports:recommendingadditionalverificationevenwhenthe evidencestronglysupportsthediagnosis(confidence=1.0,frequencymatch>0.97),whereasthefine- tunedmodelcalibratesitsrecommendationstothediagnosticcertainty.Theseresultsconfirmthat domain-specificfine-tuninginternalizesnotonlythereportstructurebutalsotheprofessionalreasoning conventionsofvibrationanalysispractice. 40 5.Conclusions Thisworkreframedthebottleneckofintelligentbearingfaultdiagnosisfromclassificationaccuracyto theverifiabilityofdiagnosticoutputs,andproposedaframeworkthatextendstheoutputofadiagnoser fromalabelâconfidencepairtoastructured,per-sampleevidencerecord.TheDiagnosticEvidence Network(DENet)augmentsanarbitrarysequenceencoderwithcharacteristicfrequencyregressionand transientimpulselocalization;pairedexperimentsacrossfourencoderarchitecturesshowedthatthis evidenceisobtainedatnostatisticallysignificantaccuracycost,whilethefrequencyheadrecovers characteristicfrequencieswithanMAEofabout6Hzfrom1,024-pointsegmentsâaregimeinwhich spectralestimationisstructurallyinapplicable. Thecentralempiricalfindingconcernsthedeviationbetweenthepredictedandtheoreticalcharacteristic frequency.Thispurelyphysicalquantity,computableatinferencetimewithoutground-truthlabels, detectsmisclassificationsnearlyaswellasthecalibratedsoftmaxconfidence(AUROC0.970and0.871 onPaderbornandJNU),stratifiesdiagnosesintorisktierswhoseerrorratesdifferbymorethananorder ofmagnitude,andâcriticallyâremainsdiscriminativewithinthehigh-confidencesubsetwhere confidence-baseddetectorsarestructurallyuninformative.Deviation-basedescalationrulestherefore provideadeployment-time,per-samplevalidationcheckthatoperatespreciselywhereexistingreliability mechanismsareblind.Complementarily,thereport-generationexperimentsprovidedquantitativegrounds foradeploymentchoice:prompt-levelprohibitionsleft10â12%ofLLM-generatedreportscontaining unsupportedclaims,whereasasix-minuteQLoRAadaptationreducedthisrateto2%andeliminated fabricatedquantities,supportinglightweightfine-tuningoverpromptengineeringwhengenerativemodels participateinmaintenancereporting. Severallimitationsdelineatethescopeoftheseconclusions.First,thefrequency-deviationcheckcurrently coversonlyfault-classpredictions;the182misseddetectionsobservedonPaderbornâfaultsamples predictedasNormal,thecostliesterrortypeinpracticeâfalloutsideitsreach,andextendingthe consistencychecktoNormalpredictions(e.g.,testingthepredictedfrequencyagainsttheshaft-frequency anchor)isanaturalnextstep.Second,nocausalclaimismadeforthedeviationmechanism:alarge deviationmayreflectanintrinsicallyambiguoussignaldegradingbothheads,whichdoesnotaffectits operationalvaluebutcautionsagainstmechanisticinterpretation.Third,allthreebenchmarksemploy artificiallyinduceddamageunderlaboratoryconditions;validatingtheevidenceframeworkonnaturally degradedbearingsandin-servicedataremainsfuturework,asdoesextendingthecharacteristic-frequency formulationtootherrotatingcomponentswiththeoreticallyderivablesignatures,suchasgearmeshing frequencies. 41 Dataavailabilitystatement Thedatawillbemadeavailableonrequest. CRediTauthorshipcontributionstatement YuntongChen:Writingâoriginaldraft,Methodology,Formalanalysis,Conceptualization.Jianyu Liu:Software,Validation,Formalanalysis.GuobinZhao:Writingâreview&editing,Formalanalysis. ZiangWang:Writingâreview&editing,Validation.ChaoChen:Originaldraft,Validation.JuHuang: Writingâreview&editing.XitianTian:Resources,Formalanalysis,Supervision.LijiangHuang: Writingâreview&editing,Supervision,Fundingacquisition,Software,Formalanalysis, Conceptualization. Declarationofcompetinginterest Theauthorsdeclarethattheyhavenoknowncompetingfinancialinterestsorpersonalrelationships thatcouldhaveappearedtoinfluencetheworkreportedinthispaper. DeclarationofgenerativeAIandAI-assistedtechnologiesinthe manuscriptpreparationprocess DuringthepreparationofthisworktheauthorusedGemini(Gemini3.1Pro)inordertoimprovethe readabilityandlanguagequalityofthemanuscript.Afterusingthistool,theauthorreviewedandedited thecontentasneededandtakesfullresponsibilityforthecontentofthepublishedarticle. References [1]R.B.Randall,J.Antoni,RollingelementbearingdiagnosticsâAtutorial,Mech.Syst.SignalProcess.25(2011) 485â520.https://doi.org/10.1016/j.ymssp.2010.07.017. [2]Y.Lei,B.Yang,X.Jiang,F.Jia,N.Li,A.K.Nandi,Applicationsofmachinelearningtomachinefaultdiagnosis: Areviewandroadmap,Mech.Syst.SignalProcess.138(2020)106587. https://doi.org/10.1016/j.ymssp.2019.106587. [3]H.Yi,D.Li,Z.Lu,Y.Jin,H.Duan,L.Hou,F.Z.Duraihem,E.M.Awwad,N.A.Saeed,VibrMamba:A lightweightMambabasedfaultdiagnosisofrotatingmachineryusingvibrationsignal,Measurement249(2025) 116881.https://doi.org/10.1016/j.measurement.2025.116881. [4]W.A.Smith,R.B.Randall,RollingelementbearingdiagnosticsusingtheCaseWesternReserveUniversitydata: Abenchmarkstudy,Mech.Syst.SignalProcess.64(2015)100â131.https://doi.org/10.1016/j.ymssp.2015.04.021. 42 [5]J.Ren,J.Wen,Z.Zhao,R.Yan,X.Chen,A.K.Nandi,Uncertainty-awaredeeplearning:Apromisingtoolfor trustworthyfaultdiagnosis,IEEE/CAAJournalofAutomaticaSinica11(2024)1317â1330. https://doi.org/10.1109/JAS.2024.124290. [6]H.Li,J.Jiao,Z.Liu,J.Lin,T.Zhang,H.Liu,TrustworthyBayesiandeeplearningframeworkforuncertainty quantificationandconfidencecalibration:Applicationinmachineryfaultdiagnosis,ReliabilityEngineering& SystemSafety255(2025)110657.https://doi.org/10.1016/j.ress.2024.110657. [7]S.Bi,M.Beer,S.Cogan,J.Mottershead,StochasticModelUpdatingwithUncertaintyQuantification:An OverviewandTutorial,Mech.Syst.SignalProcess.204(2023)110784. https://doi.org/10.1016/j.ymssp.2023.110784. [8]W.Zhang,G.Peng,C.Li,Y.Chen,Z.Zhang,Anewdeeplearningmodelforfaultdiagnosiswithgoodanti- noiseanddomainadaptationabilityonrawvibrationsignals,Sensors17(2017)425. https://doi.org/10.20944/preprints201701.0132.v1. [9]Z.Gao,Y.Wang,X.Li,J.Yao,TwinsTransformer:RollingBearingFaultDiagnosisbasedonCross-attention FusionofTimeandFrequencyDomainFeatures,(2024).https://doi.org/10.1088/1361-6501/ad53f1. [10]R.R.Selvaraju,M.Cogswell,A.Das,R.Vedantam,D.Parikh,D.Batra,Grad-CAM:VisualExplanationsfrom DeepNetworksviaGradient-BasedLocalization,InternationalJournalofComputerVision128(2020)336â359. https://doi.org/10.1007/s11263-019-01228-7. [11]S.M.Lundberg,S.-I.Lee,Aunifiedapproachtointerpretingmodelpredictions,Advancesinneuralinformation processingsystems30(2017).https://doi.org/10.48550/arXiv.1705.07874. [12]S.Li,T.Li,C.Sun,R.Yan,X.Chen,MultilayerGrad-CAM:Aneffectivetooltowardsexplainabledeepneural networksforintelligentfaultdiagnosis,JournalofManufacturingSystems69(2023)20â30. https://doi.org/10.1016/j.jmsy.2023.05.027. [13]T.Li,Z.Zhao,C.Sun,L.Cheng,X.Chen,R.Yan,R.X.Gao,WaveletKernelNet:AnInterpretableDeep NeuralNetworkforIndustrialIntelligentDiagnosis,IEEETransactionsonSystems,Man,andCybernetics:Systems 52(2022)2302â2312.https://doi.org/10.1109/TSMC.2020.3048950. [14]L.Li,C.Wei,H.Ma,Anovelwaveletconstrainedphysics-informedneuralnetworkforbearingfaultdiagnosis, JournalofAdvancedMechanicalDesign,Systems,andManufacturing19(2025)JAMDSM0032âJAMDSM0032. https://doi.org/10.1299/jamdsm.2025jamdsm0032. [15]Z.Xu,K.Zhao,J.Wang,M.Bashir,Physics-informedprobabilisticdeepnetworkwithinterpretable mechanismfortrustworthymechanicalfaultdiagnosis,AdvancedEngineeringInformatics62(2024)102806. https://doi.org/10.1016/j.aei.2024.102806. [16]L.Lin,S.Zhang,S.Fu,Y.Liu,FD-LLM:Largelanguagemodelforfaultdiagnosisofcomplexequipment, AdvancedEngineeringInformatics65(2025)103208.https://doi.org/10.1016/j.aei.2025.103208. [17]J.Chen,R.Huang,Z.Lv,J.Tang,W.Li,FaultGPT:IndustrialFaultDiagnosisQuestionAnsweringSystemby Vision-LanguageModels,IEEETransactionsonSystems,Man,andCybernetics:Systems(2026)1â14. https://doi.org/10.1109/TSMC.2026.3705972. [18]P.Wang,Y.Song,X.Wang,Q.Xiang,MD-BiMamba:Anaero-engineinter-shaftbearingfaultdiagnosis methodbasedonMambawithmodaldecompositionandbidirectionalfeaturesfusionstrategy,Measurement242 (2025)115870.https://doi.org/10.1016/j.measurement.2024.115870. [19]L.Yang,J.Wan,G.Jiang,M.Huang,H.Lu,W.Lan,WCamba:Lightweightfaultdiagnosisforaero-engine viawide-kernelconvolutionandstatespacemodeling,ResultsinEngineering28(2025)107202. https://doi.org/10.1016/j.rineng.2025.107202. [20]P.Borghesani,N.Herwig,J.Antoni,W.Wang,AFourier-basedexplanationof1D-CNNsformachine conditionmonitoringapplications,Mech.Syst.SignalProcess.205(2023)110865. https://doi.org/10.1016/j.ymssp.2023.110865. [21]T.Yan,X.Xing,T.Xia,D.Wang,Relationbetweenfaultcharacteristicfrequenciesandlocalinterpretability shapleyadditiveexplanationsforcontinuousmachinehealthmonitoring,EngineeringApplicationsofArtificial Intelligence136(2024)109046.https://doi.org/10.1016/j.engappai.2024.109046. 43 [22]Y.Li,Z.Zhou,C.Sun,X.Chen,R.Yan,VariationalAttention-BasedInterpretableTransformerNetworkfor RotaryMachineFaultDiagnosis,IEEETransactionsonNeuralNetworksandLearningSystems35(2024)6180â 6193.https://doi.org/10.1109/TNNLS.2022.3202234. [23]J.-X.Liao,C.He,J.Li,J.Sun,S.Zhang,X.Zhang,Classifier-guidedneuralblinddeconvolution:Aphysics- informeddenoisingmoduleforbearingfaultdiagnosisundernoisyconditions,Mech.Syst.SignalProcess.222 (2025)111750.https://doi.org/10.1016/j.ymssp.2024.111750. [24]J.Park,J.Yoo,T.Kim,J.M.Ha,B.D.Youn,Multi-headde-noisingautoencoder-basedmulti-taskmodelfor faultdiagnosisofrollingelementbearingsundervariousspeedconditions,JournalofComputationalDesignand Engineering10(2023)1804â1820.https://doi.org/10.1093/jcde/qwad076. [25]Z.Xie,J.Chen,Y.Feng,K.Zhang,Z.Zhou,Endtoendmulti-tasklearningwithattentionformulti-objective faultdiagnosisundersmallsample,JournalofManufacturingSystems62(2022)301â316. https://doi.org/10.1016/j.jmsy.2021.12.003. [26]H.Zhou,W.Chen,J.Liu,L.Cheng,M.Xia,Trustworthyandintelligentfaultdiagnosiswitheffective denoisingandevidentialstackedGRUneuralnetwork,JournalofIntelligentManufacturing35(2024)3523â3542. https://doi.org/10.1007/s10845-023-02221-1. [27]D.Hendrycks,K.Gimpel,ABaselineforDetectingMisclassifiedandOut-of-DistributionExamplesinNeural Networks,ArXivabs/1610.02136(2016).https://doi.org/10.48550/arXiv.1610.02136. [28]W.Liu,X.Wang,J.Owens,Y.Li,Energy-basedout-of-distributiondetection,Advancesinneuralinformation processingsystems33(2020)21464â21475.https://doi.org/10.48550/arXiv.2010.03759. [29]Y.Geifman,R.El-Yaniv,Selectiveclassificationfordeepneuralnetworks,Advancesinneuralinformation processingsystems30(2017).https://doi.org/10.48550/arXiv.1705.08500. [30]S.Zheng,K.Pan,J.Liu,Y.Chen,Empiricalstudyonfine-tuningpre-trainedlargelanguagemodelsforfault diagnosisofcomplexsystems,ReliabilityEngineering&SystemSafety252(2024)110382. https://doi.org/10.1016/j.ress.2024.110382. [31]J.Wang,T.Li,Y.Yang,S.Chen,W.Zhai,DiagLLM:multimodalreasoningwithlargelanguagemodelfor explainablebearingfaultdiagnosis,ScienceChinaInformationSciences68(2025)160103. https://doi.org/10.1007/s11432-024-4333-7. [32]Y.Yu,J.C.Ji,Z.Chen,J.Dhupia,L.Tang,Z.Liu,J.Xiang,Largelanguagemodeltoassistdataaugmentation insoftcontrastivelearningforfew-shotmachineryfaultdiagnosis,AdvancedEngineeringInformatics69(2026) 104078.https://doi.org/10.1016/j.aei.2025.104078. [33]Y.Lai,Z.Wu,M.Chen,C.Liu,H.Shao,FR-LLM:Multi-tasklargelanguagemodelwithsignal-to-text encodingandadaptiveoptimizationforjointfaultdiagnosisandRULprediction,ReliabilityEngineering&System Safety269(2026)112091.https://doi.org/10.1016/j.ress.2025.112091. [34]L.Tao,S.Li,H.Liu,Q.Huang,L.Ma,G.Ning,Y.Chen,Y.Wu,B.Li,W.Zhang,Z.Zhao,W.Zhan,W.Cao, C.Wang,H.Liu,J.Ma,M.Suo,Y.Cheng,Y.Ding,D.Song,C.Lu,AnoutlineofPrognosticsandhealth managementLargeModel:Concepts,Paradigms,andchallenges,Mech.Syst.SignalProcess.232(2025)112683. https://doi.org/10.1016/j.ymssp.2025.112683. [35]R.Yan,J.Ren,J.Wen,C.Guo,Z.Zhao,X.Chen,Largemodelsformachineryfaultdiagnosis:Current advancesandfuturedirections,ChineseJournalofMechanicalEngineering39(2026)100277. https://doi.org/10.1016/j.cjme.2026.100277. [36]Y.Artsi,E.Klang,J.D.Collins,B.S.Glicksberg,G.N.Nadkarni,P.Korfiatis,V.Sorin,Largelanguagemodels inradiologyreporting-Asystematicreviewofperformance,limitations,andclinicalimplications,Intelligence- BasedMedicine12(2025)100287.https://doi.org/10.1016/j.ibmed.2025.100287. [37]D.Wang,Y.Song,J.Xing,Y.Zhuang,J.Zhao,Y.Li,Bearingfaultdiagnosismethodbasedonmulti-level informationfusion,AdvancedEngineeringInformatics66(2025)103405.https://doi.org/10.1016/j.aei.2025.103405.