Paper deep dive
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
Yiming Gai, Junde Lu, Xuefei Huang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/9/2026, 6:08:29 AM
Summary
This paper introduces a Multi-Factor Scoring System (MFSS) to comprehensively evaluate Large Language Models (LLMs) across six dimensions: accuracy, conciseness, factual consistency, readability, coherence, and ROUGE score. Using the TruthfulQA dataset, the framework reveals that mainstream LLMs excel in reasoning tasks but struggle with complex facts and ambiguities, with Gemini-2.0-flash achieving the highest composite score of 0.6104.
Entities (10)
Relation Signals (7)
Gemini 2.0 Flash → achievesscore → 0.6104
confidence 95% · peaking at a composite score of 0.6104
Multi-Factor Scoring System → evaluates → Large Language Models
confidence 95% · This study introduces a multifactor scoring paradigm... Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths
TruthfulQA → benchmarks → Multi-Factor Scoring System
confidence 90% · Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths
Multi-Factor Scoring System → incorporatesmetric → Readability
confidence 90% · integrating accuracy, conciseness, factual consistency, readability, and coherence
Multi-Factor Scoring System → incorporatesmetric → Accuracy
confidence 90% · integrating accuracy, conciseness, factual consistency, readability, and coherence
Multi-Factor Scoring System → incorporatesmetric → Factual Consistency
confidence 90% · integrating accuracy, conciseness, factual consistency, readability, and coherence
Large Language Models → struggleswith → Misinformation
confidence 85% · pervasive limitations in navigating complex facts and ambiguities
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.
Tags
Links
- Source: https://arxiv.org/abs/2607.06940v1
- Canonical: https://arxiv.org/abs/2607.06940v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
25,661 characters extracted from source content.
Expand or collapse full text
ComprehensiveEvaluationofLargeLanguageModel Responses:AMulti-FactorScoringSystem YimingGai 12[0009-0002-0947-7182] ,JundeLu 12[0009-0003-3072-6195] ,Xuefei Huang 2( )[0000−0002−8670−1283] andYingLi 12( )[0000-0002-1545-1811] 1 SchoolofComputerScienceandEngineering,BeihangUniversity,Beijing,China gaiym,ljd2406107,liying@buaa.edu.cn 2 DataScienceandIntelligentComputingLaboratory, HangzhouInternationalInnovationInstitute,BeihangUniversity,Hangzhou,Zhejiang311115, P.R.China xuefei.huang@buaa.edu.cn Abstract.Theremarkableperformanceoflargelanguagemodels(LLMs)in linguistictasksunderscoresanurgentneedforcomprehensiveevaluationof theirresponsequality.Prevailingmethods,oftenconfinedtosingular dimensions,fallshortofcapturingthefullspectrumofmodelcapabilities.This studyintroducesamultifactorscoringparadigm,integratingaccuracy, conciseness,factualconsistency,readability,andcoherence,complementedby agraphicaluserinterface(GUI)forvisualizingoutcomes.Evaluationsonthe TruthfulQAdatasetunveilmainstreamLLMs’strengthsinreasoningtasks (peakingatacompositescoreof0.6104)alongsidepervasivelimitationsin navigatingcomplexfactsandambiguities.Transcendingthenarrowlensof traditionalmetrics,thisframeworkoffersatransparent,adaptableavenueto illuminatemodelpotentialanddeficiencies.Thoughpresentlyfocusedon Englishtasks,itshorizonsbeckontowardmultilingualdomains.Thiswork carvesanovelpathforknowledgeengineeringandmodelrefinement. Keywords:LLMEvaluation,Multi-factorScoring,LargeLanguageModels, Benchmarking,ModelComparison 1Introduction LLMshavepermeatedvariousaspectsofhumanlife,fromintelligentassistantsto academicresearch,andtheirinfluenceisubiquitous.However,ascientificand comprehensiveevaluationofLLM'sstrengthsandweaknesses,particularlyits performanceacrossmultipletestingcriteria,remainsapressingchallenge.Traditional evaluationmethodshaveprimarilyfocusedonasingledimension,suchasaccuracyor fluency.However,withthediversificationofapplicationscenarios,relyingona singlemetricisnolongersufficienttofullycapturethemodel'scapabilities.Asa result,thedevelopmentofmulti-dimensionalevaluationframeworkshasbecomea keyresearchfocus. 2YimingGai,JundeLu,XuefeiHuang,YingLi ExistingresearchhasproposedvariousmethodsforevaluatingthequalityofLLM responses.Metricsbasedonn-grammatching,suchasBLEUandROUGE,are widelyusedinmachinetranslationandtextgenerationtasks[1].However,their limitationsinsemanticunderstandingandcontextualcoherencemakethemless suitableforopen-domainquestionanswering.Recently,semanticembedding-based evaluationmethodshavegainedtraction.BERTScore[2],forexample,leveragespre- trainedmodelstocomputesemanticsimilaritybetweentexts,addressingthe shortcomingsoftraditionalmetrics.Forfactualconsistency,datasetslike TruthfulQA[3]havebeendevelopedtotestamodel’struthfulnessandreasoning abilitywithspeciallydesignedquestions.However,thesemethodsoftenoverlook userexperience-relateddimensionssuchasreadabilityandconciseness,makingit difficulttoprovideacomprehensiveevaluationperspective. Inspecializeddomains,suchashealthcareorlaw,theevaluationrequirementsfor LLMsbecomemorecomplex.Thesedomainsnotonlydemandhighfactualaccuracy butalsonecessitatetheproperuseoftechnicalterminologyandlogicalcoherence[4]. However,theapplicabilityofexistingevaluationmethodshasnotbeenfullyvalidated inthesecontexts,astraditionalmetricsandsingle-dimensionalassessmentsare inadequateforaddressingthediversepracticalneeds.Toaddressthis,thispaper proposesamulti-factorscoringsystemthatcombinesfivekeydimensions:accuracy, conciseness,factualconsistency,readability,andcoherence,whilealsointegratingthe ROUGEmetric,asshownintheFig.1.Theaimistoprovideacomprehensiveand scientificevaluationframeworkforLLMresponsequality,validatedthroughthe TruthfulQAdataset.Experimentalresultsdemonstratethatthisframeworkeffectively highlightsLLMs’strengthsinreasoningtasksandtheirpervasivelimitationsin handlingcomplexfactualscenarios. Thecontributionsofthispaperareasfollows: A.Proposingamulti-dimensionalevaluationframework.Thisstudy innovativelydesignsacomprehensivescoringsystemthatincorporatestraditional singlemetrics(suchasaccuracyorfluency).Byincorporatinguserexperience-related dimensionssuchasreadability,conciseness,andcoherence,itprovidesaholistic characterizationofLLMperformanceacrossdiversetasks. B.Enhancingevaluationapplicabilityandtransparency.Byintegratingsemantic embeddingtechniqueswiththeROUGEmetricandvalidatingtheapproachusingthe TruthfulQAdataset,thismethodisnotonlyapplicabletoopen-domainquestion answeringbutalsoprovidesascalablesolutionforthecomplexevaluationneedsof specializeddomains(suchashealthcareandlaw).Additionally,itenhancesthe visualizationandinterpretabilityofresultsthroughtheuseofagraphicaluser interface(GUI). C.Revealingmodelcapabilitiesandlimitations.Throughmulti-factoranalysis, thisstudysystematicallyrevealsthestrengthsofLLMsinreasoningtasks(suchas LogicalFalsehood)andtheirlimitationswhenhandlingcomplexfactsandambiguous information(suchasMisinformation).Thisprovidesdata-driveninsightsand theoreticalguidanceforfuturemodeloptimization. MFSSofLLM3 Fig.1.PiechartofthescoresoffiveLLMsondifferentdimensionsbasedonTruthfulQA 2Relatedwork Largelanguagemodels(LLMs),basedontheTransformerarchitectureandtrainedon extensivecorpora,excelinNLPtasksliketextgenerationandquestionanswering[5]. Deeplearning,akeydriverinAI,leveragesneuralnetworks—modeledonthehuman brain—toadvancepatternrecognition,vision,andNLP[6,7].Neuronsprocessinputs, applytransformations,andadjustoutputsviaweightedconnections[8].Byexploiting inter-layerlinks,deeplearningextractshierarchicalfeatures,enhancingdata representation[9].TheTransformer’sself-attentionmechanismhastransformedNLP featureextraction. TraditionalevaluationmetricssuchasBLEU[1]andROUGE[12]havehistorically beenemployedtoassessmachinetranslationandsummarizationmodelsby measuringn-gramoverlap.However,thesemetricsprimarilycapturesurface-level lexicalsimilaritiesandfailtoaccountfordeepersemanticrelationshipsorcontextual coherence—criticalcomponentsofhumanlanguageunderstanding.Whileeffective forsyntacticcomparisons,theseapproachesexhibitnotablelimitationsinevaluating generativemodels'comprehensionandreasoningcapabilities.Toaddressthese shortcomings,awaveofsemantic-orientedevaluationmethodshasemerged. BERTScore[2]leveragespre-trainedcontextualembeddingstomeasuretext similarity,offeringamorenuancedassessmentofsemanticalignment.TruthfulQA[3] advancesthisparadigmbydesigningadversarialquestionsthatprobethefactual consistencyandreasoningcapabilitiesofLLMs,revealingtheirrobustnessin knowledge-groundedscenarios.Anessentialcomponentunderpinningthese advancementsistokenization,particularlyBytePairEncoding(BPE)[10].BPE iterativelymergesfrequentcharacterpairs,strikingabalancebetweenvocabulary efficiencyandlinguisticexpressiveness.Thismethodsignificantlyenhancesmodel performanceinEnglish-centricbenchmarkssuchasTruthfulQA,enablingmore preciseevaluationsofLLMcapabilities.Despitetheseinnovations,existing evaluationmethodsremainconstrainedtospecificdimensions—eitherfocusingon semanticsimilarityorfactualaccuracy—therebyofferinganincompletepictureof modelproficiency. Thelimitationsoftraditionalevaluationmethodshavespurredthedevelopmentof multidimensionalframeworks.Beyondautomatedmetrics,humanevaluationoffers criticalinsightsintofluency,coherence,andinformativenessthroughdirectanalysis 4YimingGai,JundeLu,XuefeiHuang,YingLi orcomparisonwithhuman-authoredtexts[4].ROUGEvariants,suchasROUGE-N andROUGE-L,measurecontentoverlapusingn-grammatchingandlongestcommon subsequences,capturinglexicalrecallandstructuralsimilarity[13].TheF1score balancesprecisionandrecalltoassessanswerrelevanceandcompleteness.However, thesemetricsstruggletoevaluatereadability,logicalflow,anduserperception. Semanticredundancyanalysis[11]quantifiesrepetitionbutlacksadaptabilityto dynamiccontent,whilefactualconsistencyassessments,oftenreliantonkeyword matchingorexternalknowledgeretrieval[8],fallshortinopen-domainquestion answering.Evaluationcomplexityincreasesinspecializeddomainslikemedicineand law,whereprecisionandcoherencearecritical[4].Conventionalapproaches frequentlyfailinthesecontexts,revealingagapbetweenmetric-drivenassessment andpracticalutility.Consequently,multidimensionalframeworkshavegained traction,integratingdiverseperspectivestocomprehensivelyassessmodel capabilities.Motivatedbythesechallenges,thisstudyproposesamulti-factorscoring systemincorporatingaccuracy,conciseness,factualconsistency,readability,and coherence,informedbyROUGE-basedevaluation[13].LeveragingTruthfulQAasa benchmark,ourapproachtranscendssingle-metriclimitations,providingaholistic viewofLLMs’strengthsandweaknessesinreasoning,factualgrounding,anduser alignment.Thisframeworkoffersastructuredmethodologyformodelevaluation, optimization,andtask-specificadaptation,contributingtoadvancementsin knowledgeengineeringandAI-drivendecision-making. 3Methodology Tocomprehensivelyassesstheresponsequalityoflargelanguagemodels(LLMs),we proposeamulti-factorscoringsystemthatquantitativelyanalyzesresponsesacross sixdimensions:accuracy,conciseness,factualconsistency,readability,coherence, andROUGEscore.Thissystemaimstoestablishatheoreticallyrigorous, technologicallyadvanced,andpracticallyviableevaluationframework.Thissection elaboratesonthedesignprinciples,corealgorithms,metriccomputation methodologies,andcomparisonswithexistingevaluationapproaches. 3.1SystemDesignConceptandTechnicalFramework Thedesignofthissystemismotivatedbythetheoreticalnecessityofmulti- dimensionalevaluation,asasinglemetricisinsufficienttocomprehensivelycapture thediverseperformanceofLLMsacrossvarioustasks[14].Toaddressthischallenge, weadoptamodulararchitecturecomprisingadatapreprocessingmodule,asemantic embeddingmodule,andamulti-factorevaluationmodule,ensuringboth methodologicalflexibilityandsystemscalability. Thedatapreprocessingmodulestandardizesinputtext,handlesmissingvaluesand non-stringinputs,andunifiessentencesegmentationrulesusingregularexpressions ([.!?]).Thisstepleveragesestablishedtextnormalizationtechniques[13]toenhance evaluationconsistency. MFSSofLLM5 Thesemanticembeddingmoduleservesasthetechnicalcoreofoursystem, leveragingpretrainedlanguagemodelstotransformtextualresponsesintohigh- dimensionalvectorrepresentations.Weemploytheall-MiniLM-L6-v2model[15], whichisbuiltontheTransformerarchitectureandoffersbothcomputational efficiencyandmultilingualsupport.ComparedtolargermodelssuchasBERT,its lightweightdesignismoresuitableforreal-timeevaluationscenarioswhile maintainingrobustsemanticrepresentationcapabilities. Themulti-factorevaluationmodulecomputesvariousassessmentmetricsbasedon embeddingvectorsandintegratesthemintoacompositescoreusingaweighted averagingstrategy.Thismodularframeworknotonlyenhancescomputational efficiencybutalsofacilitatestheseamlessintegrationofadditionalmodelsor evaluationcriteriainfutureiterations. 3.2TheoryandAlgorithmsofEvaluationMetrics Accuracy.Accuracyquantifiesthesemanticsimilaritybetweenthemodel- generatedresponse A m andthereferenceanswer A g ,servingasafundamental evaluationmetric.Weemploycosinesimilaritytomeasurethealignmentbetweenthe embeddingvectorsofthetwotexts[16].Givenembeddingvectors E m and E g ,the accuracyscoreisdefinedas: AS=cosE m , E g = E m ⋅E g ∥E m ∥E g ∥ . (1) Toensurerobustness,inputtextundergoespreprocessing(nullvaluechecksand stringconversions)topreventinvalidcomputations.Comparedtotraditionaln-gram matchingmethods(e.g.,BLEU),thisapproachprioritizessemanticconsistencyover exactwordmatches. Conciseness.Concisenessevaluateswhetherthemodel’sresponseissuccinctand freefromredundantexpressions.Inspiredbyinformationredundancytheory[10],we designaninverseredundancymetric.Thealgorithmfirstsegments A m and A g into sentencesequencesandthencomputeseachsentence’smaximumsimilaritywiththe reference.Ifthesimilarityfallsbelowapredefinedthreshold,afull-lengthpenaltyis applied;otherwise,thepenaltyisinverselyproportionaltothesimilarityscore.For repetitivesentences,anexponentialpenalty( 2 repeat_count )isapplied,supplementedby aquadraticpenalty( count 2 ⋅0.5 ).Iftheresponse A m lengthexceedstwicethatofthe reference A g ,alength-basedscalingfactorisincorporated.Theredundancyscoreis formulatedas: R= min1.0, redundancy totallength .(2) Thefinalconcisenessscoreiscomputedas: CS=1−R.(3) Factual.Factualconsistencyassessesthealignmentbetweenkeyfactualelements inthemodel’sresponseandthegroundtruth.Weadoptaset-overlap-based approach[3],whereinboth A m and A g aretokenized,convertedtolowercase,and filteredtoremovestopwords(e.g.,“the”,“and”).Givenwordsets W m and W g ,the consistencyscoreiscalculatedas: 6YimingGai,JundeLu,XuefeiHuang,YingLi FC= ∣W m ∩ W g ∣ ∣W g ∣ .(4) If A g isempty,thescoredefaultsto1.0;ifeitherinputisempty,thescoreis0.0. Thisapproachiscomputationallyefficientandwell-suitedforEnglishdatasets.Future enhancementsmayincorporatenamedentityrecognition(NER)toimproveprecision. Readability.Readabilitymeasurestheeaseofcomprehensionofthegenerated response,basedonlinguisticreadabilitytheory[17].Weconsidertwosub-metrics: SentenceLengthScore:Deviationfromanoptimalsentencelengthof17.5words, normalized.LexicalDiversity:Theproportionofuniquewordsintheresponse(+1e-6 smoothingtopreventdivisionbyzero). Theoverallreadabilityscoreisdefinedas: RD=0.6⋅1−min1.0, ∣avg length −17.5∣ 17.5 +0.4⋅ ∣ unique words ∣ ∣words∣+10 −6 .(5) Forsingle-sentenceoremptyresponses,adefaultscoreof0.5isassignedto maintainfairness. Coherence.Coherenceevaluatesthelogicalflowwithintheresponse,groundedin discoursecoherencetheory[18].Theresponse 퐀 퐀 issegmentedintosentences 퐀= 퐀 1 ,퐀 2 ,...,퐀 퐀 ,andcosinesimilarityiscomputedbetweentheembeddingvectorsof adjacentsentences.Thecoherencescoreisobtainedbyaveragingthesesimilarity values: 퐀�= 1 퐀−1 퐀=1 퐀−1 cos (퐀 퐀 퐀 ,퐀 퐀 퐀+1 ) .(6) Forsingle-sentenceoremptyresponses,thescoredefaultsto1.0.Thismethod effectivelycaptureslocalcohesion,makingitparticularlysuitableforevaluating short-formresponses. ROUGEScore.Tocomplementsemantic-basedevaluations,weincorporate ROUGEscores[12]toquantifyn-gramoverlapbetween 퐀 퐀 and 퐀 퐀 .Specifically,we computeROUGE-1,ROUGE-2,andROUGE-LF1scores,whichareintegratedas follows: 퐀=0.5⋅퐀 1 +0.2⋅퐀 2 +0.3⋅퐀 퐀 .(7) Toenhancescoredifferentiation,thefinalscoreisscaledby1.2,withanupper boundof1.0,therebyemphasizingtheimportanceofROUGE-1andROUGE-Lin lexicalsimilarity. 3.3ComprehensiveScoringMethod ThefinalTotalScore(TS)iscomputedusingaweightedaverageofthesixevaluation metrics.Thedefaultweightassignmentsare:Accuracy(0.25),Conciseness(0.10), FactualConsistency(0.20),Readability(0.15),Coherence(0.15),andROUGEScore (0.15).Theseweightsareempiricallydeterminedbasedontaskrequirementsand inter-metriccorrelations[19]: TS= i w i ⋅M i ,(8) MFSSofLLM7 where 퐀 퐀 representstheweightand 퐀 퐀 denotestheindividualmetricscores.This frameworkallowsfordynamicweightadjustmentstoadapttodifferentevaluation scenarios. 3.4ComparisonwithExistingMethods Comparedtotraditionalevaluationapproaches,oursystemoffersseveral advantages: BLEU[21]:WhileBLEUreliesonn-gramoverlap,oursemanticembedding approachenhancessemanticsensitivityinevaluation. BERTScore[16]:UnlikeBERTScore,whichfocusesprimarilyonsemantic similarity,oursystemintroducesconciseness,factualconsistency,anduser experiencemetrics(i.e.,readabilityandcoherence)toprovideamorecomprehensive assessment. ROUGE:ByintegratingROUGE,wecomplementsemantic-basedevaluationwith surface-leveltextsimilarity,ensuringamorebalancedassessmentofbothlexicaland semanticfidelity. Thismulti-factorevaluationframeworkistheoreticallymorecomprehensiveand technicallybetteralignedwiththecomplexapplicationrequirementsofLLMs. 4ExperimentalProcessandResults Thissectiondescribeshowweapplythemulti-factorscoringsystemtoevaluatethe responsequalityofLLMs,includingtheexperimentalsetup,datapreprocessing, evaluationprocess,andresultanalysis.WeselectedtheTruthfulQAdatasettotest fivemainstreamLLMs,andthroughbothquantitativeandqualitativeanalysis,we revealtheperformancecharacteristicsofeachmodel. 4.1ExperimentalSetup DatasetSelection.WeselectedtheTruthfulQAdataset[1],whichcontains817 Englishquestionsspanningvariouscategoriessuchasscience,history,andhealth, andisdesignedtotestthemodel'struthfulnessandreasoningability.Eachquestionis pairedwithanoptimalanswer,providingastandardreferenceforevaluation.The choiceofTruthfulQAismotivatedbyitscompatibilitywiththeEnglishlanguageand theBPEencoding,whichallowsLLMstofullyleveragetheirstrengthsinEnglish- languagetasks. ModelSelection.ThefollowingfiveLLMswereevaluatedintheexperiment: qwq_plus_latest:AmultilingualmodeldevelopedbyAlibabaCloud,knownforits efficientreasoningcapabilities. deepseek_v3:Anopen-sourcemodelfocusedondeeplearningoptimization. 8YimingGai,JundeLu,XuefeiHuang,YingLi doubao-1-5-pro-32k:AconversationalmodellaunchedbyByteDance, emphasizinggenerationfluency. moonshot-v1-8k:Knownforitsstrongsemanticunderstanding,instruction following,andtextgenerationabilities. gemini-2.0-flash:Amultilingualmodelthatexcelsinmultimodalunderstanding andreasoning. DataPreprocessing.TheexperimentaldataissourcedfromtheTruthfulQACSV file,whichincludescolumnssuchas"Question"and"BestAnswer."Toensure consistencyintheevaluation,thefollowingpreprocessingstepswereperformed: DataCleaning.Thefilewasreadusingpd.read_csv,followedbychecksand removalofnullvaluesorinvalidresponses. FormatStandardization.Bothmodelresponsesandreferenceanswerswere convertedtostrings,andnon-stringinputswereprocessedtoensurecompatibility. MetadataAddition.Modelidentification(Model),questioncategory(Category), andresponsetype(Type)wereaddedtoeachrecordtofacilitategroupedanalysis. FileManagement.TheresponsesofeachmodelwerestoredasseparateCSVfiles, withfilepathsspecifiedbyparameters,andtheoutputresultsweresavedtoa designateddirectory. 4.2EvaluationProcess Theevaluationprocessisimplementedbasedontheevaluate_model_responses function,withthefollowingsteps: 1)DataTraversal.Thedatafilesarereadonebyone,extractingthequestions, modelresponses,andreferenceanswers. 2)MetricCalculation.Sixmetricsarecomputedforeachpairofresponses: Accuracy:Computesthecosinesimilarityoftheembeddingvectors. Conciseness:Calculatesredundancyandtakestheinverse. FactualConsistency:Computesthewordsetoverlaprate. Readability:Combinessentencelengthandvocabularydiversity. Coherence:Computestheaveragesimilaritybetweensentences. ROUGEScore:WeightedintegrationofROUGE-1,ROUGE-2,andROUGE-L. OverallScoring.Thetotalscoreiscalculatedusingthefollowingweights: Accuracy(0.25),Conciseness(0.10),FactualConsistency(0.20),Readability(0.15), Coherence(0.15),andROUGEScore(0.15). 3)ResultStorage.Adetailedresultstableandamodelsummarytableare generated,containingtheaveragevaluesforeachmetricandthetotalscore. Exceptionhandlingisintegratedthroughouttheprocess.Ifafileismissingora calculationerroroccurs,awarningisrecordedandtheerrorisskipped,ensuringthe stabilityoftheexperiment. MFSSofLLM9 4.3ResultAnalysis Fig.2.Category-WisePerformanceofLLMsonTruthfulQA Basedontheevaluationresults,wecandrawthefollowingconclusions,asshownin theFig.3,Table1: Table1.PerformanceMetricsofLLMs ModelAccuracyConcisenessFactualConsistencyTotalScore DeepSeek-v30.67190.56870.42280.6088 Doubao-1.5-pro-32k0.66890.54030.44450.5873 Gemini-2.0-flash0.67250.58970.39750.6104 QWQPlusLatest0.64960.51450.41260.5800 moonshot-v1-8k0.59050.49270.29360.5499 Table2.Deepseek-V3 categoryscore LogicalFalsehood0.7356 MandelaEffect0.7257 Misinformation0.2963 Confusion:People0.3259 DeepSeek-v3demonstratesanoverallperformancescoreof0.6088,reflectinga mixedcapabilityacrossvariousmetrics.ItexcelsparticularlyinhandlingLogical Falsehoodquestions,whereitsstrengthsshinethrough,butitstrugglesnotablywith Misinformationquestions,asshownintheFig.3(a),Table2. 10YimingGai,JundeLu,XuefeiHuang,YingLi Table3.doubao-1.5-pro-32k categoryscore LogicalFalsehood0.7540 MandelaEffect0.7426 Misinformation0.2497 IndexicalError:Location0.2784 Doubao-1.5-pro-32kexhibitsanoverallperformancescoreof0.5873,indicatinga variedproficiencyacrossitsevaluatedmetrics.ItshinesinaddressingLogical Falsehoodquestions,showcasingitsstrengthsinspecificreasoningtasks,yetitfalters significantlywithMisinformationquestions,asshownintheFig.3(b),Table3. Table4.gemini-2.0-flash categoryscore Misconceptions:Topical0.7628 LogicalFalsehood0.7414 Misinformation0.2950 Confusion:People0.3617 Gemini2.0Flashexhibitsanoverallperformancescoreof0.6104,showcasinga balancedyetvariedproficiencyacrossitsevaluatedmetrics.Itstandsoutin addressingMisconceptions:Topicalquestions,whereitperformsatitspeak,butit falterswhentacklingMisinformationquestions,asshownintheFig.3(c),Table4. Table5.qwq_plus_latest categoryscore MandelaEffect0.7353 LogicalFalsehood0.7059 Misinformation0.2652 Confusion:People0.3027 Theqwq-plus-latestmodelachievesanoverallperformancescoreof0.5800, indicatingamoderatelevelofeffectivenessacrossitsevaluatedmetrics.It demonstratesparticularstrengthinhandlingMandelaEffectquestions,whereits capabilitiesstandout,butitshowsarelativeweaknessinaddressingMisinformation questions,asshownintheFig.3(d),Table5. Table6.moonshot-v1-8k categoryscore Politics0.7406 Subjective0.6601 Confusion:People0.3492 IndexicalError:Time0.3607 MFSSofLLM11 Moonshot-v1-8kexhibitsanoverallperformancescoreof0.5499,indicatinga variedproficiencyacrossitsevaluatedmetrics.Itdemonstratesparticularstrengthin addressingPoliticsquestions,whereitlikelyleveragesitscapabilitieseffectively,but itshowsnotableweaknessinhandlingConfusion:Peoplequestions,asshowninthe Fig.3(e),Table6. Basedonthecomprehensiveanalysisofthemodels'performanceacrossvarious metrics,weprovidethefollowingrecommendationstailoredtospecificscenario requirements.Forscenarioswherehighaccuracyisessential,theGemini2.0Flash modelisrecommendedduetoitsexceptionalprecision.Incaseswherefactual consistencyisthetoppriority,theDoubao-1.5-pro-32kmodelstandsoutasthe optimalchoice,ensuringreliableandaccurateinformation.Additionally,for situationsthatdemandhighreadability,theDeepSeek-v3modelisthemostsuitable option,offeringclearandengagingoutput.Thesesuggestionsaredesignedtoaddress distinctneedseffectively,basedontheevaluatedstrengthsofeachmodel. Fig.3.PerformanceMetricsofLLMs 5Conclusion Thispaperintroducesamulti-factorscoringsystemtoassessLargeLanguageModels (LLMs)usingtheTruthfulQAdataset.Itevaluatesmodelsonseveralkeymetrics: accuracy,conciseness,factualconsistency,readability,coherence,andROUGE scores.Experimentalresultshighlightuniquestrengthsamongmodels:Gemini2.0 Flashstandsoutforaccuracyandconciseness,DeepSeekV3forreadabilityand coherence,andDoubao-1.5-Pro-32kforfactualconsistency.Whilethemodels performwellinreasoning,theyfacechallengeswithcomplexfactsandambiguity. Unlikesingle-metricevaluations,thisapproachoffersacomprehensiveframeworkfor modelselectionandimprovement.CurrentlyfocusedonEnglish,ithaspotentialfor futuremultilingualexpansion,multimodalapplications,andintegrationwith knowledgegraphsanduserstudies,pendingenhancementsinfactualconsistencyand weightoptimization. (a)(b)(c) (d)(e) 12YimingGai,JundeLu,XuefeiHuang,YingLi References 1.Papineni,K.,Roukos,S.,Ward,T.,Zhu,W.J.:BLEU:Amethodforautomaticevaluation ofmachinetranslation.In:40thAnnualMeetingoftheAssociationforComputational Linguistics(ACL),p.311–318.ACL,Stroudsburg(2002) 2.Zhang,T.,Kishore,V.,Wu,F.,Weinberger,K.Q.,Artzi,Y.:BERTScore:Evaluatingtext generationwithBERT.In:InternationalConferenceonLearningRepresentations(ICLR). Springer,Heidelberg(2020) 3.Lin,S.,Hilton,J.,Evans,O.:TruthfulQA:Measuringhowmodelsmimichuman falsehoods.arXivpreprintarXiv:2109.07958(2022) 4.Lewis,P.,Perez,E.,Piktus,A.,etal.:Retrieval-augmentedgenerationforknowledge- intensiveNLPtasks.In:AdvancesinNeuralInformationProcessingSystems(NeurIPS), vol.33,p.9459–9474.MITPress,Cambridge(2020) 5.Vaswani,A.,Shazeer,N.,Parmar,N.,etal.:Attentionisallyouneed.In:Advancesin NeuralInformationProcessingSystems(NeurIPS),vol.30,p.5998–6008.MITPress, Cambridge(2017) 6.Svozil,D.,Kvasnicka,V.,Pospichal,J.:Introductiontomultilayerfeed-forwardneural networks.ChemometricsandIntelligentLaboratorySystems39(1),43–62(1997) 7.Sun,Z.,Xue,L.,Xu,Y.,etal.:Asurveyondeeplearningresearch.J.Comput.Appl.Res. 29(8),2806–2810(2012) 8.Zhao,D.:Asurveyondeeplearninganddeepreinforcementlearning.ChinaNew Telecommun.21(15),174–175(2019) 9.Liu,J.,Liu,Y.,Luo,X.:Advancesindeeplearningresearch.J.Comput.Appl.Res.31(7), 1921–1930(2014) 10.Rasooli,M.,etal.:Spreadsheetsareallyouneed:ImplementingGPT-2inExcel. Availableat:[https://spreadsheets-are-all-you-need.ai/gpt2/] 11.Clark,E.,August,T.,Serrano,S.,etal.:Allthat’s‘human’isnotgold:Evaluatinghuman evaluationofgeneratedtext.In:59thAnnualMeetingoftheAssociationfor ComputationalLinguistics(ACL),p.7282–7296.ACL,Stroudsburg(2021) 12.Lin,C.Y.:ROUGE:APackageforAutomaticEvaluationofSummaries.In:Text SummarizationBranchesOut,p.74–81.ACL,Barcelona(2004) 13.Graesser,A.C.,McNamara,D.S.,Louwerse,M.M.:Methodsofautomatedtextanalysis. In:TheOxfordHandbookofComputationalLinguistics,p.375–394.OxfordUniversity Press,Oxford(2011) 14.Brown,T.B.,Mann,B.,Ryder,N.,etal.:Languagemodelsarefew-shotlearners.In: AdvancesinNeuralInformationProcessingSystems(NeurIPS),vol.33,p.1877–1901. MITPress,Cambridge(2020) 15.Reimers,N.,Gurevych,I.:Sentence-BERT:SentenceembeddingsusingSiameseBERT- networks.In:2019ConferenceonEmpiricalMethodsinNaturalLanguageProcessing (EMNLP),p.3982–3992.ACL,Stroudsburg(2019) 16.Manning,C.D.,Raghavan,P.,Schütze,H.:IntroductiontoInformationRetrieval.2nd edn.CambridgeUniversityPress,Cambridge(2008) 17.Flesch,R.:Anewreadabilityyardstick.JournalofAppliedPsychology32(3),221–233 (1948) 18.Halliday,M.A.K.,Hasan,R.:CohesioninEnglish.Longman,London(1976) 19.Manning,C.D.,Raghavan,P.,Schütze,H.:IntroductiontoInformationRetrieval. CambridgeUniversityPress,Cambridge(2008)