Paper deep dive
Machines acquire scientific taste from institutional traces
Ziqin Gong, Ning Li, Huaikang Zhou
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/22/2026, 5:44:51 AM
Summary
The paper demonstrates that artificial intelligence can acquire 'scientific taste'—the ability to evaluate the quality of research ideas—by fine-tuning language models on institutional publication records. While frontier models and human expert panels struggle to accurately categorize research pitches (achieving ~31% and ~42% accuracy respectively), fine-tuned models trained on historical journal publication decisions achieve up to 59% accuracy in management and 70% in economics, outperforming both frontier models and human experts.
Entities (5)
Relation Signals (3)
Supervised Fine-Tuning (SFT) → recovers → Scientific Taste
confidence 95% · fine-tuning language models on journal publication decisions recovers evaluative judgment
Institutional Record → contains → Scientific Taste
confidence 90% · Scientific taste was not missing from AI's reach; it was deposited in the institutional record
Frontier Models → performpoorlyon → Scientific Taste
confidence 90% · eleven frontier models... barely exceed chance, averaging 31% accuracy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Artificial intelligence matches or exceeds human performance on tasks with verifiable answers, from protein folding to Olympiad mathematics. Yet the capacity that most governs scientific advance is not reasoning but taste: the ability to judge which untested ideas deserve pursuit, exercised daily by editors and funders but never successfully articulated, taught, or automated. Here we show that fine-tuning language models on journal publication decisions recovers evaluative judgment inaccessible to both frontier models and human expertise. Using a held-out benchmark of research pitches in management spanning four quality tiers, we find that eleven frontier models, spanning major proprietary and open architectures, barely exceed chance, averaging 31% accuracy. Panels of journal editors and editorial board members reach 42% by majority vote. Fine-tuned models trained on years of publication records each surpass every frontier model and expert panel, with the best single model achieving 59%. These models exhibit calibrated confidence, reaching 100% accuracy on their highest-confidence predictions, and transfer this evaluative signal to untrained pairwise comparisons and one-sentence summaries. The mechanism generalizes: models trained on economics publication records achieve 70% accuracy. Scientific taste was not missing from AI's reach; it was deposited in the institutional record, waiting to be extracted. These results provide a scalable mechanism to triage the expanding volume of scientific production across disciplines where quality resists formal verification.
Tags
Links
- Source: https://arxiv.org/abs/2603.16659v1
- Canonical: https://arxiv.org/abs/2603.16659v1
Trouble viewing inline? Open PDF directly →
Full Text
149,697 characters extracted from source content.
Expand or collapse full text
Machinesacquirescientifictastefrominstitutionaltraces ZiqingGong 1 ,NingLi 1 ,HuaikangZhou 1 1 SchoolofEconomicsandManagement,TsinghuaUniversity,Beijing,China Correspondingauthor.Email:lining@sem.tsinghua.edu.cn Abstract Artificialintelligencematchesorexceedshumanperformanceontaskswithverifiableanswers,from proteinfoldingtoOlympiadmathematics.Yetthecapacitythatmostgovernsscientificadvanceisnot reasoningbuttaste:theabilitytojudgewhichuntestedideasdeservepursuit,exerciseddailybyeditors andfundersbutneversuccessfullyarticulated,taught,orautomated.Hereweshowthatfine-tuning languagemodelsonjournalpublicationdecisionsrecoversevaluativejudgmentinaccessibletoboth frontiermodelsandhumanexpertise.Usingaheld-outbenchmarkofresearchpitchesinmanagement spanningfourqualitytiers,wefindthatelevenfrontiermodels,spanningmajorproprietaryandopen architectures,barelyexceedchance,averaging31%accuracy.Panelsofjournaleditorsandeditorialboard membersreach42%bymajorityvote.Fine-tunedmodelstrainedonyearsofpublicationrecordseach surpasseveryfrontiermodelandexpertpanel,withthebestsinglemodelachieving59%.Thesemodels exhibitcalibratedconfidence,reaching100%accuracyontheirhighest-confidencepredictions,and transferthisevaluativesignaltountrainedpairwisecomparisonsandone-sentencesummaries.The mechanismgeneralizes:modelstrainedoneconomicspublicationrecordsachieve70%accuracy. ScientifictastewasnotmissingfromAI’sreach;itwasdepositedintheinstitutionalrecord,waitingtobe extracted.Theseresultsprovideascalablemechanismtotriagetheexpandingvolumeofscientific productionacrossdisciplineswherequalityresistsformalverification. Introduction Scienceisbuiltonverifiablefacts,butitisgovernedbyunverifiablejudgments.Artificialintelligence nowmatchesorexceedshumanperformanceontaskswithobjectivelyverifiableanswers,fromprotein structureprediction[1]toOlympiad-levelmathematics[2]tocompetitiveprogramming[3].Yetthe scientificenterprisedependsequallyonaformoftaste,notreasoning:theabilitytodiscriminatewhich untestedideasareworthpursuing.Thisscientifictasteistacit[8],sociallyconstituted[20],andtransmitted throughexposuretowhatthefieldconsidersexcellent,notthroughexplicitrules.AsAIsystemsgenerate hypothesesandmanuscriptsatunprecedentedscale[4,5],thebottleneckinsciencehasshiftedfrom productiontoevaluation[6].WhetherAIcanacquirethisformoftaste(subjective,contextual,resistantto explicitcodification)definesthenextcriticalboundaryinmachineintelligence,wherecapabilitydrops sharplyfromdomainswithclearanswerstothoserequiringdiscriminationandjudgment[7]. Evaluativejudgmentisaparadigmaticinstanceoftacitknowledge:“wecanknowmorethanwecan tell”[8].Decadesofefforttocodifywhatmakesaresearchideagood(structuredreviewrubrics,editorial frameworksdistinguishingoriginalityfromutility[9,10])havenotproducedagreement.Reviewersagree oncategoricalassignmentatbarelyabovechance[11];thegapbetweenwhatthesystemachieves collectivelyandwhatanyindividualreviewercontributesisvast.Neitherformaltraining,academic rank[13],noreditorialexperience[14]reliablypredictsreviewquality;reviewperformanceactually declineswithexperience[15].Ethnographicworkrevealsthatacademicjudgmentoperatesthrough intuitiveassessmentanddisciplinarysensibility,notruleapplication[16].Novelideas,preciselytheones thatmattermost,generatethegreatestdisagreement[17],andthesystemsystematicallyundervaluesits mostimpactfulwork[18]. Thequalitysignalisreal:itmanifestsinthelong-runsortingofpublicationsacrossprestigetiers[19]. Thatsignalisnonethelessirreduciblytacit,notinPolanyi’soriginalsenseofindividualembodiedskill, butinwhatCollinstermscollectivetacitknowledge[20]:evaluativeunderstandingembeddedin institutionalpracticesthatnoindividualparticipantcanarticulate,yetthatthesystemreliablyenacts[8,21]. InBourdieu’sterms,thisisthehabitusofascientificfield,adispositionacquiredthroughprolonged immersioninwhatthefieldrewards,notthroughrules[48].Explicitevaluativecriteriadonotimprove inter-raterreliability;theknowledgeresistscodificationevenwhencodificationisattempted[11].Yetthe institutionalsystemnonethelessproducesconsistentqualitystratificationovertime.Thisprolonged operationhasleftatrace:yearsofpublicationdecisionsacrossprestigetiersconstituteaninstitutional recordofaccumulatedscientifictaste—aformof‘darkknowledge’[38](evaluativeinformationimplicit ininstitutionaloutcomesbutabsentfromanyexplicitcriterion)thathasneverbeenexploitedasatraining signalforartificialintelligence. Thistacitcharacterexposesadeepasymmetryinartificialintelligence.Indomainswithverifiable outputs,AIsystemsachievesuperhumanperformance[1,2,3];butthisadvantagevanisheswhenthetask shiftsfromverificationtoevaluation.Whendeployedinpeerreview,asettingwheremorethanhalfof researchersalreadyreportusingAI[22,23],largelanguagemodelsconsistentlyinflatescores,recommend acceptanceatratesfarexceedinghumanpanels,andsystematicallyoverlooknovelty,thedimensionthat mostdistinguishesimportantfromincrementalwork[24,25,26].Theyevaluatehowresearchispresented, notwhatitcontributes[27].Thisleniencyisnotapromptingfailurebutapredictableconsequenceof preference-basedpost-training:reinforcementlearning(RL)fromhumanpreferencesinstillssycophantic behaviorthatintensifieswithfurtheroptimization[28,29],producingfluentmodelsthatdefaultto approvalinsteadofthediscriminationthattastedemands. Hereweconstructaheld-outbenchmarkof120article-derivedresearchpitchesinorganizational psychologyandmanagement 1 ,balancedacrossfourqualitytiersandevaluatedin2,914humanratings from48expertgatekeepersand174juniorresearchers,andreplicatetheapproachineconomicswitha 200-articlebenchmark.Wetest26AIconfigurationsspanningfrontierreasoningmodels,chatmodels, supervisedfine-tunedmodels,architecture-matchedbasecontrols,andreinforcement-learningablations, generatingover18,900independentevaluationeventsunderanexpert-derivedpromptoptimizedtofavor frontierperformance.Wealsoevaluateauxiliarypairwisediscriminationtaskstotestwhetherthis evaluativecapacitygeneralizesbeyondthetrainingformat.Frontiermodelsperformonlymarginally abovechance(31.1%versus25%;macro-F10.236).Humanevaluatorsperformsomewhatbetterand preservemorebalancedtierdiscrimination,buttheirjudgmentsremainhighlyvariable,withinter-rater agreementnearzero.NeitherpromptingthemostpowerfulAIsystemsnorassemblingexpertpanels solvestheevaluationproblem. WealignAItotheseinstitutionaltraces.Individualreviewersdisagreesubstantially,yetthe institutionalsystem,operatingthroughiteratedconsensusovertime,producesreliablequalitystratification thatindividualjudgmentsdonot[19].Itisthissystem-levelsignal,notindividualrevieweraccuracy,that supervisedlearningcanexploit[20].Byapplyingsupervisedfine-tuning(SFT)tofourbasemodelsusing historicalresearch-pitch/journal-outcomepairs,wealignthemtoimplicitevaluativecriteriathatneither writteninstructionsnorthelargestAIarchitectureshavesuccessfullytransmitted. Theresultsareclear.Allfourfine-tunedmodelsindependentlyachieve55.0–59.2%accuracy,anda simpletwo-modelpairextendsthisto60.8%,capturingapproximatelysixtimestheheadroombetween chanceandceilingrecoveredbyfrontiermodelsandmorethandoublethatoftheexpertmajorityvote,at atotaltrainingcostbelow$300.Cross-architecturereplicationconfirmsthattheeffectisnotmodel- 1 Organizationalpsychologyandmanagementisalarge,globallydistributedfield(over20,000AcademyofManagement members;1,800+rankedjournalsintheCharteredABSAcademicJournalGuide).Top-tierjournalsmaintainacceptanceratesof 4–10%,creatingawell-documentedinstitutionalhierarchy.Researchqualityisevaluatedprimarilyontheoreticalcontributionand conceptualnovelty,notresourceaccess,isolatingjudgmentsofideaqualityfromconfoundssuchaslaboratoryinfrastructure. Meta-analyticinter-raterreliabilityislow(κ=0.17across48studies[11]),consistentwithevaluationthatisgenuinelytacit. specific:eachsinglemodelexceedsboththefrontiermeanandthebestfrontiermodel.Beyondcategory prediction,thefine-tunedmodelsexhibitcalibratedself-knowledge:accuracyexceeds80%inthetop- confidencefifthandreaches100%inthetop10%ofpredictions,givingthebenchmark’sstrongest practicallyusefulsignalfordistinguishingcorrectfromincorrectjudgments.Withouttrainingonpairwise discrimination,thefine-tunedmodelstransferevaluativejudgmenttohead-to-headpitchcomparisons,a taskformatnotpresentintraining.Thelargerfine-tunedmodelfurtherretainsthisevaluativesignalwhen fullideasummariesarecompressedtoone-sentenceinputsnotencounteredduringtraining,stillexceeding thebestfrontiermodelandexpertpanelsanddemonstratinglearnedevaluativeunderstandinginsteadof memorizedtierassignments.Whentrainedoneconomicsinstitutionaltraces,thesameapproachyields 69.5%accuracy,confirmingthatthemechanismgeneralizesbeyondtheprimaryfield. Thesefindingscarryimplicationsfortheaccelerationofscience.Scientifictaste—thecapacityto judgewhichquestionsareworthasking—isthebindingconstraintonscientificdiscovery,mostacutely inthesocialandbehavioralsciences,wherethevalueofresearchquestionsresistsformalverificationand submissionvolumesfaroutstripexpertbandwidth.Cross-fieldreplicationineconomicsconfirmsthatthe mechanismisnotspecifictomanagement.Thesearepreciselythefieldswherefrontiermodelsfailmost completelyandwherescalable,calibratedevaluationwouldyieldthegreatestreturns.Themechanism shouldgeneralizetootherdomainswithsociallyjudgedoutcomesandweakexanteverification,including creativeindustries,ventureinvesting,grantallocation,hiringandpromotion,andpolicyexecution, wheneverinstitutionaldecisionhistoriesaccumulatelargerecordsofwhatlaterprovedsuccessful[19,32]. Inourbenchmark,frontier-modelperformanceclustersclosetochanceacrossthefullbreadthof proprietaryandopen-weightarchitectures,yetafine-tunedsystemtrainedonaccumulatededitorial decisionsexceedsboththemostpowerfulAIsystemsandtheexpertpanelswhosecollectivedecisions generatedthetrainingdata.ThebarriertoAIevaluationisnotrawcapabilityalone;itistheabsenceofthe righttrainingsignal. Results Abenchmarkforevaluativejudgment Peerreviewrestsonacognitivecapacitythathasresistedformalspecification:theabilitytomakeearly- stagesignificancejudgmentsbeforefullempiricalexecution,whenjournals,funders,andlabsdecide wheretoinvestscarceattentionandresources.Toisolatethiscapacity,weconstructedaheld-out benchmarkinmanagementcomprising120article-derivedresearchpitches,balancedacrossfourquality tiersandspanning15researchdomainsand17journals(ofthe19-journalsourceuniverse;seeMethods andSupplementaryMethodsSM5).ThesourcearticleswereallpublishedafterJune30,2025.Each sourcearticlewastransformedintoaresearchpitchcenteredonthecoreresearchquestionandtheoretical framing,withdetailedmethods,fullempiricalfindings,andpublicationidentifiersremoved.Articleswere assignedtofourqualitytiersbasedontheirpublicationoutlet(exceptional,strong,fair,limited;30pertier; chancebaseline25%;Fig.1).Becauseevaluatorsseeonlytheresearchquestionandtheoreticalframing (notmethods,results,orjournalidentity),tierassignmentrequiresthekindofdiscriminationclosestto taste:judginganidea’spromisefromitspresentationalone. Figure1|Studydesign:institutionaltraces,alignmentpipeline,evaluation,andmechanismtests a,Institutionaltracesofpublicationoutcomes.Journal-tierpublicationoutcomes(exceptional,strong,fair,limited)serveasthe supervisionsignalforsupervisedfine-tuning(SFT).Twotemporallydisjointtraining-dataslicesareconstructedfroma19-journal sourceuniversespanning15researchdomains:arecentslice(2020–2025;4,479research-pitch/journal-outcomepairs)usedfor theprimarySFTmodelsandanolderslice(2015–2020;3,368pairs)usedtotestthetemporalpersistenceoftheinstitutional signal.Bothslicesarefullydisjointfromthe120-pitchheld-outbenchmark.b,SFTalignmentpipeline.Institutionaloutcomesare convertedintostandardizedtrainingrepresentationsandusedtofine-tunefourbasemodelsspanningtwomodelfamiliesand multipleparameterscales(GPT-4.1,GPT-4.1-nano,Qwen3-4B,Qwen3-30B).Eachbasemodelproducesanarchitecture-matched SFTcheckpoint;checkpointoutputsarecombinedviaprobabilityaveragingintoanSFTensemblenode.c,Strictholdout evaluation.Trainingdataandtheheld-outbenchmark(N=120pitches,30pertier)arestrictlyseparated.Threeevaluatorclasses areassessedonthesameunseenbenchmarkunderacommonfour-tieroutputmapping:11frontierreasoningmodels,4SFT modelsplusensembles(ens.),andhumanreviewerscomprising48expertgatekeepers(E)and174juniorresearchers(J).All evaluatorsusethesamefrozenexpert-derivedpromptandidenticaltierdefinitions,ensuringthatperformancedifferencesreflect evaluatorcapabilityratherthantask-formatvariation.d,Mechanismtests.ThreeauxiliaryanalysesprobehowSFTacquires evaluativejudgment.Self-knowledgetest:eachmodelproducesatierpredictiontogetherwithacalibratedconfidencescore;the testaskswhetherconfidenceissystematicallyhigheroncorrectpredictionsthanonincorrectones.Pairwisediscrimination:two researchpitchesarepresentedsidebysideandthemodelidentifiesthestrongerone,ataskformatentirelyabsentfromtraining, testingwhetherSFTdevelopsgeneralizedqualityorderingthattransfersbeyondthecategoricalclassificationusedduringfine- tuning.SFTversusreinforcementlearning(RL)comparison:modelsfromthesamefamilyaretrainedviadirectsupervised alignment(SFT)orreasoning-optimizedRL(usingreasoning-enabledconfigurations),andoutputsarecomparedtotestwhether explicitchain-of-thoughtdeliberationaidsorimpedesevaluativeaccuracyontacitjudgmenttasks. Evaluatorswereaskedtoassigneachpitchtooneofthefourtiersusinganexpert-derivedassessment frameworkcenteredonoriginalityandutility[9,10].Thefullevaluatorpoolcomprisedelevenfrontier reasoningmodels,threechatvariants,foursupervisedfine-tuned(SFT)modelswitharchitecture-matched basecontrols,48expertgatekeepers,and174juniorresearchers,generating2,914humanratingsandover 16,000AIinferencesacross26modelconfigurations(includingprompt-sensitivity,pairwise discrimination,andreinforcement-learningablations),foracombinedtotalexceeding18,900independent evaluationevents.Theevaluationpromptwasselectedastheformulationthatmaximizedfrontiermodel performance,ensuringthatanyadvantageobservedforfine-tunedmodelsrepresentsaconservative estimate. Frontiermodelsandtheanatomyofpredictioncollapse Ifevaluativejudgmentwerereducibletothekindofreasoningthatfrontierlanguagemodelsexcelat (synthesizinginformation,recognizingpatternsintext,applyingcriteriatocases),thenthemostcapable modelsavailableshouldperformwellabovechanceonthisbenchmark,particularlywhengivenmaximal advantage.Wegavethemthatadvantage:anevaluationpromptincorporatingexpert-derivedassessment criteria[9,10],stress-testedacrossvariantstomaximizefrontierperformance(ExtendedDataFig.1). Itisnotenough.Acrossacohortof11frontiermodels,includingGemini3.1Pro,ClaudeOpus4.6, GPT-5.2High,Gemini2.5Pro,Qwen3.5Plus,Grok4.1Fast,andleadingopen-weightsystemssuchas KimiK2.5,DeepSeekV3.2andGLM-5,meanaccuracywas31.1%,barelyexceedingthe25%chance baselineandcapturingonly8.1%oftheavailableheadroombetweenchanceandperfectaccuracy 2 (Fig. 2a;seeMethodsforstatisticaltests).Thefailurerunsdeeperthanaccuracyalone:meanmacroF1,a measureofbalancedclassificationqualitythatpenalizesstrategiesignoringminoritytiers,was0.236, belowthe0.25expectedfromrandomguessing.Modelsachieveabove-chanceaccuracyonthetiersthey over-predictwhileproducingzerorecallontierstheyignore,astrategythatinflatesaccuracybutdestroys balanceddiscrimination. Figure2|Frontiermodelsshowstructuredpredictioncollapse 2 Gemini3.1Proreachedthehighestfrontieraccuracy(38.8%)butitsresponseoccasionallyreproducedbenchmarkcontent insteadofreturningatierlabel,raisingbenchmark-leakageconcerns;weretainitwithexplicitcaveatinginsteadoftreatingitasa cleancomparator.Toprovideacleanerwithin-familyreference,thefrontiercohortalsoincludesGemini2.5Pro,releasedinJune 2025.Evenso,nofrontiermodelreliablyseparatesfromthecohortaftermultiple-testingcorrection. a,Elevenfrontiermodelsarecomparedwithpairedaccuracyandmacro-F1barsundertheconservativeprimaryprotocol,with Gemini3.1retainedbutexplicitlymarkedforcontaminationrisk.Accuracybarsshow95%binomialconfidenceintervals(n= 120permodel),whilemacro-F1showsbootstrap95%confidenceintervalsfromper-modelpredictions.Themainpointisthatthe fullfrontiercohortremainsclosetochancedespiteusingthebest-performingpromptformulation.b,Predicted-tierdistributions (100%stacked)showheavyconcentrationinthestrongandfairtiersandsparseuseofthelimitedtier,makingthecollapse patternvisiblebeyondscalaraccuracy.c,Row-normalizedconfusionmatricesforfourrepresentativeflagshipsshowdistinct failuremodes:Gemini3.1asthebest-performingbutstillcompressedmodel,GPT-5.2Highwithmiddle-tierclustering,Claude Opus4.6withstrong-tierceilingbehavior,andQwen3.5Plusasanadditionalreferencemodelshowingthatcollapseisnot confinedtothethreehighlightedproprietaryflagships.Extendedcollapsediagnosticsacrossall11modelsareshownin SupplementaryFig.4. Thefailurefollowsthreedistinctpatternsofpredictioncollapsethatrevealhowpost-trainingalignment shapesmodelbehavioronevaluativetasks(Fig.2b).Thefirstpatternisstrong-ceilingcollapse:Claude Opus4.6assigns87%ofarticlestothe“strong”tier,producingamacroF1ofjust0.145,predicting virtuallyeverythingasabove-averagequality,usingonlytwoofthefouravailablecategories.Thesecond isleniency-to-middlecollapse:Seed2.0andGrok4.1heavilyover-assignabove-averagetiers.Thethirdis middle-tierclustering:GPT-5.2HighandDeepSeekV3.2concentratealmostallpredictionsinthestrong andfairtiers,compressingthedistributionintoanarrowband.Acrossthecohort,6of11modelsnever predict“limited”foranyarticle,producingzerorecallonthelowestqualitytier(Fig.2c;fullconfusion matricesforall11modelsinExtendedDataFig.3a;class-leveldistributiondiagnosticsinSupplementary Figs.1,4).Thesepatternsareconsistentwithreinforcementlearningfromhumanfeedbackoptimizing outputstowardagreeable,broadlyacceptableresponses[28,29],apropertythatservesconversational fluencybutsystematicallyunderminesthediscriminatingjudgmentsthatscientificgatekeepingdemands. Theevaluationpromptrepresentedthemostthoroughattempttotransmithumanevaluativecriteria throughexplicitinstruction.Itsfailureacrossall11modelssuggeststhatscientifictasteresists propositionaltransmission:itcannotbetold,onlyshown. Humanexpertiseandthelimitsofdomainknowledge Thefrontiermodels’failuremightreflectalimitationspecifictoartificialsystems.Perhapsevaluative judgmentrequiresthekindoftacitknowledgethatonlyyearsofimmersioninaresearchcommunitycan provide[8].Totestthis,werecruited48expertgatekeepers 3 :currenteditorsandeditorialboardmembers atleadingjournalsinorganizationalbehaviorandmanagement,individualswhoseprofessionalrole consistspreciselyinmakingthequalityjudgmentsthisbenchmarkdemands. Expertreviewersdidoutperformfrontiermodels,butnotbythemargintheircredentialswouldpredict. Individualexpertaccuracyaveraged36.2%,significantlyabovechancebutreflectingenormousvariability: someexpertsperformedwellbelowchancewhileothersinonecasereached100%ontheirratedarticles (Fig.3a;SupplementaryFig.2;seeMethodsfortestsofsignificance).Whenexpertratingswere aggregatedthroughmajorityvote,accuracyroseto41.6%onthe89articlesthatproducedaclearplurality (31of120articlesyieldedtiedvotesandwereexcluded),capturing22.1%oftheavailableheadroom(Fig. 3b).ExpertsachievedmacroF1valuesof0.353(individualaverage)and0.404(majorityvote),whilethe 11-modelfrontieraverageremained0.236(belowthe0.25randombaseline),demonstratingthathumans genuinelydiscriminateacrossqualitytiersinsteadofcollapsingpredictionsintoanarrowband (SupplementaryFig.6). 3 Eachofthe120benchmarkpitcheswasratedbyameanof3.2experts,matchingthetypicalreviewerloadforajournal submission(2–3independentreviewerspermanuscript).The48-raterpanelthusprovidesper-pitchevaluationdepthcomparable torealeditorialdecisionswhilecoveringthefullbenchmark. Figure3|Humanexpertisehelps,butnoiseremainsdominant a,Individualexpertandjunioraccuracydistributionsareshownwithmajority-votereferencemarkers,highlightingthatexperts outperformjuniorsonaveragebutbothgroupsremainhighlyheterogeneous;SupplementaryFig.2andSupplementaryTableST9 providethefullexpertdistributiondetails.b,Humanreferenceconfigurationsarecomparedagainstthefrontier-modelbaselineon accuracyandmacro-F1with95%uncertaintyintervals,showingthathumanraterspreservemorebalanceddiscriminationthanthe frontiercohort.c,Confidenceandtopicfamiliarityaresummarizedasrating-levelSpearmaneffectestimatesoncorrectness,with bootstrapconfidenceintervalsandpermutation-basedPvalues;thepanelmakesclearthattheseself-reportedexpertisemarkers explainlittleoftheobservedvariationinaccuracy.d,MonteCarlomatched-Njuniorsubsamplingshowshowmajority-vote accuracychangeswithpanelsize,revealingearlygainsfollowedbyaplateauratherthanunlimitedimprovement;Supplementary Fig.3andSupplementaryTableST10providethecorrespondingresamplingdetails. Thepatternofagreementamongexpertsexposesanotabledissociation.Categoricalagreement,whether expertsassignarticlestothesametier,wasnear-chance(Fleiss’kappa=0.047),barelydistinguishable fromnoise[11].Yetordinalagreement,whetherexpertsrankarticlesinroughlythesameorderofquality, wasmoderate(Krippendorff’salpha=0.307).Thisgapsuggeststhatexpertsshareacoarsesenseofwhich articlesarebetterorworsebutdisagreesubstantiallyonwheretodrawcategoricalboundaries,consistent withevidencethatexpertjudgmentisdominatedbynoisethatpersistsregardlessoftrainingor experience[30]. Wesearchedsystematicallyforexpertisemarkersthatmightpredictwhichreviewersperformwell. Noexpertisemarkerpredictedaccuracy:notcareerstage,self-reportedconfidence,nortopicfamiliarity (Fig.3c).Thispatternmatchesabroaderliteratureshowingthatneitherdomainexpertise[33],careerstage, norpublicationrecord[13]reliablypredictsevaluationaccuracy. Juniorresearchers(174doctoralandpostdoctoralscholars)performedcomparably.Individual accuracyaveraged31.7%,andfullmajorityvotereached40.8%onthe103articleswithaclearplurality. Subsamplinganalysis,repeatedlydrawingexpert-sizedpanelsfromthejuniorresearcherpool,yielded accuracynotsignificantlydifferentfromexpertmajorityvote(Fig.3d;SupplementaryFig.3).Panel-size analysisrevealedthataccuracyplateausatapproximately10reviewersperarticle(ExtendedDataFig.4), indicatingastructuralceilingonwhataggregationalonecanachieve. Theseresultsestablishthatevaluativejudgmentisgenuinelydifficult,notmerelyforAIsystems optimizedforconversationalagreeableness,butforthehumanexpertswhoseprofessionalauthorityrests onthisverycapacity.Thequalitysignalexists:institutionalpublicationoutcomesconfirmit.Butneither themostcapablereasoningsystemsinartificialintelligencenordecadesofaccumulateddomainexpertise canreliablyrecoverthatsignalfromresearchpitchesalone.Adifferentkindoflearningmightsucceed: onethatacquirestastefromtheinstitutionalrecordinsteadofreasoningoverexplicitcriteria. Fine-tuningoninstitutionaltracesrecoversscientifictaste Wehypothesizedthattheinformationneededtodiscriminatequalitytiersisnotabsentfromlanguage models’latentrepresentationsbutratherinaccessibleunderstandardpromptingandreasoningregimes.To testthis,weappliedsupervisedfine-tuning(SFT)tofourbasemodelsusing4,479historicalresearch- pitch/journal-outcomepairsfromtheprimaryrecent/newinstitutional-traceslice,thedatathatencodethe field’saccumulatedgatekeepingconsensus[19].The120benchmarkresearchpitcheswereheldout entirelyandnotseenduringtraining.SFTdoesnotinjectnewfactualknowledge;italignslatent representationstothedistributionalregularitiesthroughwhichinstitutionalqualitymanifestsinresearch pitches(seeMethodsfortrainingdetails). Beforefine-tuning,weevaluatedeachbasemodelonthebenchmarktoestablisharchitecture-matched controls.Threeoffourbasemodelsscoredatornearchancelevel(22.5–26.7%);GPT-4.1reached32.5%, modestlyabovechancebutfarbelowitsfine-tunedcounterpart(ExtendedDataTable1).Theseresults confirmthatarchitecturealonecannotextractevaluativesignal:thesmallerQwen3-4Bscorednearchance beforefine-tuningyetimprovedby32.5percentagepointsafter,whileQwen3-30B-A3Bshowedthe largestgaininthestudyat35.8percentagepoints(ExtendedDataFig.3b). Fine-tuningproducedamarkedimprovement.Allfourmodelsachieved55.0–59.2%accuracy,gains of22.5to35.8percentagepointsovertheirrespectivebasemodels(Fig.4a).Thatthesegainsemerge consistentlyacrosstwomodelfamiliesandtwoparameterscalesconfirmsthattheeffectisnotanartifact ofaparticulararchitecture,scale,ortrainingimplementation. Figure4|SFToninstitutionaltracesrecoverstierdiscrimination a,Fourarchitecture-matchedbase-to-SFTpairs(GPT-4.1,GPT-4.1-nano,Qwen3-30B,Qwen3-4B)areshownaslinkedpoints, withcirclesforaccuracyandsquaresformacro-F1.Errorbarsarebootstrap95%confidenceintervalsrecomputedfromthemodel predictionfilesusedinthisanalysis,andtheannotatedpercentage-pointgainsisolatetheeffectofsupervisedfine-tuningfrom architecturechoicealone.b,Anaccuracy-onlyevaluatorcomparisonplacestheSFT2-modelensemble,allfoursingleSFT checkpoints,expertandjuniorindividual/majorityreferences,thebestfrontiermodel,andthefrontieraverageonthesameaxis with95%confidenceintervalsfromthebenchmarksummarytables.c,Row-normalizedrecallacrossthefourqualitytiers comparestheSFTensemble,bestfrontiermodel,frontieraverage,expertmajorityvote,andjuniormajorityvote,showingthat SFTrecoversthefulltierstructureratherthancollapsingtowardthemiddle.d,Boundary-tierF1fortheexceptionalandlimited tiershighlightswhereSFTgainsarestrongestandwhytheymatterforgatekeepingdecisionsatthetwoextremesofarticlequality. Theselectedprimarypair,GPT-4.1-nano(SFT)andQwen3-30B-A3B(SFT),achieved60.8%(95%CI: 51.7–69.2%),exceedingthebestsingleQwen3-4Bcheckpointby1.7percentagepointsandcapturing 47.8%oftheheadroombetweenchanceandceiling.Thefrontieraveragecaptures8.1%;expertmajority votecaptures22.1%.TheSFTensemblethereforecapturesapproximatelysixtimestheheadroomof frontiermodelsandmorethandoublethatofexpertpanels,withamacroF1of0.607reflectingbalanced discriminationacrossallfourtiers(SupplementaryTableST3).Wesystematicallyevaluatedallsix pairwiseensemblecombinationsamongthefourSFTcheckpointsviaprobabilityaveraging;accuracies rangedfrom59.2%to60.8%,andallsixcombinationsexceededthefrontieraverage(31.1%)by28.1to 29.8percentagepoints(SupplementaryTableST4;ExtendedDataFig.3c). Comparisonsconfirmedtherobustnessofthesedifferences.Againstthefrontieraverage(31.1%),the SFTensembleoutperformedby+29.8percentagepoints(exactbinomial,p=1.74×10⁻¹);againstthe best-performingfrontiermodelunderconservativeprotocol(Gemini3.1Pro,39.1%majorityaccuracyon pairedsubset),by+20.9percentagepoints(McNemartest,p=0.001425;SupplementaryTableST8). Becausenoindividualfrontiermodelsignificantlyseparatesfromthecohortaftermultiple-testing correction,thecomparisonagainstthefrontiermeanremainsthemostrepresentative.Againstindividual evaluators,theensemblesurpassedthetypicalexpertbyalargemargin(60.8%vs.36.2%mean,t(47)=- 8.32,p<0.001,d=1.20),with83%ofexpertsand97%ofjuniorresearchersscoringbelowtheensemble (Fig.4b).Againstcollectivemajorityvotes,SFTshowedadvantagesof21.4percentagepointsoverexpert majority(p=0.00729)and21.4percentagepointsoverjuniormajority(p=0.00196). Theper-tieranalysisrevealswheretheadvantageconcentratesandexposesafundamentaldistinction inevaluationstyle(Fig.4c,d).SFTachievedhigherbalanceddetectionrates(F1)thanexpertpanelsatall fourtiers,withthelargestgainattheexceptionaltier(0.68versus0.33)andcontinuedimprovementsat thelimited(0.62versus0.40),strong(0.58versus0.46),andfair(0.55versus0.43)tiers.Humanpanels remainconservativegatekeepers:highprecisionwhentheycommittoanextremejudgment(63%for exceptional,60%forlimited),butmissingthemajorityofarticlesateachqualityend(23%recallfor exceptional,30%forlimited).TheSFTensemblepredictsmuchclosertothebenchmarkbaserate(38 exceptional,25strong,35fair,22limited),anditsfalsepositivesattheexceptionaltierremainmostly near-missesfromtheadjacentstrongtier(10of15),notcatastrophicmisclassifications.Inanysystem wherethecostofmissingexceptionalworkexceedsthecostofadditionalreview,thetypicalcondition whensubmissionvolumesfaroutstripexpertbandwidth,thisbroadercoverageaddressesthemore consequentialfailuremode. Totestwhetherthissignalreflectsdurableevaluativestructureortransientpatterns,wetrained matchedarchitecturesonanolderinstitutionaltraceslice(2015–2020,3,368pairs),introducingafive- yearlagrelativetothebenchmark.TheQwen3-30BSFTtrainedontheolderslicereached46.7%,still exceedingboththefrontieraverage(31.1%)andexpertmajorityvote(41.6%),versus58.3%foritsrecent- slicecounterpart(ExtendedDataFig.6).Whatdriftsisthequalitybar,notthesignal:theoldermodel over-predictedexceptionalarticlesby50%,reflectingcompetitivethresholdsthathavesincerisenas submissionvolumesgrewandacceptanceratesfell[44,45].Themodellearnedthestandardsofitsera faithfully;thosestandardshavesincebeenraised. Totaltrainingcostacrossallfourmodelswasunder$300,spanningtwocloudAPIfine-tuningjobs (~$200and~$10forGPT-4.1andGPT-4.1-nano,nohardwarerequired)andtwolocalGPUruns(~1and ~8A100hoursforQwen3-4BandQwen3-30B-A3B;SupplementaryTableST2).Thecriticalingredient behindthemainbenchmarkgainsisthetrainingsignal,institutionaltracesencodingaccumulated scientifictaste,notmodelscalealone.Modelscalemattersmainlyforhowrobustlythatlearnedsignal generalizesundersevereinputcompression. Themechanism:trainingsignaldrivesperformance;modelscaleshapesgeneralization TheSFTresultsposeamechanisticpuzzle.Whydoesasimpletrainingprocedureonjournal-tierlabels unlockevaluativeperformancethatneitherthelargestfrontiermodelsnorexperiencedhumanreviewers achieve?Fourlinesofevidenceclarifythemechanismanditsboundaryconditions:directexample-based alignmenttoinstitutionaltracesdrivesthebenchmarkgainsacrossarchitectures,whereasmodelscale mainlygovernshowrobustlythelearnedrepresentationgeneralizesbeyondthefullidea-summaryformat. Calibratedmetacognition.Adefiningpropertyofgenuineexpertiseisknowingwhatoneknows,the capacitytoassignhigherconfidencetojudgmentsthatprovecorrect.TheSFTensembleexhibitedthis property:itspredictionconfidencewassignificantlyhigheroncorrectpredictionsthanonincorrectones (confidencegap=+0.082,p=0.0081;expectedcalibrationerror=0.073;Fig.5a).Allfourindividual SFTmodelsshowedthesamepattern,althoughtheQwen3-4Bgapwasmuchsmaller,confirmingthat calibrationisaconsistentpropertyofthefine-tuningapproach,notanartifactofasinglemodel. Figure5|SFTmechanism:self-knowledgeandgeneralizedpairwiseranking a,MeanconfidenceforcorrectandincorrectpredictionsiscomparedfortheSFTensemble,GPT-5.2chat,experts,andjuniors. Human1-5confidenceismappedto0-1by(x-1)/4,andPvaluesareone-sidedMann-WhitneyUtestscomparingwhether correctratingsaremoreconfidentthanincorrectratingsattheratinglevel.Thepanelaskswhethereachevaluatorknowswhenit isright.b,Selective-predictioncurvesshowhowaccuracychangesaslower-confidencecasesaredeferred,withAIpredictions orderedfromhighesttolowestconfidenceandhumancurvesdefinedbyconfidencethresholds;thispanelvisualizesthepractical triagevalueofcalibratedabstention.c,Overallweightedaccuracyonthefixed300-pairpairwiseevaluationsetcomparesthe sharedfour-modelsubset:SFTGPT-4.1,Gemini3.1Pro,GPT-5.2High,andtheGPT-4.1baseline,with95%binomial confidenceintervalsandrawunadjustedexactpairedMcNemarPvaluesversusSFT.d,Ahard-boundarysummaryisolatesthe twomostinformativepairtypes,strong_exceptionalandfair_strong,showingthattheSFTpairwiseadvantageisconcentrated whererelativerankingishardest.ExtendedDataFig.2retainsthesix-pair-typeheatmapanddiscordant-pairdecompositionfor thesameplottedsubset. Thiscalibrationstandsinsharpcontrasttobothfrontiermodelsandhumanevaluators(Fig.5a;Extended DataFig.4a-c).GPT-5.2chatexhibitedextreme,undifferentiatedoverconfidence:near-maximum confidenceregardlessofwhetheritspredictionswerecorrect,despitenear-chanceaccuracy(ECE=0.657, ameasureofhowwellstatedconfidencematchesactualaccuracy;ExtendedDataFig.4b).Humanexpert confidencewassimilarlynon-discriminative.Juniorresearcherconfidenceshowedastatistically detectablebutpracticallyweakrating-levelassociationwithaccuracy.Frontiermodelsandexpertraters thereforeprovidenocomparablyusefulself-screeningsignal,andthedetectablejunioreffectremainstoo weaktosupporthigh-precisionabstention. TheSFTmodelscan.Whenpredictionswererestrictedtothetopsixthbyconfidence(16.7% coverage),accuracyreached80.0%;atthetop10.0%(12of120articles),accuracyreached100%;every high-confidencepredictionwascorrect(Fig.5b).Thisselectivepredictioncapacity,theabilitytoflag caseswherethemodel’sjudgmentishighlyreliablewhileabstainingonuncertainones,representsthe strongestpracticallyusefulmetacognitivetriagesignalinthebenchmark.Inpractice,thisenablesa confidence-basedworkflow:routehigh-confidencepredictionsdirectlyandflaguncertaincasesforhuman review. Pairwisehead-to-headdiscrimination.ToconfirmthatSFTdevelopsgeneralizedevaluative judgment,notcategory-specificpatternmatching,wetestedmodelsinapairwiseformatentirelyabsent fromtraining.Modelsreceivedtworesearchpitchessimultaneouslyandjudgedwhichrepresentedthe strongerresearchpitch.Themainfigurepairwisepanelsfocusonthesharedfour-modelsubsetused throughoutthepairwiseanalysis:SFTGPT-4.1,Gemini3.1Pro,GPT-5.2High,andtheGPT-4.1baseline. Withinthissharedsubset,SFTGPT-4.1reached84.3%(253/300),versus77.3%(232/300)forGemini3.1 Pro,78.7%(236/300)forGPT-5.2High,and76.0%(228/300)fortheGPT-4.1baseline(Fig.5c; SupplementaryTableST1).Fine-tuningoninstitutionaltracesthusexceededboththemodel’sownpre- trainingbaselineandthefrontiercomparatorsinthisauxiliarypairwiseset.Thegainswerenotuniform acrosspairtypes:thelargestmarginsappearedatthefair-strongandstrong-exceptionalboundaries,where SFTreached84.0%and64.0%versus70.0%and50.0%forGemini3.1Pro,68.0%and56.0%forGPT- 5.2High,and66.0%and54.0%fortheGPT-4.1baseline(Fig.5d).ExtendedDataFig.2providesthefull heatmapacrossallsixpairtypesandthediscordancedecompositionforthesameplottedsubset.Onthe same300paireditems,rawunadjustedexactMcNemartestsshowedsignificantSFTadvantagesversus Gemini3.1Pro(p=0.00646),GPT-5.2High(p=0.0300),andGPT-4.1baseline(p=0.000621;Fig.5c). SFTversusreinforcementlearning.Ifevaluativejudgmentistacitandresistsdecompositioninto reasoningsteps[8],trainingmethodsthatoptimizearticulatedreasoningchainsmayperformworsethan methodsthatlearndirectlyfromexamples.Cognitivesciencesupportsthisprediction:verbalarticulation ofholisticjudgmentsdegradestheirquality[35],andchain-of-thoughtpromptingyieldsnegligibleor negativegainsoutsidemathematicsandsymbolicreasoning[36,37].Wetestedthisbytraining reinforcementlearningvariantsonthesamebasearchitectures(seeMethods).RLrun-pooledaccuracy was40.3%,abovethefrontieraverage(31.1%)butbeloweverySFTsinglemodel(55.0–59.2%;Extended DataFig.5).Whenthemodel’sreasoningchaindivergedfromitsfinallabel,accuracydroppedmarkedly; whenreasoningandlabelagreed,performanceapproachedSFTlevels.Explicitdeliberationcanoverride otherwisesoundevaluativeintuitions[35,49].SFTlearnstastethroughexposuretoexemplars,not throughreasoningaboutcriteria. Transferunderinputcompression.AnaturalobjectiontotheSFTresultsisthatthemodelsmay havelearnedsuperficialpatternsintherichcontextualdescriptionsusedduringtraining,notgenuine evaluativejudgment.Wetestedthisbyevaluatingthesamefine-tunedcheckpointsonaninputformat completelydifferentfromtheoneusedintraining:one-sentenceideastatementsstrippedoftheoretical framing,methodologicaldetail,andcontributionclaims(forexamplesofpairedone-sentenceandfull- summaryinputs,seeSupplementaryMethodsSM4).ThelargerGPT-4.1fine-tunedmodelretained substantialevaluativesignalunderthiscompression,achieving49.2%accuracyonone-sentenceinputs versus55.0%onthefullsummary(ExtendedDataFig.7a).Thiscompressedinputaccuracyexceedsthe bestfrontiermodelonthefullinput(Gemini3.1Pro,38.8%)andtheexpertmajorityvote(41.6%), demonstratingthattheevaluativerepresentationlearnedthroughfine-tuningsurvivesradicalreductionof theinputsignal.Thetransferwasnotuniform:recallattheexceptionalandstrongtiersdeclinedwhile recallatthelimitedtierrosefrom60.0%to83.3%,indicatingaconservativeshiftunderuncertainty,not randomdegradation(ExtendedDataFig.7b,c). ThesmallerGPT-4.1-nanofine-tunedmodelshowedasharplydifferentpattern.Despitematching GPT-4.1SFTonthefullinput(57.5%versus55.0%),itcollapsedto33.3%undercompression,barely abovechanceandindistinguishablefromitsownbasemodel(30.8%).Thisasymmetryindicatesthat acquiringtheevaluativesignalandgeneralizingitacrossinputformatsareseparablecapabilities:thesame institutionaltrainingsignalimprovesbothmodelsonthefullbenchmark,butgreaterrepresentational capacitymakesthatlearnedsignalmorerobustunderseverecompression. Together,theselinesofevidence(calibratedself-knowledge,generalizedpairwisediscrimination, transferunderinputcompression,andtheSFTadvantageoverRL)convergeonamorespecific conclusion:supervisedfine-tuningoninstitutionaltracesextractsagenuineevaluativerepresentation.The trainingsignaldrivesthebenchmarkgainsacrossarchitectures,whilemodelscaleshapeshowrobustlythe learnedrepresentationtransfersundersevereinputcompression. Votingconsistencyandaggregationasymmetry IfindividualSFTmodelscaptureevaluativesignal,anaturalquestioniswhetheraggregatingacross modelsamplifiesthatsignalfurther,andwhetherthesameholdsforfrontiermodelsandhumanpanels. Thefinalmechanismquestioniswhetheraggregationimprovesjudgmentthroughadditionalvotesalone, orthroughthequalityanddiversityofthevoters.Fig.6presentsaunifiedcomparisonoffilteringby consensusacrossAIandhumanevaluatorclasses,revealingafundamentalasymmetryinhowaggregation functionsacrossevaluatortypes. Figure6|Aggregationasymmetryaftermetric-definitionalignment a,Cross-modelconsensuspoliciesarecomparedforthefourSFTmodelsandafixedfour-modelfrontierset.Accuracyintervals useWilson95%confidenceintervals,andcoverageindicatestheshareofarticlessatisfyingeachagreementrule.Thecontrastis substantive:stricterconsensusimprovesSFTprecisionsharply,whereasfrontiercross-modelvotingremainsweakevenathigher agreementthresholds.b,Humanconsensuspoliciescompareexpertmajority,juniormajority,junior>=50%voteshare,and expertunanimityamongarticleswithatleasttwoexpertratings.AccuracyintervalsinthispanelalsouseWilson95%confidence intervals,andlabelsreporttheassociatedcoverage,makingthehumanaccuracy-coveragetradeoffexplicit.c,Same-model8-run gainsforfrontiersystemsaresummarizedasmajority-voteaccuracyminusmeansingle-runaccuracy,showingthatrepeated samplingaloneyieldsonlymodestimprovementsanddoesnotsolvetheunderlyingjudgmentproblem. AcrossthefourSFTmodels,fullconsensus(4/4)occurson42.5%ofarticles(N=51)andreaches72.5% accuracy,thehighestprecisionachievedbyanyevaluatorconfigurationinthisstudy(Fig.6a).Withinthis unanimoussubset,accuracyissharpestatthequalityextremes:100.0%fortheexceptionaltierand66.7% forthelimitedtier,whilearticlesinthestrongtierremainhardest(53.8%;fair66.7%;Supplementary TableST12).Relaxingto>=3/4agreementexpandscoverageto80.8%at66.0%accuracy.Whenthe strongestagreementisonly2/4,accuracyfallsto34.8%,wellbelowtheunanimoussubsetandonly modestlyabovechance,indicatingthatmodeldisagreementisausefuluncertaintysignal,notrandom noise(Fig.6a). Humanvotingexhibitsasteepertradeoff.Expertunanimousagreement(amongarticleswithatleast twoexpertratings)achieves69.2%accuracyat10.8%coverage,while>=50%juniorconsensusreaches 56.0%accuracyat20.8%coverage.Withoutconsensusfiltering,juniorandexpertfull-panelplurality accuraciesare40.0%and39.2%onall120articles(Fig.6b). Thecontrastwithfrontiermodelsissharp.Cross-modelvotingusingadiversefour-modelfrontierset (Gemini3.1Pro,ClaudeOpus4.6,GPT-5.2High,GLM-5)reachedonly32.7%onthe104articleswitha clearplurality,andhigherconsensusthresholdsdidnotrecoverperformance(Fig.6a,c).Frontiermodels failtocreatemeaningfulconsensusbecausepost-trainingbiasescausecorrelatederrorsacrossthesame articles. SFTcross-modelconsensusdeliversa15-percentage-pointboostovertheaveragesingleSFTmodel (56.9%to72.0%at4/4consensus)becauseSFTmodels,despiteindependenttrainingondifferent architectures,convergeonasharedevaluativesignal(inter-modelCohen’skappa=0.50–0.60,versus humankappaapproximately0.03–0.05;SupplementaryTableST12).Thepracticalimplicationisa selectiveevaluationworkflow:routethe42%ofsubmissionswhereallfourSFTmodelsagree(at72% accuracy)directly,andconcentratereviewerattentionongenuinelyambiguouscases. Cross-fieldreplicationineconomics Theresultsaboveestablishtheinstitutionaltracemechanism,itsmechanisticunderpinnings,andits aggregationpropertiesinmanagement.Anaturalquestioniswhetherthemechanismisspecifictothat fieldorgeneralizestoothersocialsciencedisciplineswithdistinctpublicationculturesandevaluation norms.Weconstructedanadditionalbenchmark:200held-outresearchpitchesineconomics,balanced acrossthesamefourtiersanddrawnfrom2025publications(SupplementaryMethodsSM9).Wethen trainednewSFTcheckpointsonfield-specificinstitutionaltraces(5,593economicspairs)usingthesame procedure. Themechanismreplicated(Fig.7).Thebestsinglemodel(Qwen3-30B)reached69.5%accuracy,with allthreearchitecturesatleast64%(Fig.7a,b).Architecture-matchedbasemodelsremainedatchance(25– 26%),confirmingthatthegainsareattributabletofine-tuning(SupplementaryTableST15).Calibrated confidencereplicatedaswell:accuracyrosesharplywhencoveragewasrestrictedtohigh-confidence predictions(Fig.7c). Figure7|Institutional-traceretrainingreplicatesineconomics Theeconomicsvalidationreusesthesamefour-tieroutputspaceandthesamerq_with_contextrepresentationlogicasthemain benchmark,butevaluatesaheld-out2025subject-specificsample(N=200;50itemspertier).a,Architecture-matchedbase- versus-in-domain-SFTscoresineconomicsforQwen3-30B,Qwen3-4B,andGPT-4.1-nano,shownaslinkedpointswithcircles foraccuracyandsquaresformacro-F1.Horizontalwhiskersshowbootstrap95%confidenceintervalsrecomputedfromtheraw subjectpredictionfiles.Thegreydashedlinemarksthe25%chancebaseline.Thegreendashedanddottedreferencelinesmark thebestfrontier(Gemini3.1Pro)full-setaccuracyandmacro-F1ineconomics,respectively(46.0%accuracy;38.45%macro-F1), underthesamesingle-labelfrontierscoringconventionusedforthesubjectsummaries.b,Row-normalizedconfusionmatrixfor thebestsingleeconomicsmodel(Qwen3-30BSFT),showingthatthestrongestin-domaincheckpointrecoversallfourtiersrather thancollapsingintothemiddlecategories.c,Selective-predictioncurveforthatsamebesteconomicsmodel,witharticlesordered fromhighesttolowestpredictedconfidence;thecurveshowshowaccuracyrisesaslower-confidencecasesaredeferred.Together, thesepanelsshowthattheinstitutional-tracemechanismreplicatesineconomics,whileremainingclearlyseparatedfromthebest currentlyavailablefrontiercomparatorandretainingapracticallyusefulconfidencesignalunderthenewsubject-specific validation. Strikingly,eventhemanagement-trainedSFTGPT-4.1,whichneverencounteredeconomicsarticlesor journaltiersduringtraining,achieved43.5%ontheeconomicsbenchmark(p=9.2×10⁻⁹versus chance),a+14.0percentage-pointgainoveritsownbasemodel(29.5%,notsignificantlyabovechance; SupplementaryFig.7).Thiscross-fieldtransfersuggeststhatevaluativesignalssharestructureacross relatedsocialsciencedisciplines:somecomponentofthetastelearnedfrommanagementinstitutional tracesgeneralizestoeconomicsresearchevaluation.Amodeltrainedjointlyonmanagementand economicsinstitutionaltracesmaintainedstrongperformanceacrossbothdomains(Qwen3-30B:69.5% economics,61.7%management;ExtendedDataFig.8),confirmingthatevaluativesignalsfromdistinct fieldsdonotinterferewhencombined. Discussion Humanevaluatorscannotreliablyagreeonwhichresearchideasbelonginwhichqualitytier.Across nearly3,000ratings,experiencedgatekeepers(editorsandeditorialboardmembersatleadingjournals) showedcategoricalagreementbarelyabovechance(Fleiss’κ=0.047);amongalargerpanelofdoctoral andpostdoctoralresearchers,thepatternwasthesame.Thisreflectsnotindividualfailurebutthenatureof theproblem.Evaluativejudgmentinscienceiscollectivetacitknowledge[20]:noindividualreviewercan articulatewhatdistinguishesexceptionalfrommerelycompetentwork,yettheinstitutionalsystem, integratingthousandsofsuchjudgmentsoverdecades,producesreliablequalitystratification[19].The knowledgeisreal,butitisnotinanyoneperson’shead.Itresidesintheaccumulatedrecordofwhatwas selected,funded,andpublished,aninstitutionaltracethatencodesafield’saccumulatedtasteinaformno explicitrubrichasmanagedtocapture[8,11]. ThedominantpushinAIforscienceassumesthebottleneckisgeneration.Autonomoussystemscan nowproducecompletescientificmanuscriptsatscale:ProjectAPEhasgeneratedover300economics researchpapersusingfrontiermodelswithmulti-stagereviewpipelines[46],andtheAIScientistsystem hasproducedend-to-endmachine-learningmanuscriptssubmittedtopeer-reviewedvenues[47].Yetin head-to-headtournamentsagainstpublishedhumanresearch,AI-generatedpaperswinlessthan5%ofthe time[46],amargindifficulttodistinguishfromnoiseinthejudgingprocessitself.AItoolshaveexpanded scientificoutputbutmayhavecontracteditsqualityandfocus[5].Thesesystemsdemonstratethatraw generativecapabilityisnottheconstraint:frontiermodelscansynthesizeliterature,runanalyses,and writefluentprose.Whattheycannotdo,andwhatourresultsconfirm,isevaluatewhichideasdeserve pursuitinthefirstplace.Frontierreasoningmodelscapturebarely8%oftheavailableheadroomonour benchmark(Fig.2a);supervisedfine-tuningoninstitutionaltracescapturesnearlyhalf,atafractionofthe cost(Fig.4a).Theevaluativesignalwasnevermissing;itwasneverthetrainingtarget. Theseresultssuggestapracticalalternativetoautonomousscientificgeneration:anupstream evaluativefilter,trainedoninstitutionaltraces,thatscreensresearchideasbeforeresourcesarecommitted. Severalfeaturesofthefine-tunedmodelsmakesuchdeploymentrealistic.Themodelsknowwhenthey arelikelytoberight:calibratedconfidenceconcentratesthemostreliablepredictionsintoahigh-precision subsetthatcouldbeactedondirectly,whileflagginguncertaincasesforhumanreview(Fig.5a,b).The evaluativecapacityalsogeneralizesbeyondthetrainingformat.Withoutexposuretopairwise comparisonsduringtraining,themodelstransfertohead-to-headdiscriminationtasks,indicatinglearned qualityorderingandnotjustmemorizedcategoryassignments(Fig.5c,d).Cross-modelconsensus providesasecondreliabilitylever:whenindependentlytrainedmodelsagree,accuracyrisessharply; whentheydisagree,thatdisagreementitselfbecomesausefultriagesignal(Fig.6a).AndbecausetheSFT ensembleandexpertpanelserronlargelydifferentarticles,combiningthetwocouldrecoversubstantially moreoftheavailableheadroomthaneitheralone(SupplementaryFig.5).Thepracticalworkflowthis enablesisselective:routeclearcasesthroughthemodel,concentratescarcereviewerattentiononthe genuinelyambiguousremainder. Theeconomicsreplicationstrengthenstheseimplications.Thesamefine-tuningprocedure,appliedto adifferentfieldwithdistinctjournals,evaluationnorms,andintellectualtraditions,producedeven strongerresults(69.5%bestsinglemodel;Fig.7).Pooledtrainingonbothfieldssimultaneouslypreserved performanceineachdomain(ExtendedDataFig.8),suggestingthatevaluativesignalsfromdifferent disciplinescoexistwithoutinterference.Amanagement-trainedmodelthatneverencounteredeconomics articlesstillexceededchanceontheeconomicsbenchmarkby14percentagepoints(SupplementaryFig. 7),pointingtosharedevaluativestructureacrossrelatedsocialsciences.Thesecross-fieldresultsshiftthe interpretationfromafield-specificfindingtoageneralmechanism:whereverinstitutionalgatekeepinghas operatedlongenoughtoleaveatrace,thattraceislikelylearnable. Severallimitationsqualifytheseconclusions.Ourbenchmarkscovertwosocialsciencefields; whetherthemechanismextendstoSTEMdisciplineswithdifferentepistemicstructures,whereempirical reproducibilityprovidesapartialqualitysignalabsentinsocialscience,remainsuntested.Thetransfer resultsunderinputcompressionrevealanopenboundarycondition:thelargerfine-tunedmodelcarries evaluativesignalacrossradicalinputcompression,butthesmallermodeldoesnot(ExtendedDataFig.7), indicatingthatrepresentationalcapacitygovernstransferrobustness.Futureworkshouldtestwhichmodel propertiesortraining-datacharacteristicsdeterminethisthreshold.Institutionaltracesareaproxyfor quality,notanobjectivegroundtruth.Theyencodetheaccumulatedconsensusofagatekeepingsystem withwell-documentedbiases:towardincrementalovernovelwork[18],towardestablishedmethodologies, towarddominantparadigms,andtowardacompetitivebarthatrisesovertimeassubmissionvolumes grow.TheSFTmodellearnswhatthesystemhistoricallyrewarded,whichisnotidenticaltowhatis objectivelybest.Thislimitationisrealbutbounded:indomainswherequalityisirreduciblyunverifiable, institutionalconsensusovertimeisnotaproxyfortaste;itistheoperationaldefinitionoftaste[19,46]. Themodel’scalibrateduncertaintyprovidesaninternalcheckonitsownreliabilitythatthesystemit learnedfromlacks.Thesignalisalsonotstatic:modelstrainedonolderinstitutionaltracesstill outperformfrontiersystemsbutshowdriftintiercalibrationascompetitivethresholdsshift(Extended DataFig.6),implyingthatperiodicretrainingwillbenecessaryasinstitutionalstandardsevolve. Themechanismgeneralizesbeyondanysinglediscipline.Inanydomainwherecollectivehuman evaluationhasoperatedovertime,thehistoricalrecordofwhatwasselected,funded,orrewarded constitutesalearnablesignal:ventureinvesting,whereexpertpredictionisnotoriouslypoor[32];grant allocation,whererevieweragreementapproacheszero[12];andcreativeindustries,wheremarket outcomesconsistentlydefyexpertforecasts.Thesocialandbehavioralsciences,whereevaluative judgmentresistsformalverificationandsubmissionvolumesfaroutstripreviewercapacity,standto benefitmostimmediately.Theapproachisalsouniquelycost-effective:SFTrequiresonlyhistorical decisionrecordsalreadydepositedininstitutionalarchives,totaltrainingcostwasunder$300,andthe resultingmodelsprovidecalibratedconfidencescoresthatmaketheiruncertaintytransparentand auditable.Forscienceatscale,thepathforwardmaynotbeAIsystemsthatreplacehumanjudgmentor thatgenerateresearchautonomously,butsystemsthatlearnthetastehumaninstitutionshaveaccumulated overdecadesandapplyitwherehumanbandwidthcannotreach.Autorpredictedthatmachinelearning wouldovercomePolanyi’sparadox,thebarrierbetweenknowingandtelling,bylearningfromoutcome datainsteadofexplicitinstruction[39].Ourresultsconfirmthisprediction:frontiermodelsthatreceive evaluativecriteriathroughpromptsperformatchance;SFTthatlearnsfromoutcomedatarecoversnearly halftheavailableheadroominmanagementandexceedsitineconomics.Publicationrecordsfunctionas implicitfeedbackfromtheevaluativecommunity:individuallynoisybutcollectivelyinformative, encodingwhatHintonetal.termed‘darkknowledge’[38],evaluativestructureembeddedininstitutional sortingbutinvisibleinanystatedcriterion.Tastewasneveruniquelyhuman;itwasalwaysdepositedin theinstitutionalrecord,waitingforalearningproceduresimpleenoughtoextractit. Acknowledgements Wethankthe48expertgatekeepers–editorsandeditorialboardmembersatleadingjournalsin organizationalbehaviorandmanagement–andthe174doctoralandpostdoctoralresearcherswho volunteeredtheirtimetoevaluateresearchpitches.Theircarefuljudgmentsprovidedthehuman benchmarkagainstwhichallAIsystemsweremeasured,andtheirparticipationmadethisstudypossible. QiupingPengandLiyunZhangassistedwithdatacollectionandprojectmanagement. Dataavailability Thede-identifiedbenchmarkdata,modelpredictionfiles,human-ratingfiles,figures,tables,and reproducibilityscriptssupportingthisstudywillbereleasedthroughtheprojectrepositoryat https://github.com/FutureTech-OB/ai-taste.Therepositorywillalsoprovidethefine-tuningcode,training data,andinformationonaccesstotheresultingmodelweights. References 1.Jumper,J.etal.HighlyaccurateproteinstructurepredictionwithAlphaFold.Nature596,583–589(2021). 2.Hubert,T.etal.Olympiad-levelformalmathematicalreasoningwithreinforcementlearning.Nature(2025). 3.OpenAI.OpenAI2025ICPCsubmissions.GitHubhttps://github.com/openai/openai-icpc-2025(2025). 4.Si,C.etal.CanLLMsgeneratenovelresearchideas?Alarge-scalehumanstudywith100+NLPresearchers. ICLR(2025). 5.Hao,Q.etal.Artificialintelligencetoolsexpandscientists’impactbutcontractscience’sfocus.Nature649, 1237–1243(2026). 6.Karpatne,A.etal.AI-enabledscientificrevolutionintheageofgenerativeAI:secondNSFworkshopreport.npj Artif.Intell.(2025). 7.Dell’Acqua,F.etal.Navigatingthejaggedtechnologicalfrontier:Fieldexperimentalevidenceoftheeffectsof AIonknowledgeworkerproductivityandquality.HarvardBusinessSchoolWorkingPaper24-013(2023). 8.Polanyi,M.TheTacitDimension(Doubleday,1966). 9.Corley,K.G.&Gioia,D.A.Buildingtheoryabouttheorybuilding.Acad.Manag.Rev.36,12–32(2011). 10.Colquitt,J.A.&George,G.PublishinginAMJ—Part1:Topicchoice.Acad.Manag.J.54,432–435(2011). 11.Bornmann,L.,Mutz,R.&Daniel,H.-D.Areliability-generalizationstudyofjournalpeerreviews.PLoSONE5, e14331(2010). 12.Pier,E.L.etal.LowagreementamongreviewersevaluatingthesameNIHgrantapplications.PNAS115,2952– 2957(2018). 13.Callaham,M.L.&Tercier,J.Therelationshipofprevioustrainingandexperienceofjournalpeerreviewersto subsequentreviewquality.PLoSMed.4,e40(2007). 14.Black,N.,vanRooyen,S.,Godlee,F.,Smith,R.&Evans,S.Whatmakesagoodreviewerandagoodreview forageneralmedicaljournal?JAMA280,231–233(1998). 15.Callaham,M.&McCulloch,C.Longitudinaltrendsintheperformanceofscientificpeerreviewers.Ann.Emerg. Med.57,141–148(2011). 16.Lamont,M.HowProfessorsThink:InsidetheCuriousWorldofAcademicJudgment(HarvardUniv.Press, 2009). 17.Boudreau,K.J.etal.Lookingacrossandlookingbeyondtheknowledgefrontier.Manag.Sci.62,2765–2783 (2016). 18.Teplitskiy,M.,Peng,H.,Blasco,A.&Lakhani,K.R.Isnovelresearchworthdoing?Evidencefrompeerreview at49journals.PNAS119,e2118046119(2022). 19.Siler,K.,Lee,K.&Bero,L.Measuringtheeffectivenessofscientificgatekeeping.PNAS112,360–365(2015). 20.Collins,H.TacitandExplicitKnowledge(UniversityofChicagoPress,2010). 21.Nonaka,I.&Takeuchi,H.TheKnowledge-CreatingCompany(OxfordUniv.Press,1995). 22.Naddaf,M.MorethanhalfofresearchersnowuseAIforpeerreview—oftenagainstguidance.Nature649, 273–274(2026). 23.Bergstrom,C.T.&Bak-Coleman,J.AI,peerreviewandthehumanactivityofscience.Nature(2025). 24.RussoLatona,G.etal.TheAIreviewlottery:WidespreadAI-assistedpeerreviewsboostpaperscoresand acceptancerates.arXiv2405.02150(2024). 25.Zhu,C.etal.WhenyourreviewerisanLLM:Biases,divergence,andpromptinjectionrisksinpeerreview. arXiv2509.09912(2025). 26.Shin,H.etal.Mindtheblindspots:Afocus-levelevaluationframeworkforLLMreviews.Proc.EMNLP(2025). 27.Thelwall,M.CanChatGPTevaluateresearchquality?J.DataInf.Sci.9,1–21(2024). 28.Christiano,P.F.etal.Deepreinforcementlearningfromhumanpreferences.NeurIPS(2017). 29.Sharma,M.etal.Towardsunderstandingsycophancyinlanguagemodels.Proc.ICLR(2024). 30.Kahneman,D.,Sibony,O.&Sunstein,C.R.Noise:AFlawinHumanJudgment(Little,BrownSpark,2021). 31.Guo,D.etal.DeepSeek-R1incentivizesreasoninginLLMsthroughreinforcementlearning.Nature645,633– 638(2025). 32.Tetlock,P.E.ExpertPoliticalJudgment(PrincetonUniv.Press,2005). 33.Gallo,S.A.etal.Theinfluenceofpeerreviewerexpertiseontheevaluationofresearchfundingapplications. PLoSONE11,e0165147(2016). 34.Wu,Y.etal.OnthegeneralizationofSFT:Areinforcementlearningperspectivewithrewardrectification.ICLR (2026). 35.Wilson,T.D.&Schooler,J.W.Thinkingtoomuch:Introspectioncanreducethequalityofpreferencesand decisions.J.Pers.Soc.Psychol.60,181–192(1991). 36.Liu,R.etal.Mindyourstep(bystep):Chain-of-thoughtcanreduceperformanceontaskswherethinkingmakes humansworse.Proc.ICML(2025). 37.Sprague,Z.etal.ToCoTornottoCoT?Chain-of-thoughthelpsmainlyonmathandsymbolicreasoning.ICLR (2025). 38.Hinton,G.,Vinyals,O.&Dean,J.Distillingtheknowledgeinaneuralnetwork.arXiv1503.02531(2015). 39.Autor,D.H.Whyaretherestillsomanyjobs?Thehistoryandfutureofworkplaceautomation.J.Econ. Perspect.29,3–30(2015). 40.Shao,Z.etal.DeepSeekMath:Pushingthelimitsofmathematicalreasoninginopenlanguagemodels.arXiv 2402.03300(2024). 41.Wei,J.etal.Finetunedlanguagemodelsarezero-shotlearners.ICLR(2022). 42.Yu,Q.etal.DAPO:Anopen-sourceLLMreinforcementlearningsystematscale.NeurIPS(2025). 43.Qu,Y.etal.POPE:Learningtoreasononhardproblemsviaprivilegedon-policyexploration.arXiv2601.18779 (2026). 44.Gruber,M.AnalyzingAcademyofManagementJournaloperationswithartificialintelligence(2006–2022). Acad.Manag.J.68,1–10(2025). 45.Card,D.&DellaVigna,S.Ninefactsabouttopjournalsineconomics.J.Econ.Lit.51,144–161(2013). 46.Yanagizawa-Drott,D.,Awuah,K.etal.ProjectAPE:AutonomouspolicyevaluationwithAI-generated economicsresearchpapers.SocialCatalystLab,UniversityofZurichhttps://ape.socialcatalystlab.org/(2026). 47.Yamada,Y.etal.TheAIScientist-v2:Workshop-levelautomatedscientificdiscoveryviaagentictreesearch. arXiv2504.08066(2025). 48.Bourdieu,P.Distinction:ASocialCritiqueoftheJudgementofTaste(HarvardUniv.Press,1984). 49. Dijksterhuis,A.etal.Onmakingtherightchoice:Thedeliberation-without-attentioneffect.Science311,1005– 1007(2006). Methods Studydesign.Evaluativejudgmentinscienceiswidelyrecognizedastacitandinstitutionally distributed[8,16,20]:ethnographicevidenceshowsthatevenexperiencedgatekeepersrelyonintuitive assessmentanddisciplinarysensibilityratherthanruleapplication[16],anddecadesofstructuredrubric designhavenotimprovedinter-raterreliability[11].ThisstudywasdesignedtotestwhetherAIsystems canrecoversuchjudgmentfrominstitutionaldecisiontraceswhenaskedtoassessresearchideasbefore empiricalresultsareknown.Weappliedtheprotocoltotwosocialsciencefields,organizational psychology/managementandeconomics,totestboththemechanismanditsgenerality.Wecenteredthe protocolonearly-stageresearch-pitchevaluationinsteadoffull-papercritique,andimposedasharedfour- tierdecisionspaceacrossallevaluatorclasses(humanandAI)sothatperformancedifferencesreflect differencesinjudgment,nottaskformat,criterionframing,oraccesstostylisticcues.Humanevaluation, frontier-modelbenchmarking,andmechanismtests(reinforcement-learningablation,pairwise discrimination,votingconsistency)wereconductedinmanagement;economicsservedasareplicationof thecoreSFTmechanismwitharchitecture-matchedbasecontrols. Benchmarkconstructionandrationaleforlabels.Weconstructedbalancedbenchmarksintwo fields:120article-derivedresearchpitchesinorganizationalpsychologyandmanagement(30pertier,19- journalsourceuniverse;SupplementaryMethodsSM5)and200ineconomics(50pertier,38-journal sourceuniverse;SupplementaryTableST14).Allsourcearticleswerepublishedaftermid-2025.Inboth fields,thefourtiers(exceptional,strong,fair,andlimited)werederivedfromjournal-levelpublication outcomesmappedtopre-specifiedtierframeworksreviewedbydomainexperts.Operationally, exceptionaldenoteselitefield-definingjournals,strongnear-elitespecialtyoutlets,fairestablishedfield journals,andlimitedlower-prestigeornarrower-scopeoutlets.Journal-levellabelswerechosenbecause theyrepresentthemoststableinstitutionalrecordoflong-rungatekeepingdecisionscurrentlyavailableat scale[19,20]:individualreviewerjudgmentsaredominatedbynoise[30],butthesystem-levelsortingthat producesprestige-tierdifferentiationintegratesiteratededitorialconsensusovertime.Thismakesthe trainingtargetexplicit:themodellearnstorecovercollectiveinstitutionaljudgment,nottoapproximate anysinglereviewer’sopinion. Abalanceddesignwaschosentopreventmodelsfromexploitingclass-frequencypriorsandtoenable interpretableper-tiercomparisons,especiallyatthequalityextremeswhereeditorialdecisionsaremost consequential.Tierbalanceisensuredbyconstruction(30pitchespertierinmanagement;50pertierin economics). Inputstandardizationandextractionworkflow.Evaluatorcomparisonsaremeaningfulonlyifall evaluatorsseethesameinformation.Eachsourcearticlewasthereforetransformedintoastandardized research-pitchtextusingafixedextractionworkflow(SupplementaryMethodsSM4).Thesepitches presentedthecoreresearchquestionandtheoreticalframingwithoutdetailedmethods,fullempirical findings,journalidentity,orauthoridentity.TheprimaryextractorwasQwen3-235B-A22B-Instruct,a largelanguagemodelselectedforextractionqualityandformatstabilityaftercross-validationagainst alternativestrongmodels.Wegeneratedfivetextualformulations(varyinginthebalanceoftheoretical context,research-questionspecificity,andmethodologicaldetail)andselectedtheresearch-question-with- contextformastheprimaryrepresentationthroughpilotevaluationonaheld-outdevelopmentset, becauseitpreservestheoreticalmotivationwhilestrippinglater-stageevidencethatwouldtriviallyencode publicationoutcomes.Thisnormalizationisolatesidea-levelassessmentandreducesshortcutlearning. Thesameextractionpipelineandpolicywereappliedacrossbothfields,allmodelfamilies,andmaterials giventohumanraterssothatevaluatorcomparisonsremainedbehaviorallyaligned. Trainingcorpusandleakagecontrol.Foreachfield,supervisedtrainingdataweredrawnfromthe field’ssourceuniverseandfullydisjointfromitsheld-outbenchmark.Themanagementcorpuscomprised aprimaryrecent/newslice(4,479research-pitch/journal-outcomepairsfrom19journals)andanolder slicefortemporalcomparison(3,368pairs).Theeconomicscorpuscomprised5,593pairsfrom38 journals.Apooledcorpuscombiningbothfields(~10,072pairs)wasalsoconstructed.Tierlabels followedthesamejournal-levelmappingusedinevaluation.Ambiguoussampleswereexcludedduring curationtoreduceavoidablelabelnoiseinthetrainingtarget. Temporalseparationwasintroducedtocontrolcontaminationrisk:sourcearticlesunderlyingthe benchmarkpitcheswerepublishedaftermid-2025,whereasthebasemodelsusedforfine-tuningwere releasedbeforethatperiod. Becausepublicationoutcomesareinfluencedbyfactorsbeyondideaquality(executionquality, writingquality,reviewer-assignmenteffects,editorialfit)andbecausethemodelinputcapturesonlyidea- levelinformation,anoiseceilingisexpectedbydesign[30](SupplementaryMethodsSM8).Thisceiling affectsabsoluteaccuracyinterpretationwhileleavingrelativecomparisonsinternallyvalid,sincethesame informationconstraintisimposedacrossallevaluatorclasses. Evaluationcriteriaandpromptfreezingstrategy.Allevaluatorsusedthesametwo-axiscriteria frameworkcenteredonoriginalityandusefulness,consistentwitheditorialguidanceliteraturein managementresearch[9,10].AsharedrubricwasimposedonbothhumansandAIsothatscore differenceswouldreflectevaluatorbehavior,notcriterionmismatch.Foreconomics,thesameframework wasadaptedtoabroadersocial-scienceframingwhilepreservingthetierstructureandzero-shotdesign (SupplementaryMethodsSM9).TierdefinitionsandpromptvariantsareprovidedinSupplementary MethodsSM6. Threepromptvariantswerepre-specifiedbeforefinalruns:anexpert-anchoredrubric(usingthe languageandstandardsofexperiencedjournaleditors),asimplifiedplain-languagerubric,andajournal- anchoredrubric(referencingspecificpublicationvenuesasqualityanchors).Theexpert-anchoredand journal-anchoredvariantsyieldedcomparablefrontierperformanceinpre-analysischecks,butjournal- anchoredpromptsinducedmodelstoenumeratejournallistsandhuntforvenue-specificstylisticcues duringreasoning;theexpert-anchoredvariantwasthereforefrozenastheprimaryprompttoavoidthis confound,makingdownstreamcomparisonsconservativeforfrontiersystemsinsteadofoptimizedfor fine-tunedmodels.Cross-modelpromptsensitivityanalysisisreportedinExtendedDataFig.1.All evaluationsusedzero-shotpromptingtokeepconditionscomparableacrossevaluatorclasses,including modelsandhumanraters,andtoavoidnoisyexemplaranchoringfrompublicationoutcomelabels. Unresolvedmodeloutputswerecodedasincorrectforfixeddenominatorreporting(overallnon- compliancerate<1%across10,560individualruns). Humanevaluationprotocol.Thehumanstudywasapprovedbyinstitutionalreviewboardreview (ProjectNo.THU-04-2026-0034). Twopanelswereusedtoseparateexperiencedgatekeeperjudgmentfromtraineejudgmentunderthe sametaskdefinition.Theexpertpanelcomprised48raterscontributing384ratings(8researchpitchesper rater;mean3.2ratingsperpitch),recruitedthroughpersonalprofessionalnetworkstoensurehighdomain relevanceandintrinsicmotivation;twoanonymousresponseidentifierswerereconciledastwodistinct raters.Thejuniorpanelcomprised174doctoralandpostdoctoralresearcherscontributing2,530ratings (mean14.5pitchesperrater;mean21.1ratingsperpitch),recruitedviadoctoralandpostdoctoral networks;juniorraterswhospentfewerthanoneminuteperpitchwereexcludedtoremoveperfunctory responses.Unfilteredexpertdatawereprespecifiedastheprimaryexpertanalysisbecausefiltering reducespitch-levelvotecountsandincreasestiesensitivity,limitingthepitchesonwhichmajorityvoting canbecomputed.FilteredsensitivityanalysesarereportedinSupplementaryTableST7. Foreachbenchmarkpitch,ratersreportedpriorexposure(yes/no),tierassignmentusingfield-familiar labels(Top,Top-,Good,Fair,mappedtoexceptional,strong,fair,limited),confidence(5-pointLikert scale),andtopicfamiliarity(5-pointLikertscale).Theselabelsmirrorcommonjournal-evaluation shorthand,whileallanalysesconvertthemdeterministicallytotheunifiedfour-tierlabelsusedintheAI protocolsandmanuscript.Thefullsurveyinstrumentandproceduraldetailsareprovidedin SupplementaryMethodsSM7.Wecollectedconfidenceandfamiliaritybecausepriorworkshowsthat peer-reviewagreementandevaluatorqualityareweaklycoupledtoconventionalexpertisesignals[11,12]; thesevariablesalloweddirecttestingofwhetherthesamepatternheldinourdataset.Demographicand backgroundsummariesareinSupplementaryTableST5;individualexpertaccuracyvaluesarein SupplementaryTableST9;matched-NMonteCarlodetailsareinSupplementaryTableST10;anda conciseprior-exposuredescriptivesummaryisreportedinSupplementaryTableST11. AImodelfamiliesandinferenceprotocols.WeevaluatedthreeAIfamiliesunderonefrozen instructionscaffold.Thefrontiercohortincluded11reasoningmodels:Gemini3.1Pro,ClaudeOpus4.6, GPT-5.2High,Gemini2.5Pro,Qwen3.5Plus,DeepSeekV3.2,Seed2.0,MiniMaxM2.5,KimiK2.5, Grok4.1Fast,andGLM-5.Gemini3.1Proisretainedwithacontamination-riskcaveatbecausesome reasoningtracespartiallyreproducedbenchmarkcontent.Eachfrontiermodelwassampledeighttimes perpitchtoaccountfortheinherentrandomnessintextgeneration.Theprimaryfrontiermetricwaspitch- meaneight-sampleaccuracy.Fordiscretediagnostics,weusedper-pitchmajorityvotefromthesame eightrunsandreportedeffectivesamplesizeaftertieexclusion.Protocolcoverageissummarizedin SupplementaryTableST13. Weadditionallyevaluatedchat/log-probabilitytracks(GPT-5.2chat,KimiK2chat,DeepSeekChat) andarchitecture-matchedbasecontrols(thesameunderlyingmodelsbeforeanyfine-tuning).Eachbase modelwasevaluatedonthefull120-pitchbenchmarkusingthesamefrozenpromptandfour-labellog- probabilityclassificationasitsSFTcounterpart,providingarchitecture-matchedcontrolsthatisolatethe effectoffine-tuningfrompre-existingmodelcapability.Non-thinkingmodels(base,chat,andsupervised fine-tunedcheckpoints)wereevaluatedbyfour-labeltokenlog-probabilityclassification,whichyields deterministicclassprobabilitiesandavoidsfree-textparsingfailuremodes.Thesamelog-probability protocolwasusedforallfourSFTcheckpointsandtheirsixpairwiseprobability-averagingensembles (SupplementaryTableST4). ThisdualprotocolreflectsadifferenceinwhateachAPIsurfaceexposes:reasoningAPIsproduce stochasticnatural-languagecompletionsbeststabilizedbyrepeatedsampling,whilenon-reasoningAPIs exposetoken-levellog-probabilitiesthatyieldcalibratedlabeldistributionsinasingledeterministiccall, enablingbothconfidenceanalysisanddirectprobabilisticcomparison.Costandprotocolcomparisonsare summarizedinSupplementaryTableST2. Supervisedfine-tuning.Supervisedfine-tuning(SFT)isaprocedureinwhichapre-trainedlanguage modelisfurthertrainedonacurateddatasetofinput–outputexamplestospecializeitforaparticular task[41].Wefine-tunedbasemodelsspanningtwomodelfamiliesandmultiplescalesusingthesame procedureinbothfields.Inmanagement,fourmodelsweretrained:GPT-4.1,GPT-4.1-nano(asmaller, moreefficientvariant),Qwen3-4B-Instruct,andQwen3-30B-A3B-Instruct(amixture-of-experts architecturewith30billiontotalparametersbutonly3billionactiveatanytime).Ineconomics,threeof thesefourarchitecturesweretrained(GPT-4.1-nano,Qwen3-4B,Qwen3-30B);GPT-4.1wasomitted becauseitsAPIfine-tuningcostwasdisproportionatetothethreecost-effectivearchitecturessufficientto demonstratecross-architecturereplicability.Pooledmodelswereadditionallytrainedonthecombined corpusfrombothfields(~10,072pairs).Thecross-family,cross-scale,cross-fielddesigntestswhetherthe evaluativesignalisrecoverablefrominstitutionaltracesgenerallyorspecifictoaparticulararchitecture, scale,ordiscipline[34]. Eachtrainingexampleconsistedofthefrozenevaluationpromptwrappingaresearch-question-with- contextpitchasinput,withasingletier-labeltoken(oneoffour:exceptional,strong,fair,limited)asthe completiontarget.Labelsweredesignedassemanticallydescriptivesingletokenstoensure distinguishabilityunderfirst-tokenlog-probabilityclassificationandtoleveragethemodels’pre-trained qualitygradingrepresentations.Thesameprompttemplatewasusedfortrainingandevaluation, eliminatingpromptmismatchconfounds.Giveninputtextxandlabely∈1,2,3,4,trainingminimized label-tokennegativelog-likelihood,astandardobjectivethatincreasestheprobabilitythemodelassignsto thecorrecttierlabel: ℒ NLL ( � )= − 1 � �=1 � log � � � (�) ∣� (�) , withlosscomputedonlabeltokensonlyandinputtokensmaskedfromgradientupdates.Inputmaskingis critical:itpreventsthemodelfrommemorizingprompttokensandforcesthemodeltolearnexclusively themappingfromresearchcontenttoqualitytier.Trainingcorpussizesarereportedabove(Training corpusandleakagecontrol).Managementtrainingcostwasbelow$300USDacrossallfourmodels,with thelargestsinglecostarisingfromGPT-4.1APIfine-tuning(~$200);GPT-4.1-nanorequired~$10via API,whiletheopen-sourceQwen3-4BandQwen3-30B-A3Bcheckpointsrequiredapproximatelyoneand eightA100GPU-hoursrespectively;economicstrainingcostswerecomparableforthethreearchitectures used(SupplementaryTableST2).QwencheckpointsweretrainedlocallyusingTRL;GPTcheckpoints weretrainedviatheOpenAIfine-tuningAPI.Allmodelsusednear-defaultsettingswithout hyperparametersearch;fullsettingsarereportedinSupplementaryMethodsSM1.Consistentperformance gainsacrossbothaproprietaryblack-boxpipelineandafullycontrolledopen-sourcepipeline,acrosstwo fieldswithindependentjournalhierarchies,isolatethetrainingsignalandnotanyparticularoptimization decisionasthecausalfactor. Totestout-of-formatgeneralizationundersevereinformationcompression,wealsoevaluateda compressed-inputtransfersetbuiltfromthesameheld-out120articlesbutusingonlythesingle-sentence idea-statementfield(core_rq_short)extractedinSupplementaryMethodsSM4.Unliketheprimaryfull idea-summaryrepresentation(rq_with_context),thisone-sentenceversionremovesalmostalltheoretical motivationandcontextualframing,leavingonlythefocalresearchquestioninonesentence.No checkpointwastrainedonthisformat:allsupervisedfine-tuningusedthefulleridea-summaryinputonly, sothisanalysisteststransferatevaluationonly,notin-formattestperformance.ThefourGPT-family base/SFTcheckpointswerescoredwiththesamefour-labellog-probabilityclassifierusedelsewhere. Predictionswererecoveredbyargmaxoverthefourlabellog-probabilities;anymissinglabelwastreated asnegativeinfinity;andexacttieswereresolvedbyafixedlabelorder(exceptional,strong,fair,limited). Becausethebenchmarkremainsexactlybalancedacrossfourtiers,descriptivechanceperformanceis25% accuracy,andone-sidedexactbinomialtestsversus25%wereusedonlyasacompactabove-chance diagnosticforthisauxiliarytransferevaluation(ExtendedDataFig.7).Thisauxiliaryanalysiswas interpretedasatestofgeneralizationrobustnessratherthanasareplacementfortheprimaryin-format benchmarkcomparison. Atwo-modelensemblewasbuiltbyaveragingsoftmax-normalizedlabelprobabilitiesfromGPT-4.1- nano(SFT)andQwen3-30B-A3B(SFT),thenselectingthemaximumprobabilityclasswithdeterministic resolutionofanyexactclassprobabilityties.WeevaluatedallsixpairwiseSFTensemblesontheheld-out benchmarkandrankedthembyaccuracy,thenmacroF1,withafixedmodelorderconventionusedonlyif thosemetricsremainedtied;underthisruleGPT-4.1-nano+Qwen3-30B-A3Bwasretainedastheprimary pair(SupplementaryTableST4).Becauseallsixpairwiseensemblesexceededthefrontieraverageby 28.1to29.8percentagepoints,thecorefindingisrobusttoensemblecompositionandnotdependentona singlefortuitouscombination. Architecture-matchedbasecontrolsserveastheprimarycomparisonforisolatingtheSFTeffectin bothfields:becauseeachfine-tunedmodelisevaluatedagainstitsownpre-trainingcheckpointonthe identicalbenchmark,observedgainsareattributabletothefine-tuningprocedureandnottodifferencesin base-modelcapabilityorarchitecture.GPT-4.1wasadditionallyretainedforthecross-fieldtransfer evaluation(management-trainedmodeltestedontheeconomicsbenchmark;SupplementaryFig.7) becauseitdemonstratedthestrongestgeneralizationcapacityunderinputcompression(ExtendedDataFig. 7).Fulleconomicsjournal-to-tiermappingdetailsareinSupplementaryMethodsSM9(Supplementary TableST14). Reinforcement-learningablation.Totestwhetherexplicitreasoningoptimizationimprovesthistask beyondsupervisedalignment,wetrainedreasoning-enabledQwen3-4BandQwen3-32Bcheckpointswith amodifiedGRPO-styleobjective[28,31,40](SupplementaryMethodsSM2–SM3).Theimplementation removedtheKLpenalty,usedtoken-levelnormalization,andadoptedasymmetricclippingtofavor positive-advantageupdates.Wealsointroducedadaptiveprivilegedsamplingforlow-accuracyitems[42, 43]tomaintainrewardcontrastduringtraining. RLcheckpointswereevaluatedusingthesame8-runsamplingprotocolasfrontiermodels,withrun- pooledandpitch-meaneight-sampleaccuracyreportedalongsidenon-tiedmajorityvote.RLtraining inherentlyrequiresreasoning-enabledmodelconfigurationsthatgenerateexplicitchain-of-thoughtbefore afinallabel;thisisnotaconfoundbutthemechanismundertest.Thecomparisonaskswhetherexplicit deliberationaidsorimpedesevaluativeaccuracy,aquestionmotivatedbyconvergingevidencefrom cognitivesciencethatverbalarticulationcandegradeholisticjudgment[35],andfrommachinelearning thatchain-of-thoughtpromptingyieldsnegligibleornegativegainsontasksoutsidemathematicsand symbolicreasoning[36,37].Reinforcementlearningfromhumanpreferenceshasadditionallybeenshown toinstillagreement-seekingbehavior[28,29]andreduceoutputdiversitythroughmodecollapse,and GRPO-stylemethodsrepresentthecurrentstateoftheartforreasoningoptimization[31].ThisRLtrack wasthereforetreatedasamechanismtest,notadeploymentcandidate:ifexplicitchain-of-thoughtpolicy optimizationcannotmatchdirectsupervisedalignmentonthistask,theimplicationisthattheevaluative signalresidesintheinstitutionaltracesthemselvesandnotinreasoningarchitecture. Pairwisediscriminationexperiment.Wedesignedalabel-freepairwisetasktodistinguishgainsin intrinsicdiscriminationfromgainsinrubricalignment.Eachtrialpresentedtworesearchpitchessampled fromdifferenttiers,andthemodelselectedthestrongeronewithoutseeingtierlabelsorrubrictext, testingwhetheramodelcantellwhichoftwoideasisbetterevenwithoutthefour-categoryframework. Weusedafixed300-pairstratifiedsetsampledfromthe120benchmarkpitcheswithfixedrandomseed 32(150pairsattierdistance1,100atdistance2,and50atdistance3,wheredistance1meansadjacent tiersanddistance3meansthemostseparatedtiers). Thenarratedpairwisecomparisonsetwasthesharedfour-modelsubsetusedacrossthemaintext,Fig. 5,ExtendedDataFig.2,andSupplementaryTableST1:SFTGPT-4.1,Gemini3.1Pro,GPT-5.2High, andtheGPT-4.1baseline.Becausethistaskremovesabsolutecategoryassignment,improvementshere areinterpretableasimprovementsinrelativequalitydiscriminationratherthanprompt-followingalone. Fig.5carriestheheadlinepairwisecomparisonforthissamesubset:overallweightedaccuracywithexact McNemartestsandthehard-boundarysummary.ExtendedDataFig.2retainsthesix-pair-typeheatmap anddiscordancedecompositionforthesamesubset. Voting-consistencyanalysis.Consensusanalyseswereprespecifiedtoquantifyhowreliability changeswithstricteragreementthresholdsacrossevaluatorclasses.ForSFT,weanalyzed4/4,3/4,and 2/4cross-modelagreementpolicies.Forhumans,weanalyzedjuniorvotesharethresholdsandexpert unanimitysubsets.Forfrontiermodels,weanalyzedbothwithin-modelrepeated-samplingaggregation andfixedcross-modelvotingonadiversesubsetoffourtopmodels(Gemini3.1Pro,ClaudeOpus4.6, GPT-5.2High,GLM-5). Inter-raterconsistencyusedFleiss’κforhumanpanelsandpairwiseCohen’sκformodel-model agreement.Majority-voteresultsexcludedtiedpitchesandalwaysreportedeffectivenon-tiedN.Ties werenotbrokenrandomlybecauserandomresolutionintroducesavoidablevarianceandobscureswhether anevaluatorfamilyisgenuinelyindecisiveondifficultcases.Fullagreementdiagnosticsarereportedin SupplementaryTableST12. Statisticalanalysis.AllanalyseswereruninPython(NumPy,SciPy,pandas).Theprimaryendpoint wasfour-classexact-matchaccuracy.Secondaryendpointsincludedmacro-F1,per-tierprecision/recall/F1, confusionmatrices,andcalibrationdiagnostics. Ordinalinter-rateragreementwasassessedwithKrippendorff’salpha.Pairedevaluatorcomparisons usedMcNemartestsonpairedcorrectnessvectors.ForthepairwisetasksummarizedinFig.5and diagnosedfurtherinExtendedDataFig.2,significancewascomputedwiththetwo-sidedexact McNemar/binomialtestonitem-leveldiscordantpairsandreportedasrawunadjustedPvaluesforthe plottedSFTcomparisonsversusGemini3.1Pro,GPT-5.2High,andGPT-4.1baseline.Frontier-cohort heterogeneitywastestedwithCochran’sQ.Ordinalornon-GaussiananalysesusedSpearmancorrelation, Mann–WhitneyU,andKruskal–Wallistests.Individual-levelcomparisonstestedwhethertheSFT ensembleaccuracydifferedfromthedistributionofindividualhumanevaluatoraccuraciesusingaone- samplet-testwiththeensemblescoreasthereferencevalue.Confidenceintervalswereestimatedby bootstrapresampling(10,000draws).Multiple-testingcorrectionswereusedonlywhereexplicitlystated forbroaderevaluatorfamiliesandprompt/modelsweeps;thefocusedpairwiseanalysesinSupplementary TableST8andFig.5/ExtendedDataFig.2reportrawunadjustedPvalues. Thereportedsignificancetableshighlightfourcomparatorrows:frontieraverage(11models),best frontiermodelundertheconservativeprotocol(Gemini3.1Pro),expertmajorityvote,andjuniormajority vote.ThepairedprimarycomparisonsemphasizedinthemanuscriptandsupportingtablesareSFT ensembleversusexpertmajorityvoteandversusthebestfrontiermodelundertheconservativeprotocol; frontier-averageandjunior-majorityrowsarereportedasadditionalheadlinesecondarycomparisons. Additionalpairedevaluatortestswerereportedassecondaryanalyses,withrawunadjustedPvalues showninthesupportingtablesandanymultiplicity-adjustedinterpretationstatedexplicitlyintext. Headroomisdefinedas(accuracy-chance)/(100%-chance),wherechance=25%,givingthefraction ofimprovableperformancecapturedbyagivenevaluator.Humanjudgmentanalysesincorporatedthe noiseframeworkofKahnemanetal.[30]tointerpretthedissociationbetweenlowcategoricalagreement andmoderateordinalagreement.Majority-votecomparisonsinvolvereducedeffectiveNbecausetied pitchesareexcluded,whichlimitsstatisticalpowerrelativetoindividual-leveltests;thisisparticularly relevantforthecomparisonbetweenSFTandexpertmajorityvote,wheretieexclusionreducesthe evaluableset. Calibrationforprobabilisticevaluatorsusedexpectedcalibrationerror(ECE)andBrierdecomposition; selectivepredictionwasevaluatedbyplottingaccuracyasafunctionofcoverageunderconfidence thresholdsweeps.MonteCarlomatched-Nanalyses(5,000randomdraws)wereusedtocomparejunior andexpertmajorityvotingatequivalentpanelsizes,drawingexpert-sizedpanelsfromthejunior researcherpoolandcomputingmajority-voteaccuracyoneachdraw.Labelnormalizationacrosssource- articlemetadata,humansurveys,andmodeloutputsuseddeterministicmappingrules(Supplementary TableST6). Dataandcodeavailability.Allpreprocessingrules,prompts,modelinventories,hyperparameters, andsensitivityanalysesaredocumentedinSupplementaryMethodsandTables.Thebenchmarksplitwas fixedbeforefinalmodelcomparisons,andlabelnormalizationrulesweredeterministicacrosssource- articlemetadata,humansurveys,andmodeloutputs.Analysiscode,processedbenchmarkfiles,prompts, andtable-generationscriptswillbereleasedalongsidethemanuscript.Humandatawillbede-identified beforereleaseinaccordancewithapprovedethicsprocedures. ReferencescitedinMethods 8.Polanyi,M.TheTacitDimension(Doubleday,1966). 9.Corley,K.G.&Gioia,D.A.Buildingtheoryabouttheorybuilding.Acad.Manag.Rev.36,12–32(2011). 10.Colquitt,J.A.&George,G.PublishinginAMJ—Part1:Topicchoice.Acad.Manag.J.54,432–435(2011). 11.Bornmann,L.,Mutz,R.&Daniel,H.-D.Areliability-generalizationstudyofjournalpeerreviews.PLoSONE5, e14331(2010). 12.Pier,E.L.etal.LowagreementamongreviewersevaluatingthesameNIHgrantapplications.PNAS115,2952– 2957(2018). 13.Lamont,M.HowProfessorsThink:InsidetheCuriousWorldofAcademicJudgment(HarvardUniv.Press, 2009). 14.Siler,K.,Lee,K.&Bero,L.Measuringtheeffectivenessofscientificgatekeeping.PNAS112,360–365(2015). 15.Collins,H.TacitandExplicitKnowledge(UniversityofChicagoPress,2010). 16.Christiano,P.F.etal.Deepreinforcementlearningfromhumanpreferences.NeurIPS(2017). 17.Sharma,M.etal.Towardsunderstandingsycophancyinlanguagemodels.Proc.ICLR(2024). 18.Kahneman,D.,Sibony,O.&Sunstein,C.R.Noise:AFlawinHumanJudgment(Little,BrownSpark,2021). 19.Guo,D.etal.DeepSeek-R1incentivizesreasoninginLLMsthroughreinforcementlearning.Nature645,633– 638(2025). 20.Wu,Y.etal.OnthegeneralizationofSFT:Areinforcementlearningperspectivewithrewardrectification. ICLR(2026). 21.Wilson,T.D.&Schooler,J.W.Thinkingtoomuch:Introspectioncanreducethequalityofpreferencesand decisions.J.Pers.Soc.Psychol.60,181–192(1991). 22.Liu,R.etal.Mindyourstep(bystep):Chain-of-thoughtcanreduceperformanceontaskswherethinkingmakes humansworse.Proc.ICML(2025). 23.Sprague,Z.etal.ToCoTornottoCoT?Chain-of-thoughthelpsmainlyonmathandsymbolicreasoning.ICLR (2025). 24.Shao,Z.etal.DeepSeekMath:Pushingthelimitsofmathematicalreasoninginopenlanguagemodels.arXiv 2402.03300(2024). 25.Wei,J.etal.Finetunedlanguagemodelsarezero-shotlearners.ICLR(2022). 26.Yu,Q.etal.DAPO:Anopen-sourceLLMreinforcementlearningsystematscale.NeurIPS(2025). 27.Qu,Y.etal.POPE:Learningtoreasononhardproblemsviaprivilegedon-policyexploration.arXiv2601.18779 (2026). ExtendedData ExtendedDataFigures ExtendedDataFigure1|Cross-modelprompt-sensitivitylandscape a,AccuracyacrossSimple,Journal-anchored,andExpertpromptsforthesixfrontiermodelsevaluatedunderallthreeprompt formulationsintheconservative11-modelcohort:Gemini3.1Pro,GLM-5,GPT-5.2High,Qwen3.5Plus,Gemini2.5Pro,and ClaudeOpus4.6.Accuracyerrorbarsshow95%binomialconfidenceintervals.b,Macro-F1acrossthesamepromptconditions, shownasapairedrobustnessmetrictodistinguishrawaccuracyfrombalancedtierdiscrimination.c,Pooledrow-normalized confusionmatrixundertheExpertpromptbaseline.d,Pooledrow-normalizedconfusionmatrixundertheSimpleprompt.e, Pooledrow-normalizedconfusionmatrixundertheJournal-anchoredprompt.Togetherthesepanelsshowthatpromptwording changestheshapeoffrontier-modelcollapsebutdoesnotrecoverrobustfour-tierdiscrimination.BecausetheSimpleandJournal conditionsusesingle-passpromptevaluationswhereastheExpertconditionusestheconservativemajority-basedfrontierprotocol, cross-promptdifferencesshouldbeinterpreteddirectionallyratherthanasperfectlyprotocol-matchedestimates. ExtendedDataFigure2|Pairwiseintrinsicdiscriminationdetails Modelschoosethestrongerresearchpitchinpairwisehead-to-headcomparisonswithouttierlabelsorrubrictext,isolating relativequalitydiscriminationfromabsolutecategoryassignment.TheplottedED2comparatorsetisthesamesharedfour-model subsetusedinFig.5andSupplementaryTableST1:SFTGPT-4.1,Gemini3.1Pro,GPT-5.2High,andtheGPT-4.1baseline.Fig. 5nowcarriestheheadlineoverall-accuracyandhard-boundarypairwisepanels.a,Asix-pair-typeheatmap(strong_exceptional, fair_strong,fair_exceptional,limited_fair,limited_exceptional,limited_strong)showswhereperformancegapsactuallysitonce thebroaddistancebinsareunpacked.b,Netdiscordant-pairdecomposition(SFT-onlycorrectminuscomparator-onlycorrect) showsthatthepairedsignificanceisconcentratedinthehardestpairtypesratherthaninthenear-ceilingeasypairs.Raw unadjustedexactMcNemarPvaluesforSFTGPT-4.1versusGemini3.1Pro,GPT-5.2High,andtheGPT-4.1baselineare reportedindata/statistics/S14_ED2PairwiseRawPValues.json. ExtendedDataFigure3|FullconfusiondiagnosticsandSFTensembledetail a,Thecompleteconfusion-matrixatlasforall11frontierflagshipsextendstheselectedexamplesshowninFig.2andmakesthe fullrangeofcollapsemodesvisibleinoneview.b,Architecture-matchedbase-versus-SFTconfusionmatricesshowhow supervisedfine-tuningchangestierdiscriminationwithineachmodelfamily,shiftingmasstowardthediagonalratherthantoward oneortwodominanttiers.c,AccuracyrankingacrossthefoursingleSFTmodelsandsixtwo-modelensemblesshowsthatthe ensembleadvantageisrobustacrosspairchoicesratherthantiedtoasinglecombination;dashedreferencelinesmarkthebest flagshipfrontiermodelandexpert-majorityperformanceforcontext. ExtendedDataFigure4|Calibration,errorstructure,andpanel-sizereliability a,Directionalerrorprofilescompareunder-estimationandover-estimationacrosstheSFTensemble,bestfrontiermodel,frontier average,expertmajorityvote,andjuniormajorityvote,revealingthatfrontiersystemsareespeciallypronetoover-estimating articlequality.b,Thecalibrationlandscapecomparesexpectedcalibrationerror(ECE)withBrierscoreacrossevaluators,with better-calibratedsystemsappearingclosertothelower-leftcorner.c,Errorseverityisdecomposedintoexactmatches,off-by-1 errors,andoff-by-2+errors,withcompactunder-versusover-estimationsummariesprintedforeachevaluator.d,Juniorpanel- sizescalingisshowntogetherwithjuniortie-ratetrajectoriesandexpertanchorlinesforbothaccuracyandtierate,illustrating thataggregationreducesindecisionbutplateausinperformanceratherthaneliminatingthestructuralceiling. ExtendedDataFigure5|RLcheckpointdiagnostics a,RLcheckpointperformanceiscomparedwiththeSFT2-modelensemble,thebestfrontiermodel,GPT-5.2High,andthe frontieraverage,locatingRLbetweencleanfrontierbaselinesandsupervisedfine-tuning.b,RLperformanceisdisaggregated intorun-pooledeight-sampleaccuracy,pitch-meaneight-sampleaccuracy,andnon-tiedmajority-voteaccuracy,showingthatthe rankingisstableacrossreportingconventions.c,TheRLmajority-voteconfusionmatrixlocalizesresidualerrorsbytierand showsthatmiddle-tierseparationremainsthemainweakness.d,Predicted-tierdistributionscompareRLwiththebestfrontier modelandtheSFT2-modelensemble,clarifyingthatRLreducesbutdoesnoteliminatecollapserelativetothestrongestfrontier baseline. ExtendedDataFigure6|Temporalpersistenceofinstitutionaltraces a,Accuracyforthematchedrecent-trainingslice(2020-2025)andolder-trainingslice(2015-2020)onthetwoshared architectures(GPT-4.1-nano,Qwen3-30B)plustheirmatched2-modelensemble.Thebenchmarkitselfisbuiltfrompost-June- 30-2025articles,sotheoldersliceintroducesanintentionalfive-yeartraining-timelag.Accuracybarsshowbinomial95% confidenceintervals,anddashedreferencelinesmarkthefrontier-averageandexpert-majoritycomparatorsfromthebenchmark summarytables.b,Macro-F1forthesamematchedsinglesandmatched2-modelensemble,withbootstrap95%confidence intervalsandthesamebenchmarkanchors,showingthattemporaldecaypersistsonabalanceddiscriminationmetricratherthan onlyonrawaccuracy.c,Row-normalizedconfusionmatricescomparetherecentandolder-tracematchedensembles,showing thattheoldertraceretainsexceptional-tierdiagonalstructurebutlosessubstantiallymorefair-tierrecoveryandshiftsmass upward.d,Predicted-tierdistributionsplusthreetargeteddiagnosticsquantifytheolder-tracedrift:moreexceptionalpredictions, lowerexceptional-tierprecision,lowerfair-tierrecall,andmorestrong->exceptionalconfusions. ExtendedDataFigure7|Transferfromrichersupervisiontoone-sentencecore-questioninputs Allfine-tunedcheckpointsinthisfigureweretrainedonlyonthefullerresearch-ideasummaryandwerenevertrainedontheone- sentenceidea-statementformat.a,Overallperformanceontheone-sentenceideastatementbenchmarkforarchitecture-matched GPT-familybaseandfine-tunedmodels,shownasaccuracyandmacro-F1onthesame120held-outarticles.Barsshowone- sentence-inputresultswithbinomial95%accuracyconfidenceintervalsandbootstrap95%macro-F1confidenceintervals; hollowmarkersshowthesamemodel’sperformanceonthefulleridea-summarybenchmark,andthedashedlinemarksthe25% chancebaseline.b,Per-tierrecallchange(one-sentenceinputminusfullideasummary)localizeswherecompressionhurts.GPT- 4.1fine-tuningretainssubstantialsignalbutlosesrecallintheexceptional,strong,andfairtierswhilegaininglimited-tierrecall, whereasGPT-4.1nanofine-tuningcollapsesonthestrongtierentirelyundercompressedinput.c,Predicted-tierdistributions underfullerversuscompressedinputsshowthatdegradationisnotasinglegenericmiddle-tiercollapse:bothbasemodels concentrateheavilyinthemiddle,GPT-4.1fine-tuningbecomesmarkedlymoreconservativeandlimited-heavy,andGPT-4.1 nanofine-tuningstopsemittingstrong-tierpredictionsonthecompressedinput. ExtendedDataFigure8|Pooledmulti-fieldtrainingpredictsmanagementandeconomics Pooledmulti-fieldcheckpointstrainedjointlyonmanagementandeconomicsinstitutionaltracesaretestedseparatelyonthetwo held-outfieldbenchmarks.Thetestsurfacecontains120managementitemsand200economicsitems,withtierbalancepreserved ineachfieldsubset.a,AccuracyacrossthetwotestfieldsforthepooledQwen3-30BandpooledGPT-4.1-nanocheckpoints.The pooledQwen3-30Bachieves61.7%onmanagement–comparabletothe60.8%single-fieldSFTensemblereportedinthemain text–and69.5%oneconomics,matchingthebestin-domaineconomicsSFT.ThepooledGPT-4.1-nanoreaches52.5%on managementand67.0%oneconomics.Barsshowpointestimateswithbootstrap95%confidenceintervals;thedashedlinemarks the25%chancebaseline.b,Macro-F1acrossthesamefield-wiseevaluations(Qwen3-30B:0.615management,0.705economics; GPT-4.1-nano:0.489management,0.674economics),withbootstrap95%confidenceintervals,confirmingthatthelargerpooled modelmaintainssubstantiallystrongerbalanceddiscriminationthanthesmallermodelacrossbothfields.c,Predicted-tier compositionforthepooledQwen3-30Bcheckpointacrossmanagementandeconomics,shownas100%stackedbarsoverthe fourtiers.d,Predicted-tiercompositionforthepooledGPT-4.1-nanocheckpoint.Together,thesepanelsdemonstratethat evaluativesignalsfrommanagementandeconomicsdonotinterferewhencombinedintraining:thepooledQwen3-30Bmatches orexceedssingle-fieldbenchmarksinbothdomains,whilethecapacitygapbetweenthelargerandsmallerarchitectureshapes howbalancedcross-fielddiscriminationremains.GPT-4.1wasselectedforthecross-fieldtransferanalysis(SupplementaryFig.7) becauseitdemonstratedthestrongestgeneralizationcapacityunderinputcompression(ExtendedDataFig.7),retaining evaluativesignalwhenfullideasummarieswerereducedtoone-sentenceinputs;thiscapacitymadeitthenaturalcandidatefor testingwhethermanagement-trainedevaluativerepresentationstransfertoadifferentdisciplinarydomain. ExtendedDataTables ExtendedDataTable1|Basemodelcontrols ModelModel Key NAccurac y(%) Macr oF1 95% CI Lowe r(%) 95% CI Uppe r(%) Above Chanc e(p) Headroo mSkill (%) Acc exceptiona l(%) Acc stron g(%) Acc fair (%) Acc limite d(%) Pred exceptiona l(n) Pred stron g(n) Pre d fair (n) Pred limite d(n) Base (GPT- 4.1) gpt-4.112032.50.26824.240.87.510.020.060.050.00.01449570 Base (Qwen3 -4B) qwen3 -4b 12026.70.19519.234.21.72.256.746.73.30.0477120 Base (GPT- 4.1- nano) gpt- 4.1- nano 12025.00.18617.533.30.00.030.063.36.70.0308550 Base (Qwen3 -30B) qwen3 -30b- a3b 12022.50.11215.030.0-2.5-3.386.73.30.00.0972300 SupplementaryInformation Machinesacquirescientifictastefrominstitutionaltraces TableofContents SupplementaryMethods SM1:SupervisedFine-TuningHyperparametersandTrainingCorpus SM2:ReinforcementLearningObjectiveandReward SM3:RLInfrastructureandResourceConsumption SM4:Research-IdeaExtraction SM5:Journal-to-TierMapping SM6:EvaluationPromptsandZero-ShotDesignRationale SM7:HumanStudyDesign SM8:Label-NoiseCeilingAnalysis SM9:Cross-FieldValidationinEconomics SupplementaryTables ST1:PairwiseDiscriminationbyTierDistance ST2:CostandInferenceRegimeComparison ST3:CorePer-ClassMetrics ST4:AllPairwiseSFTEnsembleCombinations ST5:HumanPanelCompositionandDescriptives ST6:LabelNormalization ST7:FilteringSensitivity ST8:PairwiseMcNemarTestCompendium ST9:IndividualExpertAccuracyDistribution ST10:MonteCarloMatched-NAnalysis ST11:Prior-ExposureDescriptiveSummary ST12:AgreementandConsensusDiagnostics ST13:ModelInventoryandAccessWindow ST14:EconomicsJournal-to-TierMapping ST15:EconomicsCross-FieldValidationResults SupplementaryFigures SF1:PredictionDistributionComparison SF2:ExpertIndividualAccuracyDistribution SF3:JuniorMonteCarloSubsamplingCurve SF4:FrontierCollapse-MetricLandscape SF5:AI-HumanErrorComplementarity SF6:HumanConfusionMatrices SF7:Cross-FieldTransferfromManagementtoEconomics SupplementaryMethods SupplementaryMethods1(SM1):SupervisedFine-TuningHyperparametersandTrainingCorpus TableSM1.SFTtrainingconfiguration. ParameterQwen3-4B- Instruct Qwen3-30B-A3B-InstructGPT-4.1-nanoGPT-4.1 ArchitectureDense transformer Mixture-of-experts(30B total,3Bactive) Proprietarytransformer (undisclosed) Proprietarytransformer (undisclosed) TrainingTRL(HuggingTRL(HuggingFace)OpenAIfine-tuningAPIOpenAIfine-tuningAPI ParameterQwen3-4B- Instruct Qwen3-30B-A3B-InstructGPT-4.1-nanoGPT-4.1 frameworkFace) Training location LocalGPU cluster LocalGPUclusterOpenAIcloudOpenAIcloud Learningrate1e-42e-5API-managed-defaultAPI-managed-default Schedulercosinecosinemanualtwo-stageschedulemanualtwo-stageschedule Batchsize32323232 Epochs2233+1 OptimizerAdamWAdamWAPI-managedAPI-managed Hardware1xA1008xA100Provider-managedProvider-managed Training duration ~1hour~2hours~1hour~2hours ThefourmainSFTmodelsusedthesamecuratedrecent/newinstitutional-tracesliceandfrozen instructionscaffold.Minimalhyperparameteroptimizationwasused.TheOpenAItrackusedbatchsize32 withapragmatictwo-stagemanualschedule:3epochsatthedefaultlearningrate,followedby1epochat 0.5xthelearning-ratemultiplier.TheQwentrackusedbatchsize32,conventionaltask-tunedlearning rates,andacosineschedulerunderotherwisenear-defaultTRLsettings.Thetemporalreanalysisreused thesamescaffoldwithanolderslice. Trainingcorpus.Theprimaryrecent/newslicecomprised4,479processedresearch-pitch/journal- outcomepairs,derivedfromorganizationalbehaviorandmanagementsourcearticlesdrawnfromthe predefined19-journalsourceuniversedescribedinSM5.Amatchedoldersliceusedforthetemporal comparisoncomprised3,368pairsfromthesamesourceuniverse.Articleswereassignedtierlabelsvia thedeterministicjournal-to-tiermappingdescribedinSM5.Thedistributionwasapproximatelybalanced acrosstiers;articleswithambiguousvenueassignmentorunclearpublicationstatuswereexcludedduring curationtoreduceavoidablelabelnoise.Eachtrainingexampleconsistedofthefrozenevaluationprompt (SM6,Prompt1)wrappingaresearch-question-with-contextpitchextractedviatheSM4pipeline,witha singletier-labeltoken(oneoffour:exceptional,strong,fair,limited)asthecompletiontarget.Losswas computedonlabeltokensonly;inputtokensweremaskedfromgradientupdates,forcingthemodelto learnexclusivelythemappingfromresearchcontenttoqualitytierratherthanmemorizingprompt structure.The120benchmarkideapitches,derivedfromheld-outsourcearticles,werefullydisjointfrom bothtrainingslices. SupplementaryMethods2(SM2):ReinforcementLearningObjectiveandReward RLcheckpointsweretrainedforQwen3-4BandQwen3-32BusingamodifiedGRPO-styleobjectivewith asymmetricclippingandtoken-levelnormalization.Thedesigntestswhetherexplicitchain-of-thought policyoptimizationcanrecoverthesameevaluativesignalcapturedbydirectsupervisedalignment(SM1). Trainingobjective ℒ GRPO (�)=− 1 �=1 � |� � | �=1 � �=1 |� � | min� �,� (�)� � , clip� �,� (�), 1−�, 1+�+� higher � � where � isgroupsize, � � issampledoutput � ,and � �,� ( � )= � � (� �,� ∣� , � �,<� ) � ref (� �,� ∣� , � �,<� ) isthetoken-levelimportanceratio. Rewarddesign Weusedanordinalrewardgatedbyconsistency: �(� � ,�)=� � � label = � � reasoning ⋅�( � � label ,�) with �(� ,�)= 1, |� −� |=0 0.3, |� −� |=1 0, |� −�|≥2 Thegatesuppressesrewardforreasoning-labelmismatch;partialcreditpreservesordinalstructure. Advantagenormalization � � = � � −� � � � + � withper-groupmean� � andstandarddeviation� � . PrivilegedGRPOsamplingstrategy Toaddressadvantagevanishing(whereallsampledoutputsforagivenpromptreceivethesame reward,yieldingzeroadvantageandnopolicygradientsignal),wedevelopedasample-wiseadaptive samplingstrategy.Priortoeachrollout,thecurrentmodelperforms�diagnosticrolloutsoneachprompt �toestimateper-sampleaccuracy.Trainingmodeisthenassignedper-sample:sampleswithaccuracy belowthreshold � areroutedtoPrivilegedGRPOmode,whereahindsighthintgroundedagainstthe ground-truthlabelisprependedtotheprompt.Thismechanismcorrectsdistributionalbiasinthebase model’srollouts,ensuringthatthetraininggroupcontainssufficientrewardcontrastacrossthefulllabel spaceforstablepolicygradientupdates.Therun-levelvaluesof�and�werefixedseparatelyforeach RLexperiment;becausethisRLanalysisispresentedasamechanismtest,notahyperparametersweep, thescientificcomparisoncentersontheadaptive-samplingdesignitself,notonanysinglethreshold setting. SupplementaryMethods3(SM3):RLInfrastructureandResourceConsumption RLtrainingwasconductedonaclusterof8xA100GPUs.Qwen3-4B-Thinkingwastrainedforundera weekandQwen3-32Bforaboutaweek,withsustainedGPUutilizationexceeding95%throughout. TrainingwasbuiltonandextendedthefullyasynchronousAgentRLframeworkopen-sourcedby TsinghuaUniversity,withcustomizationstosupportourrewarddesignanddatapipeline;thefulltraining infrastructurewillbereleasedalongsidethemodelweightsandcode. TableSM3.RLtraininghyperparameters. ParameterQwen3-4B-ThinkingQwen3-32B Learningrate5x10^-51x10^-5 Batchsize3232 Adambetas0.9,0.990.9,0.99 Weightdecay0.050.05 OptimizerAdamWAdamW Clipping � 0.20.2 Asymmetricbonus � higher 0.10.1 Maxgradientnorm1.01.0 AttentionimplementationFlashAttention-2FlashAttention-2 ParallelismDDPFSDP Precisionbf16mixedprecisionbf16mixedprecision InferencebackendSGLangSGLang Hardware8xA1008xA100 Trainingduration<1week~1week SupplementaryMethods4(SM4):Research-IdeaExtraction Tostandardizeinputsacrossalltrainingandevaluationconditions,weusedQwen3-235B-A22B-Instruct (Alibaba)toextractstructuredresearchideadescriptionsfromeacharticle.Theextractionprompt instructedthemodeltoproduceastructureddescriptionincluding:(1)thecoreresearchquestion,(2) theoreticalmotivation,(3)methodologicalapproach,and(4)expectedcontribution,whileomitting methodsdetails,empiricalresults,publicationvenue,andauthoridentities. Modelselection.Theextractionpipelinewasvalidatedbycomparingoutputsacrossmultiplelarge languagemodels,includingClaudeSonnet3.5(Anthropic),Qwen3-235B-A22B-Instruct,Qwen3-32B- Instruct(Alibaba),andothers.Nosubstantivedifferenceswereobservedacrossmodels.Qwen3-235B- A22B-Instructwasselectedastheproductionextractionmodelonthebasisofhumanqualityjudgmentsof outputcoherenceandcompleteness. Extractionprompt.Theextractionpromptinstructedthemodeltoactasanobjectiveresearchpaper analyser,extractingresearchquestionsandcoreelementswithoutinterpretationorembellishment.The promptspecifiedfiveoutputversionsinJSONformat: 1.CORE_RQ_SHORT(40–60words):Distilledessentialresearchquestion(s). 2. RQ_WITH_CONTEXT(120–150words):Researchquestionwithenoughcontextforexpertevaluation, includingthephenomenon,gap,question,approach,andclaimedcontribution. 3. GAP_FOCUSED(100–130words):Whatisknown,whatremainsunknown,andhowthestudyaddresses it. 4. THEORY_AND_MODEL(100–130words):Theoreticalframework,keyvariablesandrelationships,and theoreticalcontribution. 5.CONTRIBUTION_FOCUSED(80–100words):Theoretical,empirical/methodological,andpractical contributionsasclaimedbytheauthors. ThemainbenchmarkusestheRQ_WITH_CONTEXTformat.Criticalextractionrulesrequiredfocusing ontheabstract,introduction,andtheoreticaldevelopmentsections;usingtheauthors’exactterminology forkeyconstructs;preservingtheleveloftheoreticalsophisticationintheoriginal;andavoidingany additionoftheoreticalconnections,persuasivehooks,orinferredcontributionsnotexplicitlystated. Verbatimextractionprompt(exacttext) #ROLE Youareanobjectiveresearchpaperanalyzer.Yourtaskistoextractandpresent researchquestionsandcoreelementsfromacademicpapersWITHOUTinterpretation, embellishment,orimprovement. #CRITICALPRINCIPLE:OBJECTIVITYOVERPERSUASIVENESS -PresentthepaperEXACTLYaswrittenbytheauthors -DoNOTaddtheoreticalsophisticationifit'snotthere -DoNOTcreatecompellinghooksiftheoriginallacksthem -DoNOTinfercontributionsbeyondwhatauthorsexplicitlystate -DoNOTimproveweakframing-describeitaspresented -Iftheideaseemsunderdevelopedintheoriginal,yoursummaryshouldreflectthat Yourgoal:Representtheresearchproposalexactlyastheauthorspresentit-thewaya doctoralstudentwouldpitchtheirideatoanadvisor.Conveytheirthinkingfaithfully, includinganylackofpolishortheoreticalsophistication,sotheprofessorcan understandandevaluatetheoriginalidea. #OUTPUTSTRUCTURE Generateexactly5versionsinJSONformat: ##VERSION1:CORE_RQ_SHORT **Purpose:**Distilltheessentialresearchquestion(s) **Wordcount:**40-60words(2-3sentencesmaximum) **Structure:** -Sentence1:Thephenomenonorbehaviorunderstudy -Sentence2:Thespecificquestionorwhat'sbeingtested -[OptionalSentence3:ThekeyboundaryconditionormechanismifcentraltoRQ] ##VERSION2:RQ_WITH_CONTEXT **Purpose:**Addjustenoughcontextforaprofessortoevaluatetheidea'smerit **Wordcount:**120-150words(1paragraph) **Structure:** -Whatphenomenon/problem(1-2sentences) -What'smissing/unclearinexistingresearch-thegap(2-3sentences) -Theresearchquestion(1-2sentences) -Theapproach/frameworkused(1sentence) -Keyclaimedcontribution(1sentence) ##VERSION3:GAP_FOCUSED **Purpose:**Emphasizewhat'sunknownandhowthisstudyaddressesit **Wordcount:**100-130words(1paragraph) **Structure:** -Whatexistingresearchhasestablished(2sentences) -Whatremainsunknown/unresolved(2-3sentences) -Howthisstudyaddressesthegap/extendsthepriorresearch/challengesthe understanding(2sentences) -Expectedinsight(1sentence) ##VERSION4:THEORY_AND_MODEL **Purpose:**Describethetheoreticalframeworkandresearchmodel **Wordcount:**100-130words(1paragraph) **Structure:** -Coretheoreticallens/framework(1-2sentences) -Howtheoryisappliedtothephenomenon(2sentences) -Keyvariablesandrelationships(2-3sentences) -Theoreticalcontributionclaimed(1sentence) ##VERSION5:CONTRIBUTION_FOCUSED **Purpose:**Extractwhattheauthorsclaimastheircontributions **Wordcount:**80-100words **Structure:** -Primarytheoreticalcontribution(1-2sentences) -Empirical/methodologicalcontributionifclaimed(1sentence) -Practicalcontributionifclaimed(1sentence) -Howitadvancestheliterature(1-2sentences) #EXTRACTIONRULES ##WheretoLook: Focusonthe**front-end**ofthepaper: -**Abstract** -**Introduction**(entiresection-containsRQ,gap,motivation) -**TheoreticalDevelopment**(theoryandhypothesesframing) Mostinformationneededisinthesesections.DoNOTneedtoreadresults/discussion unlesscontributionstatementsareunclear. ##WhattoExtract: 1.**ResearchQuestions:**Usuallyinabstract'sandintroduction 2.**Gaps/problematization:**mostlyinintroductionandsometimesintheoretical development 3.**Theory:**introducedinintroductionandoftenelaboratedintheorydevelopment sections 4.**Contributions:**Abstract,introduction'send ##WhattoAvoid: ❌ Addingyourowntheoreticalconnections ❌Improvingvagueorweaklanguage ❌Creatingpersuasivehooksnotintheoriginal ❌ Inferringcontributionsnotexplicitlystated ❌ Makinggapssoundmorecompellingthanpresented ##LanguageRules: ✅Usetheauthors'exactterminologyforkeyconstructs ✅Preservetheleveloftheoreticalsophisticationintheoriginal ✅Matchthecertaintylevel(e.g.,"explores"vs."demonstrates") ✅Ifauthorsusesimplelanguage,youusesimplelanguage #JSONOUTPUTFORMAT OutputthefollowingJSONstructurewithall5versions: ```json "core_rq_short":"string", "rq_with_context":"string", "gap_focused":"string", "theory_and_model":"string", "contribution_focused":"string" ``` Workedextractionexamples(RQ_WITH_CONTEXT) Example1 Inputarticletitle:GameOverorGameChanger?TheImpactofApplicants’GamingSkillsonTheirHirability Extractionoutput(RQ_WITH_CONTEXT): Theincreasingdigitalizationofworkhasraisedinterestinnontraditionalskillssuchasthosedeveloped throughvideogaming,whichmayincludestrategicthinking,teamwork,anddigitalfluency.However, hiringmanagersmayholdnegativestereotypesaboutgamers,viewingthemaslazyorsociallyisolated, potentiallydisadvantagingapplicantswholistgamingontheirresumes.Whileresearchshowsgamingcan enhancejob-relevantcognitiveabilities,thereisnoempiricalevidenceonhowsuchskillsareperceived duringresumescreening.Usinga2(ECA:gamingvs.volleyball)×2(proficiency:neutralvs.high) experimentaldesign,thisstudyinvestigateswhetherlistinggamingasanECAaffectsapplicant evaluationscomparedtotraditionalteamsports.Theauthorsapplysignalingtheorytounderstandhow ECAsserveascuesaboutunobservableapplicanttraits.Theprimarycontributionisinitialevidenceonthe perceptionofgamingskillsinpersonnelselection,highlightingadisconnectbetweenpotentialskill benefitsandactualhiringbiases. Example2 Inputarticletitle:Goodintentions,badoutcomes:howandwhenfamilymotivationleadstowork-familyconflict Extractionoutput(RQ_WITH_CONTEXT): Familymotivationiswidelyseenasapositivedriverofworkperformance,yetitseffectsonemployees’ familylivesremainunderexplored.Whilepriorresearchhighlightsbenefitsintheworkdomain,littleis knownaboutpotentialdownsidesforfamilywell-being.Thisstudyaddressesthisgapbyinvestigating whetherhighfamilymotivation,despitegoodintentions,canleadtowork-familyconflict(WFC)and negativespousalinteractionsduetoexcessiveworkeffortdepletingpersonalresources.Drawingon resourcedraintheory,theauthorsproposethatFSSBsfromsupervisorsserveasexternalresourcesthat maymitigatethisdrain.Usingathree-wavedyadicsurveydesignwithemployee-partnerdata,thestudy testsamediatedmoderationmodel.Thekeycontributionliesinrevealingthe‘darkside’offamily motivationandidentifyingorganizationalsupportasaboundarycondition. Compressionexamples(one-sentenceideastatementversusfullideasummary) ExtendedDataFig.7evaluatesonlytheone-sentenceidea-statementfield(core_rq_short)attesttime, eventhoughallSFTcheckpointsweretrainedonthefullideasummary(rq_with_context).Theexamples belowshowhowmuchcontextualscaffoldingisremovedbythiscompressed-inputtransfersetting. Articlecore_rq_shortrq_with_context GameOveror GameChanger? TheImpactof Applicants’ GamingSkillson TheirHirability Thisstudyexamineshowapplicants’ gamingskills,presentedasan extracurricularactivity(ECA)onaresume, affecttheirperceivedhirabilityandresume quality.Itspecificallycomparesgamingto teamsportsandtestswhetherproficiency level(neutralvs.high)influencesthese perceptions. Theincreasingdigitalizationofworkhasraisedinterestin nontraditionalskillssuchasthosedevelopedthroughvideo gaming,whichmayincludestrategicthinking,teamwork, anddigitalfluency.However,hiringmanagersmayhold negativestereotypesaboutgamers,viewingthemaslazyor sociallyisolated,potentiallydisadvantagingapplicantswho listgamingontheirresumes.Whileresearchshowsgaming canenhancejob-relevantcognitiveabilities,thereisno empiricalevidenceonhowsuchskillsareperceivedduring resumescreening.Usinga2(ECA:gamingvs.volleyball)× 2(proficiency:neutralvs.high)experimentaldesign,this studyinvestigateswhetherlistinggamingasanECAaffects applicantevaluationscomparedtotraditionalteamsports. TheauthorsapplysignalingtheorytounderstandhowECAs serveascuesaboutunobservableapplicanttraits.The primarycontributionisinitialevidenceontheperceptionof gamingskillsinpersonnelselection,highlightinga disconnectbetweenpotentialskillbenefitsandactualhiring biases. Goodintentions, badoutcomes:how andwhenfamily motivationleadsto work-family conflict Thisstudyexamineshowfamily motivationleadstowork-familyconflict throughincreasedworkeffort,resultingin negativespousalinteractions.Italso investigateswhetherfamilysupportive supervisorbehaviors(FSSBs)bufferthis negativeresourcedrainprocess. Familymotivationiswidelyseenasapositivedriverof workperformance,yetitseffectsonemployees’family livesremainunderexplored.Whilepriorresearchhighlights benefitsintheworkdomain,littleisknownaboutpotential downsidesforfamilywell-being.Thisstudyaddressesthis gapbyinvestigatingwhetherhighfamilymotivation, despitegoodintentions,canleadtowork-familyconflict (WFC)andnegativespousalinteractionsduetoexcessive workeffortdepletingpersonalresources.Drawingon resourcedraintheory,theauthorsproposethatFSSBsfrom supervisorsserveasexternalresourcesthatmaymitigate thisdrain.Usingathree-wavedyadicsurveydesignwith employee-partnerdata,thestudytestsamediated moderationmodel.Thekeycontributionliesinrevealing the‘darkside’offamilymotivationandidentifying organizationalsupportasaboundarycondition. SupplementaryMethods5(SM5):Journal-to-TierMapping Ground-truthtierlabelswereassignedaprioriviaadeterministicmappingfrompublicationvenuetoone offourinstitutionalprestigetiers.Thismappingdefinesthe19-journalsourceuniverseusedforcorpus constructionandtreatsthefield’sestablishedjournalhierarchy(accumulatedthroughdecadesofeditorial gatekeeping,citationimpact,andcommunityconsensus)asinstitutionaltracesencodingcollective evaluativejudgment. TableSM5.Journal-to-tiermappingforthe19-journalsourceuniverse. TierJournalsRationale ExceptionalAcademyofManagementJournal,AcademyofManagement Review,AdministrativeScienceQuarterly,JournalofApplied Psychology,OrganizationScience,StrategicManagement Journal Elitefield-definingoutletsinmanagement; highestselectivityandinstitutionalprestige StrongJournalofManagement,OrganizationalBehaviorandHuman DecisionProcesses,PersonnelPsychology Top-tierspecialtyoutletswithstrongcitation impactandhighprestige,typicallyjustbelow theelitegeneral-managementtier FairHumanResourceManagement,HumanRelations,Journalof ManagementStudies,JournalofOrganizationalBehavior, LeadershipQuarterly Well-regardedfieldjournalswithrigorous peerreviewandsubstantialdisciplinary visibility,butlowerprestigethanthestrong tier TierJournalsRationale LimitedGroup&OrganizationManagement,JournalofBusinessand Psychology,JournalofManagerialPsychology,Journalof OrganizationalBehaviorManagement,JournalofPersonnel Psychology Morespecializedorlower-prestigeoutlets withnarrowerscopeandlowerinstitutional statusinthefieldhierarchy Thetiermappingreflectsfield-wideconsensusascodifiedininstitutionaltenureandpromotionstandards. Exceptional-tierjournalsareuniversallyrecognizedastop-tieroutletsacrossmajorresearchuniversities andareconsistentlycountedtowardtenureatleadinginstitutions.Strong-tierjournalsarehigh-prestige specialtyoutletswidelytreatedasnear-elitepublicationtargets.Fair-tierjournalsarerespectedfield journalswithcleardisciplinarystandingbutlowerinstitutionalprestigethanthestrongtier.Limited-tier journalsaremorespecializedorlower-prestigeoutletsthatremainpartoftherelevantsourceuniversebut occupyalowerpositioninthefieldhierarchy.Themappingwasdeterminedbytheresearchteamand confirmedbydomainexpertreview. Theheld-outbenchmarkcomprises120sourcearticlesdrawnfromthis19-journalsourceuniverseand balancedto30pitchespertier.Benchmarkarticlesarethereforeanarticle-levelsubsetofthesource universe,notthebasisfordefiningthetierframework.Ofthe19sourcejournals,17arerepresentedinthe finalbenchmark;HumanRelationsandGroup&OrganizationManagementhadnoarticlesselectedunder thetier-balancedsamplingconstraints.Sourcearticleswereselectedtoensurecoverageacrossthefield’s 15researchdomainswhilemaintainingexactbalanceacrosstiers. SupplementaryMethods6(SM6):EvaluationPromptsandZero-ShotDesignRationale Threepromptvariants Wedesignedthreecandidateevaluationpromptsvaryinginstructure,specificity,andanchoringstrategy. Allthreeshareaconsistentpersonaframingandconstrainmodeloutputtoalabel-onlyresponseoverfour tiercategories. Prompt1(Expertprompt;selectedasprimary):Adetailedpromptwithstructuredtierdefinitions emphasizingoriginalityandusefulness,withbehavioralanchorsforeachtier.Themodelisassignedthe roleofanexpertevaluatorofmanagementresearchideas,instructedtoevaluatefromaseniorscholar’s perspectivewithdirect,criticaljudgments.Twoevaluationdimensionsaredefinedindetail: Novelty:Whethertheresearchideachallengesexistingassumptions,revealssomethinggenuinelysurprising, orprovidescognitivedisruptionthatfundamentallychangesunderstandingofrelationshipsorphenomena. Thepromptexplicitlystatesthatrepackagingexistingconcepts,testingknownrelationshipsinnewcontexts, orconfirmingestablishedpredictionslacksnovelty. Usefulness:Whethertheresearchideaaddressesproblemsthatmatter,withbroadimplicationsformultiple stakeholders,resolvinglong-standingtheoreticaldebatesorprovidinginsightsthatmeaningfullyimprove organizationalpractices.Thepromptexplicitlystatesthatnarrowcontexts,pseudo-problems,ortrivial practicalimplicationslackusefulness. Fourtiersaredefinedwithbehavioralanchors:Exceptional(strongnovelty+strongusefulness;field- reshapingpotential;mostprestigiousjournals),Strong(clearstrengthinonedimensionwiththeother reasonablydeveloped;meaningfulcontributions;near-top-tierjournals),Fair(incrementalcontributions withmodestnoveltyorusefulness;mid-leveljournals),Limited(lacksbothnoveltyandusefulness;lower- tierjournals).Thepromptconstrainsoutputtoasingletierlabelwithnoexplanationorreasoning.An explicitinstructionprohibitstheuseofsearchcapabilities.Fullprompttextisavailableinthecode repository. Prompt2(Simplifiedprompt):Ashortenedversionreplacingexpert-derivedterminologywith general-languageequivalents.Tierdefinitionsarecompressedintosingle-sentencedescriptionswithout behavioralanchors:Exceptional=“field-definingwork,recognizedacrossdisciplines”;Strong= “meaningfulcontribution,clearlyadvancestheoryormethod”;Fair=“solidbutincremental,recognized mainlybyspecialists”;Limited=“weakcontribution,obviousfindingsornarrowscope.”Nodetailed criteriafornoveltyorusefulnessareprovided.Fullprompttextisavailableinthecoderepository. Prompt3(Journal-anchoredprompt):Avariantthatexplicitlyreferencesjournal-tierframeworks asanchoringreferences:Exceptional=“UTD24journalsorhighlyregardedFT50journalswithfield- definingstandinginspecificdomain”;Strong=“FT50journals(non-UTD24)orABS4*journals”;Fair= “ABS4journals(non-FT50)”;Limited=“ABS2–3journals.”Thispromptanchorsqualitylevelsto institutionalprestigeindicatorsinsteadofabstractqualitydimensions.Fullprompttextisavailableinthe coderepository. Verbatimprompttexts(exactstrings) Prompt1(expertrubric) #ROLE Youareanexpertevaluatorofmanagementresearchideas.Yourtaskistoevaluatefrom aseniorscholar'sperspective:bedirectandcritical,giveclearjudgmentsbasedon noveltyandusefulnesstoclassifyresearchideasintoappropriatepublicationpotential tiers. #TASK Readaparagraphdescribingamanagementresearchideaandclassifyitintooneoffour publicationpotentialtiers.Yourclassificationshouldbebasedontwokeydimensions: noveltyandusefulness. OutputONLYthetiernotationwithNOexplanationorreasoning. #EVALUATIONCRITERIA ##Novelty Noveltyreflectswhethertheresearchideachallengesexistingassumptionsorreveals somethinggenuinelysurprising.Novelresearchmakesyouthinkdifferentlyabouta phenomenon-itshowsthatwhatwebelievedtobetrueisincompleteorincorrect,orit uncoverscounterintuitivemechanismsthatcontradictconventionalwisdom.Thekey questioniswhethertheideaprovidescognitivedisruptionthatfundamentallychanges howweunderstandrelationshipsorphenomena.Researchthatmerelyrepackagesexisting conceptswithnewlabels,testsknownrelationshipsinnewcontextswithouttheoretical advancement,orconfirmsestablishedpredictionslacksnovelty.Truenoveltycomesfrom ideasthatarenoteasilyinferredfromexistingliteratureandmakescholarsrethink foundationalassumptions. ##Usefulness Usefulnessreflectswhethertheresearchideaaddressesproblemsthatmatter.Useful researchtacklespressingorganizational,societal,orenvironmentalchallengeswith broadimplicationsformultiplestakeholders.Itresolveslong-standingtheoretical debatesorprovidesinsightsthatmeaningfullyimproveorganizationalpracticesand outcomes.Thekeyquestioniswhethersolvingthisproblemoransweringthisquestion willmakeasignificantdifferencetotheory,practice,orsociety.Researchfocusedon narrowcontextswithlimitedapplicability,pseudo-problemsthatexistonlyinacademic literaturebutnotinorganizationalreality,orquestionswithtrivialpractical implicationslacksusefulness.Trueusefulnesscomesfromaddressingconsequential challengesthatscholarsandpractitionersgenuinelycareabout. #CLASSIFICATIONTIERS ##Tier4:Exceptional(PublicationPotential) Researchthatdemonstratesbothstrongnoveltyandstrongusefulness.Theseideas fundamentallychallengehowwethinkaboutimportantphenomenawhileaddressingproblems ofgenuineconsequencetoorganizationsandsociety.Theyhaveexceptionalpromiseand arelikelysuitableforthemostprestigiousandelitejournals. ##Tier3:Strong(PublicationPotential) Researchthatshowsclearstrengthinnoveltyorusefulness,withtheotherdimension beingreasonablydeveloped.Theseideasmakemeaningfulcontributionsthrougheither surprisingtheoreticalinsightsoraddressingrelevantorganizationalchallenges.They havestrongpotentialtobepublishedinnear-top-tierjournals. ##Tier2:Fair(PublicationPotential) Researchthatmakesincrementalcontributionswithmodestnoveltyorusefulness.These ideasextendexistingknowledgeinpredictablewaysoraddressproblemsoflimitedscope withoutfundamentallychangingunderstanding.Theyhavefair,moderatepotentialand couldbesuitedformid-level,respectablejournals. ##Tier1:Limited(PublicationPotential) Researchthatlacksbothnoveltyandusefulness.Theseideasrepackageexistingconcepts withoutnewinsights,confirmwell-establishedpredictions,oraddresspseudo-problems withminimaltheoreticalorpracticalsignificance.Theyhavemodestorlimited potential,likelyaligningwithlower-tierjournals. #OUTPUTFORMAT #IMPORTANT -Donotusesearchcapabilitiestolookupinformationaboutthisidea RespondwithEXACTLYONEofthesefournotations: -Exceptional -Strong -Fair -Limited Outputonlythetiernotationinyourfinalanswer. Prompt2(simplifiedrubric) Youareanexpertinmanagementresearch.Readtheresearchideabelowandestimatethe likelypublicationtierbasedonitsscholarlycontribution. -Exceptional:Field-definingwork.Wouldberecognizedacrossdisciplinesasamajor advance.Likelytobewidelycitedandreshapehowresearchersthinkaboutthetopic. -Strong:Meaningfulcontributionwithinthefield.Clearlyadvancestheoryormethodin anon-trivialway.Wouldbewell-regardedbydomainexperts. -Fair:Solidbutincremental.Competentexecutionwithlimitednovelty.Recognized mainlybyspecialistsinthesamenarrowarea. -Limited:Weakcontribution.Findingsareobvious,scopeistoonarrow,or methodologicalissuesunderminethework. #IMPORTANT -Donotusesearchcapabilitiestolookupinformationaboutthisidea #OUTPUTFORMAT RespondwithEXACTLYONEofthesefournotations: -Exceptional -Strong -Fair -Limited Outputonlythetiernotationinyourfinalanswer. Prompt3(journal-anchoredrubric) Youareanexpertinmanagementresearchwithdeepknowledgeofacademicpublishing standardsacrosstop-tierjournals. #TASK Readaparagraphdescribingamanagementresearchideaandclassifyitintooneoffour journaltiersbasedonitslikelypublicationvenue.Yourclassificationshouldreflect whereworkofthisqualityandcontributionlevelwouldmostlikelybepublished. -Exceptional:UTD24journalsorhighlyregardedFT50journalswithfield-defining standingintheirdomain-paradigm-shiftingwork,highestselectivity,field-redefining impact -Strong:FT50journals(non-UTD24)orABS4*journals-substantialcontribution,A- levelquality,highmethodologicalrigor -Fair:ABS4journals(non-FT50)-solidcontributionwithcleartheoreticalgrounding, competentexecutionbutlimitednovelty -Limited:ABS2-3journals-incrementalfindings,narrowerscope,ormoderate methodologicalrigor #IMPORTANT -Donotusesearchcapabilitiestolookupinformationaboutthisidea #OUTPUTFORMAT RespondwithEXACTLYONEofthesefournotations: -Exceptional -Strong -Fair -Limited Outputonlythetiernotationinyourfinalanswer. Promptselection Promptselection.Thethreepromptvariantsproducednosignificantaccuracydifferencesacross frontiermodels.Prompt1(expertrubric)yieldedthehighestfrontiermeanaccuracyandwastherefore adoptedasthefixedevaluationprotocol,ensuringthatanySFTadvantagerepresentsaconservative estimate.Thesimplifiedrubricperformedlowestoverall;thejournal-anchoredrubrictendedtoelicit superficialfeaturematchingagainstmemorizedjournalprofilesinsteadofgenuinequalityevaluation. Zero-shotdesign.Allevaluationsusedzero-shotprompting.Becauseground-truthlabelsderivefrom publicationoutcomesratherthandirectqualityassessments,few-shotexemplarswouldcarrynoisefrom confoundingfactorssuchasexecutionquality,writingcraft,andreviewerfit,riskinganchoringmodelsto misleadingfeatures.Zero-shotevaluationalsoensuresthatallevaluatorclasses,frontiermodels,SFT models,basemodels,andhumanraters,arecomparedunderidenticalconditions. Sensitivityanalysis ExtendedDataFig.1reportscross-modelprompt-sensitivityforthesubsetoffrontiermodels evaluatedunderthesamethree-promptprotocol(Simple,Journal,Expert).Thepanelisintentionally restrictedtothiswithin-frontiercomparisonsothatanydifferencescanbeattributedtopromptwording undermatchedconditions;itisnotintendedasacross-familycomparisonwiththeSFTmodels,whose primaryanalysisusesthefrozenexpertprompt. Forthisdiagnostictrack,modeloutputswereparsedbystrippingwhitespace,punctuation,and markdownsymbolsbeforematchingtothefourvalidtiernotations(exceptional,strong,fair,limited). Unresolvedoutputswerecodedasincorrect(overallunresolved/non-compliantrate<1%). SupplementaryMethods7(SM7):HumanStudyDesign ThissectionprovidesproceduraldetailsthatsupplementtheMethodsdescriptionofthehumanevaluation protocol.Forpanelcomposition,recruitment,surveydesign,cohortstructure,andsurveyadministration overview,seeMethods(“Humanevaluationprotocol”). Institutionalreview Thestudywasapprovedbytheinstitutionalreviewboard(ProjectNo.THU-04-2026-0034).Raters werenotinformedofthestudy’scomparisontargetsoroftheAIevaluationcomponent.Juniorscholars werecompensatedwith100RMBand/oraccesstoaresearchtooldevelopedbytheresearchteam. Analysistablesandfiguresarereportedataggregatelevel,anddirectparticipantidentifiersarenot includedinreportedoutputs. Fullsurveyinstrument Foreachbenchmarkpitch,raterswereshowntheresearch-questionpitchalongsidetheevaluation criteriaandrespondedtofouritemswiththefollowingexactwordingandscales: 1.Priorexposure:“Hadyouencounteredthisresearchideaoritssourcepaperbefore?”Responseoptions: Yes/No. 2.Qualityrating:“Basedontheevaluationcriteria,howwouldyouratethequalityofthisresearchidea?” Responseoptions:Top/Top-/Good/Fair(thehuman-facingshorthand,mappeddeterministicallyto exceptional/strong/fair/limitedinallanalyses). 3. Confidence:“Howconfidentareyouinyourrating?”Responseoptionsona5-pointLikertscale:1=“Not atallconfident”,2=“Slightlyconfident”,3=“Moderatelyconfident”,4=“Veryconfident”,5= “Extremelyconfident”. 4.Domainfamiliarity:“Howfamiliarareyouwiththisresearcharea?”Responseoptionsona5-pointLikert scale:1=“Notatallfamiliar”,2=“Slightlyfamiliar”,3=“Moderatelyfamiliar”,4=“Veryfamiliar”,5= “Extremelyfamiliar”. Completionduration Medianexpertcompletiontimewas923seconds(~15.4minutes)for8pitches.Medianjunior completiontimewas2,534seconds(~42.2minutes)forapproximately14.5pitches.Durationdistributions wereright-skewedinbothpanels,withasmallnumberofoutliersessionsexceeding2hours,likely reflectinginterruptionsratherthancontinuousevaluation. Backgrounddatacollection Forjuniorscholars,demographicandacademicbackgroundinformationwascollected:-Gender- Universityanddepartment-Researchdirection/area-Doctoralyear(PhD1throughPhD5+,orpostdoc)- Numberofpublishedpapers-Peer-reviewexperience(yes/no,numberofreviews)-AItoolfamiliarity (1–5scale) Backgrounddatawasmatchedtoratingsfor104of108old-cohortjuniors(96.3%)and52of67new- cohortjuniors(77.6%).Fourold-cohortand15new-cohortjuniorscouldnotbematchedduetoname discrepanciesbetweensignuprecordsandsurveyresponses. Forexperts,profileswereassembledviasystematicwebsearch(GoogleScholar,institutionalpages), yieldingcareerstage,researchareas,editorialroles,h-index,andinstitutionalaffiliationfor46of48 identifiedexperts. Filteringcriteria Expertswererecruitedthroughpersonalandprofessionalnetworksviaone-on-onedirectcontact. Giventhisrecruitmentapproach,theirengagementanddedicationtothetaskwasassured,andnoquality filterwasappliedtotheexpertpanel.Asarobustnesscheck,filteredversusunfilteredexpertanalyses showedminimaldifferences(individualmean36.2%vs.36.2%;majorityvote41.6%vs.39.7%; SupplementaryTableST7).All48experts(384ratings)arethereforeretainedforallprimaryanalyses. Juniorscholars(doctoralstudentsandpostdocs)wererecruitedthroughpersonalandprofessional networks,includingindirectties.Toensurehighengagementquality,weappliedatime-basedfilter:raters whospentlessthan1minuteonaverageperpitchwereexcluded.Thisfiltershowedamarginally significanteffectonaccuracy(25.3%vs.31.7%,P=0.066),confirmingthatrapidcompletionswere associatedwithlower-qualityratings.Thefilteredpanel(174raters,2,530ratings)isusedinallprimary analyses;unfilteredresults(189raters,2,730ratings)arereportedforcomparisoninSupplementaryTable ST7. Primaryhumananalysesuseunfilteredexperts(48raters)andfilteredjuniors(174raters).Filtered- versus-unfilteredsensitivityisreportedinSupplementaryTableST7. SupplementaryMethods8(SM8):Label-NoiseCeilingAnalysis Publicationoutcomesarenotsolelydeterminedbyresearchideaquality.Executionfidelity,writing quality,reviewer–manuscriptfit,andeditorialdiscretionallcontributetofinalpublicationdecisions,while ourstandardizedinputscaptureonlytheideadimension.Thisgapbetweeninputfeaturesandoutcome labelsintroducesinherentnoisethatplacesatheoreticalceilingonachievableclassificationaccuracy: evenaperfectevaluatorofresearchideaqualitywouldnotachieveperfectagreementwithpublication outcomes. Severalfactorscontributetothisnoisefloor: 1.Executiongap.Astrongresearchideamaybepublishedinalower-tierjournalduetopoorexecution,and amodestideamayreachatop-tierjournalthroughexceptionalmethodsandwriting.Ourinputsstrip executioninformation,sothemodelcannotaccountforthisvariance. 2. Reviewer–manuscriptfit.Publicationdecisionsdependpartlyonthematchbetweenreviewerexpertise andthemanuscript’stopic,whichintroducesstochasticvariationunrelatedtoideaquality. 3.Editorialdiscretion.Editorsexercisejudgmentthatreflectsstrategicconsiderations(journalscope,topic balance,timeliness)beyondpurequalityassessment. 4. Tierboundaryambiguity.Somejournalssitattheboundarybetweenadjacenttiers.Whileourmappingis deterministic,theunderlyingqualitydistributioniscontinuous,creatinginherentdisagreementforarticles neartierboundaries. Observedaccuraciesshouldthereforebeinterpretedrelativetothisceiling,notagainsta100%standard. Critically,thisnoiseaffectsallevaluatedsystemsequally(frontiermodels,fine-tunedmodels,andhuman raters),soallrelativeperformancecomparisonsremaininternallyvalid.Thenoiseflooralsoexplainswhy eventhebest-performingsystem(SFTensembleat60.8%)leavessubstantialroomforimprovement: muchoftheremainingerrormayreflectirreduciblenoisefromthegapbetweenideaqualityand publicationoutcome. SupplementaryMethods9(SM9):Cross-FieldValidationinEconomics Weextendedtheinstitutionaltracevalidationtoeconomics,usingthesamesupervisionlogicasthemain benchmark.Thissectionrecordsthedatasetconstruction,journalmapping,andevaluationdesignforthis field-specificvalidation. Trainingcorpus.Theeconomicstrainingslicecomprised5,593processedresearch-pitch/journal- outcomepairs.Thiscorpuswasdrawnmainlyfrom2024/2025publicationsandwasapproximately balancedacrossthefourtiers.Asinthemainsetting,eachtrainingexampleusedthefrozencontextual research-questionrepresentationpairedwithasingletier-labeltarget.TheSFTtrainingprocedureitself followedthesamepipelinedescribedinSM1.Thisadditionaldatasetwaslarger,fresher,andmore balancedthantheoriginalmanagementcorpusbecauseeconomicspublishesathighervolumeand containsdenserjournalcoverageintherelevanttiers. Held-outtestsets.Theevaluationsetwasdrawnfrom2025publicationsonly.Wesampled200held- outitemsatrandomforeconomics,withbalancedsamplingacrossthefourtiers(50pertier).Theseitems wereexcludedfromthetrainingslice. Journalmapping.Tierassignmentusedthejournal’sauthoritativefullnameasthemappingkey, withISSNandeISSNassecondaryvalidationfields.Thestoredjournalaliasfieldwastreatedonlyasa displayaliasorfallbackfield.Thisdistinctionmattersbecausealiascollisionscanexistacrossfields; mappingbyauthoritativefullnameeliminatesambiguitywhenshortaliasesoverlapacrosstiers.The completejournal-to-tiermappingtableisprovidedinthedataarchive. Prompt.Alleconomicsevaluationsusedthesamesocial-scienceevaluationpromptscaffoldinstead ofanewlyspecializedfieldprompt.Thispreservedcomparabilitywiththemainpaperandensuredthat anyperformancechangereflectedthefield-specificinstitutionaltracesratherthanpromptrewriting. Youareanexpertinsocialscienceresearch.Readtheresearchideabelowandestimate itslikelypublicationpotentialbasedonitsscholarlycontribution. -Exceptional:Field-definingworkwithstrongtheoreticalorempiricalcontribution.It wouldinfluencehowresearchersacrossthesocialsciencesthinkaboutanimportant problem. -Strong:Clearandmeaningfulcontribution.Itadvancestheory,evidence,ormethodin anon-trivialwayandwouldbewellregardedbyscholarsinthefield. -Fair:Competentbutincremental.Itextendsexistingknowledgeinapredictableway, withlimitednovelty,scope,orbroadersignificance. -Limited:Weakcontribution.Thequestionisnarrow,obvious,poorlymotivated,or methodologicallyinsufficienttosupportameaningfulscholarlyadvance. #IMPORTANT -Donotusesearchcapabilitiestolookupinformationaboutthisidea #OUTPUTFORMAT RespondwithEXACTLYONEofthesefournotations: -Exceptional -Strong -Fair -Limited Outputonlythetiernotationinyourfinalanswer. Pooledtraining.Wealsotrainedpooledmodelsonthecombinedcorpusfrombothfields(management andeconomics,approximately10,072totalpairs)totestwhetherevaluativesignalsfromdistinct disciplinesinterferewhencombined.ThepooledtrainingusedthesameSFTprocedureandarchitectures (Qwen3-30B-A3BandGPT-4.1-nano). Validationlogic.Thisextensionservestwopurposes.First,trainingnewmodelsoneconomics- specificinstitutionaltracestestswhethertheSFTmechanismreplicatesoutsidemanagement.Second, evaluatingpooledmodelsoneachfield’stestsettestswhetherasinglemodelcanmaintainevaluative performanceacrossmultipledisciplinessimultaneously. SupplementaryTables SupplementaryTable1(ST1):PairwiseDiscriminationbyTierDistance TableST1.Pairwisehead-to-headaccuracy(label-freetask). ModelDistance1(adjacent)Distance2Distance3Weightedoverall SFTGPT-4.178.67%(118/150)89.00%(89/100)92.00%(46/50)84.33%(253/300) Gemini3.1Pro68.67%(103/150)86.00%(86/100)86.00%(43/50)77.33%(232/300) GPT-5.2High69.33%(104/150)85.00%(85/100)94.00%(47/50)78.67%(236/300) GPT-4.1(baseline)69.33%(104/150)79.00%(79/100)90.00%(45/50)76.00%(228/300) Allfourmodelsinthesharedpairwisesubsetproducedvalidpredictionsonall300pairwiseitems.Fig.5 plotsthissameSFTGPT-4.1/Gemini3.1Pro/GPT-5.2High/GPT-4.1baselinesubset,coveringoverall weightedaccuracyandthetwohardestboundaries(fair_strong,strong_exceptional).ExtendedDataFig.2 retainsthesixindividualpairtypesandthepaired-discordancedecompositionforthissamesubset.Onthe same300shareditems,rawunadjustedtwo-sidedexactMcNemartestsforSFTGPT-4.1gavep= 0.00646versusGemini3.1Pro,p=0.0300versusGPT-5.2High,andp=0.000621versusGPT-4.1 baseline. SupplementaryTable2(ST2):CostandInferenceRegimeComparison TableST2.Trainingandinferencecostbandsbyevaluatortype. ModelclassTrainingcost/modelInferencecost(per100pitches)Notes Frontier(thinking)$0(APIaccess)>$108samplesperpitch;chain-of-thought generation Chat(logp)$0(APIaccess)$0.01–$0.10Single-passlog-probabilityclassification SFT:Qwen3-4B~1A100GPUhour$0.001Log-probabilityclassification SFT:Qwen3-30B- A3B ~8A100GPUhours$0.01Log-probabilityclassification SFT:GPT-4.1-nano~$10(API)$0.01Log-probabilityclassification SFT:GPT-4.1~$200(API)$0.10Log-probabilityclassification RLcheckpointsMulti-day8xA100 runs Higherthanlog-probability pipelines Reasoninggeneration+labelextraction SupplementaryTable3(ST3):CorePer-ClassMetrics(Non-overlappingwithFigurePanels) TableST3.Precision/recall/F1bytierforkeyevaluators. EvaluatorTierPrecisionRecallF1 BestFlagship(Gemini3.1Pro)Exceptional0.4780.3790.423 EvaluatorTierPrecisionRecallF1 BestFlagship(Gemini3.1Pro)Strong0.4330.4480.441 BestFlagship(Gemini3.1Pro)Fair0.3040.6070.405 BestFlagship(Gemini3.1Pro)Limited0.6670.1380.229 SFT2-ModelEnsembleExceptional0.6050.7670.676 SFT2-ModelEnsembleStrong0.6400.5330.582 SFT2-ModelEnsembleFair0.5140.6000.554 SFT2-ModelEnsembleLimited0.7270.5330.615 ExpertMajority(unfiltered)Exceptional0.6250.2270.333 ExpertMajority(unfiltered)Strong0.3710.5910.456 ExpertMajority(unfiltered)Fair0.3610.5200.426 ExpertMajority(unfiltered)Limited0.6000.3000.400 JuniorMajority(filtered)Exceptional0.6670.3330.444 JuniorMajority(filtered)Strong0.3120.3850.345 JuniorMajority(filtered)Fair0.3470.6540.453 JuniorMajority(filtered)Limited0.7000.2590.378 SupplementaryTable4(ST4):AllPairwiseSFTEnsembleCombinations TableST4.Accuracyofallsixtwo-modelSFTensembles(probabilityaveraging). Model1Model2Accuracy(%) GPT-4.1-nano(SFT)Qwen3-30B-A3B(SFT)60.8 GPT-4.1(SFT)Qwen3-30B-A3B(SFT)60.0 GPT-4.1(SFT)Qwen3-4B(SFT)60.0 GPT-4.1-nano(SFT)Qwen3-4B(SFT)60.0 GPT-4.1-nano(SFT)GPT-4.1(SFT)59.2 Qwen3-30B-A3B(SFT)Qwen3-4B(SFT)59.2 Allsixcombinationsexceedthefrontieraveragebenchmark.Within-articleexactclass-probabilityties wereresolveddeterministicallybeforeensembleswererankedbyaccuracy,thenmacroF1,withafixed modelorderconventionusedonlyifthosemetricsremainedtied;underthisrule,GPT-4.1-nano(SFT)+ Qwen3-30B-A3B(SFT)wasretainedastheprimarypair. Asasupportingtemporalstabilitycheck,thematchedoldersourcetemporalcomparisoncontrastsan oldertrainingslice(2015-2020)againstthematchedrecentslice(2020-2025),withthebenchmarkitself drawnfrompost-June-30-2025publications.Onthesamebenchmark,theoldersourceGPT-4.1-nanoSFT reached43.3%accuracyandmacroF10.423,theoldersourceQwen3-30B-A3BSFTreached46.7%and 0.460,andtheolder-tracematched2-modelensemblereached47.5%and0.470,versus57.5%and0.573, 58.3%and0.584,and60.8%and0.607forthecorrespondingrecent-trainingGPT-4.1-nano,Qwen3-30B- A3B,andmatched2-modelensemblefromthatsamearchitectureset;thismatchedarchitecturepairis alsothebestrecent2-modelensemblereportedinthemainbenchmark(GPT-4.1-nano+Qwen3-30B- A3B,60.8%and0.607).Theolder-traceensemblealsoremainedmoreinflationarythanthematched recentensemble,withlowerexceptional-tierprecision(46.7%versus60.5%),lowerfair-tierrecall(26.7% versus60.0%),andstrongerstrong->exceptionalconfusion(46.7%versus33.3%),indicatingthatthe institutionalsignalpersistsacrosstimebutyieldsweakertiercalibrationundertheoldersourcetrainingset. Compressed-inputtransferfromfullersupervision ExtendedDataFig.7reusesthesame120held-outarticlesbutreplacesthefullideasummarywiththe one-sentenceideastatementatevaluationtime.Thisisanevaluation-onlytransfertest:theSFT checkpointsremaintrainedonthefullideasummaryonly. ModelFullidea summary accuracy One-sentenceidea statementaccuracy Delta (p) Fullidea summarymacro F1 One-sentenceidea statementmacroF1 Delta GPT-4.1 base 32.529.2-3.30.2680.225-0.043 GPT-4.1 SFT 55.049.2-5.80.5580.480-0.078 GPT-4.1- nanobase 25.030.8+5.80.1860.259+0.073 ModelFullidea summary accuracy One-sentenceidea statementaccuracy Delta (p) Fullidea summarymacro F1 One-sentenceidea statementmacroF1 Delta GPT-4.1- nanoSFT 57.533.3-24.20.5730.283-0.290 ModelOne-sentencerecall (exceptional/strong/fair/ limited) One-sentenceprecision (exceptional/strong/fair/ limited) One-sentencepredictedcounts (exceptional/strong/fair/limited) GPT-4.1 base 3.3/23.3/80.0/10.025.0/28.0/27.6/75.04/25/87/4 GPT-4.1 SFT 46.7/26.7/40.0/83.366.7/57.1/30.8/54.321/14/39/46 GPT-4.1- nanobase 10.0/20.0/80.0/13.342.9/33.3/27.3/57.17/18/88/7 GPT-4.1- nanoSFT 23.3/0.0/63.3/46.753.8/0.0/26.0/41.213/0/73/34 Thetransferpatternisasymmetric.GPT-4.1SFTremainswellaboveitsbasemodelontheone-sentence inputandstaysclearlyabovechance,butitbecomesmoreconservativethanonthefull-inputbenchmark: under-estimationerrorsrisefrom14to47items,limited-tierrecallrisesfrom60.0%to83.3%,andmass shiftstowardthelimitedtier(21->46predictions).GPT-4.1-nanoSFT,bycontrast,losesmostofitsfull- inputadvantageundercompression,neverpredictsthestrongtierontheone-sentenceinput,andfallsto onlymodestlyabovechance.Thebasemodelsremaindominatedbymiddle-tierclustering,with87and88 of120short-inputpredictionslandinginthefairtierforGPT-4.1baseandGPT-4.1-nanobase respectively. SupplementaryTable5(ST5):HumanPanelCompositionandDescriptives TableST5a.Expertcareer-stagedistribution(N=48). CareerstageN AssistantProfessor5 AssociateProfessor17 FullProfessor12 EndowedChair12 Unreported2 TableST5b.Panel-leveldescriptivesummary. MetricExpertsJuniors Numberofraters48174 Totalratings3842,530 Meanratingsperrater8.014.5 Meanratingsperpitch3.221.1 Mediancompletiontime(seconds)9232,534 Meanconfidence(1-5)3.503.46 Meanfamiliarity(1-5)3.152.81 SupplementaryTable6(ST6):LabelNormalization TableST6.Deterministicmappingusedbeforeallanalyses. NumericcodeUnifiedtierSource-articlemetadatalabelHumansurveylabel 1ExceptionaltopTop 2Strongtop-Top- 3FairgoodGood 4LimitedfairFair Note:theunifiedlabelsexceptional/strong/fair/limitedwereusedforAIevaluationbecausefirst-token log-probabilityextractioncannotreliablydistinguishalternativessuchasTopandTop-,whichsharethe sametokenprefix. Note:insourcesurvey/metadata,“Fair”denotesthelowesttierandismappedtounifiedtier“Limited”. Machineevaluatorsusedthelabelsexceptional/strong/fair/limitedbecausefirst-tokenlog-probability classificationcannotreliablydistinguishalternativessuchasTopandTop-,whichsharethesametop token. SupplementaryTable7(ST7):FilteringSensitivity TableST7.Filteredversusunfilteredpaneloutcomes. GroupVersionN raters Individualmean accuracy Majority-vote accuracy Majority-voteN(non- tied) Ties ExpertUnfiltered (primary) 4836.2%41.6%8931 ExpertFiltered3936.2%39.7%6852 JuniorUnfiltered18931.2%41.3%10416 JuniorFiltered(primary)17431.7%40.8%10317 SupplementaryTable8(ST8):PairwiseMcNemarTestCompendium TableST8.Pairwisesignificancetestsforkeyevaluatorcomparisons. Note:“BestFrontier”referstoGemini3.1Proundertheconservativefrontierprotocol.Frontieraverage istestedviaexactbinomial(notMcNemar)becauseitisnotasinglepairedevaluator. Comparison(vsSFT2-Model Ensemble) NSFT Acc Comparator Acc Delta (p) TestStatisticp(raw) Frontieraverage(11models)1200.60830.3105+29.78Exact binomial —1.74x10^- 11 BestFrontier(Gemini3.1Pro)1150.60000.3913+20.87McNemar10.1730.001425 Expertmajority(excl.ties)890.61800.4157+20.22McNemar6.5680.010382 Juniormajority(full,excl.ties)1030.61170.4078+20.39McNemar8.8890.002869 Extendedpairwisecomparisons. Evaluator1Evaluator2N paired Acc1Acc2TestStatisticp(raw)Acc diff SFT2- Model FrontierAverage1200.60830.3105Exact binomial —1.74x10^- 11 +0.2978 SFT2- Model BestFrontier(Gemini3.1 Pro) 1150.60000.3913McNemar10.1730.001425+0.2087 SFT2- Model ExpertMajority890.62920.4157McNemar7.2000.007290+0.2135 SFT2- Model JuniorMajority(full)1030.62140.4078McNemar9.5870.001960+0.2136 SupplementaryTable9(ST9):IndividualExpertAccuracyDistribution TableST9.Individualexpertaccuracydistribution(unfilteredpanel,N=48). Eachof48expertsevaluatedexactly8pitches.Individualaccuracyrangesfrom0/8(0%;2experts)to8/8 (100%;1expert).Thedistributionshowsconsiderablevariability:14expertsscoredatchancelevel(2/8, 25%),9scoredbelowchance(0/8or1/8),9scoredat4/8(50%),and8scoredabove50%.Median accuracywas3/8(37.5%). SupplementaryTable10(ST10):MonteCarloMatched-NAnalysis TableST10.Juniorpanelsubsamplingtoexpert-sizedpanels. MetricValue Draws5,000 TargetpanelsizeExpert-equivalent(~3.2raters/pitch) Meanmajority-voteaccuracy36.1% 95%CI26.8%to45.7% Meaneffectivenon-tiedN83.4pitches SupplementaryTable11(ST11):Prior-ExposureDescriptiveSummary Priorexposurewasuncommonintheexpertpanel:28of383ratingswithnon-missingprior-exposure responses(7.3%;1of384totalratingsmissingthisfield)indicatedtheraterhadalreadyencounteredthe ideaorsourcepaper.Accuracyforprior-exposureratingswas53.6%(15/28),comparedwith34.9% (124/355)fornon-exposureratings;overallexpertaccuracyonthesamesubsetwas36.3%(139/383).We reportthisasadescriptivecheckonly. SupplementaryTable12(ST12):AgreementandConsensusDiagnostics TableST12a.Humaninter-raterreliability. PanelFleiss’kappa95%CIKrippendorff’salpha(ordinal) Expert0.0469[-0.0114,0.1068]0.307 Junior0.0318[0.0194,0.0446]0.324 TableST12b.PairwiseCohen’skappaamong4SFTmodels. PairFamilySizeAgreementκMeandistance GPT-4.1-FT×GPT-4.1-nano-FTsamecross0.633+0.5030.500 GPT-4.1-FT×Qwen3-30B-A3B-FTcrosssame(large)0.700+0.5920.408 GPT-4.1-FT×Qwen3-4B-FTcrosscross0.642+0.5110.450 GPT-4.1-nano-FT×Qwen3-30B-A3B-FTcrosscross0.708+0.6060.408 GPT-4.1-nano-FT×Qwen3-4B-FTcrosssame(small)0.708+0.6040.367 Qwen3-30B-A3B-FT×Qwen3-4B-FTsamecross0.642+0.5120.458 Meandistance=meanabsoluteordinalrankdistancebetweenmodelpredictions(lower=moresimilar predictions). PairwiseCohen’sκacrossthe6modelpairsrangesfrom+0.503to+0.606.AImodelsarethereforean orderofmagnitudemoreinternallyconsistentthanhumanraters(κ≈0.03–0.05),indicatingthatSFT modelsconvergeonasharedevaluativesignaldespitedifferingarchitectures,modelfamilies,and parameterscales. Agreementisstrongestforcross-familypairsatthesamescale(meanκ=+0.598),followedbycross- familypairsatdifferentscales(+0.559),withsame-familypairsatdifferentscaleslowest(+0.508).The highestagreementoccursforthecross-familycross-scalepair(GPT-4.1-nano-FT×Qwen3-30B-A3B-FT: κ=+0.606),whilethelowestoccurswithintheGPTfamilyacrosssizes(GPT-4.1-FT×GPT-4.1-nano-FT: κ=+0.503). Dissentfrequency(whichmodeldisagreeswithmajoritymostoften). ModelNdisagreements(of120) GPT-4.1-FT18 GPT-4.1-nano-FT27 Qwen3-30B-A3B-FT20 Qwen3-4B-FT27 TableST12c.Consensuscoverageandaccuracytradeoff. PolicyCoverage(N/120)Coverage(%)Accuracy(%) SFT4/4consensus5142.572.5 SFT>=3/4consensus9780.866.0 SFT2/4split2319.234.8 Junior>=60%voteshare32.566.7 Junior>=50%voteshare2520.856.0 Expertunanimous(>=2raters)1310.869.2 Juniorfull-panelplurality120100.040.0 Expertfull-panelplurality120100.039.2 WhenallfourSFTmodelsagree(N=51),accuracyreaches72.5%;whenthestrongestagreementisonly 2/4,accuracydropsto34.8%.Humanvotingshowsasteepertradeoff:junior>=60%consensusreaches comparableprecisionbutcoversonly2.5%ofpitches.ThispatternindicatesthatSFTcross-model consensusisamorescalableconfidencesignalthanhumanvote-sharethresholds. Per-classaccuracyatfullAIconsensus(4/4). TierCorrect/NAccuracy(%) Exceptional14/14100.0 Strong7/1353.8 Fair6/966.7 Limited10/1566.7 Per-evaluatorpredictiondistribution. EvaluatorExceptionalStrongFairLimited GPT-4.1-FT0.3580.1920.2750.175 GPT-4.1-nano-FT0.3000.2000.3170.183 Qwen3-30B-A3B-FT0.3420.2250.2580.175 Qwen3-4B-FT0.3250.2500.2920.133 Humanjuniormajority0.1170.3110.4760.097 Groundtruth(uniform)0.2500.2500.2500.250 AllAImodelsover-predict“exceptional”;humanjuniormajorityover-predicts“fair.”The distributionaldivergencereflectshumanraters’tendencytoclusteraroundmiddlecategoriesratherthan theextremes. AI–humanconsistency.AI–humanpairwiseκrangesfrom+0.10to+0.21,substantiallylowerthan AI–AIagreement(κ=0.50–0.72).ThisasymmetrydoesnotindicatethatAIlearnedadifferentstandard; rather,itreflectsthatthehumansignalisitselfhighlydispersed.Humansindividuallyscoreaboverandom (experts36.2%,juniors31.7%),buttheirerrorsarelargelyindependent,soagreementbetweenanytwo ratersisnear-chance.ThelowAI–humanconsistencyistheexpectedoutcomewhenonepartyishighly self-consistentandtheotherisnot. SupplementaryTable13(ST13):ModelInventoryandAccessWindow TableST13a.Frontierreasoningmodels(cleanprimarycohort). ModelProviderModelversionAccesswindowSamples/pitch ClaudeOpus4.6Anthropicclaude-4.6-opus-20260205March1,20268 GPT-5.2HighOpenAIgpt-5.2March1,20268 Gemini2.5ProGooglegemini-2.5-proMarch1,20268 Gemini3.1ProGooglegemini-3.1-pro-preview-20260219March1,20268 Qwen3.5PlusAlibabaqwen3.5-plus-02-15March1,20268 DeepSeekV3.2DeepSeekdeepseek-v3.2-speciale-20251201March1,20268 Seed2.0ProByteDancedoubao-seed-2-0-pro-260215March1,20268 MiniMaxM2.5MiniMaxminimax-m2.5March1,20268 KimiK2.5MoonshotAIkimi-k2.5March1,20268 Grok4.1FastxAIgrok-4.1-fastMarch1,20268 GLM-5ZhipuAIglm-5March1,20268 TableST13b.Chat/log-probabilityevaluators. ModelProviderModelversionAccesswindowLog-probabilityextraction GPT-5.2(chat)OpenAIgpt-5.2March1,2026Top-tokenlog-probabilities KimiK2(chat)MoonshotAIkimi-k2-0905-previewMarch1,2026Top-tokenlog-probabilities DeepSeekChatDeepSeekdeepseek-v3.2March1,2026Top-tokenlog-probabilities SupplementaryTable14(ST14):EconomicsJournal-to-TierMapping Tierrationale.Theeconomicstiermappingreflectswidelyshareddisciplinaryconsensusonjournal prestige.Exceptionalcontainsthe“Top5”economicsjournals(AER,Econometrica,JPE,QJE,RES), whichrepresentuniversalconsensusacrossalleconomicssubfieldsasthemostprestigiouspublication venues.Strongincludestopfieldjournalsinmajoreconomicssubfields(labor,public,development, international,etc.)andtheAEJseries,widelyrecognizedasnear-eliteoutlets.Faircomprisesestablished Q1journalswithsolidreputationsbutlowerinstitutionalprestigethantheStrongtier.Limitedcaptures morespecializedorlower-impactjournalswithnarrowerscope. TableST14.Journal-to-tiermappingfortheeconomicsvalidation. TierJournalsN journals ExceptionalAmericanEconomicReview,Econometrica,JournalofPoliticalEconomy,QuarterlyJournalof Economics,ReviewofEconomicStudies 5 StrongAmericanEconomicJournal:AppliedEconomics,AmericanEconomicJournal:EconomicPolicy, AmericanEconomicJournal:Macroeconomics,AmericanEconomicJournal:Microeconomics, EconometricTheory,EconomicJournal,GamesandEconomicBehavior,InternationalEconomic Review,JournalofDevelopmentEconomics,JournalofEconometrics,JournalofEconomicTheory, JournalofInternationalEconomics,JournalofLaborEconomics,JournalofPublicEconomics,RAND JournalofEconomics,ReviewofEconomicDynamics,TheoreticalEconomics,Journalofthe 18 TierJournalsN journals EuropeanEconomicAssociation FairBrookingsPapersonEconomicActivity,EuropeanJournalofHealthEconomics,InternationalJournal ofEmergingMarkets,JournalofEconomicHistory,JournalofPopulationEconomics,NewPolitical Economy,QuarterlyReviewofEconomicsandFinance 7 LimitedAgriculturalEconomics,AmfiteatruEconomic,ASTINBulletin,JournalofCulturalEconomics, JournaloftheJapaneseandInternationalEconomies,LocalEconomy,Post-SovietAffairs, QuantitativeEconomics 8 SupplementaryTable15(ST15):EconomicsCross-FieldValidationResults TableST15.Cross-fieldvalidationresultsforeconomics. ModelNAccuracyMacroF1Role Qwen3-30BSFT(economics)20069.5%70.4%In-domainSFT GPT-4.1-nanoSFT(economics)20068.5%69.4%In-domainSFT Qwen3-4BSFT(economics)20064.0%64.8%In-domainSFT Qwen3-30Bbase20025.5%16.9%Basecontrol GPT-4.1-nanobase20025.0%15.6%Basecontrol Qwen3-4Bbase20025.0%10.0%Basecontrol Qwen3-30Bpooled(mgmt+econ)20069.5%70.5%PooledSFT GPT-4.1-nanopooled(mgmt+econ)20067.0%67.4%PooledSFT MgmtSFTGPT-4.1(cross-field)20043.5%38.9%Cross-fieldtransfer GPT-4.1base(cross-field)20029.5%–Basecomparator Pooledmodelsweretrainedonthecombinedmanagementandeconomicscorpus(~10,072pairs)using thesameSFTprocedure.Onthemanagementheld-outbenchmark(N=120),thepooledQwen3-30B reached61.7%accuracy(macro-F10.615),comparabletothe60.8%single-fieldSFTensemble,andthe pooledGPT-4.1-nanoreached52.5%(macro-F10.489).Thecross-fieldtransferrowreportsthe management-trainedSFTGPT-4.1evaluatedoneconomicswithouteconomics-specificfine-tuning(p< 10⁻⁸versuschance;SupplementaryFig.7). SupplementaryFigures SupplementaryFigure1(SF1):PredictionDistributionComparison Predictiondistributionsacrossevaluatorclasses.a,100%stackedpredicted-tiersharesforthefrontieraverage(11models),chat average(3models),SFTsingle-modelaverage(4models),SFT2-modelensemble,expertmajorityvote,andjuniormajorityvote. Thispanelshowsataglancewhichevaluatorclassescollapseintothemiddletiersandwhichusethefulllabelspace.b,Deviation ofeachpredicted-tiersharefromthebalanced25%benchmark,showingevaluator-specifictierbias.c,Normalizedprediction entropyofthefullpredicteddistribution,wherehighervaluesindicatebroaderuseofthefourtiersandlowervaluesindicate strongercollapseintoanarrowsubsetoflabels. SupplementaryFigure2(SF2):ExpertIndividualAccuracyDistribution a,Histogramofper-expertaccuracyfortheunfilteredexpertpanel(N=48;8pitcheseach).Dashedlineindicatesmeanaccuracy (36.2%);dottedlineindicateschancelevel(25%).Thepanelhighlightssubstantialheterogeneityacrossexperts,notatightcluster aroundacommonlevelofskill.b,Empiricalcumulativedistributionfunction(CDF)ofper-expertaccuracywiththesamechance baseline,makingiteasiertoseehowmuchoftheexpertpanelliesnearchanceversusinthehigher-performingtail. SupplementaryFigure3(SF3):JuniorMonteCarloSubsamplingCurve a,Majority-voteaccuracyversuspanelsizeunderrepeatedMonteCarlosubsampling(5,000draws),with95%confidenceband. Chancebaseline(25%)isshownasadottedline,andthecurveshowsthatlargerpanelshelpearlybutthenflatten.b,Marginal accuracygainperadditionalrater,showingdiminishingreturnsaspanelsizeincreasesandclarifyingwhyaggregationalonedoes notkeepimprovinglinearly. SupplementaryFigure4(SF4):FrontierCollapse-MetricLandscape Collapsediagnosticsforthe11frontiermodels.a,Shareofpredictionsassignedtothemiddletiers(strong+fair)bymodel.b, Per-tierrecallheatmapacrossthefourqualitytiers,makingvisiblewhichclassesareeffectivelyneverrecovered.c,Relationship betweenmiddle-tierconcentrationandoverallaccuracyacrossmodels.d,Normalizedpredictionentropyranking,wherehigher valuesindicatelessdistributionalcollapseandbroaderuseofthefour-tierscale. SupplementaryFigure5(SF5):AI-HumanErrorComplementarity a,OverlapdecompositionofcorrectandincorrectoutcomesbetweentheSFTensembleandexpertmajorityvoteontheexpert- comparablesubset(N=89non-tiedpitches):bothcorrect,AI-onlycorrect,human-onlycorrect,andsharederror.Thispanel showsdirectlyhowmuchofthetwosystems’successisoverlappingversuscomplementary.b,Complementarityceiling:the oracleupperboundreachedwheneitherAIortheexpertmajorityiscorrectis77.5%,comparedwithSFTalone(62.9%onthis subset)andexpertmajorityvote(41.6%),quantifyingtheremainingroomforhybridroutingstrategies. SupplementaryFigure6(SF6):HumanConfusionMatrices Row-normalizedconfusionmatricesforexpertandjuniorpanels.a,Expertindividualpooled(N=384ratings).b,Junior individualpooled(N=2,530ratings).c,Expertstrictclear-majorityvoting(N=89non-tied;31tiesexcluded).d,Juniorstrict clear-majorityvoting(N=103non-tied;17tiesexcluded).Readingpooledandmajoritypanelstogetherclarifieshowmuch disagreementissmoothedbyvotingandwhichoff-diagonalconfusionspersistevenafteraggregation.Panelcompositionis summarizedinSupplementaryTableST5,andmajority-votecountsaresummarizedinSupplementaryTableST7. SupplementaryFigure7(SF7):Cross-FieldTransferfromManagementtoEconomics Themanagement-trainedSFTGPT-4.1checkpoint(trainedexclusivelyonmanagementinstitutionaltraces)isevaluatedonthe 200-articleeconomicsbenchmarkwithoutanyeconomics-specificfine-tuning.a,Overallaccuracycomparison:GPT-4.1base (29.5%,notsignificantlyabovechance),management-trainedSFTGPT-4.1(43.5%,p<10⁻ ⁸ versuschance;+14.0percentage pointsoverbase),andthebesteconomicsin-domainSFT(Qwen3-30B,69.5%)forreference.b,Row-normalizedconfusion matrixforthemanagementSFToneconomics,showing88%exceptionalrecallbutonly10%limitedrecall,consistentwith managementqualitysignalstransferringbestatthetopofthequalitydistribution.c,Per-tierrecallcomparisonacrossallthree evaluators,showingthatthemanagementSFTcapturespartialcross-fieldsignalconcentratedintheuppertiers,whilethebase modelcollapsesintomiddle-tierpredictions. SupplementaryReferences Supplementarycitationsusethesamenumberedbibliographyasthemainmanuscript.