Paper deep dive
BAGEL: Benchmarking Animal Knowledge Expertise in Language Models
Jiacheng Shen, Masato Hagiwara, Milad Alizadeh, Ellen Gilsenan-McMahon, Marius Miron, David Robinson, Emmanuel Chemla, Sara Keen, Gagan Narula, Mathieu LauriĂšre, Matthieu Geist, Olivier Pietquin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/27/2026, 7:36:58 PM
Summary
BAGEL (Benchmark for evaluating Animal knowledGe Expertise in Language models) is a new benchmark designed to evaluate the specialized animal-related knowledge of large language models (LLMs) through a unified closed-book multiple-choice evaluation protocol. It aggregates 11,852 questions from four diverse scientific and reference sources: Wikipedia (encyclopedic knowledge), Global Biotic Interactions (ecological interaction reasoning), bioRxiv (scientific literature reasoning), and Xeno-canto (text-based bioacoustic knowledge). The benchmark allows for fine-grained analysis across various taxonomic groups and knowledge categories such as taxonomy, morphology, habitat, behavior, and vocalization, providing a testbed for studying domain-specific knowledge generalization in biodiversity-related applications.
Entities (12)
Relation Signals (10)
BAGEL â constructedfrom â Wikipedia
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â constructedfrom â Global Biotic Interactions
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â constructedfrom â bioRxiv
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â constructedfrom â Xeno-canto
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â isconstructedfrom â Wikipedia
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â isconstructedfrom â bioRxiv
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
BAGEL â isconstructedfrom â Xeno-canto
confidence 100% · BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.
Tags
Links
- Source: https://arxiv.org/abs/2604.16241v1
- Canonical: https://arxiv.org/abs/2604.16241v1
Trouble viewing inline? Open PDF directly â
Full Text
70,781 characters extracted from source content.
Expand or collapse full text
BAGEL:BenchmarkingAnimalKnowledgeExpertise inLanguageModels Jiacheng Shen 2â , Masato Hagiwara 1â , Milad Alizadeh 1 , Ellen Gilsenan-McMahon 1 , Marius Miron 1 ,DavidRobinson 1 ,EmmanuelChemla 1 ,SaraKeen 1 ,GaganNarula 1 ,MathieuLauriĂšre 3⥠, MatthieuGeist 1⥠,OlivierPietquin 1⥠1 EarthSpeciesProject 2 NYUCenterforDataScience;NYUShanghai 3 NYUShanghaiCenterforDataScience;NYU-ECNUInstituteofMathematicalSciencesatNYUShanghai Abstract Large language models (LLMs) have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle special- ized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL,aBenchmarkforevaluatingAnimalknowledGeExpertiseinLanguagemodels. BAGEL isconstructedfromdiversescientificandreferencesources, includingbioRxiv, GlobalBioticIn- teractions, Xeno-canto, and Wikipedia, using a combination of curated examples and automat- ically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animalknowledge,includingtaxonomy,morphology,habitat,behavior,vocalization,geographic distribution,andspeciesinteractions. Byfocusingonclosed-bookevaluation,BAGELmeasures animal-relatedknowledgeofmodelswithoutexternalretrievalatinferencetime. BAGELfurther supports fine-grained analysis across source domains, taxonomic groups, and knowledge cate- gories,enablingamoreprecisecharacterizationofmodelstrengthsandsystematicfailuremodes. Ourbenchmarkprovidesanewtestbedforstudyingdomain-specificknowledgegeneralizationin languagemodelsandforimprovingtheirreliabilityinbiodiversity-relatedapplications. 1Introduction Instruction-tuned LLMs and chat-style assistants have rapidly improved on a wide range of knowledge and reasoning tasks (Ouyang et al.,2022;OpenAI,2023), leading to strong per- formance on broad-domain evaluations such as MMLU ( Hendrycks et al.,2021) and science- oriented benchmarks such as ScienceQA (Lu et al.,2022). These gains have fueled interest inusing LLMs as general-purpose interfacesfor scientific information access, synthesis, and questionanswering. However, strongaggregateperformanceonbroadbenchmarksdoesnot by itself establish whether models reliably encode specialized long-tail knowledge about the naturalworld,especiallywhenansweringquestionsthatrequirespecies-levelfacts,ecological relations, or natural-history reasoning. This gap is increasingly important because language andfoundationmodelsarealreadybeingexploredforbiodiversityandanimal-relatedapplica- tions. Inecology,recentstudieshaveusedLLMstoextractstructuredecologicalinformation from scientific literature, including hostâpathogen records and large-scale species interac- tions ( Gougherty and Clipp,2024;Keck, Broadbent, and Altermatt,2025), and have begun â Theseauthorscontributedequallytothiswork â Co-supervisingauthors 1 arXiv:2604.16241v1 [cs.CL] 17 Apr 2026 BAGELBENCHMARK Benchmark/line PrimarysettingRelationtoanimals/ biodiversity KeycontrastwithBAGEL BAGELText-only, closed- book 4-option MC from four curated sourcetracks Directly centered on animal and natural- historyknowledge Reference row:combines encyclopedic animalknowledge(Wikipedia),interaction reasoning among animal species (GloBI), animal-focused scientific-literature rea- soning (bioRxiv), and bioacoustic text about animals (Xeno-canto) under one unifiedprotocol. MMLU(Hendrycks etal.,2021) BroadtextMCacross 57subjects Only incidental biol- ogy / life-science cov- erage BAGEL is animal-centered, source- heterogeneous, and reports source/do- main breakdowns rather than a broad academicaggregate. ScienceQA (Lu et al.,2022) Multimodal school- scienceMC Includes natural- science curriculum content BAGELtargetshigher-levelnatural-history content and uses text-only closed-book questionsbuiltfrombiodiversitysources. BEANS; BEANS- Zero (Hagiwara, Hoffman, et al., 2022;Robinson etal.,2024) Bioacousticau- dio tasks; audioâ languagemodeling Directlyaboutanimal sounds BAGEL removes audio and tests whether modelscananswertextualquestionsabout vocalizationpropertiesandspeciesknowl- edge. Ecological knowl- edge eval (Dorm etal.,2025) Textualevaluationof ecological LLM com- petence (mixed task types) Sharedtheme: biodiversity-related scientific knowledge inlanguagemodels BAGEL is centered onanimal natural his- toryand organism-level facts, using one closed-bookMCformatacrossfourcurated biodiversitytextsources;Dormetal.study ecological competence more broadly with aheterogeneoustextualtasksuite. EnviroExam; ELLE (Y. Huang et al.,2024;Guo, N. Li, and M. Xu, 2025) Environmental- domainexams/QA Broad environment andeco-environment coverage BAGEL is narrower but deeper on animal and natural-history knowledge, with ex- plicit species, interaction, literature, and bioacoustic-sourcestructure. PubMedQA; BLURB ( Jin et al.,2019;Gu et al., 2021) Biomedical QA / biomedicalNLP Biological,but centered on hu- man health and biomedicine BAGEL shifts the focus from biomedical language understanding to organismal di- versity, ecology, and natural-history rea- soning. Table1. RepresentativeneighboringbenchmarksandwhereBAGELdiffersmostclearly. Rowsareillus- trative rather than exhaustive; descriptions were checked against primary papers / official benchmark descriptions. to test ecological knowledge more directly, finding uneven performance across ecological tasks (Dorm et al.,2025). In animal communication and bioacoustics, prior work has intro- ducedseveralimportantresources,includingBEANS,abenchmarkcoveringabroadrangeof animal sound tasks (Hagiwara, Hoffman, et al.,2022); ISPA, a text-like transcription scheme foranimalsounds( Hagiwara,Miron,andLiu,2024);andNatureLM-audio,anaudio-language foundation model built upon text-only LLMs(Llama) for bioacoustics (Robinson et al.,2024). Morebroadly,biodiversity-focusedfoundationmodelshavealsoemergedinothermodalities, suchasBioCLIPforfine-grainedrecognitionacrossthetreeoflife( Stevensetal.,2024).Despite this momentum, most prior work emphasizes information extraction, audio understanding, orvisualrecognitionratherthanevaluatingwhethertext-onlyLLMscananswerclosed-book questions about animals. As a result, it remains unclear how well current language models generalize across the heterogeneous forms of knowledge that matter for animal expertise, such as taxonomy, morphology, behavior, habitat, vocalization, geographic distribution, and speciesinteractions. Toaddressthisgap,weintroduceBAGEL,aBenchmarkforclosed-book evaluation ofAnimal knowledGeExpertise inLanguage models. 1 BAGEL aggregates 11,852 multiple-choicequestionsderivedfromfourcomplementarysources:Wikipedia,GlobalBiotic Interactions (GloBI) ( Poelen, Simons, and Mungall,2014), bioRxiv, and Xeno-canto (Vellinga and PlanquĂ©,2015). Together they target four animal-centered skillsâencyclopedic knowl- edgeaboutanimals(Wikipedia),ecologicalinteractionreasoning(GloBI),scientific-literature reasoningaboutanimals(bioRxiv),andtext-onlybioacoustic-domainknowledgeaboutanimal vocalizations(Xeno-canto). Bydesign,BAGELtestsnotonlyoverallaccuracybutalsorobust- ness across source domains and source-specific dimensions, providing a more fine-grained viewofwhatcurrentmodelsdoanddonotknowaboutanimalsandnaturalhistory. 1 Datasetrelease:https://huggingface.co/datasets/EarthSpeciesProject/BAGEL. 2 BAGELBENCHMARK 2RelatedWork LLMsandfoundationmodelsforbiodiversityandanimal-relatedapplications.Theuse oflanguageandfoundationmodelsinbiodiversity-relevantsettingsisgrowingrapidly,butthe literature is still fragmented across application areas and modalities. In ecology, LLMs have beenexploredastoolsforextractingstructuredknowledgefromtext,includingecologicalvari- ablesfromdiseasereports(GoughertyandClipp,2024)andspeciesinteractionsfromlargesci- entific corpora (Keck, Broadbent, and Altermatt,2025). Recent evaluation work has also be- guntoprobewhethergeneral-purposeLLMspossessecologicalknowledgedirectly,reporting asubstantialgapbetweenrelativelystrongfactualortaxonomicrecallandweakerecological reasoningorconservation-orientedjudgment(Dormetal.,2025). Inparallel,animal-centered foundation-model work has expanded in bioacoustics: BEANS established a public bench- markcoveringmultipleanimal-soundtasks(Hagiwara,Hoffman,etal.,2022);ISPAproposed atext-basedrepresentationfortranscribinganimalsoundsandconnectingthemtolanguage- model-style methods (Hagiwara, Miron, and Liu,2024); and NatureLM-audio introduced an audio-language foundation model tailored to bioacoustics, with strong zero-shot generaliza- tionacrosstaxaandtasks(Robinsonetal.,2024).Outsidetextandaudio,BioCLIPdemonstrates that biodiversity-specific foundation models can substantially improve fine-grained recogni- tion across a wide taxonomic range (Stevens et al.,2024). Adjacent Earth-science domains have likewise moved toward specialized language models, including K2 for geoscience and OceanGPT for ocean science ( Deng et al.,2023;Bi et al.,2023). Together, these studies show clearmomentumtowardAIsystemsspecializedfornatureandenvironmentaldata, butthey donotdirectlyevaluateclosed-bookanimalknowledgeintext-onlyLLMs. Generalknowledgeandsciencebenchmarks.Largelanguagemodelsarecommonlyeval- uatedusingbroadknowledgebenchmarkssuchasMMLU( Hendrycksetal.,2021),whichmea- sure multitask performance across many academic subjects. In science-focused evaluation, ScienceQA (Lu et al.,2022) provides a large benchmark of science questions with associated explanationsandmultimodalcontext. Thesebenchmarksarevaluableformeasuringgeneral scientific competence, but they do not specifically target fine-grained knowledge of animals, biodiversity,ornaturalhistory. BiomedicalandscientificQAbenchmarks.Aseparatelineofworkstudiesdomain-specific evaluation in biomedicine and scientific question answering. BLURB (Gu et al.,2021) aggre- gates multiple biomedical NLP tasks into a unified benchmark, while PubMedQA ( Jin et al., 2019)focusesonquestionansweringoverbiomedicalresearchabstracts. BioASQ(Nentidiset al.,2023)hasalsoestablishedalong-runningsharedtaskcenteredonlarge-scalebiomedical semanticindexingandquestionanswering. Theseresourcesdemonstratethevalueofdomain- specificevaluation,buttheyprimarilyfocusonbiomedicalorclinicalknowledgeratherthan biodiversityandnaturalhistory. Environmental and ecological benchmarks.More recently, several benchmarks have moved closer to environmental and ecological applications. EnviroExam (Y. Huang et al., 2024)evaluatesenvironmentalscienceknowledgeoflargelanguagemodelsusingcurriculum- based questions, and ELLE (Guo, N. Li, and M. Xu,2025) proposes a QA benchmark for eco- environment applications. Work on ecological knowledge evaluation is also beginning to emerge, with recent evidence that strong general-purpose LLMs retain only partial and task- dependent ecological competence ( Dorm et al.,2025). These efforts are important adjacent steps,buttheyemphasizebroadenvironmentalscience,ecology,orsustainabilitytopicsrather thananimal-centeredexpertise. Incontrast,BAGELforegroundsanimal-centeredevaluation under one protocol: encyclopedic species facts (Wikipedia), ecological interactions among taxa(GloBI),animal-relevantscientificliteraturereasoning(bioRxiv),andbioacoustic-domain textualknowledge(Xeno-canto),withaccuracyreportedpersource. Our Contributions.BAGEL complements prior work by focusing on a distinct but under- exploredevaluationaxis: closed-bookquestionansweringaboutanimalsandnaturalhistory 3 BAGELBENCHMARK Figure1. OverviewoftheBAGELbenchmarkcurationpipeline: foursource-specificpreparationtracks feed domain prompts into a shared generator, followed by quality checks, four-option formatting, and option-ordershuffling. grounded in heterogeneous biodiversity-relevant sources. Rather than testing general aca- demicknowledgeorbiomedicalreasoningalone,BAGELisdesignedtomeasurehowwelllan- guagemodelshandlespecies-levelknowledge,ecologicalrelations,andanimal-focusedfactual generalizationacrossmultiplesourcedomains. Table 1makesthecontrastsconcretewithrep- resentativeneighbors. 3BenchmarkConstruction Figure1summarizes the end-to-end curation workflow; subsections below describe each sourcedomain. Thefourcorporaarenotintendedtoexhaustreal-worldnatural-historycom- petence; they are public, machine-accessible anchors over complementary animal-centered skills: encyclopedic knowledge about animals, ecological interaction reasoning among taxa, scientific-literature reasoning on animal-focused preprints, and text-only bioacoustic knowl- edgederivedfromanimalrecordings. 3.1Wikipedia TheWikipediasubsettargetsencyclopedicknowledgeaboutanimals:EnglishWikipediaspecies articlessupplytheevidence,andeachitemprobesclosed-bookrecalloftaxon-specificfactsa readerwouldnormallytakefromsuchanarticle,withoutaccesstothearticleattesttime. 2 Article retrieval.For each candidate taxon, structured metadata (scientific name and higher-levelclassification)isusedtoresolvetheEnglisharticletitlethroughWikidata(linking thetaxonnametoanitemandreadingitsEnglishWikipediasitelink),afterwhichthearticle plain-textextractisretrievedviatheMediaWikiAPI.Pageswithmissingtextorextractsshorter than1,000charactersareexcludedsothatstubsandminimallyinformativearticlesdonotdom- inatethebenchmark. Textpreparation.Extractslongerthan180,000charactersaretruncatedbeforegeneration, withpreferenceforparagraph-orsentence-boundarycutssothattheretainedprefixremains coherent. This cap keeps prompts within practical context-window limits of the generation model. Noadditionalmanualeditingofarticletextisperformedbeyondthiscap. Question synthesis.Questions are generated through the GPT-4o-mini API using the systemâusertemplateinAppendix A.1.1. Themodelmayemituptoeightfour-option,single- answeritemsperspecies, eachassignedtooneofeightthematicdimensions:Taxonomy;Be- 2 Portalhttps://w.wikipedia.org/; taxa are linked through Wikidata and article plain-text extracts are retrieved viatheMediaWikiAPI. 4 BAGELBENCHMARK havior(including social behavior where applicable);Communication;Morphology;Habitat; Cognition;GeographicDistribution; andDiet. Every item must be justified solely by explicit statementsinthesuppliedextract; dimensionsnotsupportedbythetextareskipped. When thearticlementionsvocalization,geographicrange,orfeeding,thepromptencouragesatleast oneiteminCommunication,GeographicDistribution,orDiet,respectively. Parsedoutputsare keptonlyiftheypasssimplestructuralchecks(validdimensionlabel,exactlyfouroptions,and adesignatedcorrectoptionthatmatchesoneoptionstringexactly). Atevaluationtime,mod- elsseeonlythestemandanswerchoices;articletextandconstructionmetadataarewithheld, consistentwiththeclosed-bookprotocolinSection4. 3.2GlobalBioticInteractions(GloBI) The GloBI subset targets ecological interaction reasoning among species. Tabular records in theGlobalBioticInteractionsexchangeformatdescribedirectedlinksbetweenasourcetaxon and a target taxon together with an interaction-type label and optional locality, coordinates, observationtime,life-stageorbody-partfields,habitat,andbibliographicprovenance. 3 Preprocessing.We read at most 10,000 rows from the record table and harmonize col- umnnamesacrosscommonGloBIexportconventions. Rowswithoutbothendpointtaxaand an interaction-type label are excluded. Each retained row is converted into a short natural- languagesummaryoftheinteractiontogetherwithoptionallocality,date,andcoordinatemeta- data,plusdataset-andreference-levelprovenance. Balancedsubsampling.Fromtheannotatedpoolweselect3,500interactions,stratifying byinteractiontypesothattheempiricaldistributionoverrelationlabelsismoreuniformthan underuniformrandomsamplingofrows. Withineachtype,rowswithrichercontextualmeta- data(locality,coordinates,dates)arepreferredwhentiesarise,usingafixedrandomseedfor reproducibility;thefinalsubsetisshuffledbeforequestionsynthesis. Multiple-choice synthesis.For each selected interaction, we prompt GPT-4o-mini throughtheOpenAIAPIwiththenatural-languageinteractionsummaryastheonlytextualev- idence. TheinstructionformatisgiveninAppendix A.1.1. Themodelmustoutputexactlyone closed-book, four-option item labeled as eitherMaskedparticipantidentificationorMasked interactiontypeinference,obeyingconstraintsthatdiscouragetrivialverbcuesandencourage ecologicallyplausibledistractors. Malformedorunparseableoutputsarediscarded. 3.3bioRxiv The bioRxiv subset targetsscientific-literature reasoning about animals. We harvest animal- related articles directly from the public bioRxiv website 4 using its month-indexed advanced searchinterface,restrictingtopostsdatedfrom2023onwardandtofoursubjectareas:Animal BehaviorandCognition , Ecology , EvolutionaryBiology ,and Zoology . Preprocessing.Title, abstract, andarticlebodyareconcatenatedintooneworkingdocu- ment. Forbodytext,wenormalizewhitespace,removeexactduplicateparagraphs,anddrop paragraphsthatbeginwithâFigureâ. Wethenapplyalighttext-cleaningpass(newline/page- number cleanup and whitespace normalization). We do not apply an additional hard length capforbioRxivdocumentssincethetextiswithinthecontextwindowthatGPT-4o-minican handle. Multiple-choicesynthesis.EachdocumentispresentedtoGPT-4o-miniviatheOpenAI API.ThepromptspecificationappearsinAppendixA.1.1. Themodelmustreturnexactlyone four-optionitemoftypeResultInterpretation: thestemmuststateenoughoftheexperimen- tal setup and findings that the item is answerable without the PDF, and all claims must be groundedinthesuppliedexcerpt. Malformedoutputsarediscarded. Thereleasedbenchmark contains2,183questions(Table 2). 3 GloBIportalandindexedinteractiondata:https://w.globalbioticinteractions.org/. 4 bioRxivserver:https://w.biorxiv.org/. 5 BAGELBENCHMARK 3.4Xeno-canto The Xeno-canto subset focuses on text-based bioacoustic knowledge concerning animals, including typical vocalization characteristics and coarse-grained acoustic structures at the specieslevel. ItisderivedfromtheXeno-cantocommunityrepository(VellingaandPlanquĂ©, 2015) 5 , which provides user-contributed recordings and associated species metadata. At test time,modelsreceiveonlythequestionstemandfouroptions;spectrogramsandaudioarewith- held,soscoresreflectvocabularyandreasoningoveranimalsoundsintext,notperceptionof aplayedwaveform.AppendixFigure3providesademonstrationoftheXeno-cantoMCQsgen- erationpipeline. Species coverage.Xeno-canto items are built only for species that already appear in the Wikipedia construction pipeline: we use the same scientific-name inventory as for the Wikipediasubset,soeverytaxonwithacandidaterecordingisoneforwhichwealsoderived species-levelquestionsfromEnglishWikipedia(subsection3.1). Recording retrieval and filtering.Recordings are retrieved from Xeno-Canto through its public query interface. For each species we request clips between 5 and 600 seconds in length, exclude material released under no-derivatives licenses, and retain at most 100 ac- ceptedfilesperspeciesbeforefurtherprocessing. Wealsoalignthespecieslisttoaninternal recording-levelmetadatainventorywithrichertaxonomicandrecordingattributesthanamin- imalspecies-levelexport, matchingonunifiedscientificnames. Audioisanalyzedat16kHz; clipsshorterthanfivesecondsorwithpeakamplitudebelow0.1aredropped,andlongerfiles aretruncatedtothefirst30secondsforfeatureextraction. Eachretainedwaveformreceivesa scalarclarityscorefromalightweightsignal-processingheuristicthatcombinesband-limited signal-to-noise(0.5â10kHz),separationofharmonicversuspercussiveenergy,meanspectral flatness, and a penalty for clipping. When many candidates exist per species, we randomize orderandkeepasmallfixednumbertodiversifyrecordingswhilecontrollingcost. Questionsynthesis.Forrecordingswithusablevocalizationlabels,free-texttypestrings arenormalized: mixedorambiguouslabels(e.g.,simultaneousâsongâandâcallâ),emptytags, orâuncertainâentriesareskipped,andotherwiseweretainasinglecanonicalphrase. Weren- deralog-frequencyspectrogramofatmost30seconds(audioresampledat44.1kHzfordisplay) andprovideittogetherwithcommonandscientificnamestoGPT-4o-minithroughtheOpe- nAIAPI.Theinstructionschema(AppendixA.1.1)asksforexactlyonefour-optioniteminone of four randomly drawn topic familiesâdominant frequency range, call or syllable duration, modulationpattern,orharmonicstructureversusbroadbandnoiseâwithrespectivesampling weights of approximately 65%, 15%, 10%, and 10%. The model should use the image only as supportforspecies-typicalacousticdescriptionsandmustnotmentionspectrograms,files,or recordingsinthequestion;thestemmustexplicitlynamethevocalizationtype(e.g.,songver- suscall). 3.5Mitigatinganswer-positionimbalance. Afterinitialconstruction,wefoundthattheindexofthecorrectoptionamongthefourordered choiceswashighlyskewedacrossdomains(AppendixF,Table9). Suchimbalanceisproblem- atic for evaluation: recent work shows that large language models are sensitive to option or- derinmultiple-choicetestsandcanexhibitsystematicselectionbiastowardparticularanswer labels independent of content (C. Zheng et al.,2024;Pezeshkpour and Hruschka,2024). To reduce the risk that models exploit positional shortcuts, we apply a random permutation of thefouroptionsforeachitemwithafixedrandomseed,updatingthestoredcorrectlabelorin- dexaccordingly.Table 10summarizestheresultingdistributionoverall11,852multiple-choice itemsinthereleasedshuffledsplit: theempiricalfrequencyofeachgoldpositionstayswithin aboutonepercentagepointofthenominal25%underabalanceddesign. Figure 2provides one representative example item from each source domain, illustrat- ing heterogeneity across encyclopedic animal knowledge (Wikipedia), interaction reasoning amonganimalspecies(GloBI),scientific-literaturereasoningaboutanimals(bioRxiv),andbioa- coustictext(Xeno-canto). Additionalexamplesforeverydomain-specificdimensionarepro- videdinAppendix B. 5 Recordingsandspeciesmetadata:https://w.xeno-canto.org/. 6 BAGELBENCHMARK bioRxiv Dimension:ResultInterpretation Level:easy Question:Inastudyoftheseaanemone Nematostella vectensis, researchers foundthatanimalskeptinenvironments with gravel substrate produced signif- icantly more clonal progeny through transverse fission compared to those without substrate. Given that substrate enhances fission rates, what can be inferred about its role in asexual repro- duction? Options:A. Substrate provides a me- chanical advantage that facilitates the physical process of fissioning. B. Sub- strate increases the genetic diversity of the clones produced during fission. C. Substrate reduces the metabolic waste that inhibits fission in high-density pop- ulations. D. Substrate alters the hor- monal balance inNematostella, promot- ingfastergrowth. Gold:A. Substrate provides a mechani- caladvantagethatfacilitatesthephysical processoffissioning. Wikipedia Dimension:Communication Level:medium Question:What is one of the distinct call types of the Andean flamingo? Options:A.WhistleB.ChirpC.GrowlD.Peep Gold:D.Peep GloBI Dimension:Maskedparticipantidentification Level:hard Question:Whichtaxonislikelytobethehostintheinteraction withPallisentisnagpurensis? Options:A. Cichlidae B. Channa striata C. Heteropneustes fos- silisD.Pallisentisnagpurensis Gold:B.Channastriata Xeno-canto Dimension:Dominantfrequencyrange Level:easy Question:Whatisthetypicaldominantfrequencyrangeofthe calloftheRedWarbler(Cardellinarubra)? Options:A.âŒ3â5kHzB.âŒ1â2kHzC.âŒ10â12kHzD.âŒ7â9kHz Gold:A.âŒ3â5kHz Figure2. Representative example items from each BAGEL domain.Levelindicates the reference diffi- cultylabelfromTable3(agreementbetweenGPT-5.4andClaudeOpus4.6onthereleasedlevelfield). The fourpanelsarechosentoincludeeasy,medium,andhardstrataacrossdomains. 4ExperimentalSetup 4.1Models Weevaluatetwofrontierclosed-sourcemodelsâGPT-5.4andClaudeOpus4.6âandaladderof open-weightmodelsfromSmolLM2-360M(BenAllaletal.,2025)upward,includingfiveQwen3 familymodels(0.6B,4B,8B,14B,and32Bwithreasoningmodeturnedoff)(A.Yangetal.,2025), Gemma 3 27B IT (Team et al.,2025) and Gemma-7B (Gemma Team,2024), Llama 3.1-8B In- struct(Grattafiorietal.,2024),Mistral-7B-Instruct(Jiangetal.,2023),andPhi-4(Abdinetal., 2024). 4.2EvaluationProtocol All models are evaluated in a closed-book multiple-choice setting. At test time, the model is givenaunifiedpromptcontainingonlythetaskinstruction,thequestionstem,andfourenu- meratedansweroptions; thesourcepassageusedduringbenchmarkconstructionisnotpro- videdatinferencetime. Weusedeterministicdecodingwithseed0andgreedygenerationfor thereportedresults,soeachmodelisrepresentedbyasinglerunratherthananaverageacross seeds. We report accuracy on each source domain and an overall score across domains. The evaluationpromptisprovidedinAppendix A.1.2. 4.3DatasetStatistics Table2summarizes the composition of BAGEL:11,852four-option, single-answer multiple- choice questions acrossbioRxiv(2,183),GloBI(3,500),Wikipedia(1,927), andXeno-canto (4,242),alongwithsource-specificdimensionsandeachcategoryâssharewithinitsdomain. Ta- ble 3reports how many items fall into each reference difficulty level (easy / medium / hard) whendefinedbyagreementbetweenGPT-5.4andClaudeOpus4.6;thesamelabelsarestored 7 BAGELBENCHMARK Domain Dimension#Questions ShareinDomain bioRxiv(#=2,183) bioRxiv ResultInterpretation2,183100.0% GloBI(#=3,500) GloBIMaskedinteractiontypeinference2,12060.6% GloBIMaskedparticipantidentification1,38039.4% Wikipedia(#=1,927) Wikipedia Behavior35318.3% Wikipedia Diet29615.4% Wikipedia GeographicDistribution28114.6% Wikipedia Taxonomy26813.9% Wikipedia Communication23412.1% Wikipedia Habitat23412.1% Wikipedia Morphology23312.1% Wikipedia Cognition281.5% Xeno-canto(#=4,242) Xeno-canto Dominantfrequencyrange2,74764.8% Xeno-canto Call/syllableduration64915.3% Xeno-canto Harmonicstructure/tonalityvs.broadband42710.1% Xeno-canto Modulationpattern4199.9% Table2. CompositionofBAGELacrosssourcedomainsandsource-specificdimensions. Thefullbench- mark contains 11,852 questions, with average question length of 180.54 characters and average option lengthof35.68characters. Domain #ItemsEasyMediumHard Wikipedia1,927 1,681(87.2%) 163(8.5%)83(4.3%) Xeno-canto 4,242 1,836(43.3%) 1,218(28.7%) 1,188(28.0%) GloBI3,500 2,549(72.8%) 439(12.5%) 512(14.6%) bioRxiv2,183 1,957(89.6%) 135(6.2%)91(4.2%) All11,852 8,023(67.7%) 1,955(16.5%) 1,874(15.8%) Table3. Referencedifficultystrata inducedbyGPT-5.4andClaudeOpus4.6underourevaluationpro- tocol:easyifbothmodelsanswercorrectly,mediumifexactlyoneiscorrect,andhardifbotharewrong. in the released data aslevel. We treat these as reference difficulty strata rather than absolute itemdifficulty,andasubsetofâhardâitemsmayalsoreflectoption-levelambiguity. Reported lengthsaremeasuredincharactersonthereleasedbenchmarkfiles: themeanquestionstem lengthis180.5andthemeanlengthofasingleoptionstringis35.7(averagedoveralloptions). 5ResultsandDiscussion 5.1Mainleaderboard Table4reportsthemainresultsonBAGEL.Threefindingsstandout. First,performancevariessubstantiallyacrosssourcedomains,evenforthestrongestmod- els. For example, frontier closed-source models perform very strongly on Wikipedia and bioRxiv,yetremainnotablyweakeronXeno-canto.ThisgapindicatesthatBAGELisnotmerely measuringbroadfactualrecall, butalsosource-sensitiveexpertiseinanimalandnaturalhis- toryknowledge. Second,open-weightmodelsspanawidebandbelowtheproprietaryfrontier. Amongthe openmodelsinTable4,Gemma327BITreachesthestrongestoverallopenscore(0.6789)under ourprotocol,followedcloselybyLlama3.1Instruct-8B(0.6522)andPhi-4(0.6511),whileQwen3- 32B is highest on Wikipedia accuracy. None of the included open models surpass GPT-5.4âs overallscore(0.7601),highlightingaremaininggaptotheproprietaryentryinthisevaluation setup. Third, thesmallestopenmodelsremainneartherandombaselineonaggregate(Table 4), whilemid-sizedmodelsalreadyreachnon-trivialaccuracyontext-heavydomainsyetcanfail onXeno-canto,indicatingunevenscalingofbiodiversity-relatedcompetence. Then we evaluate five instruct checkpoints from the same Qwen3 lineâ0.6B, 4B, 8B, 14B, and 32Bâunder an identical prompt and greedy seed-0 protocol (Table 4). Overall accuracy risessharplyfromnear-randomperformanceat0.6B(0.2464)intothe0.57â0.65bandfor4Bâ32B, and GloBI, Wikipedia, and bioRxiv scores generally increase with size up to 32B. Xeno-canto 8 BAGELBENCHMARK ModelGloBI Wiki bioRxiv Xeno-canto Overall EmpiricalRandomGuess 0.2686 0.2247 0.2652 0.2520 0.2549 Closed-SourceFrontierModels GPT-5.40.7720 0.8988 0.94410.59260.7601 ClaudeOpus4.60.81000.93050.9107 0.5601 0.7587 Open-WeightModels smolLM2-360M0.2477 0.2372 0.2510 0.2504 0.2476 Qwen3-0.6B0.2523 0.2532 0.2405 0.2414 0.2464 Qwen3-4B0.5866 0.6767 0.8901 0.3581 0.5753 Gemma-7B0.4806 0.6487 0.8543 0.2791 0.5046 Mistral-7B-Instruct-v0.3 0.6060 0.6834 0.8557 0.2949 0.5532 Llama3.1Instruct-8B 0.6386 0.7348 0.88820.50450.6522 Qwen3-8B0.6586 0.7654 0.8992 0.3932 0.6253 Phi-40.7251 0.7971 0.9244 0.3831 0.6511 Qwen3-14B0.6954 0.7940 0.91110.4125 0.6499 Gemma327BIT0.74710.8012 0.9098 0.44810.6789 Qwen3-32B0.70740.8106 0.91850.3751 0.6441 Table 4. Accuracy by domain on BAGEL. Open-weight models are ordered by increasing approximate size. Within the open-weight block, the best value in each column is shown inbold(leaders can differ bycolumn). GPT-5.4andClaudeOpus4.6areclosed-sourcereferences;betweenthosetwo,thestronger scoreineachcolumnisunderlined. behaves differently: accuracy moves from0.2414(0.6B) through0.3581(4B),0.3932(8B), and 0 . 4125 (14B),thenfallsat32B( 0 . 3751 ). ThusthelargestQwen3checkpointisnotthestrongest onXeno-cantoeventhoughitleadsonthetext-heavysources.BecauseXeno-cantocontributes manyitems,Qwen3-14BslightlyedgesQwen3-32BonOverall(0.6499vs.0.6441). Figure4inthe Appendixplotsthesamefivecheckpointswithdomain-wisecurves. Overall,theseresultsshowthatBAGELseparatesmodelsnotonlybyaggregatecapability, butalsobyrobustnessacrossheterogeneoussources. AppendixTables5â7breakdownaccu- racybysource-specificquestiontype(dimensionsmatchTable2). 5.2Discussion Construct validity and interpretation.BAGEL aggregates four complementary animal- centered tracksâencyclopedic facts (Wikipedia), pairwise species-interaction reasoning (GloBI),scientific-literature-stylestems(bioRxiv),andtext-onlyitemsgroundedinbioacoustic metadata (Xeno-canto)âunder a single closed-book, four-option MC protocol. Under this setup, reported scores reflect what models can produce from parameters and instructions alone,withoutaccesstosourcepassagesorexternalretrieval;theyshouldnotbeequatedwith field identification skills, expert ornithological practice, or perception of waveforms when audio is withheld by design (Section3). Because the four tracks emphasize different surface forms of knowledge, aggregate accuracy is best treated as a coarse summary: substantive conclusions require domain-level and, where available, dimension-level accuracy (Table 4; AppendixTables5â7). Followingstandardconcernsaboutmultiple-choiceartifacts,weapply independent random permutations of the four answer options at the item level to mitigate systematic gold-position imbalance in the released split (Section 3; AppendixF). Even after thismitigation,somemodelsexhibitunevenemissionofoptionletters,someasuredaccuracy canstillco-varywithformatsensitivityinwaysthatareorthogonaltoâcontentmasteryâinthe narrowsense. 9 BAGELBENCHMARK Heterogeneityandlimitedcross-domaintransfer.The main leaderboard underscores a recurrentpattern: modelsthatachievestrongaccuracyontext-richdomainscanremaincom- parativelyweakonXeno-canto(Table4).ThispatternisnotanartifactofasmallXenosubsetâ theXeno-cantotrackcontributesalargeshareofitemsâyetitisfrequentlythelowest-accuracy columnforfrontierandopenmodelsalike. Suchcross-trackgapsmotivatereadingBAGELas aportfolioevaluationinwhichoverallrankingissecondarytodiagnosingwhichcompetencies arepresentorabsent( Hendrycksetal.,2021). Inpracticalterms,highperformanceonencyclo- pedicorabstract-styleitemsdoesnotlicensetheconclusionthatamodelwillreliablyhandle acoustics-orientedwordingatscale. TheXeno-cantogap: itemtypes,stemstatistics,andscaling.The relative weakness on Xeno-cantodoesnotreflectasingleuniformdeficit,butrathersubstantialheterogeneityacross itsunderlyingsubtasks.AppendixTable7decomposesaccuracybythefouracoustictopicfam- iliesusedduringconstruction(dominantfrequencyrange,callorsyllableduration,harmonic structure versus broadband noise, and modulation pattern). The spread across these dimen- sionsislarge: forGPT-5.4,accuracyreachesroughly0.71ondominant-frequency-rangeitems butonlyabout0.30onmodulation-patternitems,withothermodelsexhibitingsimilarlywide within-domainvariationanddifferingstrengthsacrossdimensions.TheaggregateXeno-canto scorethereforeaveragesoversubtasksofmarkedlydifferentdifficulty,whichpartiallyexplains whyitcanlagWikipediaevenwhenbothtracksarenominallyâaboutâthesamespecieslist. To complement this item-type decomposition, Appendix Table 8analyzes the linguistic propertiesofEnglishquestionstems,usingwordfrequnigramZipfscoresasaproxyforgeneral- Englishwordfrequency(Speer,2018). RelativetoWikipediastems,Xeno-cantostemsexhibit (i)substantiallyhigherdensityoftokensdrawnfromafixedbioacoustickeywordlist(meanhits perstem:2.03versus0.05),(i)alargerfractionoflow-frequencytokens(Zipf<3:18.4%versus 11.8%),and(i)lowermeanunigramZipf(5.04versus5.33). ThesestatisticsindicatethatXeno- cantooperatesinadistinctlexicalregister, combiningdomain-specificacousticterminology withlessfrequentgeneral-Englishvocabulary. However,lexicalrarityalonedoesnotexplain performance: meanstemZipfshowsnegligiblecorrelationwithper-itemcorrectnessonXeno- cantoforGemma327BIT(SpearmanÏââ0.03),suggestingthatdifficultyarisesfromthein- teractionofdomain-specificvocabulary,reasoningdemands,anddistractordesignratherthan fromwordfrequencyalone. Finally, scaling behavior further differentiates Xeno-canto from text-heavy domains. WithintheQwen3family,Xeno-cantoaccuracyisnon-monotonicwithparametercount,with the 32B checkpoint underperforming the 14B checkpoint on this domain despite improving onWikipediaandbioRxiv(Figure 4). Thisdivergenceindicatesthatgainsingenerallanguage modeling or factual recall do not uniformly translate to improvements in bioacoustic text competence. These results are consistent with the possibility that text-based bioacoustic QA constitutesapartiallydistinctchallenge,althoughthegapmayalsoreflectdifferencesinlex- ical register, item construction, and subtask composition. Under our protocol, Xeno-canto evaluates textual reasoning about sound and should be viewed as complementary to, rather than a substitute for, audio-centric animal-sound benchmarks ( Hagiwara, Hoffman, et al., 2022;Robinsonetal.,2024). Multiple-choiceambiguityinBAGEL.OurfindingsfromthemanualauditinAppendixG align with a broader literature showing that multiple-choice evaluation can be confounded by insufficiently discriminative answer options, including semantically overlapping distrac- tors, multipleplausibleanswers, andtheforcedsingle-answerassumption( Paltaetal.,2024; Balepur, Rudinger, andBoyd-Graber,2025;W.Xuetal.,2025). Surveysandsystemslikewise emphasizedistractorqualityandoptiondistinctness( Alhazmietal.,2024;Bitewetal.,2023; Amanlou et al.,2026), while recent work argues for moving beyondstrict MC scoring toward generativeormatching-basedevaluationwhereappropriate(Chandaketal.,2025). Superficial factorssuchasoptionorderingcanalsoshiftscores(PezeshkpourandHruschka,2024). Wedo notclaimtheseissuesareuniquetoanyoneBAGELdomain:wehighlightGloBIbelowbecause interaction-centricpromptssurfacethemclearlyinourmanualreview,butanalogousriskscan arisewhereveritemsaregeneratedunderasingle-answerprotocol. 10 BAGELBENCHMARK 6Conclusionandlimitations We introduced BAGEL, an 11,852-item closed-book MC benchmark from bioRxiv, GloBI, Wikipedia,andXeno-canto. Experimentsshowlargedomainspread,agaptoaproprietaryref- erence,persistentlyweakerXeno-cantoscoresformanymodels,andcomplementarystrengths acrossopen-weightfamilies;AppendixFdiscussesMCpositionaleffects. Limitations.Resultsarereportedwithoneseedandgreedydecoding,sostochasticdecod- ing or ensembling could shift rankings. BAGELis closed-book onlyin the sense that models are evaluated without source passages or external retrieval at inference time; it does not es- tablishthatbenchmarkcontentwasabsentfrompretrainingcorpora. Becauseseveralsource domainsarepublic,scoresmayreflectamixtureofsourcefamiliarity,broaderparametricre- call, multiple-choice test-taking skill, and task-specific reasoning, which this setup does not cleanlydisentangle. ThebenchmarkisEnglish-onlyandinheritstaxonomic,geographic,and source-selection biases from its sources. MC accuracy does not measure calibration, absten- tion, oropen-endedreasoning. Corpuschoiceispragmaticandaxis-coveringratherthanex- haustive of natural-history competence; field guides, occurrence corpora, and expert exams are out of scope but complementary. The Xeno-canto-derived domain is text-only and can- notdisentangleacousticjargonfromdeepunderstanding. Atargetedauditoftheâhardâsub- setindicatesthataminorityofitemscontaininsufficientlydiscriminativeansweroptions;the âhardâsplitshouldbeinterpretedasrelativechallengeundergenerativemultiple-choicedesign ratherthanacleanestimateofbiologicalreasoningdifficultyineveryitem. Becausequestions aregeneratedautomaticallyfromsourcematerials,benchmarkqualityalsodependsongener- ationprompts,filteringrules,anddistractorconstruction,whichmayintroduceartifactsnot reducibletothesourcecontentitself. 11 BAGELBENCHMARK References Abdin,Marah,JyotiAneja,HarkiratBehl,SĂ©bastienBubeck,RonenEldan,SuriyaGunasekar, MichaelHarrison,RussellJ.Hewett,MojanJavaheripi,PieroKauffmann,etal.(2024).âPhi- 4TechnicalReport.âIn:arXivpreprintarXiv:2412.08905. Alhazmi, Elaf, Quan Z. Sheng, Wei Emma Zhang, Munazza Zaib, and Ahoud Alhazmi (2024). DistractorGenerationinMultiple-ChoiceTasks:ASurveyofMethods,Datasets,andEvalu- ation.arXiv: 2402.01512[cs.CL]. Amanlou,Mohammadetal.(2026).âKNIGHT:KnowledgeGraph-DrivenMultiple-ChoiceQues- tionGenerationwithAdaptiveHardnessCalibration.âIn:arXivpreprintarXiv:2602.20135. Balepur,Nishant,RachelRudinger,andJordanLeeBoyd-Graber(2025).âWhichofTheseBest DescribesMultipleChoiceEvaluationwithLLMs?A)ForcedB)FlawedC)FixableD)Allof theAbove.âIn:Proceedingsofthe63rdAnnualMeetingoftheAssociationforComputational Linguistics(Volume1:LongPapers).Vienna,Austria,p.3394â3418. BenAllal,Loubna,AntonLozhkov,ElieBakouch,GabrielMartĂnBlĂĄzquez,GuilhermePenedo, LewisTunstall,AndrĂ©sMarafioti,HynekKydlĂÄek,AgustĂnPiqueresLajarĂn,VaibhavSri- vastav, et al. (2025). âSmolLM2: When Smol Goes Big â Data-Centric Training of a Small LanguageModel.âIn:arXivpreprintarXiv:2502.02737. Bi, Zhen, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen (2023).âOceanGPT:ALargeLanguageModelforOceanScienceTasks.âIn:arXivpreprint arXiv:2310.02031. Bitew,SemereKiros,JohannesDeleu,ChrisDevelder,andThomasDemeester(2023).âDistrac- torgenerationformultiple-choicequestionswithpredictivepromptingandlargelanguage models.âIn:arXivpreprintarXiv:2307.16338. Chandak,Nikhil,ShashwatGoel,AmeyaPrabhu,MoritzHardt,andJonasGeiping(2025).An- swerMatchingOutperformsMultipleChoiceforLanguageModelEvaluation. arXiv: 2507. 02856[cs.CL]. Deng,Cheng,TianhangZhang,ZhongmouHe,YiXu,QiyuanChen,YuanyuanShi,LuoyiFu, WeinanZhang,XinbingWang,ChenghuZhou,ZhouhanLin,andJunxianHe(2023).âK2: AFoundationLanguageModelforGeoscienceKnowledgeUnderstandingandUtilization.â In:arXivpreprintarXiv:2306.05064. Dorm,Filip,JosephMillard,DrewPurves,MichaelHarfoot,andOisinMacAodha(2025).âLarge LanguageModelsPossessSomeEcologicalKnowledge,butHowMuch?âIn:bioRxiv. GemmaTeam(2024).âGemma:OpenModelsBasedonGeminiResearchandTechnology.âIn: arXivpreprintarXiv:2403.08295. Gougherty,AndrewV.andHannahL.Clipp(2024).âTestingtheReliabilityofanAI-BasedLarge LanguageModeltoExtractEcologicalInformationfromtheScientificLiterature.âIn:npj Biodiversity3,p.13. Grattafiori,Aaron,AbhimanyuDubey,AbhinavJauhri,AbhinavPandey,AbhishekKadian,Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. (2024). âTheLlama3HerdofModels.âIn:arXivpreprintarXiv:2407.21783. Gu,Yu,RobertTinn,HaoCheng,MichaelLucas,NaotoUsuyama,XiaodongLiu,TristanNau- mann,JianfengGao,andHoifungPoon(2021).âDomain-SpecificLanguageModelPretrain- ingforBiomedicalNaturalLanguageProcessing.âIn:ACMTransactionsonComputingfor Healthcare3.1,p.1â23. 12 BAGELBENCHMARK Guo, Jing, Nan Li, and Ming Xu (2025). âEnvironmental Large Language Model Evalua- tion (ELLE) Dataset: A Benchmark for Evaluating Generative AI Applications in Eco- EnvironmentDomain.âIn:arXivpreprintarXiv:2501.06277. Hagiwara,Masato,BenjaminHoffman,Jen-YuLiu,MaddieCusimano,FelixEffenberger,and Katie Zacarian (2022). âBEANS: The Benchmark of Animal Sounds.â In:arXiv preprint arXiv:2210.12300. Hagiwara,Masato,MariusMiron,andJen-YuLiu(2024).âISPA:Inter-SpeciesPhoneticAlpha- betforTranscribingAnimalSounds.âIn:arXivpreprintarXiv:2402.03269. Hendrycks,Dan,CollinBurns,StevenBasart,AndyZou,MantasMazeika,DawnSong,andJa- cobSteinhardt(2021).âMeasuringMassiveMultitaskLanguageUnderstanding.âIn:Inter- nationalConferenceonLearningRepresentations(ICLR). Huang,Yu,LiangGuo,WanqianGuo,ZheTao,YangLv,ZhihaoSun,andDongfangZhao(2024). âEnviroExam:BenchmarkingEnvironmentalScienceKnowledgeofLargeLanguageMod- els.âIn:arXivpreprintarXiv:2405.11265. Jiang, Albert Q., Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot,DiegodelasCasas,FlorianBressand,GiannaLengyel,GuillaumeLample,Lucile Saulnier,etal.(2023).âMistral7B.âIn:arXivpreprintarXiv:2310.06825. Jin, Qiao, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu (2019). âPub- MedQA: A Dataset for Biomedical Research Question Answering.â In:Proceedings of the 2019ConferenceonEmpiricalMethodsinNaturalLanguageProcessing(EMNLP-IJCNLP), p.2567â2577. Keck,François,HenryBroadbent,andFlorianAltermatt(2025).âExtractingMassiveEcological DataonStateandInteractionsofSpeciesUsingLargeLanguageModels.âIn:bioRxiv. Lu,Pan,SwaroopMishra,TonyXia,LiangQiu,Kai-WeiChang,Song-ChunZhu,OyvindTafjord, Peter Clark, and Ashwin Kalyan (2022). âLearn to Explain: Multimodal Reasoning via ThoughtChainsforScienceQuestionAnswering.âIn:AdvancesinNeuralInformationPro- cessingSystems(NeurIPS). Nentidis, Anastasios, Georgios Katsimpras, Anastasia Krithara, Salvador Lima LĂłpez, EulĂ lia FarrĂ©-Maduell,LuisGasco,MartinKrallinger,andGeorgiosPaliouras(2023).âOverviewof BioASQ2023:TheEleventhBioASQChallengeonLarge-ScaleBiomedicalSemanticIndex- ing and Question Answering.â In:Experimental IR Meets Multilinguality, Multimodality, andInteraction.LectureNotesinComputerScience. OpenAI(2023).âGPT-4TechnicalReport.âIn:arXivpreprintarXiv:2303.08774. Ouyang, Long, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, ChongZhang,SandhiniAgarwal,KatarinaSlama,AlexRay,JohnSchulman,JacobHilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Chris- tiano,JanLeike,andRyanLowe(2022).âTrainingLanguageModelstoFollowInstructions withHumanFeedback.âIn:AdvancesinNeuralInformationProcessingSystems(NeurIPS). Palta, Shramay, Nishant Balepur, Peter Rankel, Sarah Wiegreffe, Marine Carpuat, and Rachel Rudinger (2024). âPlausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning.â In:FindingsoftheAssociationforComputationalLinguistics: EMNLP2024.Miami,Florida,USA,p.3451â3473. Pezeshkpour,PouyaandEstevamHruschka(2024).âLargeLanguageModelsSensitivitytoThe OrderofOptionsinMultiple-ChoiceQuestions.âIn:FindingsoftheAssociationforCompu- tationalLinguistics:NAACL2024,p.2006â2017. 13 BAGELBENCHMARK Poelen,JorritH.,JamesD.Simons,andChrisJ.Mungall(2014).âGlobalBioticInteractions:An OpenInfrastructuretoShareandAnalyzeSpecies-InteractionDatasets.âIn:EcologicalIn- formatics24,p.148â159. Robinson, David, Marius Miron, Masato Hagiwara, Benno Weck, Sara Keen, Milad Alizadeh, Gagan Narula, Matthieu Geist, and Olivier Pietquin (2024). âNatureLM-audio: an Audio- LanguageFoundationModelforBioacoustics.âIn:arXivpreprintarXiv:2411.07186. Speer,Robyn(2018).wordfreq:Accesstowordfrequencydatainmanylanguages.https://github. com/rspeer/wordfreq. Stevens,Samuel,JiamanWu,MatthewJ.Thompson,ElizabethG.Campolongo,ChanHeeSong, DavidEdwardCarlyn,LiDong,WasilaM.Dahdul,CharlesStewart,TanyaBerger-Wolf,Wei- Lun Chao, and Yu Su (2024). âBioCLIP: A Vision Foundation Model for the Tree of Life.â In:Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Team,Gemmaetal.(2025).Gemma3TechnicalReport.arXiv:2503.19786[cs.CL]. URL:https: //arxiv.org/abs/2503.19786. Vellinga, Willem-Pier and Robert PlanquĂ© (2015). âThe Xeno-Canto Collection and Its Rela- tiontoSoundRecognitionandClassification.âIn:WorkingNotesofCLEF2015Conference. Vol.1391.CEURWorkshopProceedings. Xu, Weijie, Shixian Cui, Xi Fang, Chi Xue, Stephanie Eckman, and Chandan K. Reddy (2025). SATA-BENCH:SelectAllThatApplyBenchmarkforMultipleChoiceQuestions.arXiv:2506. 00643[cs.CL]. Yang, An, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, ChangGao,ChengenHuang,ChenxuLv,etal.(2025).âQwen3TechnicalReport.âIn:arXiv preprintarXiv:2505.09388. Zheng,Chujie,HaoZhou,FandongMeng,JieZhou,andMinlieHuang(2024).âLargeLanguage ModelsAreNotRobustMultipleChoiceSelectors.âIn:arXivpreprintarXiv:2309.03882. 14 BAGELBENCHMARK AAppendix A.1Prompts Weusedthefollowingpromptsinthisproject. A.1.1BenchmarkGenerationPrompt Wikipediadomainbenchmarkgenerationprompt System: You are a scientific question generator. Given a Wikipedia article about a single species, generate up to eight multiple-choice questions (no free- form). Only use information explicitly mentioned in the text. Each question must: - Fall under one of these dimensions: Taxonomy, Behavior, Communication, Morphology, Habitat, Cognition, Geographic Distribution, Diet. - For Behavior, consider social behavior (e.g., solitary vs. social, group size, social structure) when present. - Have exactly four options and exactly one correct answer. - Be factual, unambiguous, and grounded in the given Wikipedia text (no outside knowledge). - Skip dimensions not covered in the text. If the text mentions vocalization and communication, include at least one question about Communication. If the text mentions where the species is found (e.g., native range, regions, countries), include at least one question about Geographic Distribution. If the text mentions feeding or diet, include at least one question about Diet. User:Species metadata: ScientificName: ScientificName CommonNames: CommonNames Taxonomy: Kingdom=Kingdom; Phylum=Phylum; Class=Class; Order=Order; Family= Family; Genus=Genus WikipediaTitle: WikipediaTitle WikipediaText: \â\â\âWikipediaText\â\â\â Instructions: Generate up to EIGHT multiple-choice questions across relevant dimensions. Return output as a JSON list under this schema: âquestionsâ: [ âdimensionâ: â<one of Taxonomy|Behavior|Communication|Morphology|Habitat|Cognition| Geographic Distribution|Diet>â, âquestionâ: â<question text>â, âoptionsâ: [â<opt1>â, â<opt2>â, â<opt3>â, â<opt4>â], âanswerâ: â<correct option>â ] Assistant: 15 BAGELBENCHMARK GloBIdomainbenchmarkgenerationprompt System: You are creating benchmark-quality self-contained multiple-choice questions. Return only one valid JSON object User: You are generating benchmark-quality multiple-choice questions for the ecological interaction domain of an animal expertise benchmark. You will be given a short source context derived from a GloBI ecological interaction record. Your task is to generate exactly ONE self-contained multiple-choice question. The benchmark is CLOSED-BOOK: - The final test-taker will NOT see the source context. - The question must stand on its own. - The question must NOT refer to âthe passageâ, âthe textâ, âthe recordâ, âthe sourceâ, or any hidden context. - Use ONLY information supported by the provided source context. Do NOT use outside knowledge. Your goal is to produce a question that tests ecological reasoning, not keyword matching. Select exactly ONE dimension that is most suitable for the source context: - Masked interaction type inference - Masked participant identification Dimension guidance: - Masked participant identification: hide one participant and ask which taxon best fills the missing ecological slot. - Masked interaction type inference: hide the relation label and ask which interaction type best fits the interaction pattern. High-value question criteria: - The item must require at least one inference step. - The correct answer should NOT be recoverable from a single cue word alone. - The species names and ecological clues should matter to solving the question. - The question should test ecological understanding, not simple paraphrase or label recognition. Anti-leakage rules: - Do NOT use verbs or phrases that directly reveal the correct interaction type, such as: âconsumesâ, âeatsâ, âpredates onâ, âparasitizesâ, âpollinatesâ, âinfectsâ, âhostsâ, âmutualistâ, â commensalâ, âparasiteâ, âpredatorâ, âpreyâ. - Do NOT use wording that is an obvious synonym or near-restatement of the correct answer. - Do NOT write a question that can be answered correctly even if the taxa are replaced with âspecies X â and âspecies Yâ. - Do NOT make the correct answer obvious from one lexical clue in the stem. - If the source context only supports a trivial label-retrieval question, reject it. For masked interaction type inference, do not ask for the interaction label directly unless the stem requires integrating multiple clues. Prefer questions where the solver must infer the interaction from ecological roles, asymmetry, or biological consequences, not from a single action verb. Requirements for the question: - Include only the clues necessary to support a non-trivial inference. - Prefer ecological reasoning over direct restatement. - Do not simply restate the source context verbatim unless necessary. - Do not overgeneralize from a single documented interaction. - For masked-inference questions, the missing field must be reasonably inferable from the remaining clues. - Avoid questions where multiple options could plausibly be correct. - Use clear scientific language. - The question should read like an ecological reasoning problem, not like a hidden database entry. Requirements for options: 16 BAGELBENCHMARK - Exactly four options labeled A, B, C, and D. - Exactly one correct answer. - Distractors must be plausible. - Distractors should reflect realistic ecological confusions, such as: - reversed interaction direction - wrong ecological role - wrong interaction type - plausible but incorrect taxon - unsupported stronger claim Reject low-value questions such as: - direct label retrieval from a cue word - simply copying a species name with no reasoning value - masked questions where the answer is not actually inferable from the clues - items where the stem already contains the answer in nearly explicit form - items that test generic ecology vocabulary rather than animal expertise Before deciding the item is usable, internally check: 1. Can the answer be solved from a single cue word or phrase? 2. Would the item still work the same if the species names were replaced with generic placeholders? 3. Is the correct option mostly a paraphrase of the stem? 4. Does the item require ecological reasoning rather than lexical matching? If the answer to any of 1-3 is yes, strongly prefer rejection. Avoid wording such as: - âdocumentedâ - ârecordedâ - âreportedâ - âobservedâ - âin the recordâ - âaccording to the sourceâ - âthe passage statesâ Return valid JSON in one of these two formats. USABLE: âusableâ: âtrueâ, âdimensionâ: â<one of: Masked participant identification | Masked interaction type inference>â, âquestionâ: â<self-contained question>â, âoptionsâ: âAâ: â<option A>â, âBâ: â<option B>â, âCâ: â<option C>â, âDâ: â<option D>â , âcorrect_answerâ: â<A|B|C|D>â, âsupporting_evidenceâ: [ â<short evidence span 1>â, â<short evidence span 2>â ], ârationaleâ: â<brief explanation grounded in the source context>â REJECTION: âusableâ: âfalseâ, ârejection_reasonâ: â<one of: too_sparse | answer_not_self_contained | low_reasoning_value | masked_field_not_inferable | lexical_leakage>â 17 BAGELBENCHMARK SOURCE CONTEXT: Interaction record Assistant: BioRxivdomainbenchmarkgenerationprompt System: You are creating benchmark-quality self-contained multiple-choice questions. Return only one valid JSON object. User: You are creating benchmark-quality multiple-choice questions from a bioRxiv paper. The final benchmark model will see ONLY the question and answer options. Therefore, the question itself must include all necessary context for answering. Read the paper and generate exactly ONE self-contained 4-choice multiple-choice question. Use ONLY information from the paper. Do NOT use outside knowledge. Only generate the following type of dimension questions: - Result Interpretation â interpreting experimental results Requirements: - The question must be fully self-contained. - Include the minimum necessary scientific context inside the question stem so that the question is answerable without access to the paper. - The question stem may be longer than usual, but it must remain concise and focused. - The question stem must include setup and evidence when needed, but must NOT directly restate the correct answer or final conclusion in answer-like wording. - Do NOT create a question that can be answered only by having seen the paper. - The question must test scientific reasoning, not trivial word matching. - Distractors must be scientifically plausible and belong to the same conceptual category as the correct answer. - Provide exactly four options labeled A, B, C, and D. - Exactly one option must be correct. - Do not mention the dimension in the question text. - Do not include unsupported information. Question-writing guidance: - Prefer a stem that includes: 1. the study system or biological setting, 2. the relevant manipulation, comparison, or observation, 3. the key evidence or result pattern, 4. the reasoning task. - Avoid copying long sentences directly from the paper. - Lightly rewrite the source information into a natural, self-contained scientific scenario. - Do not make the stem so compressed that the correct answer becomes obvious from wording alone. Before finalizing, check: - Can a model answer the question from the stem alone? - Does the stem avoid directly revealing the correct answer? - Are the options plausible and same-category? - Is the question testing reasoning rather than recall of the paper? Return valid JSON in this schema: âdimensionâ: âResult Interpretation â, âquestionâ: â<self-contained question stem>â, âoptionsâ: âAâ: â<option A>â, 18 BAGELBENCHMARK âBâ: â<option B>â, âCâ: â<option C>â, âDâ: â<option D>â , âcorrect_answerâ: â<A|B|C|D>â, âsupporting_evidenceâ: [ â<short copied or lightly trimmed evidence span 1>â, â<short copied or lightly trimmed evidence span 2>â ], âreasonâ: â<brief explanation of why the correct answer is supported by the paper>â PAPER: paper Assistant: Xeno-cantodomainbenchmarkgenerationprompt System: You are an expert in bioacoustics. Use the provided spectrogram only as background to infer general, species-level acoustic facts. Write ONE multiple-choice question about the specified TOPIC for this âspecies vocalizations. Do NOT mention or refer to any image/recording/ spectrogram in the question. The question must read like a factual quiz item about typical acoustic properties, not about a specific recording. CRITICALLY: include the VOCALIZATION TYPE explicitly in the question text (e.g., âsongâ, âcallâ, âflight callâ). User: Animal: common_name ( scientific_name). Vocalization type: Vocalization type. Topic: topic_label. Guidance: guidance Use appropriate units (units_hint). Provide 4 options (âAD) with plausible distractors; exactly one correct. Include approximate numeric ranges where applicable. Do NOT reference any image, recording, or spectrogram. IMPORTANT: The question MUST explicitly mention the vocalization type âvtâ. Examples of reasonable value styles: examples. Assistant: A.1.2EvaluationPrompt Evaluationprompt You are answering a multiple-choice closed-book benchmark question for testing animal expertise. Choose exactly one answer. Output exactly one capital letter: A, B, C, or D. Do not output any explanation, words, punctuation, or extra text. Question: question Options: A. option_a B. option_b C. option_c D. option_d Answer: BExampleBenchmarkItems 19 BAGELBENCHMARK B.1Wikipedia Exampleitem Dimension:Taxonomy Level:easy Question:WhatisthescientificnameoftheAndeanflamingo? Options: A.Phoenicopterusandinus B.Phoenicoparrusandinus C.Phoenicopteruschilensis D.Phoenicoparruschilensis Goldanswer:Phoenicoparrusandinus Exampleitem Dimension:GeographicDistribution Level:easy Question:WhereistheAndeanflamingonativeto? Options: A.TheAmazonrainforest B.TheAndesmountainsofSouthAmerica C.ThecoastalregionsofChile D.TheplainsofArgentina Goldanswer:TheAndesmountainsofSouthAmerica Exampleitem Dimension:Diet Level:easy Question:WhatistheprimaryfeedingstrategyoftheAndeanflamingo? Options: A.Carnivorouspredation B.Filterfeeding C.Scavenging D.Grazingongrass Goldanswer:Filterfeeding Exampleitem Dimension:Behavior Level:easy Question:How do Andean flamingos adapt their foraging behavior when grouped with otherflamingospecies? Options: A.Theyforagealoneregardlessofgroup B.Theyadopttheforagingpatternsofthespeciestheyaregroupedwith C.Theyonlyforageatnight D.Theystopforagingaltogether Goldanswer:Theyadopttheforagingpatternsofthespeciestheyaregroupedwith 20 BAGELBENCHMARK Exampleitem Dimension:Communication Level:medium Question:WhatisoneofthedistinctcalltypesoftheAndeanflamingo? Options: A.Chirp B.Peep C.Whistle D.Growl Goldanswer:Peep Exampleitem Dimension:Morphology Level:easy Question:WhatdistinguishestheAndeanflamingofromotherflamingospecies? Options: A.Itsbrightredplumage B.Itsyellowlegsandthree-toedfeet C.Itslongneck D.Itsabilitytoflyathighaltitudes Goldanswer:Itsyellowlegsandthree-toedfeet Exampleitem Dimension:Habitat Level:easy Question:In which type of environment do Andean flamingos primarily live during the summer? Options: A.Forests B.Saltlakes C.Grasslands D.Deserts Goldanswer:Saltlakes Exampleitem Dimension:Cognition Level:easy Question:Whatcognitiveabilityisnotedinkearegardingproblem-solving? Options: A.Theycanmemorizesongs B.Theycansolvelogicalpuzzles C.Theycanmimichumanspeech D.Theycannavigateusingstars Goldanswer:Theycansolvelogicalpuzzles 21 BAGELBENCHMARK B.2GloBI Exampleitem Dimension:Maskedparticipantidentification Level:hard Question:WhichtaxonislikelytobethehostintheinteractionwherePearsonemaplicais anendoparasite? Options: A.Procyonlotor B.Ursusamericanus C.Canislupus D.Lynxrufus Goldanswer:Procyonlotor (VerbatimfromthereleasedBAGELGloBIJSONL;underlyingGloBIrecord:Pearsonemaplica endoparasiteofProcyonlotor.) Exampleitem Dimension:Maskedinteractiontypeinference Level:easy Question:Inanecologicalinteractionwhereatreespeciesisnegativelyaffectedbyafungal organism,whattypeofinteractionismostlikelyoccurring? Options: A.Thetreeisprovidingnutrientstothefungus. B.Thefungusisbenefitingattheexpenseofthetree. C.Thetreeisbenefitingthefunguswithoutanyharm. D.Thetreeandfungusaremutuallybenefitingeachother. Goldanswer:Thefungusisbenefitingattheexpenseofthetree. B.3bioRxiv Exampleitem Dimension:ResultInterpretation Level: easy Question:InastudyoftheseaanemoneNematostellavectensis,researchersfoundthatani- malskeptinenvironmentswithgravelsubstrateproducedsignificantlymoreclonalprogeny throughtransversefissioncomparedtothosewithoutsubstrate. Giventhatthepresenceof substrateisshowntoenhancefissionrates,whatcanbeinferredabouttheroleofsubstrate intheasexualreproductionofNematostellavectensis? Options: A.Substrateprovidesamechanicaladvantagethatfacilitatesthephysicalprocessoffission- ing. B.Substrateincreasesthegeneticdiversityoftheclonesproducedduringfission. C.Substratereducesthemetabolicwastethatinhibitsfissioninhigh-densitypopulations. D.SubstratealtersthehormonalbalanceinNematostella,promotingfastergrowth. Goldanswer: Substrateprovidesamechanicaladvantagethatfacilitatesthephysicalpro- cessoffissioning. 22 BAGELBENCHMARK B.4Xeno-canto Exampleitem Dimension:modulationpattern Level:hard Question:What common frequency modulation (FM)patternis typicallyobserved in the songoftheRedWarbler(Cardellinarubra)? Options: A.Risingglide B.Sinusoidalvibrato C.Fallingglide D.Trill Goldanswer:Risingglide Exampleitem Dimension:dominantfrequencyrange Level:medium Question:What is the typical dominant frequency range of thesongof the Red Warbler (Cardellinarubra)? Options: A.âŒ9â11kHz B.âŒ3â5kHz C.âŒ6â8kHz D.âŒ1â2kHz Goldanswer:âŒ3â5kHz Exampleitem Dimension:call/syllableduration Level:easy Question:WhatisthetypicaldurationofasinglecallfortheRedWarbler(Cardellinarubra)? Options: A.80â150ms B.200â400ms C.0.8â1.2s D.1.5â2.0s Goldanswer:80â150ms Exampleitem Dimension:harmonicstructure/tonalityvs. broadband Level:hard Question:ArethecallsoftheRedWarbler(Cardellinarubra)typicallycharacterizedastonal withharmonicsorbroadband/noisy? Options: A.Strongharmonics B.Weakharmonics C.Broadband D.Tonalwhistle Goldanswer:Strongharmonics CXeno-cantogenerationpipelinedemonstrationplot Figure3showsthepipelinetocreatetext-basedbioacousticknowledgeMCQs. 23 BAGELBENCHMARK Question: What is the typical dominant frequency range of the call of the Western Corella (Cacatua pastinator)? Options: A: 0.5â1 kHz B: 2â4 kHz C: 4â6 kHz D: 10â12 kHz Answer: B LLM (gpt-4o-mini) Figure3. Xeno-cantoquestiongenerationpipeline. Theprocessinvolvesfeedingalog-frequencyspec- trogramofabioacousticrecordingintoGPT-4o-minitogeneratestructured, multiple-choicequestions basedonvisualacousticfeatures. DAccuracybyquestiontypewithineachsourcedomain Table6showsthebreakdownaccuracyofallthemodelsacrossBehavior,Cognition,Commu- nication,Diet,GeographicDistribution,Habitat,MorphologytypequestionsontheWikipedia domain. Table7showsthebreakdownaccuracyofallthemodelsacrossDominantfrequency rangeandCall/syllableduration,Harmonicvsbroadband,andModulationtypeoftext-based bioacoustic knowledge questions in the Xeno-canto domain. Table5shows the breakdown accuracyamongMaskedinteractiontypeinferenceandmaskedparticipantidentificationtype questionsintheGloBIdomain. EQwen3familyperformanceonBAGEL Figure4showstheQwen3familyaccuracycurveonBAGEL. FAnswer-positionbiasintheinitialrelease AfterinspectingmodeloutputsontheinitialBAGELrelease(beforeoptionshuffling),wefound substantialanswer-positionskewinthebenchmarkitself.Table 9showsthatthecorrectoption wasoftenconcentratedinasinglepositionwithinadomain,despitethemultiple-choiceformat. Thisskewappearstohavecreatedashortcutthatsomesmallmodelscouldexploitundergreedy decoding. Table 10reportsthedistributionaftertheoption-shufflingmitigationdescribedat theendofSection3. We then examined whether models showed corresponding output collapse or positional preference.Table 11reportsthedistributionoffirst-optionlettersemittedbyfiverepresentative openmodelsunderthesameseed-0greedy-decodingsettingusedinthemainpaper. Qwen3- 0.6B collapses almost entirely to option A across all four domains, while smolLM2-360M col- lapsesprimarilytooptionB.Largermodelsdonotcollapseascompletely,butseveralstillex- hibitstrongpositionalskew, especiallytowardoptionBonWikipediaandXeno-canto. Their apparentaccuracyinsomedomainsthereforetracksthedominantgold-optionpositionrather thannecessarilyreflectinggenuinedomaincompetence. Thisartifactisimportantforinterpretingtheinitialleaderboard.Inparticular,strongscores fromsmallmodels onsomedomains shouldnotbe readasdirect evidenceofrobust animal- 24 BAGELBENCHMARK ModelMaskedinteractiontypeinference Maskedparticipantidentification smolLM20.2580.233 Qwen3-0.6B0.2540.250 Qwen3-4B0.8010.257 Gemma-7B0.5910.312 Mistral-7B-Instruct-v0.30.7570.374 Llama3.1Instruct-8B0.8080.378 Qwen3-8B0.8140.420 Phi-40.8800.488 Qwen3-14B0.8500.459 Gemma327BIT0.8880.530 Qwen3-32B0.8550.481 ClaudeOpus4.60.921 0.640 GPT-5.40.9090.561 Table5. AccuracybyquestiontypeontheGloBIsubsetofBAGEL.Open-weightmodelsareorderedby increasingapproximateparametercount. Bestvalueineachcolumnisunderlined. ModelBehavior Cognition Communication Diet GeographicDistribution Habitat Morphology Taxonomy smolLM20.2270.2500.2180.2360.2420.2220.2790.239 Qwen3-0.6B0.2460.1790.2390.2360.2460.2560.2920.272 Qwen3-4B0.6350.6070.6500.7200.7540.8120.6050.578 Gemma-7B0.5240.5710.5380.6930.7510.8460.5450.679 Mistral-7B-Instruct-v0.3 0.6060.6430.6790.7360.7860.8380.5410.616 Llama3.1Instruct-8B0.6770.6430.6920.7770.7900.8590.6650.705 Qwen3-8B0.7220.6430.7310.8380.8260.8970.6910.672 Phi-40.7370.857 0.7520.8450.8540.9320.7040.761 Qwen3-14B0.7220.8210.7180.8450.8750.9230.7340.750 Gemma327BIT0.7000.7500.7650.8950.8720.9230.7080.769 Qwen3-32B0.7420.6070.7740.8650.8790.9360.7120.799 ClaudeOpus4.60.870 0.8570.8970.9390.9890.9830.8930.963 GPT-5.40.8190.7860.8500.9120.9790.9700.8630.929 Table6. AccuracybyquestiontypeontheWikipediasubsetofBAGEL.Open-weightmodelsareordered by increasing approximate parameter count. Best value in each column is underlined; ties are jointly underlined. knowledgecompetencewhenthedominantoutputletteralignswiththedominantgold-answer position.Infollow-upexperiments,wethereforeshuffledansweroptionstoremovethisbench- markshortcut. GIllustrativemanualreviewonGloBI. We manually sampled GloBI items to characterize recurring construction issues; this is not a full relabeling or prevalence study, and we do not quantify rates here. We observed three il- lustrative buckets. (i)Multi-plausibleanswersandtaxonomicoverlap.For example,globi:2:0 asks for a likely host of an endoparasite without restating the host from the source interac- tion;broadhostrangescanmakemorethanonespecies-leveloptiondefensible.globi:205:0and globi:525:0canpairaconcretehostspecieswithahighertaxon,somultipleoptionsareplausi- bledependingonwhetherâtaxonâisreadatspeciesorhigherrank. globi:207:0 and globi:303:0 mixspeciesandsubspeciesnamesforthesamelineage, creatingtaxonomy-granularitycolli- sions. (i)Under-specifiedinteractionprompts.Flower-visitation templates such asglobi:21:0, globi:420:0,globi:736:0,globi:785:0, andglobi:1436:0can admit more than one interaction la- 25 BAGELBENCHMARK ModelDominantfreq.range Call/syll.dur. Harmonicvs.broadband Modulation smolLM20.2480.2570.2690.236 Qwen3-0.6B0.2370.2710.2250.239 Qwen3-4B0.3750.3420.4030.227 Gemma-7B0.2700.3140.3210.241 Mistral-7B-Instruct-v0.30.2710.3270.3770.317 Llama3.1Instruct-8B0.5840.3250.4640.305 Qwen3-8B0.4270.3220.2970.382 Phi-40.4430.3080.2860.205 Qwen3-14B0.4420.4210.3230.301 Gemma327BIT0.5210.2600.3890.325 Qwen3-32B0.4030.2900.3930.303 ClaudeOpus4.60.6710.4350.2860.310 GPT-5.40.7110.4390.3510.301 Table7. AccuracybyquestiontypeontheXeno-cantosubsetofBAGEL.Columnsfollowthedimension labelsinTable2(harmonicstructure/tonalityvs.broadband;modulationpattern). Open-weightmodels areorderedbyincreasingapproximateparametercount. Bestvalueineachcolumnisunderlined. Sourcedomain Items MeanZipf Raretok.(%) Bio-termhits/stem Wikipedia1,9275.3311.80.05 Xeno-canto4,2425.0418.42.03 Table8. LexicalstatisticsofEnglishquestionstems(optionsexcluded), computedwithwordfreq(Speer, 2018). HigherZipfindicatesmorefrequentunigramsingeneralEnglish. âRaretok.â isthemeanfraction of stem tokens with Zipf<3. âBio-term hitsâ counts tokens from a fixed bioacoustic keyword list (e.g., harmonic,khz,modulation). bel when âvisits flowersâ alone does not pin down pollination outcomes;globi:47:0mixes relationship-type wording with mechanistic phrasing so two options can appear simultane- ouslyacceptableforaparasiteâhostassociation. (i)Record-specifickeys.Somehostitemskey toaparticularinteractionrecordevenwhenalternativehostsareplausiblewithoutthesource string,soperformancepartiallyreflectsrecordrecoveryratherthanspeciesbiologyalone. Our two-modelâhardâstratumshouldthereforebereadasarelativedifficultysignalundergener- ative MCQ noise rather than a guarantee of a single objectively correct option in every case; futurereleasesshouldincludeadditionaladjudication,taxonomicde-duplication,andtighter promptconstraintswhereappropriate. 26 BAGELBENCHMARK 0.6B4B8B14B32B Qwen3 instruct checkpoint (parameters) 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy BAGEL accuracy versus Qwen3 model size (seed 0, greedy) GloBI Wikipedia bioRxiv Xeno-canto Overall Figure4. Qwen3instructfamilyonBAGEL(seed0,greedy): accuracyversusmodelsizeforeachsource domainandOverall. GloBI,Wikipedia,andbioRxivgenerallyimproveupto32B,whileXeno-cantopeaks at14Banddropsat32Bunderourprotocolâthesamenon-monotonicitydiscussedinthetext. DomainABCD bioRxiv659(30.2%) 1,143(52.4%) 325(14.9%) 56(2.6%) GloBI2,135(61.0%) 658(18.8%) 633(18.1%) 74(2.1%) Wikipedia199(10.3%) 1,244(64.6%) 445(23.1%) 39(2.0%) Xeno-canto 618(14.6%) 3,231(76.2%) 333(7.8%) 60(1.4%) Table9. Correct-answer position distribution in the initial BAGELrelease before option shuffling. Per- centagesarecomputedovertheevaluatedsubsetusedbythebenchmarkloader. DomainABCD bioRxiv519(23.8%) 556(25.5%) 551(25.2%) 557(25.5%) GloBI883(25.2%) 877(25.1%) 869(24.8%) 871(24.9%) Wikipedia487(25.3%) 479(24.9%) 490(25.4%) 471(24.4%) Xeno-canto 1,034(24.4%) 1,062(25.0%) 1,053(24.8%) 1,093(25.8%) Alldomains 2,923(24.7%) 2,974(25.1%) 2,963(25.0%) 2,992(25.2%) Table 10. Correct-answer position distributionafteroption shuffling, counting individual multiple- choiceitemsinthereleasedevaluationfiles(n=11,852). 27 BAGELBENCHMARK ModelDomainABCD Qwen3-0.6BbioRxiv2,178(99.8%)3(0.1%) 1(0.0%) 1(0.0%) Qwen3-0.6BGloBI3,500(100.0%)000 Qwen3-0.6BWikipedia 1,927(100.0%)000 Qwen3-0.6BXeno-canto 4,242(100.0%)000 smolLM2-360MbioRxiv1(0.0%) 2,168(99.3%) 14(0.6%)0 smolLM2-360MGloBI138(3.9%) 3,276(93.6%) 85(2.4%) 1(0.0%) smolLM2-360MWikipedia85(4.4%) 1,517(78.7%) 323(16.8%) 2(0.1%) smolLM2-360MXeno-canto1(0.0%) 4,225(99.6%) 16(0.4%)0 Qwen3-8BbioRxiv731(33.5%) 1,056(48.4%) 319(14.6%) 77(3.5%) Qwen3-8BGloBI1,640(46.9%) 681(19.5%) 757(21.6%) 422(12.1%) Qwen3-8BWikipedia432(22.4%) 989(51.3%) 398(20.7%) 108(5.6%) Qwen3-8BXeno-canto 1,367(32.2%) 1,965(46.3%) 741(17.5%) 169(4.0%) Llama3.1Instruct-8B bioRxiv716(32.8%) 1,065(48.8%) 331(15.2%) 71(3.3%) Llama3.1Instruct-8B GloBI1,655(47.3%) 812(23.2%) 829(23.7%) 204(5.8%) Llama3.1Instruct-8B Wikipedia453(23.5%) 1,016(52.7%) 373(19.4%) 85(4.4%) Llama3.1Instruct-8B Xeno-canto 1,131(26.7%) 2,818(66.4%) 229(5.4%) 64(1.5%) Table11.DistributionofemittedoptionlettersintheinitialBAGELreleaseforrepresentativeopenmodels underseed-0greedydecoding. Smallermodelsexhibitnear-completecollapsetoasingleoption,while strongermodelsstillshowsubstantialpositionalskewinseveraldomains,especiallyWikipediaandXeno- canto. 28