Paper deep dive
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports
Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, Lorenz K. MĂŒller, Konstantin Berestizshevsky, Lukas Cavigelli
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent framework for the automated generation of grounded, multimodal technical reports. Wyvern allows for the generation of multimodal outputs, integrating images, tables, and text with supporting references in a unified report. Additionally, a particular focus is placed on the grounding of the content, with the implementation of a claims auto-revision stage. We conduct a human evaluation study to assess the quality of our proposed framework. The results show that the figures' informativeness is perceived as superior to that of a recent baseline in 87% of cases. Furthermore, Wyvern's reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances. We also carry out automatic evaluations showing that Wyvern gains up to 2.3$\times$ in citation recall and 1.6$\times$ in citation precision with respect to the baselines.
Tags
Links
- Source: https://arxiv.org/abs/2608.14446v1
- Canonical: https://arxiv.org/abs/2608.14446v1
Trouble viewing inline? Open PDF directly â
Full Text
65,169 characters extracted from source content.
Expand or collapse full text
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports Beatrice Alessandra Motetti 1,2 , Emilien Guandalino 2 , Daniele Jahier Pagliari 1 , Alessio Burrello 1 , Lorenz K. MĂŒller 2 , Konstantin Berestizshevsky 2 , Lukas Cavigelli 2 1 Politecnico di Torino, Italy 2 Computing Systems Lab, Huawei Research, Switzerland Correspondence: beatrice.motetti@polito.it Abstract In the current artificial intelligence-driven in- novation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to syn- thesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent frame- work for the automated generation of grounded, multimodal technical reports. Wyvern allows for the generation of multimodal outputs, inte- grating images, tables, and text with supporting references in a unified report. Additionally, a particular focus is placed on the grounding of the content, with the implementation of a claims auto-revision stage. We conduct a hu- man evaluation study to assess the quality of our proposed framework. The results show that the figuresâ informativeness is perceived as su- perior to that of a recent baseline in 87% of cases. Furthermore, Wyvernâs reports are rated as more useful than those produced by three al- ternative methods in 63% to 100% of instances. We also carry out automatic evaluations show- ing that Wyvern gains up to 2.3Ăin citation recall and 1.6Ăin citation precision with re- spect to the baselines. 1 Introduction The volume of newly published content, spanning from scientific literature to general webpages, is rapidly growing, particularly in fields such as artifi- cial intelligence (Bornmann et al., 2021). Keeping up with the vast quantity of heterogeneous domain- specific information has become extremely criti- cal and complex simultaneously, especially for re- searchers (Park et al., 2023; Jones, 2009). In this context, automated synthesis and summarization tools can act as first-pass filters that retrieve, aggre- gate and condense large and relevant information into technical reports, helping researchers and end- users to stay informed and keep up in fast-evolving fields. While these tools cannot substitute criti- Web search âąInclude figures and overview table Report generation User's input topic Grounded multimodal report References [1] [2] [3] âąRevise claims wrtcited references Claims grounding yvern Figure 1: Overview of the main functionalities of Wyvern cal human judgment, they can ease the burden to efficiently retrieve and manage information. Several frameworks have been proposed re- cently to automate knowledge retrieval and its or- ganization into machine-generated textual docu- ments (Shao et al., 2024; Wang et al., 2024; Jiang et al., 2024). Furthermore, as multimodality allows for more engaging content (Fu et al., 2022), recent works have also considered the inclusion of rele- vant images in the output reports to increase their informativeness (Yang et al., 2025, 2026). A key requirement for these frameworks is the ability to ground generated statements on retrieved evidence, thereby increasing user trust. Traceability between claims and sources is fundamental, as it enables to explore with more depth the aspects of interest. Building on these observations, we propose Wyvern, a multi-agent framework for the automatic generation of multimodal technical reports on user- selected topics, with a particular focus on content groundedness. Our Wyvern framework capitalizes on a structured workflow of specialized Large Lan- guage Model (LLM)-based agents, each focused on a narrow task, such as refining the quality of previously acquired information. Organizing these agents into a coordinated framework significantly enhances reliability, modularity, and overall qual- ity of the produced content, as recent works have demonstrated (Huang et al., 2024; He et al., 2025). 1 arXiv:2608.14446v1 [cs.AI] 14 Aug 2026 As illustrated in Figure 1, Wyvern comprises var- ious modules that search and process relevant in- formation from the web, produce the multimodal report, and verify whether the textual statements are supported by the previously retrieved evidence. We evaluate Wyvern through a human evaluation study, comparing its generated reports with those produced by STORM (Shao et al., 2024), Web- Thinker (Li et al., 2025) and WikiAutoGen (Yang et al., 2025). The results show that Wyvern pro- duces reports whose figures and factuality are more positively evaluated by the end-users. We also conduct an automatic evaluation, and examine the predictive power of different evaluation means with respect to human judgment. Our key contributions can be summarized as follows: âąWe propose Wyvern, a multi-agent framework for the generation of grounded, multimodal technical reports, retrieving the most relevant information from the web. As key novelties: âwe design a module for the retrieval, se- lection, and positioning of the most infor- mative images within the textual report; âwe build a claims auto-revision stage to verify the grounding of the statements on the collected evidence. âąWe assess the quality of Wyvern-generated reports through a human evaluation study and we compare these results with automatic eval- uations, to analyze the predictive performance of the latter for human judgments. Code and data are available athttps://github .com/huawei-csl/wyvern. 2 Related Work 2.1 Automated long-form expository writing The generation of long text by LLMs is a complex task, due to their limited context length and the need to maintain semantic consistency (Wang et al., 2024; Yang et al., 2025). To overcome the limita- tions of static parametric knowledge, a common approach is to integrate Retrieval-Augmented Gen- eration (RAG) strategies (Lewis et al., 2020), often combined with web search engines (Shao et al., 2024; Li et al., 2025), to perform an up-to-date and comprehensive information retrieval. Many recent works on long-form expository writ- ing adopt a top-down approach, with an initial Method Images insertion Claims verification STORM (Shao et al., 2024)â WebThinker (Li et al., 2025)â WikiAutoGen (Yang et al., 2025)ââ Wyvern (ours)â Table 1: Overview of expository text generators. Wyvern integrates multimodality and explicit claims verification and revision into a unified framework. outline definition and a subsequent section expan- sion phase (Shao et al., 2024; Jiang et al., 2024; Yang et al., 2025; Wang et al., 2024; Yang et al., 2026); others employ a more dynamic planning strategy that builds upon recursive task decomposi- tion principles (Li et al., 2025; Xiong et al., 2025). Specifically, STORM (Shao et al., 2024) and Co- STORM (Jiang et al., 2024) employ ensembles of multi-perspective agents with a question-and- answer approach for information retrieval and out- line planning tasks, generating Wikipedia-like ar- ticles. AutoSurvey (Wang et al., 2024) tackles the problem of creating literature surveys with a four- stage framework comprising information retrieval and outline definition, and subsections drafting, in- tegration, and evaluation. WebThinker (Li et al., 2025) integrates the capabilities of performing web searches and report writing and editing within the reasoning of Large Reasoning Models. Another emerging line of works tackles the gen- eration of multimodal reports, incorporating re- trieved images within the textual body (Yang et al., 2025, 2026). In WikiAutoGen (Yang et al., 2025) an agent proposes candidate positions and descrip- tions of images, then retrieved from the web. Mul- timodal DeepResearcher (Yang et al., 2026) uses instead a structured textual representation of charts to allow for the direct generation of visualizations, by using a LLM to produce JavaScript code. We report in Table 1 an overview of the features of the main works. 2.2 Grounding Measuring and improving the factuality of text gen- erated by LLMs is an active area of research, with two main directions (Jacovi et al., 2025): one de- fines factuality as text entailment with respect to cited sources; the other verifies claims against ex- ternal ground-truth answers. Additionally, various definitions of claim unit exist, with some works identifying it at a sentence level (Gao et al., 2023; Shao et al., 2024; Balepur et al., 2023), and others 2 at a finer granularity by decomposing sentences into atomic facts (Min et al., 2023; Li et al., 2024). FActSCORE (Min et al., 2023) decomposes sen- tences into atomic facts, to overcome the chal- lenge of assigning a binary entailment label to sen- tences with multiple pieces of information. Self- Checker (Li et al., 2024) first decomposes the text into claims, and then generates queries to re- trieve relevant content to verify them. Chain-of- Verification (Dhuliawala et al., 2024) mitigates hal- lucinations in baseline LLMs responses by generat- ing a list of verification questions for the key claims, and then revising the responses via self-validation using the modelâs parametric knowledge. Gao et al. (2023) measure groundedness with re- spect to the cited sources by computing the citation recall and precision over sentence-level statements. Differently, we define the citation quality metrics on atomic claims derived from the full-sentences, avoiding cases of partial support. 3 Proposed framework Wyvern comprises distinct modules that operate sequentially on individual macro-tasks (see Fig- ure 2). Each of the modules handles a specific phase of the report generation by relying on ensem- bles of agents. In particular, Wyvern integrates: (i) a search module, that is tasked with the retrieval and processing of relevant information; (i) a report generation module, that writes the report and inte- grates it with the most informative figures from the retrieved documents, and with an overview table; (i) a grounding module, which examines all the statements of the report with the aim of improving its factuality and limiting the hallucinations and the unsupported statements. In the following, we refer to each agent involved in our framework by the subscript number next to its icon in Figure 2. 3.1 Search phase The goal of the search phase is to acquire infor- mation from the web and construct a knowledge database consisting in a set of referencesRand their corresponding content summariesR S . The search begins with the user providing Wyvern with a topic title, optionally specifying aspects of inter- est or non-interest and a list of relevant URLs. Af- terwards a LLM generates a refined topic title and a detailed description of it, and the URLs are directly incorporated into the initial references base R. Wyvernâs agent 1 then leverages the topic title and description to generatensearch queries. For each search query, it retrieves the top-k 1 URLs from a search engine. The type of webpage content determines the selection of the most suitable parser, which is used to extract the URLâs textual content into Markdown format. To assure content quality and processability, only documents whose length falls within a predefined token range are retained. This removes both empty pages and documents that are too large for the LLMâs context length. To refine the set of documents to be included in the knowledge database, Wyvern employs an ensemble of independent agents that analyze the re- trieved information. Agent 2 assigns a score on a 10-point scale to assess the completeness of the ex- tracted Markdown content for each webpage, in par- allel. Documents with a completeness score below a predefined thresholdλ comp are excluded from the references base, thereby preserving content quality by removing the documents for which the parser failed to extract meaningful text. Agent3evalu- ates, for each webpage, the need for an additional search step to retrieve additional relevant material. Similarly to agent2, it assigns a score on a 10- point scale to each document, estimating whether additional informative content can be fetched with a new search. If the score exceeds a set threshold λ search , agent4generates a candidate search query, which is used to fetch the most relevant URL. This URL is then considered for inclusion in the refer- ence base, following the parsing and completeness evaluation procedure carried out by agent2. To broaden the coverage of relevant aspects, Wyvern analyzes the embedded links within the parsed documents. This is done in a two-stage pro- cess. First, the embedded hyperlinks are extracted from each collected resource by agent 5 , and the relative links are resolved to their absolute versions. Afterwards, for each document, agent6is pro- vided with the textual content, the list of extracted hyperlinks, and the user-provided topic of interest, and is tasked with assigning a relevance score to each link. Links scoring below a set thresholdλ link are considered out of scope, while the top-k 2 by score are processed, following the same content extraction process as the initial documents. This iterative exploration process enables Wyvern to dy- namically enrich the collected knowledge base. To minimize redundancy in the reference base, a deduplication step is carried out to remove the most similar documents (e.g., earlier drafts or alternative versions of the same paper or webpage). Specifi- 3 Search Initial retrieval Search queries generation Additional search Document completeness eval. Filtering Deduplication step Outline Two-step outline generation Images Description generation Refinement Section-by-section refinement Report generation Sections Parallel section expansion Overview table Table building Retrieval via search engine & parsing Additional search eval. & search query generation Embedded links extraction & scoring Summarization Relevance filtering Grounding Text decomposition into atomic claims Claims revision Grounding check Two-stage grounding assessment wrt references Revision/removal of the unsupported claims References âąTopic title âąTopic description [optional] Figures selection & placement Insertion Placement Insertion Agent n (with Chain-of-Thought reasoning LLM as base model) n m Agent m Auxiliary tool âą Final report Report draft 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 Headings revision Headings refinement wrt the new text version 22 Figure 2: Overview of the key components of Wyvern, our proposed framework. The search module (left) retrieves relevant documents from the web. The report generation module (middle) handles of the writing of the reportâs text and of the figures insertion. The grounding module (right) verifies if the claims are supported by the cited references, and possibly revises them. cally, the Jaccard similarity index between all pairs of documents is computed. For each pair exceeding a similarity thresholdλ sim , only the longer docu- ment is kept in the final reference base R. Once the reference base is finalized, agent 7 summarizes each documentd i â R, thus creating the set of reference summariesR S that will be used for report generation. Eventually, agent 8 , based on a Chain-of-Thought reasoning LLM, conducts a final filtering step to detect semantic outliers in R S , i.e. documents that are off-topic or partially misaligned with the original topic description. The agent considers all the summaries of the references simultaneously, to infer the underlying domain, overcoming the limitation of relying only on the LLMâs static knowledge about the input topic. In this way, documents which are not aligned to the topic can be flagged as outliers and excluded from bothR S andR, allowing for more focused content. 3.2 Report generation 3.2.1 Outline creation and sections expansion Similarly to prior works (Shao et al., 2024; Wang et al., 2024), the generation of the textual content of the report follows a two-stage process, compris- ing outline creation and section expansion. The adopted approach is inspired by the exploration- exploitation trade-off. The set of summarized refer- encesR S is used by agent9to generateN o draft outlinesO 1 ,...,O N o to maximize topic cover- age. Agent10then merges the generated outline versions to produce a final outlineOincluding the most informative aspects of the explored options. Once the outlineOis produced, agent11gener- ates the text of each sectionsâ Osimultaneously, Document 1 ... Document 2 Description generation 12 ïŒSelection of the most relevant figures ïŒAssignment of each figure to the most compatible section For each section: ïŒImage positioning within the section ïŒText revision for seamless image- text integration 14 Figure 1 âąCaption âąTe x t s n i p p e t s âąTitle âąSummary Figure 1 Figure 1 âąCaption âąTe x t s n i p p e t s âąTitle âąSummary Figure N 1 Figure 1 âąCaption âąTe x t s n i p p e t s âąTitle âąSummary Figure N 2 ... 13 Description generation 12 Description generation 12 Figure 3: Structure of the figures selection and inser- tion routine. The pipeline comprises the generation of the descriptions of all the images extracted from the sources, the selection of the most relevant figures, and their integration in the report. with full access to the references summaries setR S . For each first-level heading, the agent writes the section and its associated subsections, following the outlineO. The agent is thus responsible for selecting the most salient content and placing ci- tations to the supporting references. Eventually, the concatenation of all the expanded sections is performed to obtain the draft of the report. 3.2.2 Image insertion Once the textual content of the report is generated, the workflow shifts to selecting, placing, and inte- grating figures within the report draft. The pipeline is illustrated in Figure 3. The pool of candidate images, denoted asF, comprises all the figures extracted from source doc- uments. Agent12is tasked with generating a de- scriptiond f for each candidate figuref â F. Only textual elements are used for the elaboration of the detailed depiction. Specifically, the agent retrieves the reference document title, its summaryr s â R S , 4 the caption of the figure within the full-text docu- mentr â R, and any in-text excerpt referring to the figure identifier extracted from the caption. Subsequently, reasoning agent 13 evaluates all the figuresâ descriptions together, in conjunction with the report outlineO, and selects the most rel- evant and informative images. If the length of the concatenation of the descriptions exceeds the con- text length of the LLM, this process is conducted hierarchically. The agent provides as output the set Ë F of relevant figures, and assigns each of them to the most pertinent section of the report. Then, reasoning agent14carries out the place- ment of the figures, independently for each section. Given the set of figures Ë F s â Ë Fassigned to section s, the agent has to assess their individual relevance for the specific section text, and images deemed as potentially out of scope are discarded. The key difference with respect to the selection performed a priori by agent13is that the reasoning agent has access to the section text, thus has more elements to filter out images which do not convey useful infor- mation for the textual flow. Eventually, agent 14 completes the images insertion by generating their captions and revising the sectionâs text to integrate the figures seamlessly within the narrative. 3.2.3 Overview table insertion To further enrich the reportâs content, Wyvern in- serts an overview table summarizing the key con- cepts. First, reasoning agent15analyzes the refer- ences setR S , and extracts shared and informative aspects to differentiate among the relevant works. Based on this analysis, it builds a comparative table, along with a detailed explanation of its contents. Subsequently, agent16determines the most appro- priate placement of the table within the outlineO, and outputs the name of the section in which the table should be integrated. Finally, agent17per- forms the actual table insertion. It receives the tar- get sectionâs text, the table and the accompanying explanation generated by agent15. Its tasks are identifying the optimal insertion point within the specific sectionâs text, generating a caption for the table, and editing the text of the section to ensure a meaningful content integration. 3.2.4 Refinement Following the writing of the text and the inclusion of visual content, agent18performs a final review to enhance the completeness and coherence of the report. In particular, it analyzes the content of each Figure 4: Overview of the grounding verification pipeline. The grounding of the claims is evaluated with their associated references, and then possibly revised. section, independently, to identify potential gaps or inconsistencies. When needed, it retrieves and inserts supplementary information from the refer- ences setR S to improve the depth and clarity of the section under consideration. When new content is inserted, the agent checks and possibly revises the citations. Additionally, it edits the text to ensure that the tone of the report remains neutral. 3.3 Grounding Once the report has been generated, an ensemble of agents performs a final groundedness enhancement step. The objective of this phase is to review all the statements in the report and revise or remove them if they are not supported by the cited references. This is achieved by first decomposing the text into atomic claims and then evaluating each of them independently, as illustrated in Figure 4. More in detail, the text is segmented into a setT of distinct text units, with paragraph-level granular- ity. For each text unitt i â T, the cited references constitute the references setR i S â R S for such textual element, which forms the basis for the sub- sequent grounding verification and content correc- tion. Each text unitt i is then further decomposed by agent19into a setC i of atomic textual claims, that all share the same reference set R i S . Followingthedecomposition,reasoning agent20evaluates the grounding of each atomic claimc j â C i with respect to the support reference set, independently.First, it checks the claim against the setR i S , comprising the summaries of the associated references. If the claim is flagged as potentially not supported, the agent performs a second, more detailed entailment check with respect to the full-text of the referencesR i , and the final Natural Language Inference (NLI) outcome is based on the second evaluation round. Each claim is also accompanied by a brief rationale that explains the agentâs judgment. Once all the atomic claims have been evaluated, agent21carries out the revision routine. The agent operates on each text unit independently, but takes 5 into account the broader context of the section to which the text unit belongs, in order to avoid re- dundancy or inconsistency during the refinement of the statements. Specifically, the agent considers the entire text oft i âs section, the set of unsupported claimsC unsupported i â C i to be modified withint i along with their accompanying grounding explana- tions, and the summaries of the referencesR i S . The agent then localizes each claim within the given text unit, and either revises it, in case it is only imprecise, or completely removes it. Finally, agent22operates on all the first-level sections independently, revising headings and sub- headings to ensure consistency between sectionsâ titles and content after the claims amendment. 4 Experiments 4.1 Experimental setup We implement our framework with LangChain. For each input topic, we generaten = 30search queries. We use the Serper 1 search API for the re- trieval of the top-k 1 relevant resources, withk 1 = 3. For the embedded links extraction, we setk 2 = 40. We set the threshold valuesλ comp andλ search to 5,λ link to 7, andλ sim to 0.9. For the parsing of the PDF-based documents we use Docling (Auer et al., 2024), which allows also the extraction of figures, while we use the Playwright 2 Python li- brary and Trafilatura (Barbaresi, 2021) for the other types of sources. We parse Wikipedia pages with the MediaWiki Action API 3 and the html2text 4 li- brary. We employ DeepSeek-R1 (DeepSeek-AI, 2025) to build all reasoning agents, and DeepSeek- V3 (DeepSeek-AI, 2024) for all the other agents, both with temperature equal to 0. We employ as evaluator models DeepSeek-V3 and Qwen3- 32B (Qwen Team, 2025), with temperature of 1 and 0.6 respectively. 4.2 Baselines We compare our method with STORM (Shao et al., 2024), WebThinker (Li et al., 2025) and WikiAu- toGen (Yang et al., 2025). For STORM, we use the Serper retriever with the default parameters as implemented by (Shao et al., 2024), and we re- place the proprietary models with the open-source DeepSeek-V3. This substitution was made for re- 1 https://serper.dev/ 2 https://playwright.dev/ 3 https://w.mediawiki.org/wiki/API:Main_page 4 https://pypi.org/project/html2text/ producibility reasons, as proprietary models can be continuously updated over time, hindering the pos- sibility of fair comparisons without a fixed snapshot of the modelsâ weights. For WebThinker, we em- ploy the Serper retriever and the same open-source models as in (Li et al., 2025), i.e. WebThinker- QwQ-32B, a fine-tuned version of QwQ-32B, as reasoning model, and Qwen2.5-32B-Instruct as as- sistant model. For WikiAutoGen, we use the Serper retriever and replace proprietary models with open- source alternatives, namely Pixtral Large (Mistral AI, 2024), DeepSeek-R1 and DeepSeek-V3. For all the baselines, we use as input for the report gen- eration the concatenation of the topic title and the detailed description of the aspects of interest. 4.3 Human evaluation To evaluate the quality of the reports, we conduct a user study involving 27 volunteers from diverse corners of computing systems research, including senior researchers, engineers and Ph.D. students. To have evaluations of high technical quality, we ask each participant to select a topic of interest. We then provide each reviewer with two reports, one generated using Wyvern and one using a baseline method. We obtain nine distinct comparisons of Wyvern with respect to each baseline. We do not mention the used methodologies and we randomize the order between the two reports to ensure the absence of positional bias in the aggregated results. We structure the questionnaire for the human evaluation in two main parts, i.e. the relative as- sessment and the absolute grading. The relative assessment consists of a pairwise evaluation of two reports on various rubrics, concerning the struc- ture, relevance, coverage, content presentation, fig- ures and tables informativeness, engagement, and usefulness of the report. The main objective of the relative evaluation is to directly compare our method with the baselines. Absolute grading in- volves a more detailed evaluation on the same crite- ria, but applied to a single report, whose generation method is unknown to the participant. In particular, each participant assigns a score on a 5-point Likert scale (Likert, 1932) to every statement in the ques- tionnaire. This section of the evaluation enables a more precise assessment of the quality of the report generated with our proposed methodology. 6 4.4 Automatic evaluation method 4.4.1 Grounding To measure the verifiability of the statements in the report we employ the citation recall and the citation precision, as in previous works (Gao et al., 2023; Shao et al., 2024; Wang et al., 2024). The citation recallC R measures the number of claims in the report that are supported by the cited references, i.e. givenNclaims, where each claimc i is associated to a set of cited references R i : C R = P Nâ1 i=0 f(c i , R i ) N (1) wheref(c i ,R i )is the binary output of the NLI model for claimi, that assumes value 1 if the con- catenation of the referencesR i supports the claim c i , and 0 otherwise. The citation precision measures the relevance of the citations in supporting the claims of the report. A citation r k â R i is considered relevant if: g(c i ,r k ) = [[f(c i ,r k )=1]âš[f(c i ,R i \r k )=0]] = 1 (2) where[·]denotes the Iverson bracket, which eval- uates to 1 if the enclosed logical condition is true and 0 otherwise. The overall citation precision, considering all the claims and the associated cita- tions across all paragraphs in a document can be computed as: C P = P Nâ1 i=0 P |R i |â1 k=0 f(c i , R i )â§ g(c i , r k ) P Nâ1 i=0 |R i | (3) 4.4.2 Report quality We employ the evaluation rubrics proposed by Shao et al. (2024), concerning the interest level of the report, its coherence and organization, relevance and focus, and broad coverage. We frame the eval- uation as pairwise relative assessments, as in the human evaluation study. Namely, we provide as input to the evaluator models two reports, one gen- erated by Wyvern, and the other by a baseline. We consider both order permutations, and we experi- mentally assess their effect on automated evalua- tion outcomes. 5 Results 5.1 Human evaluation 5.1.1 Pairwise relative comparisons We report in Figure 5 the results of the evalua- tion study on the pairwise comparison between Structure Content Coverage Presentation Figures Table Factuality Engagement Usefulness vs STORM vs WebThinker vs WikiAutoGen 100.00100.00100.0092.86N/A100.00100.00100.00100.00 37.5068.7575.0062.50N/A66.6775.0037.5062.50 75.0075.0087.5068.7586.9686.3687.5075.0087.50 020406080100 Preference rate for our method (%) Figure 5: Relative grading of Wyvern with respect to selected baselines in the human evaluation study. Each cell reports the preference percentage for Wyvern over the baseline on a rubric. STORM and WebThinker cannot be compared on the Figures criteria, as they produce text-only reports. the reports, extracted from 23 completed ques- tionnaires (out of the 27 distributed). Each cell contains the preference percentage of our method with respect to a considered baseline on a specific rubric. We cluster the various criteria that we ask the participants to evaluate into macro-areas, to provide a high-level overview. Additional details are given in the Appendices. Wyvern generates reports that surpass the ones produced by STORM on all the considered criteria. Additionally, it out- performs in 86.96% of cases WikiAutoGen on the figures quality (informativeness, positioning, cap- tioning). WebThinker achieves a higher preference rate (62.50%) on the questions concerning the struc- ture and the engagement level created by the report. While this highlights directions of improvement, Wyvern manages to produce reports that are quali- tatively better on all the other considered aspects. In particular, users rate Wyvernâs reports as more useful than STORMâs, WebThinkerâs and WikiAu- toGenâs ones in 100.00%, 62.50% and 87.50% of cases, respectively. 5.1.2 Absolute grading To investigate the perceived quality of our frame- work, we include in the human evaluation study a finer-grained absolute scoring task of the reports generated by Wyvern. Figure 6 shows the results of this absolute grading, on the same rubrics con- sidered in Figure 5. Wyvern exhibits the highest scores on the rubrics concerning factuality and the overview table, with an absolute score of 4.22 and 4.25 out of 5, respectively. The engagement rubric receives the lowest score overall, equal to 3.65. However, it is also visible how the standard devia- tion of this criterion is the highest, highlighting the subjectivity of the question. 7 Structure Content Coverage Presentation Figures Table Factuality Engagement Usefulness 1 2 3 4 5 Score 3.87 4.07 4.13 4.20 4.16 4.25 4.22 3.65 4.00 Figure 6: Results of the absolute grading of Wyvernâs reports from the human evaluation study OursSTORMWikiAutoGen 0 50 100 Value (%) 73.79 52.29 31.45 94.64 70.32 60.87 75.45 55.17 47.91 RecallRecall w/o uncited claimsPrecision Figure 7: Citation quality metrics computed on all the reports generated by the various methods (27 documents in total for ours, 9 for STORM and WikiAutoGen). The error bar shows the standard deviation. 5.2 Automatic evaluation 5.2.1 Grounding metrics We report in Figure 7 the citation recall and pre- cision computed over the set of generated reports for each considered baseline. We were unable to compute such metrics for WebThinker, as it does not include citations in the reportâs text. Wyvern produces reports with higher grounding metrics than STORM and WikiAutoGen. In particular, it reaches 73.79% of citation recall (+21.50 and +42.34 percentage points with respect to STORM and WikiAutoGen), and 94.64% when consider- ing only the claims with citations, thus excluding paragraphs without citations from the analysis. We further analyze the impact of the claims re- vision module on the grounding metrics computed over all the distributed reports in Table 2. Both citation recall and precision show a significant im- provement after the application of this step. In par- ticular, the citation recall increases from 60.92% to 73.79%, highlighting the positive effect of the claims revision stage on Wyvernâs final output. MetricBefore (%)After (%)Gain Cit. recall60.92 ± 6.1073.79 ± 5.441.21Ă Cit. recall w/o uncited claims 84.34 ± 4.2494.64 ± 1.551.12Ă Cit. precision67.15 ± 6.1775.45 ± 7.541.12Ă Table 2: Impact of the claims revision module on the citations quality metrics 020406080100 Win ratio [auto] (%) 0 20 40 60 80 100 Win ratio [human] (%) DeepSeek-V3 020406080100 Win ratio [auto] (%) Qwen3-32B VSBASELINE vs STORM vs WebThinker vs WikiAutoGen RUBRIC Interest Level Coherence & Organization Relevance & Focus Broad Coverage Figure 8: Comparison of human (y-axis) and automatic (x-axis) evaluation results on the four rubrics with two different evaluator models. The diagonal line represents a 1:1 correspondence between human and automatic evaluation. 5.2.2 Pairwise relative comparisons We perform an automatic evaluation, framed in a relative assessment scheme, adopting the four rubrics presented in (Shao et al., 2024), that cover the aspects of interest level, coherence and orga- nization, relevance and focus, and broad coverage. We consider as evaluator models DeepSeek-V3 and Qwen3-32B. The goal is to compare the results ob- tained with the automatic evaluation with the ones of the human evaluation study, to assess whether the first can be used as proxy for the latter, being more scalable and cost- and time-effective. Figure 8 reports the comparison between the re- sults obtained over all the reports with different evaluator models, on the four criteria of judgment, and the human evaluation scores of the respective semantically equivalent question. The automatic means of evaluation (whose scores are reported on thex-axis) achieve a good correlation degree with respect to the human-assigned scores (on the y-axis) for STORM and WikiAutoGen. However, they fail to capture the scores of WebThinker, es- pecially on the Broad Coverage and Relevance & Focus criteria. This highlights how LLM-based au- tomatic evaluation of reports on qualitative rubrics still lacks behind in capturing the trends of the human-assigned evaluations, in particular when a clear preference gap is not always present in the latter (as it emerges in Figure 5). We extend our analysis by performing experi- ments to assess whether the evaluator models ex- hibit ordering bias, i.e. if the order in which the reports are passed as input to the model influences the final score. We consider two evaluator models, DeepSeek-V3 and Qwen3-32B. Ordering bias ap- pears in 36.11% and 30.56% of the comparisons for DeepSeek-V3 and Qwen3-32B, respectively. This 8 demonstrates that there is still an ample margin of improvement for achieving robustness in this type of automated evaluation. 6 Conclusions We presented Wyvern, a multi-agent framework that enables the generation of grounded, multi- modal technical reports on given input topics. We designed different modules, composed of ensem- bles of LLM-based agents, to handle the various stages of the report generation, including figures in- clusion. Furthermore, with the integration of a final stage of text decomposition into atomic claims, ver- ified on the collected evidence, Wyvern allows for an improvement in citation recall up to 2.3Ăwith respect to the other considered methods, namely STORM, WebThinker, and WikiAutoGen. We con- ducted a human evaluation study to assess the qual- ity of the reports generated by Wyvern, showing that their usefulness is perceived as higher in 63% to 100% of the instances, depending on the refer- ence baseline. We additionally showed how the automatic means of evaluation lack accuracy and robustness in estimating human preference. Limitations Wyvern relies on auxiliary search APIs to collect relevant documents from the web. Although our proposed framework includes specific sources fil- tering and selection pipelines, the generated multi- modal report ultimately depends on the quality of the webpages retrieved by the search engine. To achieve broad coverage on the user-selected topic, Wyvern generates multiple, distinct search queries. However, this approach and the reliance on the search APIs do not always ensure a comprehensive and optimal coverage of the available information. Additionally, given the inherent characteristics of search APIs, reproducibility is not fully guaran- teed, as search engines may return different results according to the inquiry time, location, and the always evolving indexing policies. We evaluated Wyvernâs performance exclusively on topics formulated in English language. Al- though a translator module could be inserted at the beginning of the pipeline to translate a topic expressed in any language to English (e.g., with an additional LLM-based agent), the generation of search queries in English language for the in- formation retrieval penalizes (but not necessarily excludes) resources in other languages. The gen- eralization capability of the framework to multilin- gual settings should be analysed in future works. Ethical considerations Wyvern builds upon content retrieved from the web, which may inherently contain systematic biases as well as scientific inaccuracies. Furthermore, we rely on existing open-source LLMs that might not have been fully aligned with the ethical values of the reader. Currently, our framework does not in- corporate an explicit mechanism for filtering or correcting biased or erroneous information. How- ever, Wyvern does not amplify these issues beyond their presence in the original sources. Acknowledgments We thank the evaluators who participated in the human evaluation study for providing their truthful judgment of the generated reports quality. This publication is part of the project PNRR- NGEU which has received funding from the MUR â DM 118/2023. References Christoph Auer, Maksym Lysak, Ahmed Nassar, Michele Dolfi, Nikolaos Livathinos, Panos Vage- nas, Cesar Berrospi Ramis, Matteo Omenetti, Fabian Lindlbauer, Kasper Dinkla, Lokesh Mishra, Yusik Kim, Shubham Gupta, Rafael Teixeira de Lima, Valery Weber, Lucas Morin, Ingmar Meijer, Viktor Kuropiatnyk, and Peter W. J. Staar. 2024. Docling technical report. Preprint, arXiv:2408.09869. Nishant Balepur, Jie Huang, and Kevin Chang. 2023. Expository text generation: Imitate, retrieve, para- phrase. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process- ing, pages 11896â11919, Singapore. Association for Computational Linguistics. Adrien Barbaresi. 2021. Trafilatura: A web scraping library and command-line tool for text discovery and extraction. In Proceedings of the 59th Annual Meet- ing of the Association for Computational Linguistics and the 11th International Joint Conference on Nat- ural Language Processing: System Demonstrations, pages 122â131, Online. Association for Computa- tional Linguistics. Lutz Bornmann, Robin Haunschild, and RĂŒdiger Mutz. 2021. Growth rates of modern science: a latent piecewise growth curve approach to model publi- cation numbers from established and new literature databases. Humanities and Social Sciences Commu- nications, 8(1):224. DeepSeek-AI. 2024. Deepseek-v3 technical report. Preprint, arXiv:2412.19437. 9 DeepSeek-AI. 2025. Deepseek-r1: Incentivizing rea- soning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948. Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2024. Chain-of-verification reduces hallucination in large language models. In Findings of the Association for Computational Linguistics: ACL 2024, pages 3563â3578, Bangkok, Thailand. Association for Computational Linguistics. Tsu-Jui Fu, William Yang Wang, Daniel McDuff, and Yale Song. 2022. DOC2PPT: Automatic Presenta- tion Slides Generation from Scientific Documents. Proceedings of the AAAI Conference on Artificial Intelligence, 36(1):634â642. Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, pages 6465â6488, Singapore. Associa- tion for Computational Linguistics. Junda He, Christoph Treude, and David Lo. 2025. Llm- based multi-agent systems for software engineering: Literature review, vision, and the road ahead. ACM Trans. Softw. Eng. Methodol., 34(5). Dong Huang, Jie M. Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2024. Agentcoder: Multi-agent-based code generation with iterative test- ing and optimisation. Preprint, arXiv:2312.13010. Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, Carl Saroufim, Corey Fry, Dror Marcus, Doron Kuklian- sky, Gaurav Singh Tomar, James Swirhun, Jinwei Xing, Lily Wang, Madhu Gurumurthy, and 7 others. 2025. The facts grounding leaderboard: Benchmark- ing llmsâ ability to ground responses to long-form input. Preprint, arXiv:2501.03200. Yucheng Jiang, Yijia Shao, Dekun Ma, Sina Semnani, and Monica Lam. 2024. Into the unknown unknowns: Engaged human learning through participation in lan- guage model agent conversations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9917â9955, Mi- ami, Florida, USA. Association for Computational Linguistics. Benjamin F. Jones. 2009. The Burden of Knowledge and the âDeath of the Renaissance Manâ: Is Inno- vation Getting Harder? The Review of Economic Studies, 76(1):283â317. Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024a. Prometheus: Inducing fine- grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations. Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024b. Prometheus 2: An open source lan- guage model specialized in evaluating other language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4334â4353, Miami, Florida, USA. Association for Computational Linguistics. Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Hein- rich KĂŒttler, Mike Lewis, Wen-tau Yih, Tim Rock- tĂ€schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge- intensive nlp tasks. In Advances in Neural Infor- mation Processing Systems (NeurIPS), volume 33, pages 9459â9474. Curran Associates, Inc. Miaoran Li, Baolin Peng, Michel Galley, Jianfeng Gao, and Zhu Zhang. 2024. Self-checker: Plug-and-play modules for fact-checking with large language mod- els. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 163â181, Mexico City, Mexico. Association for Computational Lin- guistics. Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yu- tao Zhu, Yongkang Wu, Ji-Rong Wen, and Zhicheng Dou. 2025. Webthinker: Empowering large reason- ing models with deep research capability. In Ad- vances in Neural Information Processing Systems (NeurIPS). Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of Psychology, 140:1â55. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language mod- els use long contexts. Transactions of the Association for Computational Linguistics, 12:157â173. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettle- moyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076â12100, Singa- pore. Association for Computational Linguistics. Mistral AI. 2024. Pixtral large.https://mistral.ai /news/pixtral-large. Accessed: 2025-07-31. Michael Park, Erin Leahey, and Russell J. Funk. 2023. Papers and patents are becoming less disruptive over time. Nature, 613(7942):138â144. Qwen Team. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388. Yijia Shao, Yucheng Jiang, Theodore Kanell, Peter Xu, Omar Khattab, and Monica Lam. 2024. Assisting in writing Wikipedia-like articles from scratch with large language models. In Proceedings of the 2024 10 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6252â6278, Mexico City, Mexico. Association for Computational Linguistics. Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, Xin Zhang, Zhen Wu, Meishan Zhang, Xinyu Dai, Min zhang, Qingsong Wen, Wei Ye, Shikun Zhang, and Yue Zhang. 2024. Autosurvey: Large language mod- els can automatically write surveys. In The Thirty- eighth Annual Conference on Neural Information Processing Systems. Ruibin Xiong, Yimeng Chen, Dmitrii Khizbullin, Mingchen Zhuge, and JĂŒrgen Schmidhuber. 2025. Beyond outlining: Heterogeneous recursive planning for adaptive long-form writing with language models. In Proceedings of the 2025 Conference on Empiri- cal Methods in Natural Language Processing, pages 24678â24714, Suzhou, China. Association for Com- putational Linguistics. Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang, Xingyu Liu, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu, Bo Zhang, and Wei Chen. 2026. Mul- timodal deepresearcher: Generating text-chart inter- leaved reports from scratch with agentic framework. In Proceedings of the AAAI Conference on Artificial Intelligence. Zhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng, and Mohamed Elhoseiny. 2025. Wikiautogen: To- wards multi-modal wikipedia-style article generation. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision (ICCV), pages 15532â 15541. A Human evaluation We report in this section the complete, unaggre- gated list of questions and results of the relative assessment and absolute grading of the human eval- uation study presented in the main text of the paper. All evaluators consented to the use of aggregated anonymous statistics collected from their responses, and of portions of the generated reports, for publi- cation purposes. A.1 Relative assessment For the relative assessment in the human evalua- tion study, we provide each participant with two reports, named âReport A" and âReport B", with- out specifying the generation method. Each person receives one report generated by Wyvern, and one generated by another baseline. We divide the 27 participants into three groups of cardinality equal to 9, to have a uniform amount of comparisons over all the baselines. Furthermore, to avoid the presence of any positional bias in the participantsâ answers, we randomize the order and naming of the reports. In Table 3, the list of questions and the associated answers from the human evaluation study are given. Wyvern is systematically preferred over STORM on all the considered rubrics. WebThinker obtains a higher preference rate only on the structure, i.e. rubric (1), and on the engagement level of the re- port, namely rubric (15), where Wyvern reaches a 37.50% favor rate. Regarding the structure of the report, the main criticalities that emerged concern the redundancy of the content over the sections and the overall length of the report. Wyvern and Web- Thinker achieve the same preference rate on rubric (6), which refers to the clarity of the explanations, and on rubric (4), related to the presentation of ben- efits and limitations about the topic. Wyvern excels on the topic coverage, i.e. rubric (5), where it ob- tains a 100.00% preference rate; and on the amount of technical details and factuality of the content (rubrics (3) and (15)), where it reaches 75.00% of win percentage. When compared to WikiAutoGen, the only baseline that supports the inclusion of fig- ures within the produced reports, Wyvern achieves preference rates of 85.71%, 87.50% and 87.50% for rubrics (8), (9), and (10), which concern the quality of the positioning, the quality of the cap- tions, and the informativeness degree of the figures, respectively. On all the other rubrics, Wyvern ob- tains a preference rate spanning from 62.50% to 87.50%. A.2 Absolute grading We report in Table 4 the complete list of questions and results of the absolute grading section of the human evaluation study, before the clustering ap- plied to obtain Figure 6 of the paper. The con- sidered rubrics are the same used for the relative assessment. As discussed in the paper, the absolute grading is used to have a more detailed overview of the quality of Wyvernâs reports over distinct spe- cific rubrics. The methodology of the report for which this evaluation was asked was unknown to the participants: only a generic report index (âRe- port A" or âReport B") was given as indication of the target, hiding that it was corresponding to the one generated by Wyvern. For a performance com- parison between Wyvern and the other baselines, refer to the relative assessment results. From Table 4 it is possible to observe that the highest marks are obtained on rubrics (8) and (11), 11 Preference rate for Wyvern Question vs STORMvs WebThinkervs WikiAutoGen 1Which report has the most logical structure?100.00%37.50%75.00% 2 Which report better provides relevant information with respect to the topic? 100.00%62.50%75.00% 3 Which report presents the best amount of technical details about the topic? 100.00%75.00%75.00% 4 Which report better presents both benefits/advantages and challenges/limitations related to the topic? 100.00%50.00%87.50% 5 Which report has a broader coverage of multiple aspects of the topic? 100.00%100.00%87.50% 6Which report has clearer explanations?100.00%50.00%75.00% 7Which report has the most neutral tone?85.71%75.00%62.50% 8 Which report has the most logical figures placement (i.e., figures are assigned to a coherent section)? If not applicable (e.g., there are no figures), please select N/A. --85.71% 9 Which report has the most appropriate figures descriptions and captions? If not applicable (e.g., there are no figures), please select N/A. --87.50% 10 Which report has the most informative figures? If not applicable (e.g., there are no figures), please select N/A. --87.50% 11 Which report has the most logical overview tables placement (i.e., the overview table is assigned to a coherent section)? If not applicable (e.g., there is no overview table), please select N/A. 100.00%62.50%85.71% 12 Which report has the most appropriate overview tables descriptions and captions? If not applicable (e.g., there are no overview tables), please select N/A. 100.00%75.00%85.71% 13 Which report has the most informative overview tables? If not applicable (e.g., there are no overview tables), please select N/A. 100.00%62.50%87.50% 14 Which report has a better factuality to the best of your knowledge? 100.00%75.00%87.50% 15 Which report is more engaging and thought-provoking? 100.00%37.50%75.00% 16Which report do you find the most useful?100.00%62.50%87.50% Table 3: Results of the relative assessment section of the human evaluation study. Each percentage value represents the preference rate in the pairwise comparison between Wyvern and one of the baselines on a specific rubric. Comparisons cannot be made on criteria concerning figures when considering STORM and WebThinker as the other method, as explained in the paper. 12 QuestionScore 1The structure of the report is logical.3.87 ± 0.61 2The report provides relevant information with respect to the topic.4.26 ± 0.53 3The report presents the right amount of technical details about the topic.3.87 ± 0.85 4The report presents both benefits/advantages and challenges/limitations related to the topic.3.91 ± 0.78 5The report has a broad coverage of multiple aspects of the topic.4.35 ± 0.56 6The explanations of the report are:3.91 ± 0.88 7The tone of the report is:4.48 ± 0.65 8 The position of the figures within the report is logical (i.e., they are assigned to a coherent section). If not applicable (e.g., there are no figures), please select N/A. 4.52 ± 0.58 9 The caption and the description of the figureâs content in the text are appropriate. If not applicable (e.g., there are no figures), please select N/A. 3.87 ± 1.23 10The figures are informative. If not applicable (e.g., there are no figures), please select N/A.4.09 ± 0.72 11 The position of the overview table within the report is logical (i.e., it is assigned to a coherent section). If not applicable (e.g., there are no overview tables), please select N/A. 4.50 ± 0.58 12 The caption and the description of the overview tableâs content in the text are appropriate. If not applicable (e.g., there are no overview tables), please select N/A. 4.14 ± 0.69 13 The overview table is informative. If not applicable (e.g., there are no overview tables), please select N/A. 4.10 ± 0.75 14To the best of your knowledge, the factuality of the report is:4.22 ± 0.72 15How engaging and thought-provoking is the report?3.65 ± 1.09 16How useful do you find the generated report?4.00 ± 0.83 Table 4: Results of the absolute grading of the human evaluation study, reported with the mean and standard deviation of the given marks over the 23 questionnaires that were completed, out of the 27. Participants could assign a score in the [1,5]-range. For rubric (6), the interval extremes were associated with the adjectives âconfusing" and âclear"; for rubric (7) with âbiased" and âneutral"; and for rubric (14) with ânot satisfying" and âvery satisfying". which concern the positioning of the figures and of the overview table in the report, with an aver- age score of 4.52 and 4.50, respectively. The third highest average score, equal to 4.48, is achieved on rubric (7), related to the neutrality of the tone, which is evaluated as satisfactorily neutral. The cri- terion that receives the lowest score, equal to 3.65, is the engagement and thought-provoking level of the technical report (rubric (15) in Table 4). How- ever, it can be seen that the standard deviation is quite high, which can be explained by the highly subjective nature of the answer. Also the structure of the report, namely rubric (1), and the textual integration of the figures within the text, i.e. rubric (12), exhibit margin of improvement, reaching a score of 3.87. The factuality, the figuresâ informa- tiveness and the usefulness of the report (rubrics (14), (10) and (16), respectively) are very posi- tively evaluated, with average scores of 4.22, 4.09, and 4.00, demonstrating how Wyvernâs design suc- cessfully accomplishes the objective of producing multimodal and grounded technical reports. B Automatic evaluation We report in this section the complete results of the automatic evaluation of the reports, obtained con- Rubricsvs STORM vs WebThinker vs WikiAutoGen Interest level50.00%50.00%50.00% Coherence & Organization 50.00%50.00%50.00% Relevance & Focus 50.00%50.00%50.00% Broad Coverage 50.00%50.00%50.00% Table 5: Results of the automatic relative assessment, performed using Prometheus 2-7B as evaluator model sidering the same four criteria presented in (Shao et al., 2024), which concern the interest level, the coherence and organization, the relevance and fo- cus of the content, and the broad coverage of the topic. We consider both a relative assessment setup, discussed in the paper, and an absolute grading one. B.1 Relative assessment In Figure 8 of the paper we presented the results of the pairwise relative evaluation, obtained using Qwen3-32B and DeepSeek-V3 as evaluator mod- els. As discussed in the paper, for each pair of reports (one generated by Wyvern and the other by one of the baselines) we conduct the relative as- sessment twice, swapping the order of the reports, and assess the impact on the output of evaluators. 13 As shown in Table 2 of the paper, the two models did exhibit ordering bias, i.e. the outcome of the evaluation turned out to depend on the ordering of the reports under assessment. We carried out but omitted from the paper the same experiments using as evaluator Prometheus 2-7B (Kim et al., 2024b), a model specifically trained on a relative ranking task. Table 5 presents the results of this analysis, which shows that Prometheus 2-7B struggles even more to capture the human-assigned scores, giving the same preference rate to all the methods. By a manual inspection of the outcome, it was revealed that the evaluator always identifies as best report the first of the two options. Thus, since we con- sider both orderings when we perform the evalua- tion, the preference rate is always equal to 50.00%, regardless of both method and criterion under con- sideration. The information position hence alters the final output of the model, as already shown in the literature (Liu et al., 2024). We hypothesize that the cause of this behavior could lie in the fact that Prometheus 2-7B was trained with a limited maximum sequence length (4,096 tokens) (Kim et al., 2024b), potentially affecting its capability to handle longer inputs properly. B.2 Absolute grading We conduct an automatic evaluation in an absolute grading schema, similarly to (Shao et al., 2024). In these experiments, we provide only a single re- port to the evaluator model, which is prompted to assign a score in the [1,5]-range for each of the four considered criteria, i.e. Interest Level, Co- herence and Organization, Relevance and Focus, and Broad Coverage. We employed as evalua- tor models Prometheus-13B (Kim et al., 2024a), Prometheus 2-7B (Kim et al., 2024b), Qwen3- 32B, and DeepSeek-V3. For Prometheus-13B, we employ the same iterative trimming strategy as in (Shao et al., 2024) to reduce the length of text. However, we identify this evaluation setup as sub- optimal, as for longer reports the majority of the content gets removed, impeding a fair evaluation. For this reason, we did not employ Prometheus- 13B for the relative assessment, but only on the absolute grading with the same evaluation setup used in (Shao et al., 2024). Table 6 presents the results of the experiments, that show that there is no uniform agreement between the evaluator mod- els. For each rubric, distinct evaluators identify different optimal methods. This highlights how the choice of the evaluator model can highly affect the results and ranking between various method- ologies, making it hard to obtain a fair and robust quantitative interpretation of the evaluations. B.3 Claims evaluation To verify whether automatic means of evaluation are effective for the grounding assessment, we con- ducted a manual evaluation comparing LLMs and human judgments on a sample of claims. Specifi- cally, we randomly sampled 60 claims and verified whether the human annotations and the LLM judg- ments were consistent. The agreement rate was 93.33%, indicating strong alignment between the two in this specific task. C Cost analysis We report the breakdown of the input and comple- tion tokens of each module and agent of Wyvern in Figure 9 and Figure 10, respectively. Statis- tics are computed on a sample of 6 reports, using DeepSeek-V3 as base model and DeepSeek-R1- 0528 as reasoning model. Agents2-7in the search module consume approximately the same number of input tokens, as they all analyze the collected documents individually. Within the re- port generation module, agents 11 and 18 have high input and completion tokens usage as they are responsible for the reportâs text generation and re- finement, based on the input references. Agent12 produces a significant amount of completion tokens due to the generation of image descriptions. The grounding module (especially agent20) required lengthy inputs mainly due to the two-stage ground- ing check, which may lead to considering many references (also in their full-text form); while the long outputs can be attributed to the long reasoning trace needed for the grounding assessment of each claim. Regarding the parsing of the resources, which falls outside of the tokens usage analysis, we note that it can contribute significantly on the end-to- end latency of the framework, especially when a high number of long documents with many figures must be processed. D Failure cases While the overall evaluation of our framework is positive, as detailed in Section 5.1, the human eval- uation study also served to identify the main weak- nesses and directions of improvement. 14 Score RubricsEvaluator model STORMWebThinkerWikiAutoGenWyvern Prometheus-13B4.00 ± 0.004.56 ± 0.504.00 ± 0.004.07 ± 0.47 Prometheus 2-7B4.22 ± 0.924.44 ± 0.834.67 ± 0.674.85 ± 0.45 DeepSeek-V33.22 ± 0.634.44 ± 0.833.44 ± 0.684.85 ± 0.52 Interest Level Qwen3-32B3.33 ± 0.474.33 ± 0.673.67 ± 0.673.78 ± 0.50 Prometheus-13B4.78 ± 0.424.44 ± 0.504.33 ± 0.473.96 ± 0.58 Prometheus 2-7B4.67 ± 0.474.67 ± 0.474.22 ± 0.424.70 ± 0.46 DeepSeek-V34.89 ± 0.314.67 ± 0.474.22 ± 0.924.85 ± 0.36 Coherence & Organization Qwen3-32B4.56 ± 0.685.00 ± 0.00 4.22 ± 0.634.52 ± 0.63 Prometheus-13B4.89 ± 0.314.89 ± 0.314.78 ± 0.634.11 ± 0.92 Prometheus 2-7B5.00 ± 0.005.00 ± 0.005.00 ± 0.005.00 ± 0.00 DeepSeek-V34.78 ± 0.424.56 ± 0.504.00 ± 0.944.81 ± 0.47 Relevance & Focus Qwen3-32B4.56 ± 0.505.00 ± 0.004.11 ± 0.314.59 ± 0.56 Prometheus-13B5.00 ± 0.004.89 ± 0.315.00 ± 0.004.19 ± 1.54 Prometheus 2-7B4.44 ± 0.50 4.22 ± 0.424.00 ± 0.004.41 ± 0.49 DeepSeek-V34.44 ± 0.504.33 ± 0.474.00 ± 0.004.37 ± 0.48 Broad Coverage Qwen3-32B4.44 ± 0.684.89 ± 0.314.67 ± 0.474.93 ± 0.26 Table 6: Results of the automatic absolute grading, performed with different evaluator models, namely Prometheus- 13B, Prometheus 2-7B, DeepSeek-V3, and Qwen3-32B. The results are reported in terms of the mean and standard deviation of the scores over all the reports generated by a given method. Underlined, the method achieving the highest average score with a specific evaluator on a given criterion. SearchReport gen. Grounding 10 5 10 6 10 7 Number of tokens 2.58e6 7.54e5 3.15e6 4.20e4 7.22e4 3.50e5 Input tokensCompletion Tokens Figure 9: Input and completion tokens used by each module of Wyvern. The main challenge noted by a small number of evaluators was the reportsâ excessive length and redundancy, sometimes resulting in a fragmented structure. A few evaluators also identified an un- even balance between breadth and depth, with too much focus given to peripheral topics. Issues involving images, e.g. imprecise place- ment or explanation, instead stem mainly from in- sufficient textual descriptions in the input sources, which may lead to overly general figure descrip- tions by agent 12 . This in turn may lead to the placement of figures in sections not fully relevant to the figureâs topic, or to overly vague in-text expla- nations of their content. Additionally, the parsing tool occasionally crops images imprecisely, result- ing in minor inconsistencies between the imagesâ content and their captions. Regarding claims grounding, instances of only partially correct explanations and confusing termi- nology were also reported. Interestingly, we also came across an instance of a clearly false state- ment. After manually inspecting the cited refer- ence, we found that the claim was indeed faithful to the source, and it was the source itself that con- tained the factual inaccuracy. As discussed in the Ethical considerations section, we emphasize on this matter that while Wyvern targets the grounded- ness of the claims, it does not incorporate a mecha- nism to filter out erroneous information retrieved from the web. 15 10 1 10 2 10 3 10 4 10 5 10 6 10 7 10 8 Number of tokens 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Agent Search Report gen. Grounding Input tokensCompletion tokens Figure 10: Input and completion tokens used by each agent of Wyvern. 16