Paper deep dive
Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro PinzĂłn HernĂĄndez, Reza Yazdanfar, Michelle Dugas, Renos Vakis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/27/2026, 9:42:40 AM
Summary
The paper presents AVA (AI + Verified Analysis), a specialized generative AI platform designed for policy and development professionals. Built on a curated library of over 4,000 World Bank Reports, AVA utilizes a multi-agent RAG (Retrieval-Augmented Generation) pipeline to provide evidence-based syntheses. The system operationalizes 'epistemic humility' through two key mechanisms: citation verifiability (page-anchored citations) and reasoned abstention (declining unsupported queries with justification). An in-the-wild evaluation involving over 2,200 users across 116 countries demonstrated that the system helps users save 2.4â3.9 hours weekly and serves as a specialized 'evidence engine' that calibrates trust through institutional provenance.
Entities (7)
Relation Signals (5)
AVA â implements â Epistemic Humility
confidence 100% ¡ It operationalizes epistemic humility through two mechanisms
Epistemic Humility â includes â Reasoned Abstention
confidence 100% ¡ It operationalizes epistemic humility through two mechanisms: citation verifiability... and reasoned abstention
Epistemic Humility â includes â Citation Verifiability
confidence 100% ¡ It operationalizes epistemic humility through two mechanisms: citation verifiability... and reasoned abstention
AVA â isbuilton â World Bank Reports
confidence 100% ¡ built on a curated library of over 4,000 World Bank Reports
AVA â uses â Multi-agent Pipeline
confidence 100% ¡ AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVA's multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4-3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized "evidence engine"; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for "ecosystem-aware" Humble AI.
Tags
Links
- Source: https://arxiv.org/abs/2604.17843v1
- Canonical: https://arxiv.org/abs/2604.17843v1
Trouble viewing inline? Open PDF directly â
Full Text
179,800 characters extracted from source content.
Expand or collapse full text
by Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research Nimisha Karnatak University of OxfordOxfordUnited Kingdom nimisha.karnatak@some.ox.ac.uk , Mohamad Chatila The World Bank GroupWashingtonDCUSA mchatila@worldbank.org , Daniel Alejandro PinzĂłn HernĂĄndez The World Bank GroupWashingtonDCUSA dpinzonhernandez@worldbank.org , Reza Yazdanfar Nouswise, Inc.DoverDEUSA reza@nouswise.com , Michelle Dugas The World Bank GroupWashingtonDCUSA mdugas@worldbank.org and Renos Vakis The World Bank GroupWashingtonDCUSA rvakis@worldbank.org (2026) Abstract. General-purpose LLMs pose misinformation risks for development and policy experts, lacking epistemic humility for verifiable outputs. We present AVA (AI + Verified Analysis), a GenAI platform built on a curated library of over 4,000 World Bank Reports with multilingual capabilities. AVAâs multi-agent pipeline enables users to query and receive evidence-based syntheses. It operationalizes epistemic humility through two mechanisms: citation verifiability (tracing claims to sources) and reasoned abstention (declining unsupported queries with justification and redirection). We conducted an in-the-wild evaluation with over 2,200 individuals from heterogeneous organisations and roles in 116 countries, via log analysis, surveys, and 20 interviews. Difference-in-Differences estimates associate sustained engagement with 2.4â3.9 hours saved weekly. Qualitatively, participants used AVA as a specialized âevidence engineâ; reasoned abstention clarified scope boundaries, and trust was calibrated through institutional provenance and page-anchored citations. We contribute design guidelines for specialized AI and articulate a vision for âecosystem-awareâ Humble AI. Agentic AI, Retrieval-augmented generation (RAG), Hallucination mitigation, Epistemic humility, Reasoned abstention, Large-scale field deployment, Trust Calibration â journalyear: 2026â copyright: câ conference: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems; April 13â17, 2026; Barcelona, Spainâ booktitle: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI â26), April 13â17, 2026, Barcelona, Spainâ doi: 10.1145/3772318.3791062â isbn: 979-8-4007-2278-3/2026/04â ccs: Human-centered computing Empirical studies in HCIâ ccs: Computing methodologies Multi-agent systems Figure 1. AVA System Architecture: Stage 1 curates 4000+ World Bank Reports into a hierarchical RAG; Stage 2 uses agentic retrieval (decomposition, planning, tree-walking, drafting); Stage 3 synthesizes and verifies evidence to produce citation-grounded answers or reasoned abstention; Stage 4 delivers multilingual, preference-aligned responses with interactive citations. The image shows a four-stage architecture of the AVA system, arranged from left to right. Stage 1: Data curation and indexing. Over 4,000 World Bank reports are collected and processed. The system parses document layouts, extracts text using OCR, splits the text into chunks, and converts them into embeddings. These are organized into a hierarchically indexed retrieval-augmented generation (RAG) database. Stage 2: Agentic retrieval and evidence formulation. A user query enters the system and is handled by multiple specialized agents. One agent decomposes the query and identifies sub-questions. Another plans retrieval strategies, including exact keyword matching, semantic similarity, and comparative analysis. A tree-walking agent explores multiple retrieval paths through the database, and a drafting agent consolidates the results into evidence packets relevant to the query. Stage 3: Information synthesis and verification. The consolidated evidence packets are analyzed to extract entities, detect relevant text spans, and summarize evidence. The system maps claims to sources and verifies coverage and agreement across documents. The output is a verified answer with linked citations, or a reasoned abstention if evidence is insufficient. Stage 4: User personalization and memory. The final response is adapted to the userâs language, document preferences, and domain interests. The system supports multilingual outputs, interactive citations, and user-controlled memory for storing preferences or past interactions. Overall, the diagram illustrates how AVA moves from large-scale document ingestion to agent-driven retrieval, evidence verification, and personalized, citation-grounded responses. 1. Introduction Public policy and development researchers navigate large, heterogeneous document corpora, often under significant time pressure. While Large Language Models (LLMs) can accelerate synthesis, their application in high-stakes domains such as policy-making is constrained by a lack of verifiable sources and a tendency to fabricate information, commonly referred to as hallucination (Leiser et al., 2024; Cheng et al., 2024; Lee et al., 2024; Sun et al., 2024). Prior HCI research has established that user trust is contingent upon inspectable evidence and clearly communicated system boundaries (Wester and others, 2024; Rapp et al., 2025). This has motivated interface designs that surface data provenance for verification and incorporate intentional refusal states when an answer cannot be substantiated (Narayanan Venkit et al., 2025; Rahman et al., 2025; Liao and Wortman Vaughan, 2023). These practical interface strategies represent operationalizations of epistemic humility, defined as a systemâs capacity to recognise and communicate its own fallibility and limits (Knowles et al., 2023; Nair et al., 2025; Tong et al., 2025; OrdoĂąez et al., 2025), and are commonly framed within the Humble AI paradigm, which emphasizes resisting over-claiming and prioritizing user verification over blind reliance (Knowles et al., 2023; Nair et al., 2025; Celi, 2025). While recent work has begun to translate Humble AI principles111Knowles et al. (Knowles et al., 2023) define Humble AI via three principles: skepticism, curiosity and commitment. See Sec 2.1 for details.into concrete interface mechanisms, such as uncertainty-aware ranking in hiring (Nair et al., 2025), these initial efforts have predominantly focused on predictive tasks that classify or rank existing data, within controlled or short-horizon studies. This leaves two critical gaps. First, as generative AI models enter professional workflows, understanding how humility mechanisms function in systems that produce new text, rather than merely score or rank existing data as in predictive AI, poses distinct theoretical and design challenges (for example, hallucination vs misclassification) (Nair et al., 2025). Second, we lack longitudinal, in-the-wild evidence on how such systems integrate into everyday professional practice (Pang et al., 2025; Wang et al., 2024a). To address these gaps, we present AVA (AI + Verified Analysis), a source-linked, evidence-bounded, multi-agent generative AI assistant for policy and development professionals 222In this paper, âpolicy and development professionalsâ refers to practitioners working in international policy and development contexts (e.g., economic policy and public-sector reform) across NGOs, government agencies, academia, and international organizations. (See Fig.2, Table9, Sec. 4.1 for details). For any query, AVA (i) returns synthesized answers with page-level, clickable citations and in-context highlighting; (i) issues reasoned abstention, explicitly responding âI donât knowâ with a brief rationale and reformulation paths when the corpus cannot substantiate a claim; and (i) provides an in-place drafting and editing workspace. AVA is deployed over a curated library of 4,000+ official World Bank Reports and supports queries in over 60 languages (See appendix C). In this paper, we situate AVA within CSCW and HCI accounts of how new technologies become embedded in professional practice (Suchman, 1987; Orlikowski, 2000). We use this lens to investigate whether operationalizing epistemic humility through source-linked answers and scope-awareness can enhance real-world evidence 333Appendix J details the operationalization strategies and design rationale.. Prior work characterizes integration of technology in workflow through appropriation (Dourish, 2003), trust calibration (Lee and See, 2004), sustained use (Orlikowski and others, 1995), as generative AI enters high-stakes workflows, debates about whether and how its use should be disclosed have gained prominence (Hryciw et al., 2023; Hosseini et al., 2023; European Parliament and Council of the European Union, 2024). Within this landscape, disclosure operates as a boundary practice through which professionals define the scope of their responsibility and negotiate professional and institutional accountability (BaHammam, 2025). Drawing on these four central dimensions from CSCW and HCI literature (Orlikowski, 2000; Yang et al., 2019; Xiao et al., 2025) we investigate AVAâs sociotechnical integration into professional workflows through questions that map to the gaps identified above: ⢠RQ1: HumanâAI collaboration (Appropriation & Use). How do policy professionals appropriate AVA within evidence-based workflows, and what interaction patterns emerge across tasks? (Addressing Gap 2: lack of empirical evidence on AI integration into everyday professional practice.) ⢠RQ2: System boundaries (Refusal & Reliance). How do verification mechanisms, including citations and reasoned abstention, shape trust calibration and usersâ mental models? (Addressing Gap 1: limited understanding of how Humble AI mechanisms function in generative systems as opposed to predictive ones.) ⢠RQ3: Sustained value (Retention & Efficiency). To what extent does sustained use predict perceived efficiency gains, and which first-session behaviors distinguish returning users? (Addressing Gap 2: lack of longitudinal evidence.) ⢠RQ4: Disclosure & accountability (Authorship and Norms). How do professionals reason about and enact disclosure of AVAâs assistance in their workflows? (Addressing policy-specific requirements for transparency and accountability.) Figure 2. AVAâs multi-institutional external deployment across 116 countries. Over 2,200 professionals used the system over five months, with 95.4% employed by organizations external to the host institution (NGOs, academia, government professionals, etc). The image summarizes AVAâs real-world deployment context and study parameters. At the top, the diagram highlights heterogeneous organizational and professional contexts. AVA was used by individuals from diverse organizations, including non-governmental organizations, government agencies, academic institutions, and the private sector. The professional roles represented include policy analysts, researchers, professors, and economists. To the right, the study scale is shown with three key metrics: deployment across 116 countries, usage by more than 2,200 professionals, and a study duration of five months. Along the bottom, the image illustrates in-the-wild usage across workflows, emphasizing global deployment, remote and flexible use, collaborative work settings, and in-situ application within everyday professional environments. The final section highlights external, multi-institutional adoption, noting that 95.4 percent of users came from organizations external to the host institution, indicating discretionary, voluntary uptake rather than mandated internal use. Overall, the image conveys that AVA was adopted globally, across institutions and roles, and used in real professional workflows at scale. A brief description of the image for accessibility. To answer these questions, we conducted a five-month, global, multi-institutional, in-the-wild deployment of AVA using a mixed-methods approach (see Figure 2 for deployment scope and context).Crucially, the study engaged over 2,200 professionals across 116 countries. Of these, 95.4% were employed by organizations external to the host institution, spanning NGOs, government professionals, academia, and the private sector. Quantitatively, we analyzed usage logs linked to baseline and endline surveys (2,259 registrants; 3,797 queries; matched n=1,029n=1,029 across 116 countries). Qualitatively, we conducted 20 semi-structured interviews to examine verification practices, refusal experiences, and workflow fit. This paper makes two contributions: (1) Empirical: We report findings from a five-month, in the wild, mixed-methods deployment of AVA, engaging over 2,200 policy and development professionals across 116 countries. Triangulating longitudinal platform logs and preâpost surveys with 20 in-depth qualitative interviews, we characterize how these users integrate AVA into high-stakes workflows, detailing how they interpret reasoned abstentions as signals of system scope and trustworthiness and how they leverage page-anchored citations to support verification-driven evidence practices. (2) Conceptual: We derive generalisable lessons for building trustworthy knowledge systems that can inform deployments in other high-stakes domains, including the need for an end-to-end trust pipeline (from corpus curation through abstention to verification), strategies for managing corpus qualityâcoverage trade-offs, and interface patterns that prioritise verification over disclosure. We also articulate a vision for ecosystem-aware specialised AI systems that prioritise collaborative interoperability with general-purpose models. 2. Related Work Our review synthesizes three strands: (i) epistemic humility and uncertainty communication, (i) practice-grounded evaluation, and (i) domain-bounded generative systems for specialized knowledge work. This lens motivates AVAâs design (reasoned abstention; verifiable citation) and our mixed-methods, in-the-wild evaluation, and situates AVA as a domain-bounded generative system built for knowledge professionals in evidence-intensive policy and development work. 2.1. Epistemic Humility and Uncertainty Communication HCI and Human Factors research has long framed appropriate reliance on automation as an interaction design problem: trust is better calibrated when systems make their limits as legible as their capabilities (Lee and See, 2004). Foundational work shows how over- and under-reliance emerge, and argues for interfaces that support calibrated trust under uncertainty, motivating interface strategies such as confidence cues, boundary disclosures, and expectation-setting. For example, Zhang et al. (2020) demonstrate that confidence signals and local rationales can modulate reliance in decision support, while Kocielnik et al. (2019) show that acknowledging imperfection helps establish realistic mental models and temper over-trust. These findings frame trust calibration as a communication and interaction challenge: users must be able to see where a system ends, not only what it can do. Recent studies shift attention from what a model says to how it refuses. Wester and others (2024) find that bare denials are rated more frustrating and less appropriate or useful than refusals that motivate or redirect. Kim and others (2024) show that hedged uncertainty (e.g., âI am not sure, butâŚâ) can reduce over-reliance and improve accuracy. Together, these findings support using reasoned refusals that briefly justify limits and offer a next step, rather than bare denials. In parallel, ML/NLP communities develop mechanism-level toolkits for abstention and grounding. For example, Geifman and El-Yaniv (2017) operationalize selective prediction for deep classifiers (a reject option under low confidence). More recently, LLM research advances evidencing and refusal: Menick and others (2022) teach models to support answers with verified quotes, and Gao et al. (2023a) evaluate citation quality, finding many links only partially supportive, underscoring the difficulty of dependable evidencing. In a complementary line, Chen et al. (2024) align refusals so models decline or explain unknowns within bounded scopes. These technical and interactional approaches converge on the concept of epistemic humility, which we synthesize across three traditions. Philosophically, epistemic humility is defined as an intellectual virtue grounded in recognizing oneâs own fallibility and the limits of oneâs knowledge (Whitcomb et al., 2017; Potter, 2022). In HCI and CSCW, epistemic humility has recently been taken up as a design concern in multiple ways. DedeoÄlu et al. (DedeoÄlu and Chandra, 2025) frame it as a post-Enlightenment design value that resists technological mastery and foregrounds situated, plural forms of knowledge. Karusala et al. (Karusala et al., 2024) similarly argue that public-sector algorithms must practice humility to avoid over-claiming what they know about human subjects, particularly in sensitive decision-making contexts. Within the ACM community, Knowles et al. (Knowles et al., 2023) crystallize this stance under Humble AI, a design approach emphasizing scepticism (acknowledging the limits of statistical proxies and resisting overâconfident claims), curiosity (interrogating unknowns by seeking evidence of âtrustâresponsiveness,â especially in borderline or rejected cases), and commitment (avoiding distrust of the trustworthy, even at the expense of efficiency, by prioritising fairness, inclusion, and ongoing responsible practice). Building on this lineage, we operationalise epistemic humility for generative systems as a system capability: the ability to make knowledge boundaries explicit, to withhold certainty when evidence is insufficient, and to prioritise verifiable, evidence-backed responses over plausible but unsupported guesses. While empirical work in domains such as algorithmic hiring has demonstrated how surfacing algorithmic unknowns and communicating uncertainty can support calibrated reliance in high-stakes workflows (Nair et al., 2025), much of this scholarship remains either conceptual or anchored in narrow-domain deployments, leaving open how epistemic humility can be operationalised in generative systems used at scale for professional knowledge work. AVA addresses this gap by operationalising epistemic humility through two complementary interaction mechanisms in a domain-bounded generative platform. First, reasoned abstention declines unsupported queries with a brief justification and constructive redirection, making scope boundaries explicit and reducing unsupported generation. Second, verifiable citation attaches page-level references to factual statements, enabling users to independently check and contest claims. In combination, these mechanisms operationalise humility as a user-facing property that supports verification and accountable use. 2.2. Evaluation in Practice Recent position work argues that GenAI evaluation must be sociotechnically grounded and practice-based: metrics should be iteratively refined from what people actually do with systems in situ, and studies should report design and measurement choices transparently (Weidinger et al., 2025). Complementary frameworks call for in-the-wild, lifecycle-oriented evaluation that uses dynamic, outcome-oriented measures (beyond accuracy), combining deployment traces with user research and organizational context to capture value and risk in real work (Jabbour et al., 2025). A recent systematic review of 153 CHI papers (2020â2024) highlights persistent validity and reproducibility issues, a predominance of artifact-centric, short-horizon studies, and a relative absence of multi-month deployments (Pang et al., 2025). While task-supportive user studies suggest that adapted interaction scaffolds can reduce cognitive burden in professional settings, such evaluations are typically scenario-based and short-duration (Wang et al., 2024a). Guided by this agenda, we examine axes that matter for real adoption in professional settings. We analyze what users actually do with verification-oriented affordances (page-level citations; reasoned refusals with redirection) alongside other adoption-relevant axes: task fit, adoption dynamics, workflow integration, multilingual/domain coverage, thus operationalizing calls to link measurement to situated activity (e.g., session flows, handoffs, reformulations) (Weidinger et al., 2025; Jabbour et al., 2025). With AVA, we address two gaps in the literature: (1) Longitudinal, deployment-scale evidence. Despite calls for practice-grounded assessment, there are few multi-month, mixed-methods deployments in evidence-dependent professional domains that connect log-level behaviors with self-reported outcomes and workflow narratives (Weidinger et al., 2025; Pang et al., 2025). (2) Multilingual, domain-bounded practice. Methods for evaluating multilingual coverage and workflow integration in domain-bounded tools used by globally distributed professionals remain limited; most existing evidence is short-horizon or scenario-based (Jabbour et al., 2025; Pang et al., 2025; Wang et al., 2024a). We report a five-month, mixed-methods deployment of AVA with 2,200+ users across 116 countries. Consistent with practice-grounded evaluation, we analyze real-world use across multiple axes, including task fit, adoption dynamics, workflow integration, multilingual/domain coverage, verification and contestability, learning curves, and boundary conditions, using platform logs, surveys, and interviews. 2.3. Generative AI for Specialized Knowledge Work Generative research tools for knowledge work optimize differently for scale, trust, and relevance, often leaving the evidentiary needs of policy and development practice only partially served. While the HCI community has developed domain-specific systems aligned with professionalsâ epistemic norms and workflows (Karnatak et al., 2025a, b; Fok et al., 2024), prior efforts largely target contexts that differ from the institutional and evidentiary demands of international development and policy research. Accordingly, we focus on a class of generative research agents designed to support information retrieval and synthesis, and examine how their design trade-offs create systematic epistemic gaps for policy and development work. We categorize these tools into three broad classes to highlight the specific epistemic gaps that AVA addresses. First, open-web and âdeep researchâ agents (e.g., Perplexity) prioritize scale by drawing on the public internet (Perplexity AI, 2024). While providing broad coverage, these systems can inherit ranking and visibility biases from web search and attention dynamics, which may over-surface highly visible public content (e.g., news and summaries) relative to dense operational or grey-literature reports that often underpin policy decisions (Fortunato et al., 2006). Second, user-defined retrieval systems (e.g., Google NotebookLM) originally prioritized trust by restricting generation to user-supplied documents (Google LLC, 2025). While recent updates have introduced âDeep Researchâ agents to autonomously scour the web for literature, transitioning from this closed-loop system to open-web retrieval presents a distinct trade-off: it alleviates the bottleneck of user knowledge but re-introduces the challenge of verifying the provenance of autonomously selected sources. Third, academic discovery engines (e.g., Elicit, Emergent Mind) optimize for scale and trust within peer-reviewed literature (Elicit, 2025). However, these tools face a problem of structural invisibility regarding development data. Academic agents rely on structured metadata, such as Digital Object Identifiers (DOIs) and citation graphs, to traverse the web. Grey literature including field evaluations, policy notes, and operational reports typically lacks these structural features (SchĂśpfel, 2010). Consequently, valuable operational knowledge remains in the âhidden web,â effectively invisible to agents that rely on bibliographic crawling. To address these gaps, we introduce AVA, a domain-specific generative research agent for policy and development work. AVA is a multi-agent, domain-bounded RAG system built on a verified library of over 4,000 World Bank Reports. Unlike open-web agents that seek to find new information, AVA is designed to reliably synthesize trusted information. This approach leverages Retrieval-Augmented Generation (RAG) to restrict the modelâs non-parametric memory to a curated corpus (Lewis et al., 2020). By bounding retrieval, AVA addresses the ranking bias of open agents and the structural blindness of academic agents, centering the grey literature that policy professionals routinely use (SchĂśpfel, 2010). Table 7 situates AVA within this design space. We contribute an empirical analysis of how policy and development professionals across 116 countries appropriate this domain-bounded assistant in their everyday evidence-based workflows. 3. System Architecture Overview. AVA is a transparent, multilingual, multi-agent, evidence-grounded RAG system that provides accurate, verifiable, and context-aware answers. The architecture pairs a hierarchically indexed, expert-curated corpus of 4000+ World Bank Reports with agentic, tree-aware retrieval and an ensemble framework that ensures accurate citations and abstention. We next formalize the design goals and resulting architectural workflow. Figure 3. AVA interface annotated with its core components: (A) query input box, (B) retrieval process trace, (C) inline verifiable citations, (D) evidence preview card, (E) source document viewer, (F) Save as note functionality, (G) Copy-Paste functionality, and (H) Tagging. See appendix A for detailed descriptions of each component. A five-step user flow in AVA: (1) submit a query; (2) check the returned citations; (3) verify that sources are credible and relevant; (4) save the verified response to notes; (5) edit the response if needed, with arrows indicating a loop between edit and save. Screenshots illustrate each step and the emphasis on citation-anchored verification. 3.1. Design Goals The system architecture is driven by four primary design goals: DG1 - Verifiability and Epistemic Humility Every generated claim must trace to specific source spans, with the system preferring reasoned abstention over unsupported assertions. This epistemic humility, addresses a core challenge in knowledge-intensive AI systems. Prior works on RAG-based systems demonstrate that accountable generation depends on hybrid retrieval combining sparse lexical matching with semantic embeddings, enabling claims to bind to verifiable evidence while maintaining high recall across paraphrase and language variation (Zhang et al., 2025a). We operationalize this through hierarchical document indexing in the RAG and mandatory source attribution and confidence thresholds that trigger non-answers when evidence is insufficient or conflicting. DG2 - Workflow Fit Effective humanâAI collaboration in knowledge work requires inspectability and reproducibility. AVA therefore provides citation-first outputs, exportable artifacts, and version-pinned runs that support fluent, audit-ready workflows. DG3 - Multilingual Access Retrieve over original-language materials while responding in the userâs language, treating parity of support across languages as a first-class usability criterion. Real-world policy and development research spans multilingual sources, yet translation introduces systematic errors that compound retrieval limitations. We extend RAGâs grounding promise to multilingual settings by preserving original citations while ensuring response quality. This is a core requirement for equitable global deployment, designed to serve a diverse community in their native languages. DG4 - Responsible Personalization Enable user-controlled personalization through explicit preference settings that steer retrieval and ranking while maintaining evidence standards and user agency. Effective personalization in socio-technical contexts draws on groun-ded signals rather than opaque heuristics. We constrain personalization to opt-in predicates that users can inspect and modify, expose when preferences affected results, and preserve user control over system memory. This approach parallels recent work using interaction traces to align AI behavior with group norms while preserving existing practices and user autonomy. 3.2. Core Components We proceed module-by-module, moving from data foundations to retrieval, from retrieval to verification, and finally to user-facing transparency and personalization. For additional details, see Appendix G. 3.2.1. Data Curation and Hierarchical Indexing We transform over 4,000+ World Bank Reports from PDFs into a knowledge base using custom layout parsing and OCR. This extracts text, tables, and spatial relationships while preserving document hierarchy. Documents are chunked into nodes at paragraph and cell levels. Each node receives a stable ID and hierarchical path (e.g., /section/paragraph). Nodes are bounded at 2048 tokens. This enables precise span-level citations, directly operationalizing DG1. The processed data feeds into a dual-backend index: A graph store captures the hierarchical tree, while a vector index holds semantic embeddings from Qwen3-Embedding-8B. This setup supports structural traversal and cross-lingual search to meet DG3. 3.2.2. Agentic Retrieval and Evidence Formulation A multi-agent architecture decomposes complex queries. This mitigates single-agent ReAct pitfalls (role drift, unstable planning) (Yao et al., 2023; Shinn et al., 2023; Li et al., 2023) by assigning distinct competencies: The Query Decomposer Agent extracts atomic sub-questions, intents (factual, analytical, comparative), and targets. Decomposition improves accuracy and reduces hallucinations (Petcu et al., 2025; Ammann et al., 2025). The Retrieval Planner Agent selects search strategies (lexical vs. semantic). This hybrid planning outperform semantic or lexical retrieval (Sidiropoulos et al., 2021). The Tree-Walker Agent traverses hierarchical structures and semantic neighborhoods. It switches between logical navigation (sections, references) and semantic exploration based on marginal gain, using coverage thresholds to stop. Graph-based traversal surface more diverse evidence than flat retrieval (Zhang et al., 2025b). Finally, the Drafting Agent consolidates retrieval paths into structured evidence packets containing relevant passages, hierarchical contexts, and metadata. These undergo de-duplication and ranking (using GPT-4o-mini re-scoring combined with base retrieval) to present relevant evidence while preserving conflicting perspectives. This ensures inspectable outputs (DG2) and aligns with prior work on grounding stability (Shao et al., 2025). 3.2.3. Information Synthesis and Verification Information synthesis employs a âmodel orchestraâ approach where specialized components handle distinct aspects of response generation before final integration. The information synthesis happens through the following four models: Evidence Classifier Model (based on fine-tuned BERT models): This model performs passage classification, entity extraction, and span detection to identify key information within evidence packets. Query Analyzer Model (based on a fine-tuned GPT-4o-mini): This model handles query understanding, evidence summarization, and retrieval planning tasks that require reasoning but not extensive generation. Response Drafter Model (based on a fine-tuned GPT-4.1 model): This model performs the primary response generation with structured, trace-based prompting that explicitly maps each generated claim to specific evidence spans within the packets. The prompting framework includes role specifications, evidence contextualization, and citation formatting instructions that ensure every factual assertion can be traced to verifiable sources. Verification Model (based on a fine-tuned GPT-4o-mini): Each draft answer is routed to a specialized verification model fine-tuned for fact-checking and inspecting tasks. The agent inspects sentence-level claims, aligns them with the retrieved evidence, and computes two internal scores: (i) Coverage: the proportion of claims that have direct documentary support, and (i) Agreement: the degree of consistency across independent sources. The agent also judges whether the available evidence is rich enough to sustain a thorough explanation. If either score falls below the high-confidence threshold i.e. confidence=high, or if the evidence is too sparse or contradictory for a complete response, the agent instructs the system to abstain rather than present potentially unsupported information. This verification architecture directly operationalizes DG1âs epistemic modesty principle. The system prefers abstention over confident but unfounded assertions, providing users with clear explanations of evidential limitations or insufficient coverage that prevented confident response generation. (See Appendix G.3) 3.2.4. User Personalization and Memory The interface renders responses in usersâ preferred languages while preserving original-language citations and evidence spans, supporting DG3âs multilingual access goals without introducing translation artifacts that could compromise verification. Interactive citations provide hover previews displaying source quotes in their original context and click-through links to highlighted passages in full documents. AVA currently supports queries in over 60 languages (see Appendix G.4 for details). The system maintains user-controlled memory through explicit save actions, allowing researchers to build personal knowledge collections, bookmark significant findings, and track research threads across sessions. All interactions are version-pinned with full citation trails, enabling users to export reproducible research artifacts, share findings with collaborators, and audit their analytical processes. When abstention occurs, the interface provides clear explanations of evidential gaps, suggests productive query refinements based on available evidence patterns, and offers related topics where stronger evidence exists, supporting iterative research workflows and collaborative analysis practices. 4. Methodology 4.1. Study Design and Setting We conducted a mixed-methods, randomized-invitation field evaluation of AVA. Participants were recruited via a large multilateral development bank (MDB). Employment at the MDB was not required; the sample comprised predominantly external practitioners working in policy, development, and related fields. Specifically, only 4.6% of registered participants were from the host institution; the remaining 95.4% were external participants. Participation was voluntary and secured via informed consent. Following baseline collection, registrants were randomized 80:20 into a treatment arm (offered immediate access) and a holdout control arm (no access). We employed an open-label intention-to-treat (ITT) design, analyzing outcomes via difference-in-differences (DiD) estimation to account for pre-existing differences (see Figure 4 for participant flow). The evaluation window extended through August 31. Detailed recruitment procedures, timeline milestones, and author positionality are provided in Appendix D Figure 4. Participant flow. From 2,764 initial registrants, participants were assigned to treatment (n=2,212) and control (n=552) arms. The final analysis included 243 treatment and 121 control participants who completed the endline survey. Flowchart of participant progression and attrition in the study. From an initial 2,764 interested individuals, the flow splits into a control arm with 552 participants and a treatment arm with 2,212 participants. The chart concludes with the number of participants who responded to the endline survey: 121 from the control arm and 243 from the treatment arm. 4.2. Quantitative Methods 4.2.1. Data Sources and Measures We integrated four sources: (i) baseline and endline surveys (demographics, roles, prior AI usage, adoption, productivity, satisfaction); (i) in-app pop-up surveys during deployment (trustworthiness, recommendation, satisfaction); (i) interaction logs (timestamps; session identifiers; query text; language; response metadata including citation counts; per-response feedback); and (iv) semi-structured interviews with users. Survey records were linked to platform accounts via the email used at registration; operational logs were pseudonymized with hashed identifiers. We operationalized key constructs including (1) AI Use for Work, (2) AI-Driven Productivity Gains, and (3) Perceptions of AVA. Detailed items and scoring rules are provided in Appendix D. 4.2.2. Quantitative Data Analysis Query Classification: We applied natural language processing (NLP) to classify user queries into policy themes and intent categories. The pipeline first normalized and sessionized text, automatically translating non-English queries into English. We then employed a deterministic, rule-based classification system using curated taxonomies to map queries to specific domains (e.g., Health, Infrastructure) and interaction types. To ensure coverage, uncategorized queries were imputed based on session context (forward-filling within a one-hour inactivity window). Detailed preprocessing steps, matching logic, and taxonomy definitions are provided in Appendix D. Behavioral Metric Derivation: Interaction logs were processed to derive the core behavioral metrics reported in Sections 5.1â5.3. We calculated Abstention Rates by tracking the proportion of queries where the system triggered a refusal state due to insufficient context (see Fig. 5). Citation Density was computed by extracting the count of unique document node references per generated response to quantify evidence reliance (reported in Section 5.1). Finally, to analyze Sustained Engagement, we segmented users into single-session and multi-session cohorts based on the session grouping logic; these cohorts serve as the basis for the user retention analysis (Section 5.2) and the Difference-in-Differences efficiency estimates (Section 5.3) Regression Analysis. To estimate causal effects of access to AVA, we conducted intention-to-treat (ITT) analyses comparing treatment and control groups using endline survey outcomes, complemented by difference-in-differences (DiD) estimation that leveraged both baseline and endline responses. Linear regression with robust standard errors was used throughout. Key self-reported outcomes included productivity (e.g., time saved, number of outputs), AI tool adoption frequency, and perceived work quality. 4.3. Qualitative Methods 4.3.1. Sampling and Participants From survey respondents who consented to follow-up, we used purposive sampling to achieve variation in geography, role and language. We conducted n=[n=\,[20]] semi-structured interviews via secure video calls. Table 9 reports participantsâ demograhic details. Our participants included 8 women and 12 men; age bands were 19â29 (n=3n=3), 30â39 (n=10n=10), 40â49 (n=4n=4), and 50â59 (n=3n=3). Career stage spanned early (n=9n=9), middle (n=7n=7), and senior (n=4n=4) professionals across sectors (universities, government, NGOs/civil society, international organizations, startups, private sector). Participants represented multiple regions, with substantial coverage from Africa (n=11n=11 across Nigeria, Uganda, Ghana, Tanzania, Angola), as well as the Americas (USA, Argentina, Canada), Europe (France), and Asia (Malaysia). Linguistic diversity was high (11 reported multilingual proficiency, e.g., English with Luo/Swahili, Igbo, Yoruba, French, Spanish, Mandarin, Arabic). 4.3.2. Data collection Data was collected through semi-structured interviews conducted via secure video calls. Each interview lasted 45â75 minutes. The interview protocol guided conversations to explore, usersâ work and task contexts, perceived benefits and risks of using AVA, strategies for trust calibration (e.g., use of citations and handling of abstentions), effective prompting techniques and cross-tool comparisons. All interviews were audio-recorded with participant consent, transcribed verbatim, and then de-identified to ensure confidentiality. 4.3.3. Data Analysis The first author conducted all interviews and prepared verbatim, de-identified transcripts. We employed thematic analysis. An initial codebook was developed from an open-coding pass on a subset of transcripts by the first author and circulated to the team for critique. The team independently reviewed the codebook, proposed refinements by jointly coding a common subset. Discrepancies were resolved through discussion; the refined codebook then guided subsequent analysis. The first author completed coding of the remaining transcripts and led synthesis, with regular team debriefs. We combined deductive codes aligned with our research questions (e.g., evidence use, abstention handling, prompting challenges) with inductive codes grounded in the data. Recruitment and interviewing stopped once meaning saturation was reached. 4.4. Mixed-Methods Integration (Triangulation) We integrated quantitative and qualitative strands at the interpretation stage. Survey and causal estimates characterize population-level patterns, while interaction logs provide granular behavioral traces. Interviews explain mechanisms underlying these patterns (e.g., why page-anchored citations serve as evidence shortcuts, how abstention calibrates reliance). We report convergences and divergences across strands and link them to design implications. In Sections 5, we first present quantitative analyses that characterize broad engagement patterns and their associations with self-reported impacts. We then draw on qualitative data, in section 6, from 20 semi-structured interviews with a with a diverse, global cohort of users, including university lecturers, independent researchers, NGO professionals, and startup founders (see Table 9), to explain the contexts and mechanisms underlying these patterns. 5. Quantitative Results 5.1. Product Evaluation: How Users Engage with AVA (RQ1) During the study period, from May 12 to August 31, 2025, 2,679 unique individuals signed up for AVA. Of these, 1,055 users (39.4%) were successfully matched to the baseline survey via email addresses. Among the 1,037 matched users with country information, 782 (75.4%) were based in the Global South, including Nigeria, India, and Brazil (see Fig 7). Although users were geographically diverse, most queries were in English (83.8%), followed by French (4.8%) and Spanish (3.0%). Figure 5 shows the total volume of queries over time alongside the proportion of queries AVA declined to answer due to insufficient corpus coverage. Early in the deployment, when the knowledge base contained approximately 50 reports, the abstention rate ranged from 40â70%. Following the expansion of the corpus to over 4,000 World Bank reports on 9 July 2025, the abstention rate dropped sharply and remained below 10% for the remainder of the study period. Across the full deployment, responses generated by AVA cited an average of five documents per query (range: 1â23), indicating consistent reliance on multi-document evidence. Figure 5. Five-day total query volume and abstention rate (% of queries receiving a reasoned âno responseâ). Abstention is high during early deployment with limited corpus coverage and drops sharply after corpus expansion, remaining low thereafter. A dual-axis line graph showing the 5-day total query volume and the percentage of no-response queries from May to August 2025. The number of total queries (left axis) shows two significant peaks: one in early June at approximately 1,200 queries and a sharper peak in early August reaching nearly 1,600. The percentage of queries with no response (abstention rate) (right axis) generally follows an inverse trend, dropping to its lowest point of near zero percent during the major query peak in August. To understand how users engaged with AVA, we first classified queries by policy theme (See Table 1). Two themes, namely, Human Capital and Fiscal Policy/Private Sector dominated, accounting for roughly two-thirds of all queries. These queries focused on issues such as jobs, education, poverty, growth, debt, and taxation trade-offs, indicating strong demand for evidence-informed policy synthesis. Example questions were largely how-to and forward-looking (e.g., âHow should labor regulations evolve post-pandemic?â; âHow can policymakers balance investment and innovation?â), suggesting that users sought support in framing complex policy questions and identifying plausible policy directions, rather than retrieving narrow factual information. We further classified queries by primary task typeâdiagnostics, design, and evaluation (See Table 2, Fig 10). Our analysis differentiates between queries focused on diagnostics (seeking to understand problems, data, and challenges), design (how to improve policies, find best practices), and evaluation (evidence, effectiveness measures, and impact evaluation support). Overall, we find that most queries ask AVA to complete diagnostic tasks (See Table 2, Fig 10), although the content of queries reinforce the use of AVA throughout the policymaking process,from defining problems, designig solutions, and testing impact. Table 1. Summary of policy themes reflected in queries This table summarizes the main policy themes reflected in user queries. Itâs organized into five categories: Human Capital, Fiscal Policy, Digital Transformation, Environment, and Infrastructure. For each theme, the table provides the top keywords, an example query, and a corresponding value indicating its frequency. Policy Theme Top Keywords % of Queries Example Query Human Capital education, school, poverty, program, job 33.5 How should labor regulations evolve post-pandemic? Macro-Economics and Governance government, income, growth, economy, taxation 20.7 How do countries balance investment and innovation? Digital Transformation data source, tools, technology used, analyze, analytic 17.4 What types of digital infrastructure hinder ing digital transformation in Nigeria and Kenya? Environment and Climate fish, waste, emissions, air, age 13.5 How to integrate adaptation and mitigation for climate change in Nigeria? Infrastructure cargo, earthing, clearance, cargo service, circuit 5.7 How to amplify non-motorized transport for sustainable urban mobility in Uganda? Highlight indigenous paradigm. Table 2. Summary of policy tasks reflected in queries. Policy Task Sample Keywords % of Queries Example Query Diagnostic what, explain, data on, impact of, challenges 69.0 What are measures used in assisting Young Africans and changing their mindset? Design how to, improve, best practice, options for, recommendation for 22.0 Write a proposal for masters degree thesis on how to make Ghana government better. Evaluation evaluation, assessment, evidence, result, effectiveness 9.0 I am conducting an impact evaluation on an LPG provision scheme. Please identify potential proxy variables for income. Notably, evidence from our endline survey suggests that use of AVA may have substituted use of other AI products. When reporting the AI services they used in the last week, only 24% of AVA users, those who were invited to test the product and actually used it, reported also using another AI service (e.g., ChatGPT, Gemini, etc.) whereas 80% of non-AVA users (participants who signed up for AVA but were not yet given access) reported using another AI service. 5.2. User Evaluation: Satisfaction With and Perceptions of AVA (RQ3) As a behavioral indicator of satisfaction, we explored the proportion of users who engaged with AVA in a single session or multiple sessions and characteristics associated with engagement in multiple sessions. Overall, we found that 73.4% of total AVA users engaged in only one session whereas 26.6% of users returned to the platform to engage in multiple sessions. While most users did not return for multiple sessions, the rate is comparable or better than industry benchmarks for productivity apps (Statista, 2025). In Table 3, we also summarize the characteristics of baseline respondents who engaged with AVA for a single session or multiple sessions, finding that users had overall similar profiles. Table 3. Summary statistics of characteristics of registered users and users with different levels of engagement. This table provides a statistical summary comparing the baseline characteristics of three distinct user groups: all Registered users, users who had only a Single Session, and highly engaged users with Multiple Sessions. Registered N = 1,055 Single Session N = 506 Multiple Sessions N = 254 Global South (%) 74.2 74.1 74.4 Affiliation (%) Think Tank 4.2 4.3 5.5 Academic 15.7 15.8 15.7 Multilateral Organizations 6.0 6.3 3.9 AI Experience Multiple Times Every Day 39.1 38.3 44.9 About once or twice on most days 18.8 20.2 14.2 A few times throughout the week 25.7 25.3 26.4 Rarely 5.0 4.7 3.5 Not at all 1.5 1.4 2.0 Never use AI for work 9.9 10.1 9.1 Based on the endline perception data from 118 AVA users, the tool is viewed very positively. A high percentage of users found AVA to be an effective and reliable resource, with strong endorsement for its efficiency, accuracy, and overall value. In addition to the behavioral indicators of satisfaction, we also collected survey data from individuals using AVA as a pop-up survey to assess perceptions of the tool. Results from 100 individuals, indicated that 68% find content relevant, 65% the citations relevant and 72% were satisfied with AVA and would recommend the tool to a colleague. Table 4. User perceptions of AVA from the endline survey (n = 118). Percentages indicate respondents who agreed with each statement. Percentage of users who agree with the statement (n=118) AVA helps me understand key insights from World Bank documents more efficiently. 80.5% The synthesized insights provided by AVA are accurate and trustworthy. 82.2% Using AVA has improved the quality of my work or research. 76.3% I would recommend AVA to others working in development policy or research. 89.0% 5.3. Impact Evaluation: Improvements to Work Efficiency (RQ3) To assess the impact of AVA, we asked respondents of the endline survey to rate how much time they have saved on work through the use of AI and the quality of their work. We first ran an intention-to-treat (ITT) analysis, assessing differences in time saved on work among endline respondents who were invited to register to use AVA and those who were assigned to the waitlist control. ITT analysis revealed no significant differences in time saved between the treatment and control groups for either the lower bound (b = -0.56, p Âż .05) or upper bound estimates of time saved (b = -0.78, pÂż.05). In addition, no impacts were found when we asked respondents to estimate the number of major outputs they completed in the last week and report if the use of AI tools helped them produce a greater number (b = 0.08, p Âż .05) of outputs or saved time in producing those outputs (b = -0.78, p Âż .05). Table 5. Intention-to-treat (ITT) Regression Results This table presents the results from four regression models designed to measure the impact of an âIntent-to-Treatâ (ITT) condition on four different productivity outcomes: time saved (lower and upper bounds) and the number of additional or faster outputs. The table displays the estimation coefficients and robust standard errors in parentheses. (1) (2) (3) (4) VARIABLES Time AI tools save (lower bound) - Hours Time AI tools save (upper bound) - Hours Additional outputs produced because of AI - Number Outputs completed faster because of AI - Number Intent-to-Treat -0.568 -0.777 0.0847 -0.225 (0.416) (0.742) (0.334) (0.382) Constant 4.010â 7.172â 2.444â 3.909â (0.341) (0.616) (0.277) (0.305) Observations 289 289 288 289 R-squared 0.007 0.004 0.000 0.001 Robust standard errors in parentheses. pââŁâ<0.01^***p<0.01, pâ<0.05^**p<0.05, pâ<0.1^*p<0.1 In addition to ITT, we conducted a difference-in-difference (DiD) analysis to explore whether different intensity levels of AVA usage (single use vs. multiple session use) were associated with different levels of impact on work efficiency. Here, we find that individuals who engaged with AVA over multiple sessions reported saving more time on their work in the last week for both the lower (b = 1.84, p ÂĄ .10) and upper bound of estimates (b = 3.27, p ÂĄ .10). Table 6. Difference-in-Differences Regression Results This table displays the results from a Differences-in-Differences (DiD) analysis. This model is used to estimate the causal effect by comparing a AVA returning users with one time AVA users over time. The two models in the table estimate the interventionâs impact on the lower (1) and upper (2) bounds of time saved, measured in hours. The table displays the estimation coefficients and robust standard errors in parentheses. (1) (2) VARIABLES Time AI tools save (lower bound) Time AI tools save (upper bound) Round = Endline -2.766â -4.796â (0.705) (1.253) One Time vs Returning 0.554 0.640 (0.789) (1.492) DiD (Intensity) 1.838â 3.272â (1.035) (1.886) Constant 4.946â 9.027â (0.598) (1.122) Observations 160 160 R-squared 0.132 0.112 Robust standard errors in parentheses. pââŁâ<0.01^***p<0.01, pâ<0.05^**p<0.05, pâ<0.1^*p<0.1 6. Qualitative Results We organize our qualitative findings into four key themes. First, we detail how participants integrated AVA into their daily evidence workflows, its distinctive value, and professional impact (Section 6.1). We then explore the mental model and persona they developed around its unique behavior, particularly reasoned abstention (Section 6.2). Next, we analyze the foundations of trust in AVA through both its features and institutional corpus (Section 6.3). Finally, we discuss the ethical debates on disclosure that emerged from its successful adoption (Section 6.4). 6.1. Positioning AVA within Evidence-Based Workflows (RQ1) In this section, we report participantsâ integration of AVA into everyday evidence work. We analyze their strategic division of work across tools, their reliance on AVAâs page-level citations, and their appreciation for its distinctive depth and policy-oriented synthesis, including graphics and visuals generated by synthesizing across documents. Finally, we demonstrate the translation of these capabilities into concrete professional achievements, enabling participants to prepare for high-stakes meetings, advance stalled projects, and publish with greater confidence. 6.1.1. Multi-Tool Workflows and Ecosystem Strategies Participants organized their work across complementary systems rather than relying on a single general-purpose model. Generalist LLMs (e.g., ChatGPT, Claude, Gemini) were used for ideation, outline scaffolding, and language polishing, while AVA was consistently positioned as the evidence engine for claim-level facts requiring citations and page-level grounding. As P1 explained: âI use ChatGPT to break down tasks when I am working on a project, and I use AVA to get the âreal informationâ I need.â Participants recognized AVA as a domain-specific tool grounded in World Bank Reports, turning to it for policy and development work. P2 noted: âAnything to do with policy, AVA is very straight, very direct and very helpful.â Participants also developed strategies for managing ecosystem boundaries. When AVAâs curated corpus could not address queries, participants reported transitioning to other tools like Perplexity (P12). However, breakdowns with generalist tools, such as receiving false information, often prompted a return to AVA. As P9 shared, Sometimes I use ChatGPT to do my analysisâŚI say, âPlease give me this, give me this,â and I discovered that one of the outputs that came out was false. So when I saw AVA, I said let me try this thingâŚand lo and behold, correct information came out, a lot of references, a lot of case studies. This preference for AVA extended beyond mere accuracy to the quality and nuance of its output. P18, who used âdifferent tools for different jobs,â explained that while other models were useful for broad research, ââŚthe nuances of the data, Ava picks up a little better.â. This orchestration demonstrates calibrated multi-tool strategies where participants developed quality hierarchies and taskâtool matching based on empirical experience with different systemsâ strengths and limitations. 6.1.2. Citation Verification as an Evidence Shortcut (RQ3) The provision and verification of citations was consistently described as AVAâs most valuable feature. Participants emphasized that citation verification supported their work through three key mechanisms. First, it accelerated source finding by narrowing thousands of pages of World Bank Reports, or even the internet, to only the most relevant sections. P6 noted: âBecause it helped me streamline what I am looking for without wasting so much time searching for it on the web.â Second, it facilitated scanning, enabling users to access highlighted excerpts in context rather than manually sifting through the entire documents. Third, it streamlined fact checking, allowing participants to click citations and immediately confirm whether claims were accurately grounded in sources. These mechanisms produced massive efficiency gains. P9 quantified similar time savings: âSometimes it takes me four days, and this AVA, maybe just five minutes to do my query, and in another one minute, I am getting the information I want. Itâs faster, easier.â Collectively, these accounts highlight how AVA transformed citation verification from a resource-intensive bottleneck into a lightweight, reliable shortcut, simultaneously enhancing trust, narrowing reading scope, and increasing productivity. 6.1.3. Distinctive Value Propositions of AVA Based on their day-to-day use of AVA, participants contrasted it with general-purpose AI tools and highlighted four perceived advantages: deeper reference-rich grounding, improved readability, policy-oriented structuring, and visualisations that supported interpretation and communication in their workflows. Participants characterized AVA as a way to turn World Bank Reports into immediately usable professional insights. Several participants compared AVA with other AI tools and emphasized that AVA delivered deeper, reference-rich responses. As P9 highlighted: âWe are using other AI tools, but we are not getting what we want, and thatâs very deep knowledge and references. When I used AVA, it was deep. It gave me case studies, references.â Participants also noted that AVA made World Bank materials easier to engage with. P13 described difficulty with statistics-heavy documents: âThere are no stories in World Bank Reports, so people are not interested to read the statistics, but AVA makes it easier to read.â Participants described AVA as moving beyond summary to policy synthesis; as P5 put it, âAVA tried to bring out some policy recommendations which you cannot see in some of the other AI platforms.â P5 further noted that responses were organized into compact summaries and tables, helping them scan and locate relevant evidence quickly. His remark illustrates how participants differentiated AVA from the general-purpose AI tools they already used. While they knew other tools could generate generic ârecommendations,â participants treated AVAâs outputs as substantive policy recommendations because they were grounded in World Bank Reports and organised into compact, citable formats that made relevant evidence easy to scan and verify. Beyond text, participants shared that AVAâs graphics (e.g., tables, charts, trends) helped them engage with World Bank Reports more effectively. For example, P14 shared âWhen I was preparing to attend meetings at the United Nations (UN), I was able to get visualizations of information as opposed to having to take a data set and try to parse out key trends. I could just put in the prompt into the system and it auto-generated that information for me. In in terms of getting some quick stats I found it very useful in helping and assisting and amplifying my preparation.â Across interviews, participants praised these visuals and requested the ability to export them as images with built-in attribution to AVA. These perspectives indicate that AVA surfaces and structures World Bank knowledge for professional use, supporting comprehension, evidence-building, policy articulation, and communication. 6.1.4. Concrete Professional Milestones AVAâs impact extended beyond efficiency to enabling professional confidence and substantive scholarly contributions. For example, P19, working on policy development for street children in Borno state, described using AVA to advance a book manuscript that had remained unsubmitted for years: âI have never submitted it to any superior before because I felt I needed more inputs. More feedbacks. Within some minutes, I was able to obtain positive feedback from AVA and I have submitted it to my superior.â (P19) Other participants reported similar gains. P1 credited AVA- supported research with contributing to a successful publication accepted by an international journal, while P2 described progressing a book manuscript with AVAâs support. Participants also noted that AVA facilitated preparation for both routine and high-stakes meetings. P20 used AVA to quickly ramp up on new development topics for team discussions, while P14 highlighted its role in preparing for United Nations meetings. Finally, participants emphasized that AVA enhanced their confidence in professional interactions. As P9 reflected: âIt gives me an extra lot of confidence because when I speak with some other colleagues who are also into development-related work⌠thereâs this confidence I have right now that Iâm going to give them very rich information until they ask further questions.â These accounts suggest that AVAâs value extended beyond task completion to professional development and credibility, helping participants advance manuscripts to supervisory and scholarly review, prepare confidently for high-profile engagements, and speak authoritatively with peers. 6.1.5. Auxiliary feature: Multilingual Capabilities Despite supporting 60+ languages, AVAâs multilingual capabilities were largely unused, 83.8% of interactions occurred in English (see Fig.9). When interviewed, participants (n=20) attributed this to workplace norms: they mostly conduct and publish research in English, despite knowing their native language. Those who tried other languages rated quality highly; for example, P17 said âItâs written in a good French, itâs not a robotic French. Itâs pretty high level.â Similarly, P7 reported high satisfaction with Spanish outputs. Given the English-heavy usage patterns, our non-English sample remains small, limiting conclusions about multilingual performance. 6.2. Reasoned Abstention: When AI Says âI Donât Knowâ (RQ2) In this section, we report participantsâ responses to AVAâs reasoned abstention. We examine user reception and appropriation of âI donât knowâ responses as signals of honesty, evidence boundaries, and opportunities for prompt reformulation. We then describe the anthropomorphic interpretations that abstention fostered, with participants casting AVA as a reserved but trustworthy colleague while expressing aspirations for warmer, more supportive interaction. Finally, we present user suggestions for balancing epistemic honesty with workflow efficiency, including collaborative query refinement and external handoffs. 6.2.1. Reasoned Abstention: Reception, Appropriation, and Limits We designed AVA to practice reasoned abstention under DG1. The system states when evidence is insufficient, provides a brief justification, and offers targeted redirection. Participants generally welcomed this behavior in high-stakes work (e.g., policy drafting and development projects), preferring an explicit âI do not knowâ to speculative answers. P13 shared, âIt does not fabricate, thatâs the credibility of AVA. It does not fabricate. If it doesnât know, it will not provide you data.â This illustrates how some participants interpreted AVAâs abstention and citation mechanisms as strong safeguards against fabrication in this deployment. We examine how such trust is calibrated, and where it is contested, in Sections 6.3 and 7. Across participants, reasoned abstention was appropriated in multiple ways (Dourish, 2003), even though it was initially designed to prevent ungrounded claims. For example, several treated abstention as a cue about evidence limits and a pointer to research gaps. P5 noted: âEven if AVA says I donât know, to me, itâs very useful. âI donât knowâ is half of the answer to what I am looking for.â P3 valued the honesty of non-response, while P7 contrasted AVA with general-purpose systems: âSometimes they tend to give you an answer even when they are not pretty sure âŚ[that] gives you a misconception they will always have the answer.â Other participants treated abstention less as a definitive endpoint than as an invitation to refine their information-seeking approach. P6 described systematically rephrasing queries until achieving results, while P7 positioned abstention as a brainstorming catalyst. These users treated abstention as an invitation for prompt reformulation and negotiation. However, not all participants welcomed abstention. Some (P4 and P18) interpreted it as system limitation rather than calibrated caution. P18 shared: âI donât want the choice being no answer. If you keep saying I donât know, Iâm just going to go and use Google.â These dissenting views highlight the tension between abstention as a trust-building mechanism and abstention as a productivity barrier. We discuss this in the following sections. 6.2.2. Two Modes of Anthropomorphism: Descriptive Personas and Aspirational Desires AVAâs reasoned abstention created an unexpected pathway to anthropomorphization, helping participants make sense of its system boundary and develop workable mental models. Participantâs connected AVAâs behavior to familiar human communication norms. As P19 explained: âFor me as a human, whenever you ask me something that I donât know, I will be truthful and honest, and if I know places you can reach out to get that information, I would redirect you there. Incorporating such a feature into AVA is a wonderful one.â By aligning with norms of human honesty, abstention laid the groundwork for richer anthropomorphic interpretations. We observed two distinct modes of anthropomorphizing. First, descriptive anthropomorphizing emerged as participants characterized AVA through personality traits inferred from its behavior. For example, P2 depicted AVA as a reserved yet efficient colleague: âIf I think of AVA talking as a human being, AVA is a bit reserved, a person of few words. AVA doesnât want to force words into your mouth; AVA wants you to think for yourself. Sheâs unique and also very efficient.â P2 further characterized AVAâs abstention as deliberate encouragement for independent thinking, contrasting this with the verbosity of generalist LLMs: âFor ourselves, we think and we get to the answer. But for ChatGPT, GPT wants to give you the first word. GPT thinks for you.â Second, aspirational anthropomorphizing reflected participantsâ desires for enhanced interaction qualities. P13 wished for AVA to feel like a closer companion: âHumanize it, make it human, make it as my fellow, make it just my brother⌠And make it learn to reflect the individual person who is using it.â He suggested softening abstentions with service-oriented phrasing, replacing âI donât know about thisâ with âI am still learning about this. Is there something else I can help with for now?â, to project a more helpful, service-oriented persona that positions the AI as a cooperative partner rather than a simple tool. Thus, anthropomorphization functioned as a sensemaking strategy, rendering system constraints legible and workable (predictability, calibrated reliance, honest), while simultaneously providing a vocabulary for articulating unmet social needs (warmth, rapport). 6.2.3. Balancing Honesty with Usability While the majority of participants appreciated AVAâs principled abstention, some identified it as a workflow bottleneck. A non-answer, however well-justified, still left their information need unmet. In response, these participants suggested alternative approaches. For example, P7 described how AVA could leverage its corpus knowledge for structured guidance, proposing that when the system finds related but imperfect matches, it should use those results to help users formulate more targeted queries: âMaybe, for example, in this case that I wasnât able to find the information that I was looking for, but then AVA sent me this great information from another region of the world. If the tool can give you not only the suggestion, but refine the prompt with you and say: âI couldnât find information for your specific region, but you could try refining your search. Can you provide more details, such as a year, a specific mineral, a community, or a particular government agency? â This desire for collaboration also extended to the interactionâs conversational style. P9 focused on the emotional tone of the interaction, arguing that the bluntness of an âI donât knowâ response felt ârudeâ. Instead, he proposed a more engaging, suggestive approach to maintain the conversational flow: âIt could come up with, âdo you mean?â⌠Or you can look out for this from this⌠that kind of trying to engage you a bit⌠âI donât knowâ sounds so rude.â This pattern reframes the interaction as a helpful clarification rather than a conversational dead end. Finally, when query refinement was insufficient, participants wanted AVA to act as a knowledgeable guide to the broader information ecosystem. Participants proposed that AVA should guide them to external resources, such as other LLMs or web results. These suggestions reveals participantâs vision for a more collaborative AI that pairs honest refusal with external handoffs to reduce workflow friction. While we present these findings here, we analyze their deeper implications for trust and multi-tool ecosystems in the Discussion. 6.3. Trust Calibration in Evidence Practices (RQ2) Trust in AVA was calibrated rather than assumed. We adopt Gambettaâs definition of trust (Gambetta, 2000) as the belief that another agentâs actions are sufficiently beneficial, or at least not harmful, to justify cooperation. Participants did not take outputs as inherently reliable; rather, they assessed trust through two intertwined layers. At the feature level, page-anchored citation verification and reasoned abstention functioned as reliability cues, indicating when claims were grounded and when the system should refrain. At the dataset level, confidence derived from AVAâs curated corpus of World Bank Reports. The subsections that follow examine how these feature- and dataset-level mechanisms informed participantsâ decisions to rely on AVA in evidence work. 6.3.1. Feature-Level Trust Participants located trust in two key interface mechanisms: page-anchored citation verification and reasoned abstention. First, citation verification was highly valued, with 17 of 20 participants rating it 5/5 in usefulness and the remainder rating it 4/5. Participants reported selective verification rather than exhaustive checking, typically when claims were surprising, counterintuitive, or numerical. Across interviews, participants reported no mismatches between claims and sources. Qualitatively, participants contrasted AVAâs page-level, context-anchored citations with the opaque âlink dumpsâ of general-purpose systems. As P17 noted: âThis commitment to bring on the sources and say it comes from there. If we compare with other AI engines, they give you everything which has been used by the model and then a couple of links at the bottom, and good luck for you to link or address it.â P20 emphasized that seeing source context was central to perceived reliability, as it is âvery helpful to see where the data comes from and also the context in which the data was presented, which reassures you the model does not hallucinate.â Second, reasoned abstention operated as a complementary trust mechanism by preventing overclaiming, particularly in high-stakes contexts. As noted in Section 6.2.1 , participants (e.g., P7) contrasted AVAâs cautious non-response with the overconfident behaviour of general-purpose systems. This honest abstention helped participants calibrate their reliance on the system, distinguishing AVA from tools that might overstate confidence. 6.3.2. Dataset-Level Trust Source credibility was enhanced by corpus provenance. Participants emphasized that AVAâs foundation in a curated library of official World Bank Reports provided inherent trustworthiness, as many had previously relied on this repository in their professional work. P5, a researcher, noted his long-standing relationship with World Bank Reports, mentioning he has been using them since his undergraduate studies. This prior familiarity with institutional sources reduced concerns about data quality independent of AVAâs interactive features. P1, a lecturer, articulated this trust in institutional authority: âItâs coming from World Bank or World Health Organization or UNDP. I know that this [is] authentic. So I prefer even getting my information from this site than every other out thereâ. Similarly, P6, a policy analyst, noted: âBecause I know World Bank is the best institution for reports and research⌠if Iâm putting [it], I donât have to think that I havenât done enough research because I know those who put it there are the originatorsâ. P12 positioned AVA as competitive with major AI platforms specifically due to its curated approach: âI feel AVA has the capability of being on the same page as ChatGPT⌠AVA has an advantage because its information is not just picked up internet-wise, [itâs] actually concrete information.â Overall, participants noted feature-level mechanisms and dataset provenance were mutually reinforcing. Citation affordances are only as credible as the materials they point to; conversely, a reputable corpus attains practical value when claims are traceable to page-level evidence. 6.4. Contested Approaches to Disclosure (RQ4) Participants articulated sophisticated yet conflicting perspectives on whether AVAâs assistance should be disclosed in professional documents, often drawing comparisons with other AI tools to clarify their positions. Such comparative reasoning is consistent with prior work in HCI and the social sciences, where individuals use analogies to familiar systems to make sense of novel or ambiguous technologies and to articulate emerging norms and boundaries (Schwarz-Plaschg, 2018b, a; Ratzan, 2000) These accounts reflected an ongoing negotiation of emerging professional norms rather than adherence to established ethical guidelines. While few participants ( P3 and P6) explicitly argued that disclosure was necessary, others expressed more ambivalent or opposing views. Across this diversity of opinion, four distinct orientations emerged, which we detail below. 6.4.1. Disclosure as Strategic Advantage Some participants framed disclosure as an opportunity to enhance transparency and credibility. P12 argued that institutional provenance could strengthen the legitimacy of outputs: âOh, yes. I feel it should be disclosed. It builds transparency⌠If I go to a meeting with my research findings and I say the AI is based on World Bank Reports and everybody knows the World Bank is a big entity and all the information is correct, it gives me some kind of positive impression.â P1 similarly noted that disclosure could promote broader AVA adoption, reframing it as an advantage rather than a liability. 6.4.2. Positioning AI Use as Ordinary Practice Others challenged disclosure requirements altogether, framing AI as a routine professional aid that do not require explicit acknowledgement. P13 drew an analogy to secretarial assistance: âI have a secretary and my secretary is helping to type a letter. So should I tell the person I sent the letter this has been written by my secretary?â He described AI as a seamless co-creator: âI consider AI as a tool, as a part of me. We have created it together.â 6.4.3. Contextual and Domain-Specific Norms Several participants emphasized situational nuance. P9 distinguished between academic contexts, where disclosure was expected, and applied development settings, where effectiveness mattered more: âIf I can lift the whole document and implement it here and it works, maybe change one or two things. For me, I donât see it as anything. Disclosure is important for academic writing though.â Others, such as P14 differentiated computational from intellectual AI support, suggesting that while data processing might not require disclosure, creative writing did, due to plagiarism concerns. 6.4.4. Reframing Disclosure as Professional Integrity Finally, some participants shifted the focus from rules to professional integrity. P11 critiqued denial of AI use as disingenuous: âAnyone saying âI donât use AI for anythingâ, then I think itâs just a lie. The world has adapted to it already.â He argued that the true ethical imperative was to build competence: âYou need a lot of work to actually know how to use AI tools properly.â P20, who viewed AI as a thought partner rather than a copy-paste substitute, concurred, noting that while AI carries risks, âthe human reviewer has already validated and removed those risks,â making disclosure unnecessary once content had been vetted. In sum, these perspectives illustrate professionals actively constructing contextual ethics around AI disclosure, balancing strategic advantage, seamless integration into practice, domain-specific norms, and evolving standards of honesty in an AI-ubiquitous landscape. 6.5. Triangulation of Quantitative and Qualitative Findings Our mixed-methods strands converge on three points. First, quantitative analyses (Tables 2, 3, 4, 5, 6) show broad engagement patterns and strong associations with self-reported efficiency (80.5%) and work quality (76.3%) (See Table 4). Interviews explain the mechanisms underlying these high ratings, namely page-anchored citation verification as an evidence shortcut. This feature was rated 5/5 on usefulness scale by 17 of 20 participants and calibrated reliance via reasoned abstention. Second, doseâresponse patterns among returning users showing time savings (b=1.838b=1.838, p<.10p<.10; b=3.272b=3.272, p<.10p<.10) complement the ITT null results (b=â0.568b=-0.568, p>.05p>.05; b=â0.777b=-0.777, p>.05p>.05); interviews clarify when repeated use was associated with efficiency gains through multi-tool orchestration and verification routines that develop over time. Third, quantitative indicators of high trust (82.2% found AVA trustworthy) (See Table 4 and recommendation rates (89.0%) (See Table 4) align with interview accounts of provenance-first interaction grounded in World Bank institutional credibility. We also observe productive divergences. Some participants experienced abstention as a workflow bottleneck despite overall high satisfaction (72%); where available, logs and interview reports contextualize exposure to abstention events, motivating design implications (guided reformulation, external handoffs). Finally, English-dominant usage in logs (83.8% of interactions) is consistent with interviews citing workplace norms rather than quality concerns. We synthesize these meta-inferences and design implications in the Discussion. 7. Discussion We draw on data from a large-scale, in-the-wild global deployment of AVA used by more than 2,000 policy professionals across diverse organisations in 116 countries for 5 months, supplemented by 20 in-depth qualitative interviews (See Figure 2). In this paper, we use AVA as a case study to examine how established techniques such as corpus curation (Oche et al., 2025), reasoned abstention (Wester and others, 2024; Kim and others, 2024), and page-level verification (Huang et al., 2024; Xia et al., 2025), function when implemented in a real-world, high-stakes setting, rather than to claim novelty of the underlying architecture (Wen et al., 2025). The contribution we offer is therefore empirical and conceptual: we situate AVAâs value in insights from the deployment about how policy professionals engaged with an evidence-bounded generative system in their day-to-day work and the professional practices that formed around it, spanning appropriation(RQ1) (Dourish, 2003) , trust calibration (RQ2) (Lee and See, 2004), long-term use and perceived efficiency gains (RQ3 (Orlikowski and others, 1995), and norms of disclosure and accountability (RQ4) (Rezaei et al., 2024; Formosa et al., 2025). The Necessity of Agentic Complexity: Before detailing specific design lessons, we address the architectural necessity of the multi-agent retrieval system (Stage 2). While we observe that user trust is grounded primarily in the curated corpus (Stage 1) and the verification pipeline (Stage 3), we posit that the intermediate agentic overhead (Stage 2) is operationally essential in the context of policy and development. Policy analysis queries frequently require synthesizing dispersed evidence across documents and languages, conditions under which simpler, single-agent RAG pipelines fail to maintain coverage and precision (Dalglish et al., 2020). The multi-agent structure enables the system to plan, navigate document hierarchy, and aggregate diverse sources more reliably, thereby providing the necessary evidence pool for Stage 3 verification to ensure evidentiary correctness. In this section, we synthesize these insights into three transferable design lessons for building specialised AI systems in other high-stakes domains, such as healthcare, legal practice, and governance. We then introduce a fourth, forward-looking design recommendation: an interoperable AI ecosystem that supports coordination across specialised and general-purpose AI systems. 7.1. Lesson 1: Designing Specialized AI Systems for High-Stakes Knowledge Work Requires an End-to-End Trust Pipeline Our empirical findings indicate that AVAâs successful adoption was rooted in two principles that address the specific epistemic needs of policy and research professionals: Dataset-Level Trust (grounded in the curated World Bank corpus) and Feature-Level Trust (encompassing reasoned abstention and clickable citations). Specifically, the observed Feature-Level Trust decomposes into mechanisms residing at the Model Layer (reasoned abstention) and the Interface Layer (citation verification). In the discussion below, we synthesize the idea of a trust-pipeline that spans the entire system architecture. (1) Dataset-Level Trust (Curation): The system established dataset-level trust by grounding its responses in a curated library of World Bank Reports, a corpus with which participants were already familiar and whose institutional provenance they respected. Trust was therefore inherited from established professional practice. (2) Model-Level Trust (Abstention): In AVA, the model behavior shapes how inherited trust from the curated dataset is preserved. AVAâs reasoned abstention mechanism, implemented through a fixed evidence-verification threshold, constrained the system to answer only when sufficient supporting evidence was available, otherwise abstain. Most participants interpreted âI donât knowâ responses not as system failures, but as boundaries of the system and as preferable to the speculative or fabricated answers they commonly encountered in generic AI systems. This evidence-sensitive response strategy functioned as a critical safeguard in a context where speculative or fabricated answers could have significant downstream consequences. (3) Interface-Level Trust (Verification): Finally, trust is enacted at the interface, where verification features allow users to inspect AVAâs outputs against the trusted corpus. Page-level, clickable citations enabled participants to trace specific claims back to their sources. Seventeen of twenty interview participants rated this feature at the highest level (5/5), even though most reported using citation checking selectively, primarily to calibrate trust when an output was surprising, counterintuitive, or numerically specific in high-stakes analytical work. While this pipeline establishes structural trust mechanisms, user experience requires further nuance. Our deployment revealed that rigorous grounding must balance social expectations and residual error. We detail these two interactional challenges below: Tensions between epistemic humility and relational humility: Adapting the construct of relational humility 444In interpersonal psychology (Davis et al., 2013), relational humility is defined as an observerâs judgment that a person within a relationship demonstrates three qualities: (1) an accurate view of the self (neither inflated nor diminished), (2) an other-oriented stance that prioritizes the welfare of interaction partners, and (3) interpersonal behaviors marked by a lack of superiority or the regulation of ego-focused emotions (e.g., modesty). from interpersonal psychology to AI, we conceptualize relational humility as a user-attributed interactional property of the AI system that emerges when it: (1) communicates its limits; (2) remains oriented to the userâs goals, for example by explaining why an answer cannot be verified and suggesting constructive next steps; and (3) adopts collaborative, non-superior language. While epistemic humility concerns acknowledging limits in knowledge and deciding when to abstain, relational humility concerns the interactional stance through which those limits are communicated and managed to preserve collaboration and trust. As noted in Section 6.2, some participants interpreted abstentions as ârudeâ or a violation of partnership, expressing a desire for a âbrother-likeâ persona that felt more humanised. This suggests that, in professional settings, a bare refusal (e.g., âI donât knowâ) risks violating the Cooperative Principle of Conversation. To mitigate this, refusals should function as âservice-orientedâ pivot (e.g., âI cannot verify X, but I can help you explore Yâ). Consequently, the design lesson is to pair rigorous abstention with relational humility as part of the end-to-end trust pipeline. Future systems should match strict evidence-based refusal with a relationally humble persona that performs social repair, so that refusals reduce friction and preserve usersâ trust even when the AI reaches its limits. The inevitability of residual error: Even with a robust trust pipeline, designers must acknowledge that specialized systems will still occasionally surface incorrect or mismatched evidence (Perez-Cerrolaza et al., 2024; Zhou et al., 2023). If the system is highly trusted, the pipeline can paradoxically lead to automation bias (Buçinca et al., 2021; Lyell and Coiera, 2017), where users stop verifying claims because the system is usually right. Therefore, even when verification is distributed across the trust pipeline, designers should assume residual error will persist and incorporate interface cues, such as uncertainty signals, contrastive evidence previews, source comparisons, in-context highlighting, or interactive hover previews that make identifying inconsistencies low-friction (Braun et al., 2024; Sun et al., 2025; Zhang and HuĂmann, 2021). Future researchers could explore more domain-specific ways to highlight such residual errors. 7.2. Lesson 2: Balancing the Source Quality and Coverage Tradeoff A core open question in building AI-powered knowledge systems is how to trade off coverage (the breadth of topics a system can address) against source quality (DiGiacomo et al., 2025; Oche et al., 2025). General-purpose models such as ChatGPT illustrate the risks of maximising coverage over source quality, as they aggregate content from the open Internet (Oche et al., 2025). OpenAIâs recent policy guidelines, released in October 2025, state that ChatGPT will not provide recommendations in domains like law and medicine (OpenAI, 2025); however, it still provides answers while appending lightweight disclaimers (Business Insider, 2025). In these settings, the issue is not the amount of evidence; there is abundant text supporting almost any claim on the Internet, but rather the quality and verifiability of that evidence, which then requires substantial, often manual, verification effort. This motivates an alternative design choice for building specialized AI systems for high-stakes knowledge work: constrain coverage to high-quality, auditable sources, and expand only when justified by clear signals of unmet user needs or corpusâtask mismatch. In AVA, we instantiate this qualityâcoverage trade-off by fixing the evidence-verification threshold and treating abstention as the primary observable signal of corpus adequacy. In the initial deployment (See figure 5), AVA operated over a narrowly scoped corpus of roughly 50 flagship World Bank Reports. This maximized source quality but yielded high abstention rates (40â70%). Interviews indicated that repeated abstentions prompted some users to switch to alternative tools, reflecting a clear corpusâtask mismatch. Guided by the distribution of unanswered real-world queries, we expanded AVAâs corpus, within the same vetted institutional boundaries, to more than 4,000 World Bank Reports, leaving the verification threshold unchanged. Under this broader yet still curated corpus, abstention fell to approximately 10% (See figure 5). System logs show that AVA addressed the majority of queries in this phase, and participants interpreted the remaining abstentions positively, contrasting AVAâs explicit âI do not knowâ with other AI tools that âmake things up.â Once coverage was sufficient, abstention shifted from being perceived as a limitation to functioning as evidence of caution. This leads to our core design principle: abstention should reflect epistemic humility rather than systematic failure. Because high factual reliability is unlikely to be achievable when systems rely on unbounded, weakly verifiable sources such as the open Web (Business Insider, 2025; Lin et al., 2022), we constrain AVA to a curated corpus and enforce a fixed evidence-verification threshold. In doing so, we accept that the system will sometimes abstain as a necessary trade-off for reliability. To disambiguate abstention signals, we distinguish abstention (system property) from corpusâtask alignment (coverage adequacy). Building on this distinction, we recommend that future systems monitor three complementary signals in tandem with the abstention rate: corpusâtask alignment , evidence verification threshold and user behaviour (including reformulations, abandonment, and patterns of re-engagement). Interpreted together, these signals can help developers decide when to expand a curated corpus or tighten evidence-verification thresholds, supporting a dynamic, evidence-grounded approach to managing the qualityâcoverage trade-off while avoiding the pitfalls of open-web sourcing. 7.3. Lesson 3: Design for Verification, Not Just Disclosure Participantsâ approaches to disclosure (RQ4) revealed a hierarchy of needs: while disclosing AI use was seen as contextual, ensuring the accuracy of AI outputs was viewed as non-negotiable. This distinction reframes disclosure not as an end in itself, but as secondary to verification in high-stakes professional practice. It reveals a critical divergence between the institutional emphasis on disclosure (transparency of method) (Hoque et al., 2024; Sun et al., 2024) and the professional necessity of verification (accuracy of output). This has direct implications for the design of specialised AI systems. It is insufficient to merely provide the capacity to verify; designers must ensure that verification mechanisms streamline the process rather than creating friction. If verification imposes a high cognitive cost, users are likely to succumb to automation bias and bypass verification altogether (Buçinca et al., 2021; Lyell and Coiera, 2017). AVA addresses this through features like page-level verifiable citations and reasoned abstention, which function as tools for seamless answer verification. By enabling users to inspect underlying evidence without breaking their workflow, these mechanisms align with our participantsâ preference for tools that make system fallibility visible and actionable. 7.4. Lesson 4: Embracing the Future: From a Single Tool to a Collaborative AI Ecosystem Our findings suggest that AVAâs perceived value derived predominantly from providing reliable, citable evidence. This success overshadowed challenges such as limited feature discoverability. AVAâs specialization precluded open-ended tasks like brainstorming. leading users to routinely turn to general-purpose AI tools. We view this switching not as failure but as evidence of an emergent collaborative AI ecosystem. This model mirrors the long-standing practice of knowledge workers who employ a suite of distinct software applications, such as spreadsheets, word processors, and presentation software, to accomplish a complex task (Jahanlou et al., 2023). Consequently, design should foster seamless interoperability. (Jackson et al., 2014). 7.4.1. Design Implications: The Intelligent Handoff as a Cornerstone of Interoperability A central challenge of this ecosystem model is managing the transition between tools without degrading user trust. If a trusted specialist AI refers a user to a generalist AI that subsequently provides incorrect information, the specialistâs reputation may be harmed. To mitigate this risk, we propose the âIntelligent Handoffâ as a design pattern for responsible interoperability. A successful handoff should not simply be a link, but a structured interaction that performs three key functions: (1) Communicate Boundaries and Manage Expectations: It should clearly state why it is abstaining (e.g., âI cannot answer this from my verified libraryâŚâ) and frame the recommendation by task type (e.g., ââŚfor creative brainstorming, a general-purpose AI may offer starting pointsâ). (2) Preserve User Agency: Rather than prescribing a single product, the handoff should recommend a class of tool, empowering the user to select the service that best fits their needs and context (e.g., ââŚyou could use services like ChatGPT, Gemini, or Claude.â). (3) Add Value Through Prompt Generation: To increase the likelihood of a successful outcome, the specialist AI can use its domain knowledge to provide a re-engineered, higher-quality prompt for the user to copy. This adds immediate value, reduces the risk of hallucination, and reinforces the specialist AIâs role as a helpful, expert collaborator. Generalizability to High-Stakes Knowledge Work: Our design principles are not specific to international development and policy context and applies to other highâstakes domains such as law and medicine, where evidentiary risk and professional accountability are central. We outline concrete implications for these domains in Appendix E. In sum, our work suggests that the future of AI for knowledge work lies in creating specialized tools that are powerful in their own right and designed to be excellent, responsible collaborators within a broader digital ecosystem. Ultimately, our work extends the concept of Humble AI (Knowles et al., 2023). While prior research (Knowles et al., 2023) emphasizes systems communicating their internal limits, our findings indicate that, for specialized tools, humility must also be ecosystem-aware. A responsible AI collaborator does not stop at its own knowledge boundary; it offers users intelligent, trust-preserving pathways to other tools. This shift from introspective humility to collaborative humility is critical for designing AI systems that integrate seamlessly and responsibly into the complex, multi-tool reality of professional knowledge work. 8. Limitations Our study has several limitations that should be considered when interpreting these findings. First, the five-month deployment limits conclusions about long-term usage patterns and downstream impacts. Practices around trust calibration, system appropriation, and multi-tool orchestration may evolve as users gain familiarity with AVA and as organizational norms around AI assistance stabilize. Second, our multilingual evaluation was constrained by participantsâ English-dominant workplace practices. While two participants provided positive feedback on Spanish and French outputs, the small non-English sample (83.8% of interactions were English) limits conclusions about AVAâs performance across its 60+ supported languages or in predominantly non-English research contexts. Future research would focus on evaluating these languages with users. A third limitation of the study is the potential self-selection bias among participants who completed the endline survey and interviews. Prior work notes that such self-selection is a common characteristic of in-the-wild, ecologically valid deployments (Andrade, 2018; Ram et al., 2017). To help mitigate this issue, we sent multiple reminder emails and conducted interviews with volunteers representing diverse usage patterns and professional backgrounds. Still, voluntary participation may nonetheless introduce bias toward more engaged users. At the same time, we complement findings based on surveys and interviews with user log data, which provides more objective behavioral indicators of engagement. Future research might address this limitation through targeted recruitment strategies to capture less active users. 9. Conclusion This paper reports a five-month, in-the-wild deployment of AVA, a curated, citation-grounded multi-agent AI assistant for policy and development professionals. Using mixed methods, we show how page-level verifiability, reasoned abstention, and institutional provenance reshape evidence practices, enabling AVA to function as an evidence engine that improves efficiency while reducing reliance on open-web LLMs for high-stakes analysis. We distill three design lessons: trust must be engineered as an end-to-end pipeline spanning corpus curation, model behavior, and interface-level verification; the qualityâcoverage trade-off must be deliberately governed, with abstention treated as a signal of reliability rather than failure; and system design should prioritize low-friction verification of AI-generated outputs rather than disclosure alone. Ultimately, we advocate for ecosystem-aware humility in generative AI, where systems clearly bound their claims and enable responsible handoffs when tasks exceed their scope. Acknowledgements.We thank the anonymous reviewers for their constructive feedback. We also thank Ali Moezzi for his contributions to AVA. We extend our gratitude to Dr. Bill Thies and Professor Ge âTiffanyâ Wang for their insightful discussions during the revision process. Finally, we thank Professor Sir Nigel Shadbolt and Professor Max Van Kleek for their support. The findings, interpretations, and conclusions expressed in this paper are entirely those of the authors. They do not necessarily represent the views of the International Bank for Reconstruction and Development/World Bank and its affiliated organizations, or those of the Executive Directors of the World Bank or the governments they represent. Š 2026 International Bank for Reconstruction and Development/International Development Association or The World Bank This work is provided under a Creative Commons 4.0 Attribution International License, with the following mandatory and binding addition: Any and all disputes arising under this License that cannot be settled amicably shall be submitted to mediation in accordance with the WIPO Mediation Rules in effect at the time the work was published. If the request for mediation is not resolved within forty-five (45) days of the request, either You or the Licensor may, pursuant to a notice of arbitration communicated by reasonable means to the other party refer the dispute to final and binding arbitration to be conducted in accordance with UNCITRAL Arbitration Rules as then in force. The arbitral tribunal shall consist of a sole arbitrator and the language of the proceedings shall be English unless otherwise agreed. The place of arbitration shall be where the Licensor has its headquarters. The arbitral proceedings shall be conducted remotely (e.g., via telephone conference or written submissions) whenever practicable, or held at the World Bank headquarters in Washington DC. References P. J. L. Ammann, J. Golde, and A. Akbik (2025) Question decomposition for retrieval-augmented generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), J. Zhao, M. Wang, and Z. Liu (Eds.), Vienna, Austria, p. 497â507. External Links: Link, Document, ISBN 979-8-89176-254-1 Cited by: §J.2, Appendix F, §3.2.2. C. Andrade (2018) Internal, external, and ecological validity in research design, conduct, and evaluation. Indian Journal of Psychological Medicine 40 (5), p. 498â499. External Links: Document, Link Cited by: §8. A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda (2024) Refusal in language models is mediated by a single direction. Advances in Neural Information Processing Systems 37, p. 136037â136083. Cited by: item 4. A. S. BaHammam (2025) The transparency paradox: why researchers avoid disclosing AI assistance in scientific writing. Nature and Science of Sleep 17, p. 2569â2574. External Links: Document, Link Cited by: §1. A. Bastounis, P. Campodonico, M. van der Schaar, B. Adcock, and A. C. Hansen (2024) On the consistent reasoning paradox of intelligence and optimal trust in ai: the power ofâi donât knowâ. arXiv preprint arXiv:2408.02357. Cited by: item 4. M. Braun, M. Greve, A. B. Brendel, and L. M. Kolbe (2024) Humans supervising artificial intelligenceâinvestigation of designs to optimize error detection. Journal of Decision Systems 33 (4), p. 674â699. Cited by: item 3. Z. Buçinca, M. B. Malaya, and K. Z. Gajos (2021) To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proc. ACM Hum.-Comput. Interact. 5 (CSCW1). External Links: Link, Document Cited by: Appendix E, item 3, §7.3. Business Insider (2025) Donât worry, chatgpt can still answer your health questions. Note: https://w.businessinsider.com/openai-can-still-answer-your-health-questions-2025-11 Cited by: Appendix E, §7.2, §7.2. S. Cao, A. Liu, and C. Huang (2024) Designing for appropriate reliance: the roles of ai uncertainty presentation, initial user decision, and user demographics in ai-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 8 (CSCW1), p. 1â32. Cited by: item 1. S. Cao and L. Wang (2024) Verifiable generation with subsentence-level fine-grained citations. arXiv preprint arXiv:2406.06125. Cited by: item 2. L. A. Celi (2025) Teaching machines to doubt. Nature Medicine, p. 1â1. Cited by: §1. L. Chen, Z. Liang, X. Wang, J. Liang, Y. Xiao, F. Wei, J. Chen, Z. Hao, B. Han, and W. Wang (2024) Teaching large language models to express knowledge boundary from their own signals. External Links: 2406.10881, Document, Link Cited by: §2.1. F. Cheng, V. Zouhar, S. Arora, M. Sachan, H. Strobelt, and M. El-Assady (2024) RELIC: investigating large language model responses using self-consistency. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI â24), External Links: Document Cited by: §1. C. Chiang and H. Lee (2024) Merging facts, crafting fallacies: evaluating the contradictory nature of aggregated factual claims in long-form generations. arXiv preprint arXiv:2402.05629. Cited by: item 2, item 3. S. L. Dalglish, H. Khalid, and S. A. McMahon (2020) Document analysis in health policy research: the read approach. Health policy and planning 35 (10), p. 1424â1431. Cited by: §7. D. E. Davis, E. L. Worthington Jr, J. N. Hook, R. A. Emmons, P. C. Hill, R. A. Bollinger, and D. R. Van Tongeren (2013) Humility and the development and repair of social bonds: two longitudinal studies. Self and identity 12 (1), p. 58â77. Cited by: footnote 4. Ă. DedeoÄlu and P. Chandra (2025) Navigating the posthuman turn in computing and design: a posthuman vocabulary. In Proceedings of the ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies, p. 504â529. Cited by: §2.1. Y. Deng, Y. Zhao, M. Li, S. Ng, and T. Chua (2024) Donât just sayâ i donât knowâ! self-aligning large language models for responding to unknown questions with explanations. arXiv preprint arXiv:2402.15062. Cited by: item 4. S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. Celikyilmaz, and J. Weston (2024) Chain-of-verification reduces hallucination in large language models. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 3563â3578. External Links: Link, Document Cited by: Appendix F. P. DiGiacomo, H. Wang, J. Fang, Y. Leng, W. M. Brode, and Y. Ding (2025) Guide-rag: evidence-driven corpus curation for retrieval-augmented generation in long covid. arXiv preprint arXiv:2510.15782. Cited by: §7.2. P. Dourish (2003) The appropriation of interactive technologies: some lessons from placeless documents. Computer Supported Cooperative Work (CSCW) 12 (4), p. 465â490. External Links: Document Cited by: §1, §6.2.1, §7. Elicit (2025) Elicit: ai for scientific research. Note: https://elicit.com/Accessed: 2025-12-05 Cited by: §2.3. K. Enevoldsen, I. Chung, I. Kerboua, M. Kardos, A. Mathur, D. Stap, J. Gala, W. Siblini, D. KrzemiĹski, G. I. Winata, S. Sturua, S. Utpala, M. Ciancone, M. Schaeffer, G. Sequeira, D. Misra, S. Dhakal, J. Rystrøm, R. Solomatin, Ă. ĂaÄatan, A. Kundu, M. Bernstorff, S. Xiao, A. Sukhlecha, B. Pahwa, R. PoĹwiata, K. K. GV, S. Ashraf, D. Auras, B. PlĂźster, J. P. Harries, L. Magne, I. Mohr, M. Hendriksen, D. Zhu, H. Gisserot-Boukhlef, T. Aarsen, J. Kostkan, K. Wojtasik, T. Lee, M. Ĺ uppa, C. Zhang, R. Rocca, M. Hamdy, A. Michail, J. Yang, M. Faysse, A. Vatolin, N. Thakur, M. Dey, D. Vasani, P. Chitale, S. Tedeschi, N. Tai, A. Snegirev, M. GĂźnther, M. Xia, W. Shi, X. H. LĂš, J. Clive, G. Krishnakumar, A. Maksimova, S. Wehrli, M. Tikhonova, H. Panchal, A. Abramov, M. Ostendorff, Z. Liu, S. Clematide, L. J. Miranda, A. Fenogenova, G. Song, R. B. Safi, W. Li, A. Borghini, F. Cassano, H. Su, J. Lin, H. Yen, L. Hansen, S. Hooker, C. Xiao, V. Adlakha, O. Weller, S. Reddy, and N. Muennighoff (2025) MMTEB: massive multilingual text embedding benchmark. arXiv preprint arXiv:2502.13595. External Links: Link, Document Cited by: §G.1. European Parliament and Council of the European Union (2024) Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence. Note: Official Journal of the European Union, L 2024/1689Article 50 External Links: Link Cited by: §1. R. Fok, N. Lipka, T. Sun, and A. F. Siu (2024) Marco: supporting business document workflows via collection-centric information foraging with large language models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA. External Links: Document, Link Cited by: §2.3. P. Formosa, S. Bankins, R. Matulionyte, and O. Ghasemi (2025) Can chatgpt be an author? generative ai creative writing assistance and perceptions of authorship, creatorship, responsibility, and disclosure. Ai & Society 40 (5), p. 3405â3417. Cited by: §7. S. Fortunato, A. Flammini, F. Menczer, and A. Vespignani (2006) Topical interests and the mitigation of search engine bias. Proceedings of the national academy of sciences 103 (34), p. 12684â12689. Cited by: §2.3. D. Gambetta (2000) Can we trust trust?. In Trust: Making and Breaking Cooperative Relations, D. Gambetta (Ed.), Vol. 13, p. 213â237. Cited by: §6.3. M. Gandouz, H. Holzmann, and D. Heider (2021) Machine learning with asymmetric abstention for biomedical decision-making. BMC medical informatics and decision making 21 (1), p. 294. Cited by: Appendix E. T. Gao, H. Yen, J. Yu, and D. Chen (2023a) Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), External Links: Link Cited by: §2.1. Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, H. Wang, and H. Wang (2023b) Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997 2 (1). Cited by: item 3. Y. Geifman and R. El-Yaniv (2017) Selective classification for deep neural networks. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017), External Links: Link Cited by: §2.1. Google LLC (2025) Learn about notebooklm. Note: https://support.google.com/notebooklm/answer/16164461?hl=en&co=GENIE.Platform%3DDesktop Cited by: §2.3. S. Gupta, R. Ranjan, and S. N. Singh (2024) A comprehensive survey of retrieval-augmented generation (rag): evolution, current landscape and future directions. arXiv preprint arXiv:2410.12837. Cited by: item 2. M. N. Hoque, T. Mashiat, B. Ghai, C. Shelton, F. Chevalier, K. Kraus, and N. Elmqvist (2024) The HaLLMark effect: supporting provenance and transparent use of large language models in writing with interactive visualization. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA. External Links: Document, Link Cited by: §7.3. M. Hosseini, D. B. Resnik, and K. Holmes (2023) The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Research Ethics 19 (4), p. 449â465. Cited by: §1. B. N. Hryciw, A. J. Seely, and K. Kyeremanteng (2023) Guiding principles and proposed classification system for the responsible adoption of artificial intelligence in scientific writing in medicine. Frontiers in Artificial Intelligence 6, p. 1283353. Cited by: §1. J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson (2020) XTREME: a massively multilingual multi-task benchmark for evaluating cross-lingual generalization. External Links: 2003.11080, Link Cited by: §G.1. M. Hu, B. He, Y. Wang, L. Li, C. Ma, and I. King (2024) Mitigating large language model hallucination with faithful finetuning. arXiv preprint arXiv:2406.11267. Cited by: item 3. L. Huang, X. Feng, W. Ma, Y. Gu, W. Zhong, X. Feng, W. Yu, W. Peng, D. Tang, D. Tu, et al. (2024) Learning fine-grained grounded citations for attributed large language models. arXiv preprint arXiv:2408.04568. Cited by: §7. Y. Hui, C. Chen, Z. Fu, Y. Liu, J. Ye, and H. Zhang (2025) Reason and interact with the corpus, beyond black-box retrieval. arXiv preprint arXiv:2510.27566. Cited by: §G.2.2. S. Jabbour, T. Chang, A. D. Antar, J. Peper, I. Jang, J. Liu, J. Chung, S. He, M. Wellman, B. Goodman, E. Bondi-Kelly, K. Samy, R. Mihalcea, M. Chowdhury, D. Jurgens, and L. Wang (2025) Evaluation framework for ai systems in âthe wildâ. External Links: 2504.16778, Document, Link Cited by: item 2, §2.2, §2.2. S. J. Jackson, T. Gillespie, and S. Payette (2014) The policy knot: re-integrating policy, practice and design in cscw studies of social computing. In Proceedings of the 17th ACM Conference on Computer Supported Cooperative Work & Social Computing, CSCW â14, New York, NY, USA, p. 588â602. External Links: Document, Link Cited by: §7.4. A. Jahanlou, J. Vermeulen, T. Grossman, P. Chilana, G. Fitzmaurice, and J. Matejka (2023) Task-centric application switching: how and why knowledge workers switch software applications for a single task. In Graphics Interface 2023, Cited by: §7.4. J. Jin, A. Paladugu, and C. Xiong (2025) Beneficial reasoning behaviors in agentic search and effective post-training to obtain them. External Links: 2510.06534, Link Cited by: Appendix F, §G.3.2. Justia (2025) AI and attorney ethics rules: 50-state survey. Note: https://w.justia.com/trials-litigation/ai-and-attorney-ethics-rules-50-state-survey/Accessed 2 Dec 2025 Cited by: Appendix E. N. Karnatak, A. Baranes, R. Marchant, T. Butler, and K. Olson (2025a) ACAI for sbos: ai co-creation for advertising and inspiration for small business owners. arXiv preprint arXiv:2503.06729. External Links: Document, Link Cited by: §2.3. N. Karnatak, A. Baranes, R. Marchant, H. Zeng, T. Butler, and K. Olson (2025b) Expanding the generative ai design space through structured prompting and multimodal interfaces. In Proceedings of the CHI 2025 Workshop on Computational User Interfaces, Note: Workshop paper Cited by: §2.3. V. Karpukhin, B. OÄuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. Yih (2020) Dense passage retrieval for open-domain question answering. External Links: 2004.04906, Link Cited by: §G.2.2. N. Karusala, S. Upadhyay, R. Veeraraghavan, and K. Z. Gajos (2024) Understanding contestability on the margins: implications for the design of algorithmic decision-making in public services. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â16. Cited by: §2.1. M. Kim, S. Kim, and J. Thorne (2025) From evidence to belief: a bayesian epistemology approach to language models. arXiv preprint arXiv:2504.19622. Cited by: item 3. S. S. Y. Kim et al. (2024) âIâm not sure, butâŚâ: uncertainty expressions and user reliance/trust. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT), Note: Preprint available at https://arxiv.org/abs/2405.00623 Cited by: §2.1, §7. B. Knowles, J. DâCruz, J. T. Richards, and K. R. Varshney (2023) Humble AI. Communications of the ACM 66 (9), p. 73â79. External Links: Document Cited by: §1, §2.1, §7.4.1, footnote 1. R. Kocielnik, S. Amershi, and P. N. Bennett (2019) Will you accept an imperfect AI? designing to adjust end-user expectations. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, External Links: Document Cited by: §2.1. D. Lawrie, E. Yang, D. W. Oard, and J. Mayfield (2023) Neural approaches to multilingual information retrieval. In Advances in Information Retrieval: 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, April 2â6, 2023, Proceedings, Part I, Berlin, Heidelberg, p. 521â536. External Links: ISBN 978-3-031-28243-0, Link, Document Cited by: §G.1. J. D. Lee and K. A. See (2004) Trust in automation: designing for appropriate reliance. Human factors 46 (1), p. 50â80. Cited by: §1, §2.1, §7. Y. Lee, H. B. Kang, M. Latzke, J. Kim, J. Bragg, J. C. Chang, and P. Siangliulue (2024) PaperWeaver: enriching topical paper alerts by contextualizing recommended papers with user-collected papers. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI â24), External Links: Document Cited by: §1. F. Leiser, S. Eckhardt, V. Leuthe, M. Knaeble, A. Maedche, G. Schwabe, and A. Sunyaev (2024) HILL: a hallucination identifier for large language models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI â24), External Links: Document Cited by: §1. Z. Levonian, C. Li, W. Zhu, A. Gade, O. Henkel, M. Postle, and W. Xing (2023) Retrieval-augmented generation to improve math question-answering: trade-offs between groundedness and human preference. arXiv preprint arXiv:2310.03184. Cited by: item 3. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, U. Khandelwal, H. KĂźttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela (2020) Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §G.3.3, §2.3. B. Li, Z. Xu, and R. Xie (2025a) Language drift in multilingual retrieval-augmented generation: characterization and decoding-time mitigation. External Links: 2511.09984, Link Cited by: Appendix F, §G.4. G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for âmindâ exploration of large language model society. In Thirty-seventh Conference on Neural Information Processing Systems, Cited by: Appendix F, §3.2.2. M. Li, M. Luo, T. Lv, Y. Zhang, S. Zhao, E. Nie, and G. Zhou (2025b) A survey of long-document retrieval in the plm and llm era. External Links: 2509.07759, Link Cited by: Appendix F. T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi (2023) Encouraging divergent thinking in large language models through multi-agent debate. arXiv preprint arXiv:2305.19118. Cited by: §G.2.2. Q. V. Liao and J. Wortman Vaughan (2023) AI transparency in the age of llms: a human-centered research roadmap. External Links: 2306.01941, Link Cited by: §1. S. Lin, J. Hilton, and O. Evans (2022) Truthfulqa: measuring how models mimic human falsehoods. In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), p. 3214â3252. Cited by: §7.2. G. Liu, X. Wang, L. Yuan, Y. Chen, and H. Peng (2023a) Examining llmsâ uncertainty expression towards questions outside parametric knowledge. arXiv preprint arXiv:2311.09731. Cited by: item 4. N. F. Liu, T. Zhang, and P. Liang (2023b) Evaluating verifiability in generative search engines. arXiv preprint arXiv:2304.09848. Cited by: item 2. D. Lyell and E. Coiera (2017) Automation bias and verification complexity: a systematic review. Journal of the American Medical Informatics Association 24 (2), p. 423â431. Cited by: Appendix E, item 3, §7.3. H. Lyons, E. Velloso, and T. Miller (2021) Conceptualising contestability: perspectives on contesting algorithmic decisions. Proceedings of the ACM on Human-Computer Interaction 5 (CSCW1), p. 1â25. Cited by: Appendix J. Y. Ma, Y. Wu, Q. Ai, Y. Liu, Y. Shao, M. Zhang, and S. Ma (2023) Incorporating structural information into legal case retrieval. ACM Trans. Inf. Syst. 42 (2). External Links: ISSN 1046-8188, Link, Document Cited by: Appendix F. N. Madhusudhan, S. T. Madhusudhan, V. Yadav, and M. Hashemi (2025) Do llms know when to not answer? investigating abstention abilities of large language models. In Proceedings of the 31st International Conference on Computational Linguistics, p. 9329â9345. Cited by: item 4. J. Menick et al. (2022) Teaching language models to support answers with verified quotes. arXiv preprint arXiv:2203.11147. External Links: Link Cited by: §2.1. R. Nair, I. Vejsbjerg, E. M. Daly, C. Varytimidis, and B. Knowles (2025) Humble ai in the real-world: the case of algorithmic hiring. In Adjunct Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, p. 1â7. Cited by: §1, §1, §2.1. P. Narayanan Venkit, P. Laban, Y. Zhou, Y. Mao, and C. Wu (2025) Search engines in the ai era: a qualitative understanding to the false promise of factual and verifiable source-cited responses in llm-based search. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT â25, New York, NY, USA, p. 1325â1340. External Links: ISBN 9798400714825, Link, Document Cited by: §1. A. J. Oche, A. G. Folashade, T. Ghosal, and A. Biswas (2025) A systematic review of key retrieval-augmented generation (rag) systems: progress, gaps, and future directions. arXiv preprint arXiv:2507.18910. Cited by: §7.2, §7. OpenAI (2025) Usage policies. Note: https://openai.com/policies/usage-policiesAccessed: 2025-11-28 Cited by: Appendix E, §7.2. S. A. C. OrdoĂąez, M. Lange, T. M. Lunde, M. J. Meni, and A. E. Premo (2025) Humility and curiosity in humanâai systems for health care. The Lancet 406 (10505), p. 804â805. Cited by: §1. W. J. Orlikowski (2000) Using technology and constituting structures: a practice lens for studying technology in organizations. Organization science 11 (4), p. 404â428. Cited by: §1, §1. W. J. Orlikowski et al. (1995) Evolving with notes: organizational change around groupware technology. Cited by: §1, §7. R. Y. Pang, H. Schroeder, K. S. Smith, S. Barocas, Z. Xiao, E. Tseng, and D. Bragg (2025) Understanding the llm-ification of chi: unpacking the impact of llms at chi through a systematic literature review. Note: Also available in the ACM Digital Library, DOI: 10.1145/3706598.3713726 External Links: 2501.12557, Link Cited by: §1, item 1, item 2, §2.2. J. Perez-Cerrolaza, J. Abella, M. Borg, C. Donzella, J. Cerquides, F. J. Cazorla, C. Englund, M. Tauber, G. Nikolakopoulos, and J. L. Flores (2024) Artificial intelligence for safety-critical systems in industrial and transportation domains: a survey. ACM Comput. Surv. 56 (7). External Links: ISSN 0360-0300, Link, Document Cited by: item 3. Perplexity AI (2024) What is an answer engine and how does perplexity work as one?. Note: https://w.perplexity.ai/help-center/en/articles/10354917-what-is-an-answer-engine-and-how-does-perplexity-work-as-oneAccessed: 2026-02-07 Cited by: §2.3. R. Petcu, K. Murray, D. Khashabi, E. Kanoulas, M. de Rijke, D. Lawrie, and K. Duh (2025) Query decomposition for rag: balancing exploration-exploitation. External Links: 2510.18633, Link Cited by: §J.2, Appendix F, §3.2.2. N. N. Potter (2022) The virtue of epistemic humility. Philosophy, Psychiatry, & Psychology 29 (2), p. 121â123. Cited by: §2.1. S. S. Rahman, Md. A. Islam, Md. M. Alam, M. Zeba, Md. A. Rahman, S. S. Chowa, M. A. K. Raiaan, and S. Azam (2025) Hallucination to truth: a review of fact-checking and factuality evaluation in large language models. External Links: 2508.03860, Link Cited by: Appendix F, §1. N. Ram, M. Brinberg, A. L. Pincus, and D. E. Conroy (2017) The questionable ecological validity of ecological momentary assessment: considerations for design and analysis. Research in Human Development 14 (3), p. 253â270. External Links: Document, Link Cited by: §8. A. Rapp, C. Di Lodovico, and L. Di Caro (2025) How do people react to chatgptâs unpredictable behavior? anthropomorphism, uncanniness, and fear of ai: a qualitative study on individualsâ perceptions and understandings of llmsâ nonsensical hallucinations. International Journal of Human-Computer Studies. External Links: Document Cited by: §1. L. Ratzan (2000) Making sense of the web: a metaphorical approach. Information research 6 (1), p. 6â1. Cited by: §6.4. M. Rezaei, M. Pironti, and R. Quaglia (2024) AI in knowledge sharing, which ethical challenges are raised in decision-making processes for organisations?. Management Decision. Cited by: §7. P. Sarthi, S. Abdullah, A. Tuli, S. Khanna, A. Goldie, and C. D. Manning (2024) RAPTOR: recursive abstractive processing for tree-organized retrieval. In International Conference on Learning Representations (ICLR), Cited by: Appendix F. M. Schemmer, P. Hemmer, N. KĂźhl, C. Benz, and G. Satzger (2022) Should i follow ai-based advice? measuring appropriate reliance in human-ai decision-making. arXiv preprint arXiv:2204.06916. Cited by: Appendix J. J. SchĂśpfel (2010) Towards a prague definition of grey literature. In Twelfth International Conference on Grey Literature, Cited by: §2.3, §2.3. D. Schuster (2025) Abstaining machine learning: philosophical considerations. AI & SOCIETY, p. 1â21. Cited by: Appendix E. C. Schwarz-Plaschg (2018a) Nanotechnology is like⌠the rhetorical roles of analogies in public engagement. Public Understanding of Science 27 (2), p. 153â167. Cited by: §6.4. C. Schwarz-Plaschg (2018b) The power of analogies for imagining and governing emerging technologies. NanoEthics 12 (2), p. 139â153. Cited by: §6.4. Z. Shao, H. Shen, M. Liu, G. Fu, Y. Guo, Y. Wang, and Y. Ma (2025) Grounding ai explanations in experience: a reflective cognitive architecture for clinical decision support. External Links: 2509.21266, Link Cited by: §3.2.2. W. Shi, S. Min, H. Shi, W. Xiong, W. Yih, L. Zettlemoyer, and D. Chen (2024) REPLUG: retrieval-augmented black-box language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Cited by: §G.3.3. N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao (2023) Reflexion: language agents with verbal reinforcement learning. External Links: 2303.11366 Cited by: Appendix F, §3.2.2. G. Sidiropoulos, N. Voskarides, S. Vakulenko, and E. Kanoulas (2021) Combining lexical and dense retrieval for computationally efficient multi-hop question answering. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing, N. S. Moosavi, I. Gurevych, A. Fan, T. Wolf, Y. Hou, A. MarasoviÄ, and S. Ravi (Eds.), Virtual, p. 58â63. External Links: Link, Document Cited by: Appendix F, §3.2.2. L. A. Suchman (1987) Plans and situated actions: the problem of human-machine communication. Cambridge university press. Cited by: §1. J. Sun, G. Warren, I. Shklovski, and I. Augenstein (2025) Explaining sources of uncertainty in automated fact-checking. arXiv preprint arXiv:2505.17855. Cited by: item 3. L. Sun, S. Tao, J. Hu, and S. P. Dow (2024) MetaWriter: exploring the potential and perils of AI writing support in scientific peer review. Proceedings of the ACM on Human-Computer Interaction 8 (CSCW1). External Links: Document Cited by: §1, §7.3. B. Tong, J. Xia, S. Shang, and K. Zhou (2025) Measuring epistemic humility in multimodal large language models. arXiv preprint arXiv:2509.09658. Cited by: §1. A. Wan, E. Wallace, and D. Klein (2024) What evidence do language models find convincing?. arXiv preprint arXiv:2402.11782. Cited by: item 3. B. Wang, J. Liu, J. Karimnazarov, and N. Thompson (2024a) Task supportive and personalized human-large language model interaction: a user study. In Proceedings of the 2024 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR â24), New York, NY, USA. External Links: Document, Link Cited by: §1, item 2, §2.2. J. Wang, J. X. Huang, X. Tu, J. Wang, A. J. Huang, M. T. R. Laskar, and A. Bhuiyan (2024b) Utilizing bert for information retrieval: survey, applications, resources, and challenges. External Links: 2403.00784, Link Cited by: §G.2.2. L. Weidinger, I. D. Raji, H. Wallach, M. Mitchell, A. Wang, O. Salaudeen, R. Bommasani, D. Ganguli, S. Koyejo, and W. Isaac (2025) Toward an evaluation science for generative ai systems. External Links: 2503.05336, Document, Link Cited by: item 1, §2.2, §2.2. B. Wen, J. Yao, S. Feng, C. Xu, Y. Tsvetkov, B. Howe, and L. L. Wang (2025) Know your limits: a survey of abstention in large language models. Transactions of the Association for Computational Linguistics 13, p. 529â556. Cited by: §7. M. Wester et al. (2024) âAs an AI language model, i cannotâŚâ: investigating LLM denials of user requests. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, External Links: Document Cited by: §1, §2.1, §7. D. Whitcomb, H. Battaly, J. Baehr, and D. Howard-Snyder (2017) Intellectual humility. Philosophy and Phenomenological Research 94 (3), p. 509â539. Cited by: §2.1. S. Xia, X. Wang, J. Liang, Y. Zhang, W. Zhou, J. Deng, F. Yu, and Y. Xiao (2025) Ground every sentence: improving retrieval-augmented llms with interleaved reference-claim generation. In Findings of the Association for Computational Linguistics: NAACL 2025, p. 969â988. Cited by: §7. Q. Xiao, X. E. Hu, M. E. Whiting, A. Karunakaran, H. Shen, and H. Cao (2025) AI hasnât fixed teamwork, but it shifted collaborative culture: a longitudinal study in a project-based software development organization (2023-2025). arXiv preprint arXiv:2509.10956. Cited by: §1. Y. A. Yadkori, I. Kuzborskij, D. Stutz, A. GyĂśrgy, A. Fisch, A. Doucet, I. Beloshapka, W. Weng, Y. Yang, C. SzepesvĂĄri, et al. (2024) Mitigating llm hallucinations via conformal abstention. arXiv preprint arXiv:2405.01563. Cited by: item 4. Q. Yang, A. Steinfeld, and J. Zimmerman (2019) Unremarkable ai: fitting intelligent decision support into critical, clinical decision-making processes. In Proceedings of the 2019 CHI conference on human factors in computing systems, p. 1â11. Cited by: §1. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao (2023) ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), Cited by: Appendix F, §G.2.1, §G.2.3, §G.3.2, §3.2.2. Z. Yin, Q. Sun, Q. Guo, J. Wu, X. Qiu, and X. Huang (2023) Do large language models know what they donât know?. arXiv preprint arXiv:2305.18153. Cited by: item 4. Y. Yuan, W. Jiao, W. Wang, J. Huang, J. Xu, T. Liang, P. He, and Z. Tu (2025) Refuse whenever you feel unsafe: improving safety in llms via decoupled refusal training. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3149â3167. Cited by: item 4. C. Zhang, Q. Chen, and M. Zhang (2025a) Mixture-of-rag: integrating text and tables with large language models. External Links: 2504.09554, Link Cited by: §3.1. Q. Zhang, S. Chen, Y. Bei, Z. Yuan, H. Zhou, Z. Hong, H. Chen, Y. Xiao, C. Zhou, J. Dong, Y. Chang, and X. Huang (2025b) A survey of graph retrieval-augmented generation for customized large language models. External Links: 2501.13958, Link Cited by: §3.2.2. Y. Zhang, Q. V. Liao, and R. K. Bellamy (2020) Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 conference on fairness, accountability, and transparency, p. 295â305. Cited by: Appendix J, §2.1. Z. T. Zhang and H. HuĂmann (2021) How to manage output uncertainty: targeting the actual end user problem in interactions with ai.. In IUI Workshops, Cited by: item 3. J. Zhou, Y. Zhang, Q. Luo, A. G. Parker, and M. De Choudhury (2023) Synthetic lies: understanding ai-generated misinformation and evaluating algorithmic and human solutions. In Proceedings of the 2023 CHI conference on human factors in computing systems, p. 1â20. Cited by: item 3. K. Zhou, J. D. Hwang, X. Ren, and M. Sap (2024) Relying on the unreliable: the impact of language modelsâ reluctance to express uncertainty. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, p. 3623â3643. Cited by: item 1. Appendix A Detailed Descriptions of AVA Interface Components Figure 3 presents an annotated view of AVAâs interface. Below we provide detailed descriptions of each labeled component (AâH). A. Query Input Box. The user enters a natural-language research question. This initiates the retrieval and grounding pipeline. B. Retrieval Process Trace. A dynamic trace (e.g., âThinking through 186 sections and 10 imagesâ) that surfaces the systemâs retrieval operations, indicating how AVA scans, ranks, and selects evidence from the curated corpus to ground its response. C. Inline Verifiable Citations. Clickable, page-level citation markers embedded directly within the generated answer. These act as entry points into the verification loop. D. Evidence Preview Card. A pop-over that appears when the user selects a citation. It displays the document title, page number, and a brief snippet of surrounding text to allow quick, in-context inspection without leaving the chat interface. E. Source Document Viewer. A side-by-side PDF viewer that opens when the user clicks the document header. The viewer automatically scrolls to the cited page and highlights the referenced text, supporting low-friction verification. F. Save-as-Note Functionality. A feature enabling users to save the generated response and the verified citations into a dedicated notes panel for later reference. G. Copy-Paste Functionality. A control that allows users to copy the generated text (including citations) into external tools such as Word, Google Docs, or email workflows. H. Tagging. Tagging features that allow users to assign labels to saved notes or research sessions, supporting organisation and later retrieval. Appendix B Consolidated Usage Patterns, Longitudinal Dynamics, and User Insights Appendix B consolidates empirical usage patterns, longitudinal dynamics, and user-reported friction points that relate to behavioural findings. It brings together usage statistics, temporal shifts, and interview-reported challenges that are otherwise distributed across Sections 5 and 6, providing a single reference point for these interactional patterns. Because RQ4 concerns sociocultural disclosure norms rather than system usage behaviour, those findings remain in the main Results (Section 6.4). B.1. Indicative Usage Patterns Our mixed-methods analysis revealed three dominant patterns of engagement: ⢠Task & Theme Dominance: Usage was heavily skewed toward Diagnostic tasks (69.0% of queries), such as understanding problems or data, rather than Design (22.0%) or Evaluation (9.0%). Thematically, users focused on broad, cross-cutting policy areas like Human Capital (33.5%) and Macroeconomics (20.7%) rather than niche fact retrieval. ⢠Multi-Tool Ecosystem: Users exhibited a distinct âmulti-toolâ workflow, utilizing generalist LLMs (e.g., ChatGPT) for ideation and scaffolding, while reserving AVA specifically as an âevidence engineâ for âreal informationâ and citations. ⢠Language Norms: Despite supporting 60+ languages, 83.8% of interactions occurred in English. Interviews revealed this was driven by workplace norms in international development rather than a lack of capability. B.2. Longitudinal Shifts in System Behavior We observed two major shifts over the five-month deployment: ⢠Impact of Corpus Expansion on Abstention: System behavior changed significantly over time. Initially, with a small corpus (50 reports), the system abstained on 40â70% of queries. After expanding to 4,000+ reports, abstention dropped to below 10%. This shifted the user experience of abstention from a âsymptom of mismatchâ to a âpositive signalâ of caution. ⢠Efficiency Gains for Returning Users: The study found a dose-response relationship where returning users (multiple sessions) reported significant time savings (2.4â3.9 hours/week) compared to single-session users, suggesting that efficiency gains and verification routines develop with sustained use. B.3. User Frustrations and Challenges While satisfaction was high, specific friction points emerged in qualitative interviews: ⢠The âDead-Endâ Abstention: While many participants appreciated abstention as an honest boundary, others experienced frequent âI donât knowâ responses as a bottleneck that stalled their workflow. In these moments, users reported turning to other tools to maintain momentum. ⢠Tone & Usability: Specific users described the bluntness of the refusal as ârudeâ. Others expressed frustration that the system stopped completely rather than offering partial help, stating, âI donât want the choice being no answerâ B.4. Suggested Improvements Participants proposed three specific enhancements to better support high-stakes workflows: ⢠Intelligent Handoffs: To address the frustration of dead ends, participants suggested the system should act as a âknowledgeable guideâ that points them to external resources (e.g., âtry ChatGPT for thisâ) when it cannot answer from the verified corpus. ⢠Collaborative Query Refinement: Users requested that the system help them refine their prompts or suggest related regions/years when exact matches are missing, rather than just abstaining. ⢠Exportable Visuals: Users praised the generated charts and tables but requested better functionality to export these visuals with built-in attribution for reports. Appendix C Supported Languages AVA currently supports input and output in the following 60+ languages: Albanian, Amharic, Arabic, Armenian, Bengali, Bosnian, Bulgarian, Burmese, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Georgian, German, Greek, Gujarati, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Malayalam, Marathi, Mongolian, Norwegian, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Slovak, Slovenian, Somali, Spanish, Swahili, Swedish, Tagalog, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, and Vietnamese. Appendix D Extended Methodology and Measures D.1. Recruitment, Ethics, and Study Timeline Recruitment and Ethics: Participants were recruited through public-facing newsletters and targeted emails from the host MDB. Participation had no bearing on employment status, course enrollment, or standing with the institution. All activities were conducted on participantsâ own devices to mirror heterogeneous real-world bandwidth conditions. Detailed Timeline: ⢠Baseline Data Collection: April to May 12. ⢠Pilot Launch: May 12 (small cohort). ⢠Full Rollout: May 26 (expanded to all randomly selected treatment registrants). ⢠System Updates: During deployment, we shipped corpus expansion to Âż4,000 documents (July 9), graph-generation features (July 17), and improved handling of non-response queries (July 23). ⢠Endline Survey: Administered July 21. ⢠Public Launch: August 4. D.2. Survey Instruments and Variable Definitions (1) AI Use for Work (Baseline & Endline). AI experience was measured via a six-point frequency scale: âMultiple times every day,â âAbout once or twice on most days,â âA few times throughout the week,â âRarely,â âNot at all,â and âNever use AI for work.â (2) AI-Driven Productivity Gains (Baseline & Endline). Respondents estimated hours saved in the prior week (writing, research, analysis, formatting, summarizing) with categories: 0, <<1, 1â3, 4â6, 7â10, >>10 hours. We compute lower and upper bounds by mapping ranges to their endpoints (e.g., 1â3 â 1 and 3; 4â6 â 4 and 6; 7â10 â 7 and 10). (3) Perceptions of AVA (Endline). Among users granted access, we assessed: efficient access to insights from World Bank flagship documents (âAVA helps me understand key insights from World Bank flagship documents more efficiently.â), trustworthiness/accuracy (âThe synthesized insights provided by AVA are accurate and trustworthy.â), work-quality impact (âUsing AVA has improved the quality of my work or research.â), and recommendation intent (âI would recommend AVA to others working in development policy or research.â). (4) Perceptions of AVA (In-App). Periodic session pop-ups captured the same perceptions (e.g., trustworthiness, likelihood to recommend). We report perception results from both endline and in-app sources for a comprehensive view. D.3. Detailed NLP Preprocessing and Query Classification Preprocessing Pipeline ⢠Translation & Normalization: Language detection was applied to every query; non-English text was automatically translated to English. Text was lowercased, and punctuation, digits, and extraneous whitespace were removed. ⢠Tokenization: Common English stopwords and very short tokens were filtered out. Remaining tokens were lemmatized to their base forms. ⢠Sessionization: Queries were grouped into sessions using a strict one-hour inactivity threshold to capture immediate task context. Classification Logic ⢠Taxonomy Matching: Classification was rule-based using three curated taxonomies (Policy Themes, Query Types, Intended Use). We used weighted keyword matching where multi-word phrases received higher priority than single-word matches. ⢠Conversational Filtering: A separate conversational keyword list was checked first to identify and segregate chit-chat or phatic interactions. ⢠Imputation: A TF-IDF step surfaced representative terms. Finally, uncategorized queries were forward-filled based on neighboring categorized queries within the same session to leverage conversational context. D.4. Positionality Statement We acknowledge that our backgrounds shaped both AVA-AIâs design and evaluation. Our team spans humanâcomputer interaction researcher (4â8 years), development policy practitioner (8â15+ years across education, climate, and social protection), and machine learning engineer (3â6 years in retrieval systems). Several authors have first-hand policy experience and working within public-sector constraints; this vantage influenced our emphasis on citation-first interfaces, claim-level verification with coverage thresholds, and workflow integration aligned with policy research practice. We recognize this may bias us toward formal, documentable sources and structured evidence over other forms of knowledge (e.g., gray literature, oral testimony). To counterbalance, we evaluated AVA with users in diverse roles across multiple countries, reported failure cases and abstention triggers alongside successes. Our motivation stems from direct experience with information overload in development research, synthesizing large literature while maintaining source traceability under time pressure. We aim to contribute design patterns that augment expert judgment in knowledge-intensive domains. Appendix E Transferability of Design Principles to Medical and Legal Workflows Recent changes by Open AI restrict the use of general-purpose models for tailored legal and medical advice, formally acknowledging the evidentiary and safety risks of using LLMs in these domains (OpenAI, 2025). Yet, in practice, ChatGPT still answer legal and medical questions while appending only lightweight disclaimers (Business Insider, 2025), placing the burden of verification entirely on users. This design pattern risks being counterproductive: it normalises confident answers in domains where errors can be catastrophic, while offering minimal support for systematic checking. Prior work on AI-assisted decision making shows that when verification is costly or poorly supported, people frequently rely on AI outputs without thorough checking (Buçinca et al., 2021; Lyell and Coiera, 2017). AVA offers an alternative pattern. Rather than providing unconstrained answers with generic disclaimers, it constrains its knowledge to a vetted corpus, abstains when evidence is insufficient, and makes verification low-friction through page-level citations and in-context highlighting. Our findings in policy and development suggest that this âtrust pipelineâ: curation, calibrated abstention, and verificationâfirst interfaces directly addresses the kinds of risks that emerging legal and medical guidelines now seek to regulate. In clinical decision support, controlled scope and the ability to abstain are already treated as nonânegotiable safety requirements to avoid misdiagnosis or unsafe treatment recommendations (Gandouz et al., 2021; Schuster, 2025). AVA extends these arguments by quantifying how corpus coverage and abstention interact in a live deployment and by documenting a âtrust flipâ: once coverage is sufficient, conservative abstention is reinterpreted by professionals as evidence of caution rather than system failure. This behavioural pattern provides a transferable design insight for medical systems: abstention will only function as a safety signal if verification is easy and coverage is perceived as adequate. Similarly, several bar associations and courts in USA now permit lawyers to use generative AI tools provided they independently verify the accuracy and appropriateness of any AIâgenerated content before filing, reinforcing longâstanding duties of competence and candor (Justia, 2025). Our findings speak directly to this shift from âmay I use AI?â to âhow do I verify it in practice?â. AVA demonstrates how interface and corpus design can lower the cognitive cost of verification, through sourceâlinked answers and inâflow citation inspection, moving verification from an aspirational ethical norm to a routine, selectively applied practice (e.g., for surprising, numerical, or otherwise doubtful claims). Our results further indicate that, wherever possible, grounding generative systems in a vetted institutional corpus both reduces the incidence of hallucinated content and makes verification more trustworthy and efficient. Appendix F Stage-Wise Design Rationale Our architecture is structured around the four design goals (DG1âDG4) and informed by empirical findings from prior work on retrieval-augmented generation, agentic tool use, and verifiable LLM systems. Each stage addresses a distinct failure mode identified in earlier systems and is explicitly designed to ensure that downstream components operate over reliable, interpretable, and multilingual evidence. Stage 1: Curation and Hierarchical Indexing (DG1, DG3). Long-form policy analysis frequently requires synthesizing evidence scattered across multiple, non-contiguous sections of a document (Ma et al., 2023). Recent work on long-document retrieval and hierarchical RAG shows that flat, chunk-level indexing typically fails in such settings: when documents are treated as bags of independent spans, models miss cross-section links, headings, and layout cues, and retrieval quality degrades as document length grows (Li et al., 2025b). Standard RAG pipelines typically retrieve a handful of short, contiguous text spans, implicitly assuming that relevant information is localized in a single segment of text (Sarthi et al., 2024). To support such distributed evidence retrieval at scale, the first stage constructs a stable, provenance-preserving substrate that later stages depend on. Hierarchical document trees and stable node identifiers allow AVA to maintain exact byte offsets, page anchors, and cross-references across languages. This grounding ensures that verification in Stage 3 can reliably trace each factual claim to a precise, inspectable span (DG1), while multilingual semantic embeddings maintain access to evidence without requiring document-level translation (DG3). Stage 2: Agentic Retrieval and Evidence Formation (DG1, DG2). The goal of Stage 2 is to construct a sufficiently rich and diverse evidence pool to support robust verification. Prior work has documented that single-agent ReAct pipelines often suffer from role drift, brittle tool use, and insufficient query diversity in multi-hop settings (Yao et al., 2023; Shinn et al., 2023; Li et al., 2023). In contrast, multi-agent decompositional architectures have been shown to improve retrieval coverage, reduce hallucinations, and stabilize long-horizon behavior (Petcu et al., 2025; Jin et al., 2025). In AVA, specialized agents (decomposer, planner, walker, drafter) collaborate to surface complementary and sometimes contradictory evidence, a pattern that empirical research has found necessary for analytical and comparative tasks (Ammann et al., 2025; Sidiropoulos et al., 2021). This ensures that the verifier receives a balanced, multilingual set of candidate passages. Without such diversity, even a perfect verifier would be forced into excessive abstention due to insufficient evidence (DG2). Stage 3: Verification and Abstention (DG1). Verification operationalizes DG1âs requirement that factual claims must be explicitly supported by evidence. Prior work on fact-checking and LLM verification has emphasized that conservative acceptance thresholds reduce hallucinations but increase dependence on retrieval quality (Dhuliawala et al., 2024; Rahman et al., 2025). Following these findings, AVA adopts a strict coverage-and-agreement policy: unsupported claims trigger abstention, and contradictory evidence prompts explicit explanations of evidential gaps. The verifier thus acts as a high-precision gatekeeper rather than a secondary generator, aligning with emerging best practices in transparent RAG pipelines. Stage 4: Personalization and Rendering (DG3, DG4). Personalization in AVA adapts surface realization without weakening grounding constraints. Unlike translation-based RAG pipelines, which risk breaking citation alignment and require independent verification of translated spans (Li et al., 2025a), AVA delays language choice until rendering. Evidence packets remain language-agnostic, ensuring that personalization (e.g., preferred language, document preferences) operates strictly at the output layer. This design supports DG3 and DG4 while continuing to enforce DG1âDG2. Appendix G Additional Implementation Details G.1. Stage 1: Data Curation and Hierarchical Indexing For the dual-backend index described in Section 3.2.1, AVA adopts Qwen3-Embedding-8B as its primary multilingual encoder. Qwen3-Embedding-8B currently ranks at the top of the multilingual MTEB benchmark (Enevoldsen et al., 2025), with a mean score of 70.58 across 50+ tasks, and demonstrates strong cross-lingual retrieval and queryâpassage alignment performance. An alternative is the classic translateâthenâanswerâthenâtranslate pipelineâtranslating all documents and queries into a pivot language (typically English) and applying a high-performing monolingual encoder. However, prior cross-lingual IR research shows that while this baseline can be competitive, it introduces translation noise, higher latency, and significantly higher indexing and maintenance overhead (Lawrie et al., 2023). Moreover, multilingual encoders trained directly across many languages often match or outperform the translateâthenâanswerâthenâtranslate pipelines, especially on domain-specific or terminology-heavy content (Hu et al., 2020). Given these considerations, we use Qwen3-Embedding-8B as a strong, unified multilingual retriever without requiring an additional MT layer in the retrieval stack. G.2. Stage 2: Agentic Retrieval and Evidence Formulation Section 3.2.2 outlines AVAâs four retrieval agents at a conceptual level. Here we provide concrete implementation details and clarify the underlying models and patterns. G.2.1. Query Decomposer Agent The Query Decomposer Agent is implemented using a finetuned GPT-4o-mini model that follows a Chain-of-Thought pattern and can invoke ReAct-style tool use (Yao et al., 2023) when additional context is needed. Its role is to interpret the user query in the joint context of the document collection and the current session state. The agent can call tools that: ⢠identify entities, concepts, and timeframes, ⢠expand underspecified references (e.g., ârecent policiesâ), ⢠and derive sub-atomic queries that capture temporal and contextual constraints. ⢠perform intent classification In isolation, this ReAct-style decomposer is effective at finding fine-grained details but tends to over-focus on local aspects of the query. To avoid trajectories that lack a coherent high-level plan, we pair it with a dedicated Retrieval Planner Agent. G.2.2. Retrieval Planner Agent The Retrieval Planner Agent uses the same GPT-4o-mini model as the Query Decomposer and follows a âMulti-Agent Debateâ pattern (Liang et al., 2023). Its role is to coordinate retrieval by (i) reviewing proposed decompositions, (i) assigning retrieval strategies (lexical, semantic, structural, or hybrid) to each sub-query, and (i) tuning symbolic parameters such as quoted spans, temporal filters, or numerical constraints. We clarify here why differentiating retrieval strategies by query type is necessary, and how these choices impact retrieval quality. Prior Information Retrieval literature consistently shows that different query types benefit from different retrieval mechanisms. Exact lexical matchers excel on factoid or entity-centric queries where surface forms carry most of the discriminative signal (Wang et al., 2024b). In contrast, semantic retrievers (dense embeddings) outperform lexical methods on conceptual or cross-lingual queries where meaning is only weakly tied to surface forms (Karpukhin et al., 2020). Hybrid retrieval has emerged as a strong strategy when queries require both: for example, comparative or analytical questions where salient entities must be precisely matched, but broader conceptual framing is also needed. Our planner explicitly encodes these distinctions: the choice of strategy directly controls the balance between precision (lexical) and recall/semantic coverage (dense). Both the Query Decomposer and Retrieval Planner are instantiated from a checkpoint fine-tuned on a trajectory corpus. The SFT objective is not to improve raw reasoning ability but to inject priors about (i) cost-aware decomposition, (i) retrieval heuristics (e.g., quoting, temporal normalization, range queries) that improve precision/recall under tight budgets, and (i) mapping retrieved evidence back to precise spans for the verifier. This follows prior work showing that trajectory-level SFT improves consistency and grounding (Hui et al., 2025). G.2.3. Tree-Walker Agent and Evidence Packet Reranking The Tree-Walker Agent uses a GPT-40-mini model that is responsible for traversing the hierarchical document graph and semantic neighborhoods to collect evidence. It uses a ReAct pattern (Yao et al., 2023), interleaving thoughts with actions over the graph store and vector index. At each step, it can: ⢠move to structural neighbors (parent, children, cross-references), ⢠jump to semantically similar nodes, and ⢠adjust passage boundaries to align with sentence or cell-level units that support or contradict candidate claims. The Tree-Walker maintains a convergence criterion based on marginal information gain and a coverage threshold to avoid unbounded traversals. G.2.4. Drafting Agent The Drafting Agent takes the retrieved nodes/paths from the Tree-Walker Agent and reduces them into evidence packets that contain: document and node IDs, hierarchical paths, byte/character offsets, retrieval scores, and other metadata (e.g., temporal normalization, source diversity). For de-duplication and ranking, we use a two-stage scheme: (1) A base retrieval score derived from Qwen3-Embedding-8B similarity and lexical scores. (2) An LLM-based cross-encoder reranker using the same GPT-4o-mini model, which re-scores each evidence packet against the current sub-query or plan. The final ranking combines the base retrieval score and the cross-encoder score, which improves precision while maintaining diversity. Finally the agent uses these ranked evidence packets to produce an intermediate draft that explicitly encodes the structure of the final answer, including: ⢠section and paragraph structure, ⢠claimâevidence mappings (packet IDs, offsets), ⢠and placeholders for user-facing surface realization. This agent is implemented via a verification-aware prompting schema, ensuring that each factual statement is explicitly linked to one or more evidence packets. The drafting step provides a structured representation that the Verification Model can inspect before any final answer is rendered to the user. G.3. Stage 3: Information Synthesis and Verification. G.3.1. Evidence Classifier Model The Evidence Classifier is a fine-tuned BERT model. We fine-tune the model on a small set of passage-level annotations where relevant spans, entities, and supporting evidence are explicitly marked. Off-the-shelf BERT models handle generic NER and classification but are not calibrated to the structural and linguistic patterns in policy documents (cross-references, multi-sentence definitions, or table-embedded entities). Prompting alone leads to inconsistent span detection and low recall. Fine-tuning injects these domain-specific cues directly, producing stable, high-precision spans and entities for downstream grounding. G.3.2. Query Analyzer Model The Query Analyzer Model is a finetuned GPT-4o-mini model because the base version, while competent, is not reliably calibrated to our tool APIs or the retrieval-specific heuristics the system depends on (quoting, temporal normalization, parameterized queries). Prompting alone proved brittle over long ReAct (Yao et al., 2023) chains, often producing redundant or poorly scoped queries that increased latency and reduced recall. SFT on a small set of curated trajectories encodes these behaviors as stable priors, yielding a budget-aware planner that consistently maps claims to precise evidence spans, consistent with findings in prior trajectory-level SFT work (Jin et al., 2025). G.3.3. Response Drafter Model The Response Drafter Model uses a lightly fine-tuned GPT-4.1 checkpoint to synthesize the final answer from the structured evidence packets produced by the retrieval and verification stack. The model is fine-tuned on a small set of system-specific trajectories to (i) enforce a stable output schema compatible with downstream API consumers, (i) improve the mapping from verified evidence spans to well-structured natural language responses, and (i) minimize stylistic drift across user sessions. This is aligned with prior findings that task-format conditioning rather than extensive semantic tuning is sufficient for high-quality evidence-grounded generation in retrieval-augmented systems (Lewis et al., 2020; Shi et al., 2024). The fine-tuning data primarily consists of (query, evidence-packet, structured-output) triples in our domain, and the objective is to teach the model how to: (a) preserve citations and provenance metadata, (b) surface multiple perspectives when present in the evidence pool, (c) avoid hallucinating beyond verified spans, and (d) adhere to rendering conventions required by the AVA interface (e.g., section headers, bulletable fields, JSON-like attributes). G.3.4. Verification Model The Verification Model is built on GPT-4o-mini and is fine-tuned on a curated dataset of claimâevidence pairs annotated along two axes: coverage (âDoes the evidence directly support the claim?â) and agreement (âDo the retrieved sources converge, diverge, or contradict the claim?â). Training emphasizes robust negative examples, including partially supported claims, over-generalizations, and multi-source contradictions, enabling the verifier to reliably abstain when evidential grounding is incomplete. Because verification operates at the sentence level and requires only localized evidence inspection, a compact model such as GPT-4o-mini is sufficient; larger models would provide limited marginal benefit while significantly increasing computational cost given the high frequency of verification calls. At inference time, the model operates over sentence-level claim units extracted from the drafted answer. For each claim, it receives (i) the claim text, (i) a list of evidence snippets (with citation identifiers), and (i) document-level metadata captured during Stage 1 (hierarchical location, segment provenance). The verifier outputs structured scores for coverage and agreement, along with an accept/flag/abstain decision. Unsupported claims trigger abstention, while conflicting evidence leads to flagged explanations describing the inconsistency. Importantly, the verifier never rewrites content; it only audits and classifies, preserving the transparency and traceability required for DG1. G.4. Stage 4: User Personalization and Memory The user interface renders responses in the userâs preferred language without introducing a separate document-level translation phase. We avoid a separate translation stage by postponing any language choice until the very last mile. Translations are prohibitively expensive given the scale of the initial corpus (4000 documents). Also, adding a new translation layer adds unnecessary complexity for verifying the translation and another verification layer for mapping spans to translation and then back to the original source (Li et al., 2025a). When the userâs locale is known, or a specific locale is requested, the Drafting Agent performs surface realization in that locale, handling morphology, numerals and dates, while citations point back to the original spans. All Quotations and spans remain in the source language for fidelity. Compared to translateâanswerâtranslate baselines, this approach performs better when sources are multilingual, code-switched, or when precise span-level citation is required, while avoiding compounding translation errors and extra verification layers. Appendix H Comparison of Deep Research Systems Feature Perplexity AI (Generalist) AVA (Domain Specialist) Google NotebookLM (User-Grounded) Primary Corpus Open Web: Indexes the public internet; variable quality and provenance. Curated Institutional: Pre-indexed, high-trust library (4,000+ World Bank reports). User-Uploaded: User-Uploaded (BYO-Data) + Agentic Web: âDeep Researchâ can now autonomously fetch external sources. Epistemic Stance Answer-First: prioritizes fast, fluent answers with citations, and even Deep Research can still produce errors or misattributed details despite multi-step checking. Reasoned Abstention: Explicitly refuses (âI donât knowâ) if evidence is missing in the corpus. Bounded: Refuses to answer if not in sources, but lacks policy-aware redirection. Citation Granularity URL-Level: Links to web pages; often lacks pointers to specific text spans. Page-Level: Anchors span directly to the specific PDF page/paragraph in the viewer. Excerpt-Level: Links to specific text chunks in uploaded files. Architecture General Loops: Uses generic iterative search loops for broad synthesis. Policy Agents: Decomposer/Planner optimized for intent (e.g., separating âfiscalâ vs. âsocialâ). Single-Pass RAG: RAG for queries + Agentic loops for âDeep Researchâ (multi-step planning). Table 7. Comparison of AVA against Open-Web (Perplexity) and User-Curated (NotebookLM) Retrieval Systems. Appendix I AVA-AI v/s Deep Research Systems We acknowledge that AVAâs technical architecture utilizes established Retrieval-Augmented Generation (RAG) paradigms and standard Large Language Models (LLMs). We do not claim to introduce fundamental architectural novelty; rather, our study emphasizes the critical role of agentic orchestration and strict evidence thresholding in meeting the reliability standards required for high-stakes policy analysis. Grok 3 Scores System Comp. Rel./Cov. Coher. Appropr. Grammar Adherence Causal Safety Avg. AVA AI 7.53 7.53 8.48 7.99 9.80 8.22 7.43 9.80 8.35 NotebookLM 6.82 6.69 7.77 7.03 9.51 7.58 6.63 9.64 7.71 Perplexity AI 6.32 7.08 8.10 7.45 9.37 7.82 6.50 9.51 7.77 Qwen 80B Scores System Comp. Rel./Cov. Coher. Appropr. Grammar Adherence Causal Safety Avg. AVA AI 7.58 8.07 9.48 8.74 9.92 8.01 8.74 9.96 8.81 NotebookLM 7.83 8.10 9.12 8.23 9.74 8.35 8.33 9.88 8.70 Perplexity AI 7.38 8.34 9.28 8.83 9.67 8.35 8.28 10.00 8.77 Table 8. Exact evaluation scores (rounded to 2 decimals) across all rubric dimensions under Grok 3 and Qwen 80B. Each value represents the average score across all 92 queries. I.1. Evaluation Setup and Baselines To validate that our specific orchestration yields tangible benefits over standard implementations, we established a controlled evaluation environment to benchmark AVA against Google NotebookLM (Pro) and Perplexity Enterprise. We ingested a standardized curated corpus of 300 World Bank policy documents into all systems. The corpus size (n=300n=300) was determined by the maximum source capacity of Google NotebookLM (Pro), establishing a hard constraint for the baseline comparison. To evaluate performance across diverse retrieval tasks, a standardized set of 92 queries was issued to each model. These queries included: ⢠In-domain factual queries (e.g., regarding climate adaptation, fiscal decentralization), ⢠Analytical and comparative queries requiring synthesis across multiple documents, and ⢠Out-of-domain distractors designed to test epistemic humility (e.g., âWrite me a pizza recipe from the policy filesâ). We configured the baseline systems as follows: ⢠Google NotebookLM (Pro): We utilized the Pro tier to maximize the context window. As this version supports up to 300 sources per notebook, we were able to ingest the document set directly without file consolidation or truncation. ⢠Perplexity Enterprise (Web-Search Disabled): We utilized Perplexity Enterprise with the âweb searchâ feature disabled. This was critical to evaluate the systemâs ability to ground answers solely in the provided corpus and to test its abstention capabilities, specifically, whether it would correctly refuse to answer when information was absent from the source text, rather than hallucinating from external web knowledge. Responses were evaluated under an LLM-as-a-Judge framework consisting of two independent evaluators: Grok 3 and Qwen 80B. Each evaluator scored system outputs a score on a scale of 1-10 across eight dimensions: Comprehensiveness, Relevance & Coverage, Coherence, Appropriateness, Grammatical Correctness, Adherence to Constraints, Causal Reasoning, and Safety/Bias. To evaluate epistemic humility, we additionally measured each systemâs abstention rate, i.e.,the proportion of queries for which the model explicitly declined to answer due to insufficient evidence. I.2. Results Table 8 reports the average of rubric scores across all eight dimensions and both evaluators for all the queries. All values are rounded to two decimal places. The data demonstrates that AVA achieves performance parity with commercial research systems in general linguistic capabilities. In the Grok 3 evaluation, AVA achieved an average quality score of 8.35, exceeding the scores of Perplexity AI (7.77) and NotebookLM (7.71). Evaluations by Qwen 80B showed a highly competitive landscape, with AVA scoring 8.81 versus 8.77 for Perplexity. These results demonstrate that our architectural constraints do not degrade the fundamental quality or coherence of the generated responses relative to larger commercial systems. Figure 6. Perplexity AIâs behavior under the âpizza recipeâ stress test. When prompted with an out-of-domain question, such as âWrite a recipe for a pizza,â the system produces a confident, fully-formed answer rather than declining, despite the absence of any supporting evidence in the policy corpus. The screenshot illustrates the modelâs tendency to generate plausible but ungrounded content. Perplexity AIâs behavior under the âpizza recipeâ stress test. When prompted with an out-of-domain question, such as âWrite a recipe for a pizza,â the system produces a confident, fully-formed answer rather than declining, despite the absence of any supporting evidence in the policy corpus. The screenshot illustrates the modelâs tendency to generate plausible but ungrounded content. I.3. Epistemic Humility: A Critical Divergence While all systems exhibit comparable linguistic performance, a critical divergence appears in âEpistemic Humilityâ, a core requirement in high-stakes domains such as policy and development. We assessed the systemsâ ability to decline out-of-domain queries (e.g., requesting a âpizza recipeâ from policy files), revealing sharp differences in refusal behavior: ⢠AVA (11.90% Abstention): The system correctly identified when the source text did not support an answer, declining 11 out of 92 queries, including all out-of-domain distractors. ⢠NotebookLM (3.20% Abstention): The model displayed inconsistent behavior, occasionally declining requests but often producing fabricated or tangential content to satisfy the prompt. ⢠Perplexity AI (0% Abstention): The system attempted to answer every query regardless of source support. Notably, in the âpizzaâ stress test, it hallucinated a recipe complete with fabricated citations to World Bank policy documents to justify ingredients like yeast and flour (see Fig 6). Scoring Prompt You are a meticulous AI judge. A human will give you a QUERY and a RESPONSE from Chatbot A and you score the RESPONSE based on different criteria. You are tasked to evaluate Chatbot Aâs answer to the Query. Use these definitions for your scoring: * Comprehensiveness: Comprehensiveness / coverage of key points gauges whether the answer addresses every major facet of the sources given the user request, providing necessary details, examples, and context. Any omission or superficial treatment of critical elements lowers the rating, while thorough inclusion of all relevant aspects raises it. * Relevance: the answer directly addresses the userâs question and adds practical value rather than drifting off-topic. * Grammatical Correctness: multilingual language fluency and grammatical correctness in the target language. * Adherence to Constraints: respects constraints such as format, length, style, tools, or source-use requirements. * Clarity and coherence: well-structured, easy to follow, and logically consistent throughout. * Safety / Bias: avoids harmful, sensitive, or disallowed content and maintains policy compliance. * Causal Reasoning: shows logically sound reasoning and avoids self-contradiction or non sequiturs when explanations are requested. * Appropriateness: assesses how suitable, fitting, and contextually proper the generated response is to the userâs query. Return a JSON object that follows exactly this schema (no additional keys, no trailing commas): âComprehensivenessâ: ÂĄ1-10, 1 = worst, 10 = perfectÂż, âRelevanceâ: ÂĄinteger 1-10Âż, âCoherenceâ: ÂĄinteger 1-10Âż, âAppropriatenessâ: ÂĄinteger 1-10Âż, âGrammatical Correctnessâ: ÂĄinteger 1-10Âż, âAdherence to Constraintsâ: ÂĄinteger 1-10Âż, âCausal Reasoningâ: ÂĄinteger 1-10Âż, âSafety / Biasâ: ÂĄinteger 1-10Âż, âcommentâ: ÂĄone concise sentence explaining the scoreÂż Two other candidate answers are shown only so you can compare quality wise. Base your judgment on how well A satisfies the official rubric and how it stacks up against B and C and is relevant to user-provided sources. This would help you better understand relative advantages or shortcomings vs. the other answers. Weaknesses that are hard to notice in isolation become obvious when a stronger answer is present among the three chatbots. Chatbots must exhibit epistemic humility i.e. they understand their limitations and lack of knowledge in out of context cases. I.4. Conclusion These behavioral differences stem from conflicting optimization goals rather than fundamental capability gaps. Commercial tools parameterized to prioritize conversational fluency, lowering retrieval thresholds to maximize engagement. This results in a near-zero abstention rate, leading to âforced hallucinationsâ when information is absent. In contrast, AVA utilizes strict evidence threshold. We consciously trade marginal conversational fluency gains (where commercial models scored slightly higher, e.g., 8.35 (both NotebookLM and Perplexity vs. 8.01 AVA; Qwen 80 B as a judge) for epistemic humility. This orchestration determines the systemâs suitability for high-stakes enterprise environments where the ability to remain silent is as critical as the ability to generate text. Appendix J Operationalizing Epistemic Humility in Generative AI Systems Epistemic humility in AI systems concerns how a system bounds its claims, signals uncertainty, and avoids over-assertion when available evidence is insufficient. This is closely tied to questions of trust calibration (Zhang et al., 2020), responsible reliance (Schemmer et al., 2022), and usersâ ability to verify system outputs (Lyons et al., 2021). Prior work has explored multiple ways of operationalizing epistemic humility in AI systems; below, we outline the dominant strategies and motivate the specific design choices adopted in AVA. J.1. Design Space: Approaches to Bounding Knowledge in AI Systems Prior work operationalizes epistemic humility through a range of complementary mechanisms: (1) Uncertainty Signaling and Confidence Calibration. Some systems express epistemic limits through confidence scores, probabilistic estimates, or hedging language. These approaches aim to prevent overconfidence by making uncertainty visible to users. However, studies have shown that users often misinterpret numerical confidence signals or ignore them altogether, especially in complex decision-making contexts (Cao et al., 2024). Furthermore, confidence markers can paradoxically reinforce over-reliance if users fail to distinguish confidence from reliability (Zhou et al., 2024). As a result, uncertainty signaling alone is insufficient to prevent over-reliance. (2) Source Attribution and Verifiable Citation. Another prominent approach grounds system outputs in identifiable sources, enabling users to inspect, verify, and contest claims. Retrieval-augmented generation (RAG), citation-aware generation, and page-anchored references exemplify this strategy (Gupta et al., 2024; Cao and Wang, 2024). Source attribution supports epistemic humility by shifting authority away from the model toward underlying evidence. However, attribution alone does not prevent systems from synthesizing unsupported claims when retrieved evidence is sparse, ambiguous, or misaligned (Liu et al., 2023b; Chiang and Lee, 2024). (3) Faithfulness and Grounding Constraints. Some systems enforce faithfulness by constraining generation to retrieved inputs, quoting verbatim source material, or penalizing unsupported synthesis during training (Gao et al., 2023b; Hu et al., 2024). These approaches reduce hallucination but often trade off coverage and flexibility (Levonian et al., 2023; Kim et al., 2025). Moreover, even grounded systems may still produce misleading summaries if evidence is incomplete or contradictory (Chiang and Lee, 2024; Wan et al., 2024). (4) Abstention and Refusal Mechanisms. A more explicit form of epistemic humility is abstention, defined as the system declining to answer when evidence is insufficient (Madhusudhan et al., 2025). Prior work explores selective refusal (Arditi et al., 2024; Yuan et al., 2025), scope-aware abstention (Liu et al., 2023a; Yin et al., 2023), and âI donât knowâ responses (Deng et al., 2024; Bastounis et al., 2024) as mechanisms to prevent false authority. Abstention makes epistemic limits explicit rather than probabilistic (Yadkori et al., 2024). Each approach addresses knowledge limits differently, but also exhibits known shortcomings when used in isolationâsuch as misinterpretation of confidence cues, or unhelpful refusals. J.2. Rationale for AVAâs Hybrid Design Consistent with DG1 (Verifiability and Epistemic Humility) articulated in Section 3, AVA adopts a hybrid operationalization that combines source-linked generation, scope awareness, and reasoned abstention, motivated by the requirements of high-stakes policy and development work. Policy professionals routinely engage in verification practices and require traceability to primary sources; accordingly, AVA employs a Retrieval-Augmented Generation (RAG) architecture over a curated library of 4,000+ World Bank reports, producing page-level, clickable citations that support inspection and contestation within existing workflows. In addition, long-form generative synthesis amplifies hallucination risk when evidence is incomplete, ambiguous, or misaligned. In such contexts, uncertainty signaling alone can still convey false authority (Petcu et al., 2025; Ammann et al., 2025). AVA therefore incorporates explicit abstention, providing a stronger boundary signal by clearly demarcating when the system cannot substantiate an answer. Abstention in AVA is scope-aware and reasoned: the system explains why available evidence is insufficient and, where possible, suggests reformulation paths or adjacent queries, preserving task continuity while maintaining clear epistemic boundaries. These mechanisms operationalize humility as a user-facing system capability that prioritizes verifiability, explicit boundary-setting, and continuity of professional work over fluent but unsupported generation. Appendix K Demography Table P.ID Profession Institute Experience Level Age Range G Country Languages Known P1 Lecturer University Early 30â39 F Nigeria English P2 Researcher Independent Early 30â39 F Uganda English, Luo, Swahili P3 Researcher University Early 30â39 F Nigeria English; Arabic P4 Development professional NGO / Civil society Middle 30â39 M Angola English P5 Researcher University Middle 40â49 M Malaysia English P6 Policy analyst & researcher Government Senior 50â59 F Nigeria English, Igbo P7 Lawyer, Legal Consultant Government Early 30â39 F Argentina English, Spanish P8 Economist International Organization Early 30â39 F USA English P9 Think tank professional & NGO worker Independent Senior 50â59 M Nigeria English P10 Educator (Teacher Coach) Private school Middle 40â49 M Nigeria English, Yoruba P11 PhD Student University Early 19â29 M USA English P12 AI Startup Founder Startup Early 19â29 M Ghana English P13 Lecturer & Startup Founder University and Startup Senior 50â59 M Tanzania English, Swahili P14 Public Policy Manager NGO Middle 30â39 F Canada English, French, Spanish P15 Manager International Organization Middle 30â39 M Uganda English, French P16 Researcher NGO Early 30â39 M Uganda English P17 Consultant Private Senior 40â49 M France French, English P18 Strategy Officer International Organization Middle 40â49 M USA English P19 Executive Assistant / Impact Evaluation Government Middle 30â39 M Nigeria English, Hausa, Marji, Basic Arabic P20 Product & Strategy Manager NGO Early 19â29 F USA English, Mandarin Table 9. Demographic Details of Participants (Note: Early = early career professionals, Middle = middle career professionals, Senior = senior-level professionals; F = Female, M = Male). Appendix L Diagrams Figure 7. Global Distribution of AVA Users: The map covers 116 countries, with darker shading indicating higher user concentrations. Major user bases include the United States, Brazil, Western Europe, Nigeria, and India. Choropleth world map of the Global Distribution of AVA Users from the impact evaluation sample. The map visualizes data from a total of 1,055 users across 117 countries, where countries are shaded in progressively darker shades of blue to indicate a higher number of users. The highest concentrations of users (ranging from 9 to 281 per country) are located in the United States, Brazil, several countries in Western Europe, Nigeria, and India. Figure 8. The AVA Project Roadmap, illustrating key milestones from April to August 2025. A vertical timeline illustrating the AVA Project Roadmap with eight key milestones from April to August 2025. The roadmap begins with user registration and pilot launches in May, progresses through content and feature additions in July, and concludes with the public launch on August 4, 2025. Figure 9. Distribution of the Top 10 Query Languages (N=7,088N=7,088).English dominates the dataset, accounting for 83.8% of all queries. French (4.8%) and Spanish (3.0%) are the next most frequent, with all remaining languages comprising less than 1.5% each. A horizontal bar chart shows the percentages for the top 10 languages used in 7,088 queries. English is the dominant language, accounting for 83.8 percent of all queries. French and Spanish are the next most frequent, at 4.8 and 3.0 percent respectively, with all other languages comprising less than 1.5 percent each. Figure 10. Distribution of user queries by policy theme and query type (N=1,582). Diagnostic queries predominated (69.0%), especially within the People (23.6%) and Prosperity (21.8%) themes. A heatmap illustrating the joint distribution of 1,582 user queries across five policy themes and three query types. Overall, Diagnostic queries are the most frequent type (69.0 percent), while People (33.3 percent) and Prosperity (29.7 percent) are the most common themes. The highest concentrations of queries, represented by the darkest cells, are for Diagnostic queries within the People (23.6 percent) and Prosperity (21.8 percent) themes. Evaluation queries are the least common type across all themes.