Paper deep dive
Don't Measure Once: Measuring Visibility in AI Search (GEO)
Julius Schulte, Malte Bleeker, Philipp Kaufmann
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 4:06:14 AM
Summary
This paper investigates the reliability of Generative Engine Optimization (GEO) metrics by analyzing the visibility and stability of brand mentions and cited sources in AI-based search engines (ChatGPT, Gemini, Google AI Mode, and Perplexity). The authors demonstrate that AI search results exhibit significant stochastic volatility, both over time and within simultaneous repeated runs, rendering single-point measurements unreliable. They propose that GEO performance should be characterized as a distribution rather than a single value, utilizing Jaccard similarity and Rank-Biased Overlap (RBO) to quantify visibility stability.
Entities (7)
Relation Signals (3)
Jaccard similarity → measures → Visibility
confidence 95% · Jaccard similarity (Jaccard, 1901) measures the set overlap of cited sources (or detected brands) between two observations
GEO → requires → Repeated Measurements
confidence 95% · our findings underscore the need for repeated measurements to assess a brand’s GEO performance
ChatGPT → exhibitsstochasticvariation → Visibility
confidence 90% · The inherent probabilistic nature of AI search changes this paradigm. Answers can vary across runs, prompts, and time
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As large language model-based chat systems become increasingly widely used, generative engine optimization (GEO) has emerged as an important problem for information access and retrieval. In classical search engines, results are comparatively transparent and stable: a single query often provides a representative snapshot of where a page or brand appears relative to competitors. The inherent probabilistic nature of AI search changes this paradigm. Answers can vary across runs, prompts, and time, making one-off observations unreliable. Drawing on empirical studies, our findings underscore the need for repeated measurements to assess a brand's GEO performance and to characterize visibility as a distribution rather than a single-point outcome.
Tags
Links
- Source: https://arxiv.org/abs/2604.07585v1
- Canonical: https://arxiv.org/abs/2604.07585v1
Trouble viewing inline? Open PDF directly →
Full Text
57,545 characters extracted from source content.
Expand or collapse full text
DON’T MEASURE ONCE: MEASURING VISIBILITY IN AI SEARCH (GEO) Julius Schulte ∗ Global Center for Entrepreneurship and Innovation University of St. Gallen 9000 St.Gallen, Switzerland julius.schulte@unisg.ch Malte Bleeker Marketing und Customer Insight University of St. Gallen 9000 St.Gallen, Switzerland malte.bleeker@unisg.ch Philipp Kaufmann Marketing und Customer Insight University of St. Gallen 9000 St.Gallen, Switzerland philipp.kaufmann@unisg.ch April 10, 2026 ABSTRACT As large language model-based chat systems become increasingly widely used, generative engine optimization (GEO) has emerged as an important problem for information access and retrieval. In classical search engines, results are comparatively transparent and stable: a single query often provides a representative snapshot of where a page or brand appears relative to competitors. The inherent probabilistic nature of AI search changes this paradigm. Answers can vary across runs, prompts, and time, making one-off observations unreliable. Drawing on empirical studies, our findings underscore the need for repeated measurements to assess a brand’s GEO performance and to characterize visibility as a distribution rather than a single-point outcome. Keywords AI Visibility·Generative Engine Optimization (GEO)·Search Engine Optimization (SEO)·Information Retrieval 1 Introduction With the emergence of Large Language Models (LLMs), the way consumers retrieve information is undergoing a paradigm shift. This transformation poses a direct challenge to the traditional search ecosystem (Wen et al., 2025), which remains the foundational marketing communication channel for most companies (Stürze et al., 2022). Signs of a market restructuring are evident: Google, having historically monopolized search in the Western hemisphere, experienced its first market share drop in a decade in late 2024 (Goodwin, 2025). In stark contrast, generative search (or LLM-based search) has seen explosive growth, exemplified by ChatGPT increasing its weekly active user base from≈ 100 million (Jan. 2024) to≈780 million (Sept. 2025) (Chatterji et al., 2025). Some studies estimate that ChatGPT will surpass Google in four years (Krisinger, 2025). This trend reflects a broader change in user behavior, as consumers increasingly rely on AI agents not just for information retrieval, but for executing complex tasks such as shopping, planning, and coding (Economist, 2025). This paradigm shift from traditional search to generative search fundamentally alters the mechanisms of performance monitoring. While digital marketers have historically benefited from high levels of data transparency in Search Engine Optimisation (SEO), which is facilitated by first-party utilities such as Google Search Console (GSC), the ∗ Affiliated with Aurora Intelligence. arXiv:2604.07585v1 [cs.IR] 8 Apr 2026 Don’t Measure Once: Measuring Visibility in AI Search (GEO) transition toward Generative Engine Optimisation (GEO) has introduced a significant observability gap (Aggarwal et al., 2024). Unlike traditional search engines, the providers of LLMs do not currently offer native, proprietary monitoring tools equivalent to GSC. Consequently, foundational metrics such as which specific search queries are used and the corresponding query volumes are no longer directly observable within the GEO ecosystem. In the absence of ground-truth data, marketers must adopt alternative methodologies to evaluate performance. The emerging industry standard is to measure visibility, which is defined as the frequency and prominence of brand mentions within generated responses (Rejón-Guardia et al., 2025). However, relying on static visibility alone is insufficient due to the stochastic nature of generative models. This study introduces stability as a critical, complementary performance dimension. Even as novel third-party tools are developed to quantify AI-specific visibility (Rejón-Guardia et al., 2025), marketers must avoid drawing premature conclusions from these snapshot metrics. Without accounting for the stability of mentions over time and across varying prompt iterations, visibility data remains prone to volatility and misinterpretation. 2 Literature Recent empirical research provides behavioural evidence highlighting the shift from traditional search to generative search. Using click-stream data, Padilla et al. (2025) show that adoption of AI-based search engines leads to a gradual but substantial decline in traditional search. Specifically, traditional search queries fell by more than 20% after generative search adoption, with particularly strong reductions in informational and question-based searches (Padilla et al., 2025). This suggests that generative search increasingly acts as a substitute for traditional search engines, particularly for complex or knowledge-seeking queries. This shift towards AI search requires marketers to change their performance measurement. Research offers marketers insights into how GEO performance can be measured. Aggarwal et al. (2024) indicate that GEO performance can be evaluated using visibility (impression) metrics, such as position-adjusted citation prominence and subjective relevance of sources within generative engine responses. This aligns with Chen et al. (2025), who argue that GEO visibility can be assessed through domain (brand) presence and citation-based source visibility within AI-generated responses. Therefore, a shift from click-based metrics to visibility occurs (Rejón-Guardia et al., 2025). However, quantifying visibility as a GEO performance metric presents inherent challenges due to the intransparent architectures of generative search engines. These models function as "black boxes", restricting a firm’s ability to predict precisely when or how its brand will be referenced. Whereas SEO visibility typically oscillates along a deterministic ranking spectrum, GEO visibility is subject to far greater instability. LLM-generated responses often exhibit a binary inclusion-exclusion dynamic, where a source is either prominently integrated or omitted entirely. As Wen et al. (2025) articulate, "unlike SEO, which competes for ranked link positions, GEO focuses on inclusion and prominence within LLM-generated answers" (Wen et al., 2025, 2). This phenomenon is driven by the probabilistic nature of token generation and retrieval-augmented evidence selection processes that compress information from diverse sources into a constrained answer space, thereby increasing visibility volatility (Aggarwal et al., 2024). As prior research calls for improved tracking of LLM search outputs and brand visibility within generative engines (Wen et al., 2025), this study extends existing work by focusing not only on the measurement of GEO visibility but also on the stability and consistency of this visibility across prompts and verticals. Thereby, it expands previous GEO visibility papers (e.g., (Aggarwal et al., 2024)) that do not explicitly measure or empirically quantify GEO visibility fluctuation. 3 Methods & Datasets For the analysis, two datasets were created. The first dataset contains the daily results of four AI search engines across four Swiss-German campaign verticals, collected over a 45–46-day window (Jan 24 – Mar 20, 2026). The second contains results of repeated prompts submitted simultaneously on the same day to isolate stochastic variation from temporal drift (see Section 5). The prompts were derived from high-search-volume SEO keywords. These keywords were entered into Google, and the “People Also Ask” feature was subsequently used to identify and generate relevant prompts. Eight prompts per campaign were selected, approximating real user search behaviour, as 70% of AI-powered search users ask top-of-funnel questions to learn more about products and services (Silliman et al., 2025), and users are increasingly moving away from simple keywords toward a conversational tone (Martins, 2025). The study covers four verticals, namely Telecommunications, Real Estate Sales, Sporting Goods, and Consumer Electronics, representing frequently searched domains in Swiss 2 Don’t Measure Once: Measuring Visibility in AI Search (GEO) SEO environments. 2 For each campaign, eight prompts were entered into four engines: ChatGPT, Gemini, Google AI Mode, and Perplexity. The full list of prompts in both German (original) and English is provided in Appendix G. Data coverage is summarised in Table 1; rationale for restricting to this single period is given in Section 6. Table 1: Data coverage by campaign and search engine, Jan 24 – Mar 20, 2026 (collection days with≥ 1 result). CampaignQueriesDaysChatGPTGeminiGoogle AI ModePerplexity Consumer Electronics84543224343 Real Estate Sales84539264343 Sporting Goods84638234444 Telecommunications84540234342 Note: Gemini had sporadic gaps in this period. Jan 30, 2026 excluded (citation volume≈ 2× the daily average). Similarity Metrics The stability of AI-based search engine results is measured along two dimensions: differences in cited sources and differences in mentioned brands across repeated prompts. Two complementary metrics are used. Jaccard similarity (Jaccard, 1901) measures the set overlap of cited sources (or detected brands) between two observations, defined asJ(A, B) = |A ∩ B|/|A ∪ B|(see Appendix D). It is rank-agnostic but intuitive and easy to interpret. Rank Biased Overlap (RBO,p=0.9) addresses Jaccard’s rank insensitivity by weighting items at the top of the ranked list more heavily than those further down (Webber et al., 2010), making it appropriate when source position reflects relevance (see Appendix E). The non-extrapolated minimum-bound variant(4)is used, truncating the weighted sum at k = min(|S|,|T|). The unified policy for handling empty source and brand sets is detailed in Appendix C. 4 Results 1: Source Visibility over Time Across all four campaigns over the 45–46-day observation window (Jan 24 – Mar 20, 2026), the day-to-day Jaccard similarity for cited sources averages between 0.34 and 0.42 (see Table 2 and Figure 1). A Jaccard value of 0.35 implies that, on average, only about 35% of the cited sources overlap between two consecutive days — meaning roughly 65% of all sources change from one day to the next. The RBO scores are consistently lower than Jaccard (0.21–0.26), indicating that not only do the source sets change, but so does the rank order in which they appear. These values confirm that day-to-day source instability is a persistent characteristic of AI search, not a transient artifact. Table 2: Day-to-day source similarity by campaign (Jan 24 – Mar 20, 2026). CampaignJac. MeanJac. SDRBO MeanRBO SD Consumer Electronics0.3360.2430.2060.163 Real Estate Sales0.3780.2930.2560.219 Sporting Goods0.3550.2690.2240.183 Telecommunications0.4230.2440.2530.180 Note: 4,044 consecutive-day pairs aggregated across all queries and engines. RBO at p=0.9. Edge-case policy: see Appendix C. Beyond source-level instability, we additionally examine brand-level visibility. For each response we detect brands mentioned in the answer text using a campaign-specific lexicon (32–51 canonical brands per vertical; see Appendix F for the full list). Before computing brand similarity, we apply a quality filter: campaigns are included in the brand analysis only if their mean brand-detection rate across all runs exceeds 70%. This threshold was computed on the temporal dataset (Jan 24 – Mar 20, 2026); the same campaign qualifications were then applied to the simultaneous-run analysis. The Real Estate Sales campaign falls below this threshold (mean detection rate: 53.6%), driven by several generic tax- and investment-oriented queries (e.g., "Wie viel ist kapitalertragssteuerfrei?") for which LLMs answer without citing any specific brand. It is therefore excluded from brand similarity analyses; its brand lexicon is retained in the appendix for completeness. Among the three qualifying campaigns (Telecommunications, Sporting Goods, Consumer Electronics), the resulting brand similarity scores, reported in Table 3, are markedly higher than source similarity — Jaccard values of 0.45–0.59 — reflecting that brand mentions are somewhat more stable than individual source citations. Nevertheless, substantial day-to-day variation remains: RBO scores of 0.19–0.30 indicate that the implicit ordering of brands within responses 2 Henceforth, these are called campaigns. The original German campaign labels are listed in Appendix A. 3 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 1: Day-to-day Jaccard similarity (left) and Rank-Biased Overlap (right) for cited sources across all four campaigns (Jan 24 – Mar 20, 2026). Each box aggregates all 4,044 consecutive-day pairs. Median values annotated above each box. Source sets overlap by only 34–42% on average. also shifts considerably over time. Sporting Goods shows the lowest brand Jaccard (0.45), likely because the large pool of substitutable sports-shoe brands means the model draws from a wide set across days. Table 3: Day-to-day brand similarity by campaign (Jan 24 – Mar 20, 2026; campaigns with mean brand-detection rate ≥ 70%). CampaignJac. MeanJac. SDRBO MeanRBO SD Consumer Electronics0.5570.2370.2890.184 Sporting Goods0.4530.3260.1870.182 Telecommunications0.5890.2110.3040.179 Note: Finance and Real Estate Sales excluded. 2,924 consecutive-day pairs (non-NaN). Edge-case policy: see Appendix C. Taken together, these results show that both the specific sources cited and the brands mentioned in AI-generated responses fluctuate substantially from day to day across a 45–46-day window. Source instability appears to be a persistent property of the generative search process rather than an occasional glitch. Source citation inequality is also notably high: a small number of domains account for the vast majority of citations across all campaigns and engines. Figure 3 shows Gini coefficients per campaign and engine for Jan 24 – Mar 20, 2026 (after filtering theimages.openai.comCDN artifact; see Section 6). The mean Gini across all campaigns and engines is 0.715. Google AI Mode exhibits the highest citation concentration (Gini = 0.782), while Perplexity shows the lowest (0.671). Across campaigns, Telecommunications reaches the highest Gini (0.750) and Sporting Goods the lowest (0.680). These values imply a highly unequal citation landscape in which a handful of domains captures most of the AI-generated visibility. The numerical breakdown by campaign and engine and the formula with a worked example are provided in Appendix H–I. What is cited in one run today is not necessarily cited in the same run tomorrow. The drivers of this instability may be external — algorithmic updates, changes in domain authority, index freshness — or they may be time-independent, arising from the inherently probabilistic nature of the LLM’s output distribution. The simultaneous re-run analysis in Section 5 isolates these contributions. 4 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 2: Day-to-day brand similarity (Jaccard, left; RBO, right) for the three campaigns meeting the≥ 70%brand- detection threshold. Brand stability is higher than source stability but still far from perfect, with broad interquartile ranges indicating substantial response-to-response variation within campaigns. Median values annotated above each box. 5 Results 2: Source Visibility Under Simultaneous Re-Runs The temporal analysis in Section 4 establishes that sources and brands change from day to day. However, this variation could in principle be driven by changes external to the model itself — algorithmic updates or index freshness. To isolate the contribution of the model’s inherent stochasticity, we examine cases in which the same prompt was issued multiple times on the same calendar day. This removes temporal drift as a confounder: any variation observed within a single day is very likely to originate from the probabilistic nature of the LLM’s output or the output system of the respective AI-based search engine. We draw on a dedicated simultaneous-run collection, which contains 8 prompts per campaign queried up to 10 times to all 4 engines in succession. Because the collection spanned multiple calendar days for some engines (ChatGPT: Mar 21–23, 2026; Perplexity: Mar 23; Gemini: Mar 22–23; Google AI Mode: Mar 24–25), runs are filtered so that only pairs with a timestamp difference of at most 24 hours are compared, ensuring that temporal drift cannot confound the results. For source-overlap analysis a second quality filter is applied: only runs in which the engine returned at least one extracted citation are included, removing zero-citation responses that would otherwise inflate the false-zero Jaccard scores (75.4% of runs pass this filter; ChatGPT is lowest at 42.2%, reflecting its tendency to suppress web search on definitional queries). After both filters, 3,409 pairwise source comparisons remain across the four campaigns, with up to 10 runs per engine–prompt group. The similarity edge-case policy described in Section 3 applies here as well. Under classical search-engine assumptions one would expect near-perfect source overlap for identical queries issued within minutes of each other. As Table 4 shows, the actual pairwise Jaccard similarity for sources averages between 0.32 and 0.43 across campaigns — values in the same range as the day-to-day figures in Section 4, confirming that intra-day stochastic variation alone accounts for most of the observed instability. We repeat the same pairwise analysis for brand mentions, restricting to the three campaigns that surpass the 70% detection-rate threshold (Telecommunications, Sporting Goods, Consumer Electronics; see Section 4). Brand similarity uses all runs with a non-empty response within the 24-hour window (no citation requirement, since brands are extracted from response text). The brand-level Jaccard values (Table 5) are higher than source-level values for Consumer Electronics and Telecommunications (0.46–0.48), confirming that brand mentions are somewhat more stable than individual cited sources even within a single day. Sporting Goods shows a lower brand Jaccard (0.33), reflecting the wide interchangeable pool of running-shoe brands from which the model draws. All three campaigns show high 5 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 3: Source citation Gini coefficient by campaign and engine, Jan 24 – Mar 20, 2026. Values close to 1.0 indicate that a single domain captures nearly all citations; values around 0.6–0.7 indicate moderate concentration. All campaigns and engines exhibit high Gini values (mean 0.715). Table 4: Pairwise source similarity across repeated runs within a 24-hour window (Jaccard and RBO; runs with at least one extracted citation only). CampaignPairsMax RunsJaccard MeanJaccard SDRBO Mean Consumer Electronics830100.3270.2200.156 Real Estate Sales805100.3910.2590.225 Sporting Goods886100.3210.2190.149 Telecommunications888100.4340.2380.272 Note: Only pairs with|∆t|≤ 24 h compared. Runs with zero extracted citations excluded from source analysis. 4 engines: ChatGPT, Perplexity, Gemini, Google AI Mode. Up to 10 reps per engine–prompt. Both-empty source pairs excluded (NaN policy). RBO at p=0.9. within-campaign variance (SD≈ 0.30): some prompts yield near-perfect brand consistency across runs while others change almost entirely, consistent with the prompt-level heterogeneity documented in Figure 5. Figure 4 visualises the full distribution of pairwise Jaccard scores across campaigns. The consistently low median values and broad interquartile ranges demonstrate that LLM output variation is not reducible to external temporal factors: a substantial fraction of observed instability originates from the model’s stochastic generation process itself. The practical implication is direct: if a marketer queries an AI search engine once on a given day, the resulting brand-visibility snapshot may differ substantially from a second query executed minutes later under identical conditions. Characterising true GEO visibility therefore requires aggregating over multiple runs rather than relying on a single observation. Figure 5 shows per-prompt mean Jaccard and RBO values for both source and brand similarity across all campaigns. The top row covers source similarity for all four campaigns; the bottom row covers brand similarity for the three 6 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Table 5: Pairwise brand similarity across repeated runs within a 24-hour window (Jaccard and RBO; campaigns with mean detection rate≥ 70%). CampaignPairsMax RunsJaccard MeanJaccard SDRBO Mean Consumer Electronics1235100.4770.2980.229 Sporting Goods1027100.3270.2980.175 Telecommunications1233100.4630.3010.220 Note: Finance and Real Estate Sales excluded (detection rate < 70%). Pairs with|∆t|≤ 24 h only (Mar 21–25, 2026; 10 reps). Brand detection uses a campaign-specific lexicon (see Appendix F). Both-empty pairs excluded per edge-case policy (Section 3). Table 6: Mean pairwise similarity within 24 hours by engine: sources (cited runs only, all 4 campaigns) and brands (Telecommunications, Sporting Goods, Consumer Electronics). SourceBrand EngineJac. MeanRBO MeanJac. MeanRBO Mean ChatGPT0.2330.0880.4370.192 Perplexity0.2820.1020.4920.202 Gemini0.5050.2300.4090.196 Google AI Mode0.3180.2540.3750.238 Note: Source columns include only runs with≥ 1 extracted citation and pairs with|∆t|≤ 24 h. Brand columns restricted to campaigns meeting the≥ 70% detection-rate threshold. RBO at p=0.9. qualifying campaigns (Telecommunications, Sporting Goods, Consumer Electronics). Within each campaign, prompts vary substantially in their similarity levels, suggesting that query specificity — rather than engine behaviour alone — determines how consistently a prompt is answered. Specific product queries (e.g., “Welche Sportschuhe sind die besten?”) tend to attract more consistent source and brand sets than broad, generic queries. As Figure 5 illustrates, variability across prompts is substantial: some prompts yield consistently high similarity (Jaccard > 0.8) while others remain persistently low (< 0.2). This prompt-level heterogeneity implies that single-prompt- based visibility assessments are unreliable, and that monitoring strategies must account for query-level variation in addition to temporal and simultaneous-run variation. Together with the temporal results, this strongly motivates a repeated-measurement framework for GEO monitoring. 6 Limitations Several data-quality and methodological limitations should be noted when interpreting the results. ChatGPT data collected containedimages.openai.com— an OpenAI image-delivery CDN — as a spurious domain data (889 occurrences, 5.8% of ChatGPT P2 citations). Because mixing API and interface data would create a methodological inconsistency for ChatGPT specifically, all analyses in this paper are restricted to the January 24 – March 20, 2026 window, where collection is consistent across all four engines. Theimages.openai.comdomain is additionally filtered from all calculations. Future studies should pin collection method and model version across the full observation window. Swiss-server context. All data were collected from servers located in Switzerland. Prompts are therefore served with Swiss IP addresses and locale settings, which may affect geo-personalised index selection, language weighting, and citation patterns. Results may not generalise to other regional or linguistic markets. Simultaneous-run collection window. The geo-brand-monitor simultaneous collection spanned multiple calendar days for some engines (ChatGPT: Mar 21–23; Perplexity: Mar 23; Gemini: Mar 22–23; Google AI Mode: Mar 24–25, 2026). The raw dataset additionally contains responses from Google AI Overviews (“Google AIO”, Mar 22–23), a distinct Google product that generates AI-summary snippets within regular search results rather than operating as a dedicated AI search interface. Because Google AIO differs fundamentally in interaction mode and citation behaviour from the four dedicated AI search engines in the study, it is excluded from all analyses; the study focuses on ChatGPT, Gemini, Google AI Mode, and Perplexity. The analysis further controls for multi-day collection by including only pairs whose timestamps are within 24 hours of each other; all cross-day pairs exceeding this window are discarded. Additionally, ChatGPT activates web search only for specific queries, leaving 57.8% of its runs with zero citations; the source-similarity analysis therefore excludes zero-citation runs across all engines to ensure comparisons reflect genuine source overlap rather than collection failures. 7 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 4: Pairwise Jaccard and RBO for sources (top row) and brands (bottom row) across repeated runs within a 24-hour window (geo-brand-monitor collection, Mar 21–25, 2026; 4 engines, up to 10 reps). Source boxes include only runs with≥ 1extracted citation. Source overlap averages 32–43%; brand overlap ranges from 33% (Sporting Goods) to 48% (Consumer Electronics). Median values annotated above each box. Brand-detection coverage. Brand detection relies on substring matching of a fixed lexicon. Brands cited via synonyms, abbreviations, or paraphrases are missed; conversely, generic terms that are substrings of brand names may produce false positives. The 70% detection-rate threshold used to qualify campaigns for brand similarity analysis mitigates the worst of this, but does not eliminate the problem. 7 Conclusion The research demonstrates that visibility in AI search is inherently unstable and cannot be treated like traditional SEO rankings. In contrast to SEO, where results may shift in position but typically remain present in the ranking set, generative search operates on an inclusion–exclusion dynamic in which brands or sources may appear in one response and disappear entirely in another. Even when identical prompts are executed simultaneously under controlled conditions, cited sources and brand mentions vary substantially. Source sets overlap by only 34–42% between consecutive days; brand sets somewhat more, at 45–59%, but with wide variance. As a result, single observations of AI visibility are misleading and risk over- or underestimating true brand presence. 8 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 5: Per-prompt mean Jaccard (dark) and RBO (light) for source similarity (top row, all campaigns) and brand similarity (bottom row, campaigns with detection rate≥ 70%). Prompts are sorted by ascending Jaccard. Vertical dashed line at 0.5 for reference. Instead of relying on snapshot metrics, marketers must conceptualize GEO performance as the probability of being mentioned across repeated runs. This shift from deterministic rankings to probabilistic visibility fundamentally changes how marketing performance in AI search should be monitored and managed. Several concrete implications follow: • Minimum run count. A single daily query cannot provide a reliable estimate of true visibility. A bootstrap convergence analysis on the 10-run simultaneous dataset (Appendix J) shows that the standard error of the estimated per-brand detection rate drops below 0.10 atn = 7runs (95% CI±0.158) and below 0.08 atn = 8 runs (95% CI±0.121). For source coverage the convergence is slower: SE< 0.10requiresn = 8runs, reflecting higher source-level stochasticity. Practitioners should therefore use at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters. •Multi-prompt coverage. Prompt-level Jaccard scores vary widely within campaigns (from below 0.2 to above 0.8). Monitoring based on one or two prompts will reflect the idiosyncrasies of those prompts rather than campaign-level visibility. A large prompt portfolio of diverse queries is advisable. • Sustained observation windows. Day-to-day source instability (≈ 65%turnover) means short observation windows, e.g. days to a week, are insufficient to distinguish signal from noise. A per-brand rolling-window convergence analysis on the temporal dataset (Appendix K) shows that the standard error of ad-day per-brand detection rate estimate drops below 0.10 atd = 10days and below 0.05 atd = 24days (95% CI±0.105at d = 21;±0.065atd = 28). Short windows that appear precise at the campaign level are substantially noisier when tracking individual brands. Rolling aggregation over two to four weeks is therefore recommended to obtain per-brand estimates that are both statistically stable and representative of sustained visibility rather than momentary snapshots. •Campaign-specific benchmarking. Citation concentration (Gini≈ 0.71on average) varies meaningfully across both campaigns and engines. Google AI Mode concentrates citations most strongly; Perplexity distributes them most evenly. Marketers should set engine-specific visibility baselines rather than applying a single threshold across all AI search products. •Brand vs. source monitoring. Brand-level day-to-day stability (Jaccard 0.45–0.59) exceeds source-level stability (0.34–0.42), suggesting that brand presence aggregated over a campaign is a more reliable KPI than tracking individual cited URLs. Source-level monitoring remains valuable for understanding which content pieces drive inclusion. 9 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Future work should examine whether the instability patterns documented here hold across other languages and regional markets, and whether targeted GEO interventions (e.g., structured content, authoritative backlink profiles) can shift a brand’s inclusion probability in a measurable and durable way. 10 Don’t Measure Once: Measuring Visibility in AI Search (GEO) References P. Aggarwal, V. Murahari, T. Rajpurohit, A. Kalyan, K. Narasimhan, and A. Deshpande. GEO: Generative Engine Optimization, 2024. URL http://arxiv.org/abs/2311.09735. A. Chatterji, T. Cunningham, D. J. Deming, Z. Hitzig, C. Ong, C. Y. Shan, and K. Wadman. How people use chatgpt, 2025. M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba. Evaluating Large Language Models Trained on Code, 2021. URL https://arxiv.org/abs/2107.03374. M. Chen, X. Wang, K. Chen, and N. Koudas. Generative Engine Optimization: How to Dominate AI Search, 2025. URL http://arxiv.org/abs/2509.08919. T. Economist. How AI is disrupting shopping. 2025. ISSN 0013-0613. URLhttps://w.economist.com/ business/2025/12/08/how-ai-is-disrupting-shopping. B. Efron and R. Tibshirani. An Introduction to the Bootstrap. Chapman and Hall/CRC, 0 edition, 1994. ISBN 978-0- 429-24659-3. doi:10.1201/9780429246593. URL https://w.taylorfrancis.com/books/9781000064988. C. Gini. Measurement of Inequality of Incomes. 31(121):124, 1921. ISSN 00130133. doi:10.2307/2223319. URL https://w.jstor.org/stable/10.2307/2223319?origin=crossref. D. Goodwin. Google’s search market share drops below 90% for first time since 2015, 2025. URLhttps:// searchengineland.com/google-search-market-share-drops-2024-450497. P. Jaccard. Étude comparative de la distribution florale dans une portion des Alpes et des Jura. 37:547–579, 1901. J. Krisinger. Ai visibility: Seo, aeo & geo für digitale sichtbarkeit | deloitte deutschland, 2025. URLhttps: //w.deloitte.com/de/de/services/consulting/perspectives/ai-visibility.html. J. Leskovec, A. Rajaraman, and J. D. Ullman. Finding similar items. In Mining of Massive Datasets, pages 68–122. Cambridge University Press, 2014. G. Lior, E. Habba, S. Levy, A. Caciularu, and G. Stanovsky. ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments, 2025. URL https://arxiv.org/abs/2505.22169. J. M. G. P. d. O. Martins. The Evolution of SEO in the Age of Generative Search Engines. 2025. URLhttp: //hdl.handle.net/10362/190409. M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, and G. Stanovsky.State of What Art?A Call for Multi-Prompt LLM Evaluation.12:933–949,2024.ISSN 2307-387X. doi:10.1162/tacl_a_00681. URLhttps://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00681/ 123885/State-of-What-Art-A-Call-for-Multi-Prompt-LLM. N. Padilla, H. T. Lam, A. Lambrecht, and B. Hollenbeck. The Impact of LLM Adoption on Online User Behavior, 2025. URL https://papers.ssrn.com/abstract=5393256. F. Rejón-Guardia, S. Molinillo, and R. Anaya-Sánchez. Generative engine optimization: How search engines integrate AI-generated content into conventional queries. In Encyclopedia of Artificial Intelligence in Marketing, pages 1–8. Springer, 2025. E. Silliman, J. Boudet, and K. Robinson.Winning in the age of AI search | McKinsey, 2025. URLhttps://w.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/ new-front-door-to-the-internet-winning-in-the-age-of-ai-search?utm_source=chatgpt.com. S. Stürze, M. Hoyer, C. Righetti, and M. Rasztar. Agile Marketing Performance Management. 2022. W. Webber, A. Moffat, and J. Zobel. A similarity measure for indefinite rankings. 28(4):1–38, 2010. ISSN 1046-8188, 1558-2868. doi:10.1145/1852102.1852106. URL https://dl.acm.org/doi/10.1145/1852102.1852106. Y. Wen, N. Zhang, H. Yuan, X. Chen, H. Zhang, and H. Guo. Position: On the Risks of Generative Engine Optimization in the Era of LLMs, 2025. URLhttps://w.techrxiv.org/users/1010717/articles/ 1370817-position-on-the-risks-of-generative-engine-optimization-in-the-era-of-llms? commit=d9d611ded6e17d7b0dc93e855718f984775127d. 11 Don’t Measure Once: Measuring Visibility in AI Search (GEO) A Campaign Names The four study verticals are reflecting German terminology. Table 7 maps the English names used throughout this paper to the original German labels. Table 7: English campaign names used in this paper and corresponding original German labels. English (used in paper)Original German label TelecommunicationsTelekom Real Estate SalesImmobilienverkauf Sporting GoodsSportartikel Consumer ElectronicsElektronik B Dataset — Visibility over Time Table 8: Full data coverage by campaign and engine, Jan 24 – Mar 20, 2026 (collection days with≥ 1result). Finance excluded; all data collected via web-interface scraping. CampaignQueriesDaysChatGPTGeminiGoogle AI ModePerplexity Consumer Electronics84543224343 Real Estate Sales84539264343 Sporting Goods84638234444 Telecommunications84540234342 Note: Gemini had sporadic gaps. Jan 30, 2026 excluded (citation volume≈ 2×daily average).images.openai.comfiltered from all calculations. C Similarity Edge-Case Policy The following policy is applied uniformly to both source and brand similarity calculations throughout the paper. •Both lists empty: the pair is excluded from aggregation (assigned NaN). When both runs return the same empty state — no sources cited or no brands detected — there is agreement, but it carries no information about which items are stably cited. Including such pairs would inflate mean similarity scores, particularly for campaigns or queries with low detection rates. •One list empty, the other non-empty: Jaccard= 0.0, RBO= 0.0(maximum disagreement). One run produced citations or brand mentions; the other did not. This is treated as the most severe form of instability. • Both lists non-empty: standard Jaccard and RBO as defined in Appendix D and E. This policy is especially important for brand similarity: campaigns with low detection rates (notably Real Estate Sales, mean 53.6%) would otherwise exhibit inflated Jaccard values driven by many (empty, empty) run-pairs. The 70% detection-rate threshold for including a campaign in brand similarity analyses further mitigates this issue. D Jaccard Similarity The Jaccard Similarity (also known as the Jaccard Index) measures the similarity between two finite sample sets (Jaccard, 1901). It is defined as the size of the intersection divided by the size of the union of the sample sets. It is widely used in information retrieval and biology to compare the overlap of two unweighted sets. Given two sets A and B, the Jaccard coefficient J(A, B) is defined as (Jaccard, 1901; Leskovec et al., 2014): J(A, B) = |A∩ B| |A∪ B| = |A∩ B| |A| +|B|−|A∩ B| (1) Where: 12 Don’t Measure Once: Measuring Visibility in AI Search (GEO) • 0≤ J(A, B)≤ 1. • If the sets are disjoint (A∩ B =∅), J = 0. • If the sets are identical (A = B), J = 1. E Rank Biased Overlap (RBO) Rank Biased Overlap is a similarity measure for indefinite rankings. Unlike the Jaccard index, RBO is designed for ranked lists rather than sets. It weights items at the top of the list more heavily than those at the bottom and can handle lists of different lengths or lists that are not conjoint (do not contain the same items). RBO calculates similarity based on the overlap at each depthd, weighted by a geometric decay determined by a persistence parameter p. The general definition sums to infinite depth (Webber et al., 2010): RBO(S, T, p) = (1− p) ∞ X d=1 p d−1 · A d (2) Where: • S and T are the two ranked lists. • p is the persistence parameter (0 < p < 1). A higherpindicates a stronger interest in the lower-ranked items (the “tail” of the list). • d is the rank depth. • A d is the agreement (overlap) at depth d, calculated as: A d = |S :d ∩ T :d | d (3) Here, S :d and T :d denote the sets of items present in lists S and T up to rank d. Implementation note. Because both source and brand lists are finite, the infinite sum must be truncated in practice. This study uses the non-extrapolated (minimum-bound) variant, in which the sum is truncated atk = min(|S|,|T|)— the length of the shorter list — and A d = 0 is assumed for all d > k. This gives: RBO min (S, T, p) = (1− p) k X d=1 p d−1 · A d , k = min(|S|,|T|)(4) For identical lists of lengthk, this yields1−p k rather than 1.0, because the geometric series is not summed to infinity. At p = 0.9andk = 5, for instance,RBO min = 1− 0.9 5 ≈ 0.41for perfectly matching lists. This is a conservative lower bound on the true RBO: any unobserved overlap beyond depthkwould only increase the score. The minimum-bound variant is appropriate here because the lists under comparison (AI-generated source citations or brand detections) vary in length across runs, and extrapolating beyond what was actually observed introduces assumptions that are not warranted. Duplicate items are removed from each list before computation to satisfy the package’s requirement for unique elements. While Jaccard is set-based and order-agnostic, RBO is rank-sensitive. Jaccard is appropriate when the presence of an item is the only factor, whereas RBO is appropriate when the position of the item signifies importance (e.g., search engine results). RBO scores are computed using the Python implementation by Changyao Chen. 3 F Brand Lexicon The brand lexicon can be found in Table 9. 3 The respective Python package can be found here. 13 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Table 9: Complete brand lexicon used for brand-mention detection, by campaign. Finance is excluded from all analyses. Real Estate Sales is excluded from brand similarity (det. rate 53.6%<70% threshold) but its lexicon is listed for completeness. CampaignBrands tracked (canonical names) Telecommunications (51 brands) 1&1, ALDI, Alao, Besteabos, CH Mobile, Comparis, Congstar, Coop, Deinabo, Digital Republic, Dschungelkompass, Freshnet, GGA Maur, Galaxus Mobile, Gigamobile, Handyabo-Vergleich, Init7, Jio, Lebara, Leucom, Lidl, Lycamobile, MTEL, Migros, Moneyland, Monzoon, Mucho, Net+, Netplus, Netzwoche, O2, Peoplefone, Post Mobile, Quickline, Sak Digital, Salt, Solnet, Spusu, Sunrise, Swisscom, Swype, TalkTalk, Teleboy, Toppreise, Ubigi, VTX, Vodafone, Wingo, Yallo, gomo, iWay Sporting Goods (43 brands) 21run, ASICS, Adidas, Altra, Berg-Freunde, Birkenstock, Brooks, Bächli, De- cathlon, HOKA, Idealo, Inov-8, Intersport, KURU Footwear, Karhu, Kiprun, La Sportiva, Lauf-bar, Merrell, Mizuno, New Balance, Nike, Norda, OOFOS, Ochsner- sport, On, Puma, Reebok, Runnersworld, Running Point, Runningxpert, Salomon, Saucony, Scott, Shop4runners, Skechers, Sportscheck, The North Face, Topo Ath- letic, Transa, UGG, Under Armour, Vivobarefoot Consumer Electronics (47 brands) AMD, ASUS, Acer, Alienware, Alternate, Amazon, Apple, Back Market, Brack, CHUWI, Conrad, Cyberport, Dell, Digitec, Dynabook, ERAZER, Framework, Fujitsu, Fust, Galaxus, Gigabyte, HP, Honor, Huawei, Idealo, Intel, Interdiscount, LG, Lenovo, Logitech, MSI, Mediamarkt, Medion, Microsoft, Microspot, NVIDIA, Notebookcheck, Panasonic, Preisvergleich, Razer, Samsung, Sony, Steg Electron- ics, Toppreise, Toshiba, VAIO, XMG Real Estate Sales (excl.; det. rate 53.6%; 32 brands) AXA, Acheter-Louer, BEKB, Baloise, Beobachter, Blick, CBRE, Comparis, En- gel & Völkers, Finanztip, Helvetia, Homeday, Homegate, ImmoScout24, Im- moverkauf24, Immowelt, JLL, Livit, Mobiliar, Neho, Newhome, PostFinance, Privera, Properti, RE/MAX, Raiffeisen, Sotheby’s, Swiss Life, UBS, Wincasa, Wüst & Wüst, ZKB Note: Detection is substring-based on lower-cased answer text. Multiple search patterns may map to the same canonical brand name (e.g., m-budget, mbudget→ Migros; aldi mobile→ ALDI). Brand counts per vertical: Real Estate Sales 32, Sporting Goods 43, Consumer Electronics 47, Telecommunications 51. G Campaign Prompts The following tables list all eight prompts per campaign as originally issued to the search engines in German, alongside their English translations. Prompts were derived from high-search-volume Swiss SEO keywords using Google’s “People Also Ask” feature. Table 10: Prompts used for the Sporting Goods campaign. German (original)English (translation) Auf was muss man beim Laufschuhkauf achten?What should you pay attention to when buying running shoes? In welchen Schuhen läuft man wie auf Wolken?In which shoes do you run as if on clouds? Was ist die 80%-Regel beim Laufen?What is the 80% rule in running? Welche Marke sind gute Laufschuhe?Which brand makes good running shoes? Wie finde ich die richtigen Laufschuhe für mich?How do I find the right running shoes for me? Wie teuer muss ein guter Laufschuh sein?How expensive does a good running shoe have to be? Wie viel kosten sehr gute Laufschuhe?How much do very good running shoes cost? Wie viel sollte ein Anfänger für Laufschuhe ausgeben?How much should a beginner spend on running shoes? 14 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Table 11: Prompts used for the Consumer Electronics campaign. German (original)English (translation) Wann ist die beste Zeit, um einen Laptop zu kaufen?When is the best time to buy a laptop? Was ist der Unterschied zwischen einem Notebook und einem Laptop? What is the difference between a notebook and a laptop? Welche Laptops sind zurzeit die besten?Which laptops are currently the best? Welche Marke ist für Laptops am besten geeignet?Which brand is best suited for laptops? Welche Marke ist gut bei Laptops?Which brand is good for laptops? Welcher Laptop wird am meisten gekauft?Which laptop is purchased most often? Wie viel sollte ein guter Laptop kosten?How much should a good laptop cost? Worauf sollte man beim Kauf eines Laptops achten?What should you look for when buying a laptop? Table 12: Prompts used for the Telecommunications campaign. German (original)English (translation) In der Schweiz: Welche Handy-Abos sind weltweit unlimi- tiert? In Switzerland: Which mobile phone plans are unlimited worldwide? In der Schweiz: Welcher Anbieter hat unlimited Datenvolu- men? In Switzerland: Which provider offers unlimited data volume? In der Schweiz: Welcher Internetanbieter ist zurzeit der beste? In Switzerland: Which internet provider is currently the best? Was ist das schnellste Internet in der Schweiz?What is the fastest internet in Switzerland? Welche Anbieter in der Schweiz bieten ein Internetabon- nement ohne Vertrag an? Which providers in Switzerland offer an internet subscription without a contract? Welcher Anbieter hat das beste Internet in der Schweiz?Which provider has the best internet in Switzerland? Welcher Handyanbieter ist der beste in der Schweiz?Which mobile phone provider is the best in Switzerland? Wie bekomme ich unbegrenzte Daten in der Schweiz?How do I get unlimited data in Switzerland? Table 13: Prompts used for the Real Estate Sales campaign. German (original)English (translation) Ist es aktuell sinnvoll, ein Haus zu verkaufen?Is it currently advisable to sell a house? Was muss ich alles bezahlen, wenn ich mein Haus verkaufe?What do I have to pay when I sell my house? Was muss ich beachten, wenn ich mein Haus verkaufen möchte? What do I need to consider if I want to sell my house? Wie berechne ich den Gewinn einer Immobilie?How do I calculate the profit from a property? Wie hoch ist der langfristige Kapitalgewinn beim Verkauf einer Immobilie? How high is the long-term capital gain when selling a property? Wie lange muss ich mein Haus besitzen, um es steuerfrei zu verkaufen? How long do I have to own my house to sell it tax-free? Wie viel Steuer muss ich zahlen, wenn ich mein Haus verkaufe? How much tax do I have to pay when I sell my house? Wie viel ist kapitalertragssteuerfrei?How much is exempt from capital gains tax? H Source Citation Inequality (Gini) The heatmap of Gini coefficients by campaign and engine is shown in Figure 3 in the main text (Section 4). All campaigns and engines show high Gini values (0.63–0.83), confirming that a small number of domains receive the vast majority of citations. The overall mean Gini is 0.715. Google AI Mode exhibits the highest concentration (0.782) and Perplexity the lowest (0.671). Across campaigns, Telecommunications reaches the highest Gini (0.750) and Sporting Goods the lowest (0.680). Tables 14 and 15 provide the numerical breakdown by campaign and engine, respectively; the formula and a worked example follow below. 15 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Table 14: Mean Gini coefficient by campaign (mean across engines), Jan 24 – Mar 20, 2026. CampaignGini Consumer Electronics0.713 Real Estate Sales0.718 Sporting Goods0.680 Telecommunications0.750 Table 15:Mean Gini coefficient by search engine (mean across campaigns), Jan 24 – Mar 20, 2026. images.openai.com CDN artifact excluded (see Section 6). Search EngineGini ChatGPT0.684 Gemini0.723 Google AI Mode0.782 Perplexity0.671 I Gini Coefficient: Formula, Example, and Implementation I.1 Formula The Gini coefficient (Gini, 1921) measures inequality in a distribution:G = 0means all domains receive equal citations; G = 1means one domain receives all citations. For a finite set ofnnon-negative values, sorted in ascending order y 1 ≤ y 2 ≤·≤ y n , it is computed as: G = 2 P n i=1 i· y i n P n i=1 y i − n + 1 n (5) wherey i is the citation count of domainiandiits ascending rank. This is the standard rank-weighted computational form, algebraically equivalent to the classical Lorenz-curve definition (G = 2×area between the Lorenz curve and the line of equality). It requires only one pass through the sorted data. I.2 Worked Example Consider five domains with citation counts [1, 2, 3, 4, 10] (already sorted ascending). 1. Assign ranks i = [1, 2, 3, 4, 5]. 2. Compute weighted sums: P y i = 1 + 2 + 3 + 4 + 10 = 20 P i· y i = 1·1 + 2·2 + 3·3 + 4·4 + 5·10 = 80 3. Apply Equation (5): G = 2× 80 5× 20 − 6 5 = 1.6− 1.2 = 0.4 G = 0.4indicates moderate inequality: the dominant domain (10 citations) accounts for 50% of all citations while the four others share the remaining 50% unevenly. J Convergence Analysis: How Many Runs Are Sufficient? J.1 Motivation and Method The stochastic nature of LLM outputs means that a single query yields only a noisy snapshot of a brand’s true visibility. This is analogous to the pass@k problem in code generation, where Chen et al. (2021) show that estimating the probability of a correct solution requires multiple independent samples. Lior et al. (2025) formalise this for general LLM evaluation via the method of moments, deriving the number of repeated runs needed for reliable evaluation. 16 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Mizrahi et al. (2024) similarly find that single-prompt LLM evaluation is unreliable and recommend multi-prompt, multi-run designs. We apply the same logic to GEO measurement, asking: how many repeated runs are needed to estimate brand-mention probability reliably? The minimum-run-count recommendation in the Conclusion (Section 5) is derived empirically from the 10-run simultaneous dataset. The full dataset contains 128 engine–prompt groups (4 engines×8 prompts×4 campaigns) with all 10 runs available. Brand analysis is restricted to the three qualifying campaigns (96 groups; see Section 4); source coverage analysis uses all 128 groups. We treat the 10-run mean as the best available proxy for the “true” detection probability. Method: subsampling without replacement. For each group and subsample sizen ∈ 1, . . . , 9, we draw 2,000 random subsamples of sizenwithout replacement (Efron and Tibshirani, 1994) and record the mean binary detection indicator for each individual brand (1 if that specific brand was detected in the response, 0 otherwise). Rather than collapsing to a campaign-level “any brand” indicator, we treat each canonical brand as a separate series. Brands that are never detected across all 10 runs are excluded (their SE is trivially zero and not informative). This yields 1,216 per-brand series across the three qualifying campaigns. The standard deviation of the 2,000 subsample means is the estimated SE of ann-run estimate for that brand; we report the mean SE across all 1,259 series. Note that sampling without replacement introduces a finite population correction: atn = 9fromN = 10runs, the SE is mechanically smaller than it would be for truly independent additional runs. The reported SE values therefore represent a lower bound; actual SE from fresh data would be somewhat higher at large n. Source coverage. We measure how well ann-run union of cited domains approximates the 10-run reference union, using Jaccard similarity between the two. We draw 2,000 subsamples per group and report SE of this Jaccard across all 128 groups. J.2 Results Figure 6 shows both curves. For per-brand detection rate (left panel), SE falls below 0.10 atn = 7runs (95% CI±0.158) and below 0.08 atn = 8runs (95% CI±0.121). The curve is steep between 1 and 5 runs, and flattens thereafter: moving from 8 to 9 runs only reduces SE from 0.062 to 0.041. For source coverage (right panel), convergence is similar: SE remains above 0.10 untiln = 8runs (SE = 0.096, 95% CI±0.187), reflecting the higher stochasticity of which specific URLs are cited. Figure 6: Subsampling standard error of the estimated per-brand detection rate (left) and source-coverage Jaccard (right) as a function of the number of independent runs per prompt. Each point is the mean SE across all 1,216 per-brand series (left) or 128 engine–prompt groups (right), using 2,000 subsamples without replacement. The red dashed line marks SE = 0.10. Per-brand monitoring reaches SE < 0.10 at n = 7 runs; source coverage requires n = 8 runs. 17 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Table 16: Subsampling SE and 95% confidence interval half-width for per-brand detection rate estimation by number of runs (averaged across 1,216 per-brand series). Runs (n)SE95% CI (±) 10.3700.724 20.2460.483 30.1880.369 40.1510.296 50.1230.241 60.1010.197 70.0810.158 80.0620.121 90.0410.081 Note: SE = std of 2,000 subsamples without replacement per brand series. “True” rate proxied by the 10-run mean for that brand. 96 engine–prompt groups (4 engines× 8 prompts× 3 qualifying campaigns)× multiple brands per group = 1,216 per-brand series (brands never detected across all 10 runs excluded). SE atn = 9is subject to finite population correction (FPC) and underestimates the SE of truly independent runs. J.3 Interpretation A single run (SE = 0.370) is essentially uninformative: a true per-brand detection rate of 50% could appear anywhere from−22%to+122%in a nominal 95% interval (clipped to [0,1] in practice). Atn = 7runs SE drops to 0.081, giving a 95% CI of±0.158—adequate for detecting large differences (e.g., a brand detected in 80% vs. 20% of runs) but insufficient for fine-grained ranking of brands with similar visibility. Atn = 8runs SE falls to 0.062 (±0.121), and source coverage reaches comparable precision (SE = 0.096). The per-brand framing is more actionable than a campaign-level “any brand” indicator because it surfaces which specific brands are consistently absent and which are reliably cited. These thresholds assume intermediate detection probabilities; for brands that are either always or never cited, fewer runs suffice. This empirical result aligns with the formal framework of Lior et al. (2025), who derive minimum sample requirements for reliable LLM evaluation from first principles. K Temporal Convergence: How Long an Observation Window Is Sufficient? K.1 Motivation and Method The sustained-observation-window recommendation in the Conclusion is derived empirically from the temporal dataset (Jan 24 – Mar 20, 2026). Mirroring Appendix J, we use a per-brand approach: for each canonical brand in the three qualifying campaigns, we build a daily binary series (1 if that brand was detected in the day’s response, 0 otherwise). Brands never detected across the entire observation period are excluded. This yields 1,726 per-brand series (3 qualifying campaigns×4 engines×8 prompts×multiple brands, filtered to those with≥ 1detection), spanning 40–46 days per series. We ask: as the rolling window lengthdincreases, how precisely can a practitioner estimate the underlying per-brand detection probability from a d-day mean? For each series and each possibled-day consecutive window within its observation period, we compute the mean detection rate. The standard error (SE) across all such window means, averaged over all 1,726 series, quantifies estimation uncertainty as a function of window lengthd. This follows the same logic as rolling-window volatility estimation in time-series analysis. K.2 Results Figure 7 shows the mean SE as a function of window lengthd(left panel, linear scale; right panel, log scale). Table 17 lists key thresholds. K.3 Interpretation The per-brand convergence is considerably slower than a campaign-level “any brand” indicator, reflecting the higher variability of individual brand detection rates. SE falls below 0.10 atd = 10days and below 0.05 atd = 24days (95% CI±0.105atd = 21;±0.065atd = 28). A 14-day window still leaves SE at 0.080 (±0.157)—sufficient for 18 Don’t Measure Once: Measuring Visibility in AI Search (GEO) Figure 7: Mean standard error of thed-day rolling window per-brand detection rate (averaged across 1,726 per-brand series) as a function of window lengthd. Left: linear scale; right: log scale. The red dashed line marks SE= 0.05; the purple dashed line marks SE = 0.02. SE falls below 0.05 at d = 24 days and below 0.02 at d = 34 days. Table 17: Rolling-window SE and 95% confidence interval half-width for per-brand detection rate estimation by window length (averaged across 1,726 per-brand series). Window (d days)SE95% CI (±) 10.3220.631 20.2380.467 30.2010.393 50.1600.313 70.1350.264 100.1070.210 140.0800.157 210.0530.105 280.0330.065 Note: SE computed as the standard deviation of all possible d-day window means within each per-brand series. Series: 1,726 per-brand indicators across 3 qualifying campaigns× 4 engines× 8 prompts (brands never detected excluded). Temporal dataset: Jan 24 – Mar 20, 2026. directional monitoring but not for fine-grained brand comparison. This result must be interpreted in the context of AI search dynamics. AI search engines undergo regular algorithmic updates and index refreshes that can shift brand inclusion probabilities substantially over days to weeks. Short windows—even when statistically tight at the campaign level—may produce estimates that are unrepresentative of the longer-run visibility level for specific brands. A two-to-four-week rolling window is recommended because it (a) reduces per-brand SE below 0.05–0.08, within practical precision requirements for brand monitoring, and (b) averages over short-lived fluctuations introduced by minor model updates, thereby providing a more durable and actionable estimate of sustained per-brand visibility. L Code and Data Availability All analysis code and the datasets used in this study are publicly available at: https://github.com/jatlantic/DONT-MEASURE-ONCE-MEASURING-VISIBILITY-IN-AI-SEARCH The repository contains the two analysis scripts (paper_analysis_v9.py,temporal_brand_v9.py), the shared similarity utility (similarity_functions.py), the simultaneous-run dataset (live_20260321_042355.jsonl), the brand lexicon, and the processed temporal data files derived from the Aurora Intelligence export. 19