Method
One run is an anecdote. Keoph runs experiments.
What the published evidence says about AI citations, how noisy the measurement is, and how Keoph designs, measures and reports an experiment. Every figure comes from a source listed at the end. Where a result is unreviewed or only correlational, we say so.
01
The measurement is noisy
Run the same prompt twice on the same day and the cited sources barely overlap. Schulte, Bleeker and Kaufmann measured the Jaccard similarity of cited sources between two same-day runs:
About 65% of cited sources change from one day to the next, and the authors put the monitoring floor at 7–8 runs per prompt.
Wording matters as much as timing. SparkToro collected 142 human phrasings of a single intent; their average similarity was 0.081, and the same list of brands came back in under 1% of runs.
Organic rank is not a stand-in either. Ahrefs found that 8% of URLs cited by ChatGPT and 28.6% of those cited by Perplexity rank in Google’s top 10.
A score built from one run per prompt moves with this noise whether or not anything on your site changed.
02
What the evidence says moves citations
Much less than the optimisation market claims, and not consistently across engines.
- The “up to 40%” figure
- comes from the original GEO paper: a within-context share shift on GPT-3.5 with five fixed sources, and a Perplexity test in which sources were uploaded as files. A 2026 re-test across ten current engine families found the same levers moved citation on none of them (not yet peer reviewed).
- Content scores barely predict citation
- A query-blind content score correlates at ρ ≈ 0.11 with real citation; conditioning on the query lifts that to about 0.4 (same unreviewed study).
- C-SEO Bench, NeurIPS 2025
- 3 of 54 method-and-setting cells were positive, none in question answering. “Add statistics” hurt in 19 of 24 settings. Where a page sits in the model’s context beat every rewrite, and gains fell to zero once competitors used the same method.
- SAGEO Arena, KDD 2026
- Rewriting only the body cut retrieval by 9% and citation by 6%; adding structural fields raised retrieval by 22%. Retrieval and citation are separate stages, and a change can help one and hurt the other.
- E-GEO
- 10 of 15 heuristics were neutral or negative. Optimised pages converged on summary, headings, bullets and FAQ for a gain of about 0.3 rank positions; manipulation got flagged, and preferences differed by engine.
- The weak levers that remain
- Extractable evidence (correlational only) and visible freshness dates (causal in rerankers). There is no causal evidence for schema.org markup, and Q&A formatting on its own correlated negatively. Engines favour earned media over brand-owned pages.
There is no recipe to buy. The useful question is whether a specific change, on your pages, moved your citations. That takes an experiment.
03
How a Keoph experiment works
-
Randomize groups of pages
You choose the change: a doc template, a freshness date, a restructured reference page. We group comparable docs and pages and assign groups at random to receive the change or hold as they are, so the comparison isn’t confounded by which pages you’d have picked.
-
Fix the prompts before anything ships
The prompt set covers the intents those pages answer, with several phrasings per intent, because paraphrase moves results as much as time does. It is set before the change goes live.
-
Measure every stage, from the engines’ own reports
- Google Search Console, generative-AI report: AI impressions per page, country, date and device. It has no queries or clicks, and exports only from the UI.
- Bing Webmaster Tools, AI Performance: per-page Copilot citations and the grounding queries behind them, as sampled by Bing.
- Repeated sampling: Google AI Mode through SERP APIs, the most faithful sampled surface, run at or above the 7–8 runs per prompt floor. Other engines’ APIs are reported as API results, not as the consumer app: OpenAI’s API search shares 25.6% of its sources with the ChatGPT app, and returns sources on 26% of prompts against 84% in the app.
- Organic guardrail: the same pages’ organic search performance, so a change that wins AI citations but costs organic traffic is caught. It happens: SearchPilot has published a change rejected for −6.5% organic.
-
Run long enough to count
The effective sample is prompts × days, not API calls. Detecting a 10-point lift at 80% power takes roughly 350 effective observations per arm. Citation also lags publication, and many new pages are never cited at all, so an experiment runs for weeks rather than days.
-
Report an effect with an interval
Each outcome gets an estimate and an interval, and one of three verdicts. Inconclusive is a result: it says the effect is smaller than this experiment could detect.
Lift Inconclusive Harm no effectIllustration, not data. Square: estimate. Bar: interval.
04
What we don’t claim
- AI-driven clicks and conversions. Google doesn’t separate AI clicks from organic clicks, and GA4’s “AI Assistant” channel misses visits that arrive without a referrer.
- The consumer ChatGPT app. It can only be sampled by scraping against OpenAI’s terms, so it is out of scope.
- A “citability” score. No published study links LLM-judged citability to measured citation. Keoph’s outcome is measured citation.
- One engine standing in for another. Preferences are engine-specific, so a result on AI Mode is reported for AI Mode.
05
Sources
- Schulte, Bleeker, Kaufmann, “Don’t Measure Once”, arXiv 2604.07585, April 2026. Same-day overlap, daily churn, run floor.
- arXiv 2609.07559, 2026, not peer reviewed. Content-score correlation and the re-test across ten engine families.
- Aggarwal et al., “GEO: Generative Engine Optimization”, arXiv 2311.09735. Origin of the “up to 40%” figure.
- C-SEO Bench, NeurIPS 2025.
- SAGEO Arena, KDD 2026.
- E-GEO, arXiv 2511.20867.
- SparkToro, consistency of AI brand recommendations across runs and phrasings.
- Ahrefs, overlap between AI-cited URLs and Google’s top 10.
- SearchPilot, published page-level GEO tests with an organic guardrail.
- Google Search Console generative-AI performance report (global from 2026-08-31); Bing Webmaster Tools AI Performance (2026-02-10).