Skip to content
AnswerGeo

Methodology v1.0 · not yet published · change log at the end

How the numbers are made, in the order a sceptic checks them.

7 things a buyer of AI-visibility data can read here in writing. How each engine is collected, which model, from where and in which language. How many runs, what counts as a citation, what changed and when, how to export everything. Numbers we do not own yet are printed as pending, not hidden.

2 Sep 2026 · methodology v1.0

What we measure

Each term below carries its formula, and every metric label on the site links to the one it means.

  • A project is 1 business: a domain, brand aliases and up to 5 competitor domains.

  • A market is 1 country plus 1 language, such as DE · de or CH · fr. Every query in a market is written in that language.

  • A query is 1 question a buyer would ask an assistant, in the market's language, such as "Spedition für Umzug nach Portugal". A query set has 50 queries per market on Starter and above, 10 on the free snapshot.

  • A sample is 1 query × 1 engine × 1 market × 1 run. It stores the full answer text, the cited pages in order, the sub-queries the engine exposed, the brands named with the quoted sentence, the model id and the timestamp.

  • Mention rate

    = samples where the brand appears anywhere in the answer ÷ all samples.

  • Citation rate

    = samples where one of the brand's domains is cited as a source ÷ all samples.

  • Top-1 rate

    = samples where the brand is the first recommendation ÷ all samples.

  • Share of AI voice

    = the brand's mentions ÷ all brand mentions in the market × engine for the period.

  • Sub-query = a search the engine ran on its own while answering, as exposed by the engine. Stored with the count of runs in which it appeared.

  • Sub-queries covered

    = sub-queries where Bing or Google lists one of the brand's pages in its first 20 results ÷ sub-queries looked up on both. A sub-query not looked up on both indexes is left out of the count and shown beside the figure as not checked.

  • Daily core = the queries you mark as core, measured once a day on Growth and above (20 on Growth, 100 on Pro).

  • Panel set = 30 queries per market, fixed for the month, drawn from the query list: your core queries first, then the next by citation count in the last 4 weeks. On a plan without a daily core, the 30 with the highest citation count.

  • API vs panel gap = per engine and market, the mention rate from automated samples minus the mention rate from panel samples for the same panel set. Published monthly.

What the API measures, and what it cannot see

Every request is made from a fixed location in the country, in the market's language. It carries no browsing history, no memory and no signed-in account. That is the reproducible and comparable half of what a buyer sees.

A buyer with 3 years of chat history sees something else. No tool on this market can measure that, and anyone who claims otherwise is measuring an API too.

What we do about it: the model id and the run number sit on every cell. Paid plans run every query at least twice. The location, model and parameters per engine are printed in Table M1.

You can read the rest of this page, or measure your own domain first: 10 queries, 1 country, no account.

2 Sep 2026 · methodology v1.0

Collection method per engine

2 layers: the engine APIs and a SERP provider for Google, on every plan; a human panel in the real interfaces from Growth. The gap between them is printed, not smoothed.

2 layers. The automated layer uses each engine's official API where one exists and a SERP provider for Google's AI surfaces; it runs every query in every plan, and counts 5 engines, with Google AI Overviews and AI Mode as 1 engine with 2 surfaces. The human panel runs the panel set in the real, logged-in interfaces from inside the country, on Growth and above. Nothing on this page automates a consumer chat interface.

updated 2 Sep 2026
  • ChatGPT

    Capture method
    API with user_location with counsel
    Execution origin
    market country, market language
    Runs per query per week
    1 (Starter) / 2 (Growth, Pro)
    Model version
    first and last seen id stored per sample
    Language handling
    prompt written in the market language, sent unchanged
    Sub-queries exposed
    yes
    Known limits
    –
  • Gemini

    Capture method
    API with Google Search grounding, location passed as prompt context with counsel
    Execution origin
    validated against the panel
    Runs per query per week
    1 / 2
    Model version
    stored per sample
    Language handling
    sent unchanged
    Sub-queries exposed
    yes
    Known limits
    cited links are redirects; the source domain is read from the chunk title and the page URL is not available
  • Perplexity

    Capture method
    API with user_location with counsel
    Execution origin
    market country, market language
    Runs per query per week
    1 / 2
    Model version
    stored per sample
    Language handling
    sent unchanged
    Sub-queries exposed
    partial
    Known limits
    internal sub-queries are not exposed
  • Claude

    Capture method
    API with user_location with counsel
    Execution origin
    market country, market language
    Runs per query per week
    1 / 2
    Model version
    stored per sample
    Language handling
    sent unchanged
    Sub-queries exposed
    yes
    Known limits
    –
  • Google AI Overviews and AI Mode

    Capture method
    SERP provider with the market's location code with counsel
    Execution origin
    reflects Google's output at fetch time
    Runs per query per week
    1 / 2
    Model version
    not applicable
    Language handling
    not applicable; rank data instead
    Sub-queries exposed
    no
    Known limits
    AI Mode is available in fewer markets than AI Overviews
  • Copilot

    Capture method
    human panel only with counsel
    Execution origin
    real people in the country
    Runs per query per week
    – (Starter) / weekly (Growth, Pro)
    Model version
    not applicable
    Language handling
    the panelist types the query
    Sub-queries exposed
    no
    Known limits
    not in the automated layer
Table M1 · Collection per engine · methodology v1.0

The human panel

Real people who live in the country, recruited task by task, working on AnswerGeo-issued accounts so that no screenshot holds a panelist's own data. Personalisation and memory are off. Panelists give separate consent for running the queries, storing their country and language, and being contacted for the next round; each consent can be withdrawn from every e-mail. The consent wording is with counsel.

The number of people per market and the date of the last round are printed in Table M2; the first round has not run.

Weekly on Growth and Pro, on the same weekday and at the same local hour. The Free snapshot and Starter have no panel and use the published gap instead.

The panel set of the market, 30 queries fixed for the month, typed into the logged-in interfaces of ChatGPT, Gemini, Perplexity, Claude and Copilot, and Google Search for AI Overviews and AI Mode.

The public share link of each answer and a screenshot, kept 12 months. Panelists are paid at or above the floor rate of the recruitment service.

No date is set for the first round.

API vs panel gap

pending · not published yet
Not published yet; the table structure is final.
EngineMarketPanel queriesAPI mention ratePanel mention rateGap (pts)Panel sizeMonth
ChatGPTDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
GeminiDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
PerplexityDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
ClaudeDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
Google AI OverviewsDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
Google AI ModeDE30not measured yetnot measured yetnot measured yetnot measured yetnot measured yet
Not published yet; the table structure is final.

This table never prints a gap without its panel size and month.

External context, not our number: a University of Hamburg and Leibniz HBI study of 24,013 ChatGPT responses over 5 weeks in 2025 found that the web interface's sources overlapped a reference list of German news outlets 45.5 % of the time against 27.3 % for the API. The gap moves by surface, engine and month, which is why ours is dated, carries n, and is never smoothed across months.

2 Sep 2026 · methodology v1.0

Model version

Every sample stores the model id the engine returned. A model change restarts the baseline and is labelled on every chart that crosses it.

Every sample stores the model id the engine returned, unchanged, such as gpt-5-mini. When an engine's default model changes, the baselines for that engine are recomputed from that date, every comparison that crosses the change is labelled, the change appears as a numbered footnote on the trend chart and as an entry in the change log below, and the weekly digest names it. The first and last seen id per engine are printed in Table M1.

2 Sep 2026 · methodology v1.0

Where and in which language each query runs

Each query runs with the engine's location set to the market and is written in the market's language. Sets are generated per market, never translated.

On ChatGPT, Perplexity and Claude, each query runs with the engine's location parameter set to the market's country. The queries are generated by a language model prompted in the market's language, then reviewed and edited by you before the first run; a paid native-speaker review is planned as an option. A query is sent as written: its text is never changed. Beside it, each engine that runs a model is given the market in its instruction, which says where the user is located. Gemini has no location parameter in the grounded API, so that instruction is where its location is passed, with the country's capital and a request to prefer sources available there, and its rows are checked against the panel. Google AI Overviews and AI Mode are fetched with the market's location code through a SERP provider. Queries are never translated between markets: a Spanish set is generated for Spain and a separate one for Mexico.

Engines often research in English even for a German query; the sub-queries you see include those English searches, which is one reason a page in the market language alone may not be enough.

Coverage

Countries with ≥ 1 engine · 39 · All 5 automated engines · 39 · AI Mode available · 39

updated 2 Sep 2026
Which engines answer in your countries.
MarketChatGPTAPIGeminiAPI · contextPerplexityAPIClaudeAPIGoogle AI (1 engine, 2 surfaces)Microsoft Copilotpanel
Google AI OverviewsSERPGoogle AI ModeSERP
AT Austriaavailableavailableavailableavailableavailableavailableavailable
AU Australiaavailableavailableavailableavailableavailableavailableavailable
BE Belgiumavailableavailableavailableavailableavailableavailableavailable
BG Bulgariaavailableavailableavailableavailableavailableavailableavailable
BR Brazilavailableavailableavailableavailableavailableavailableavailable
CA Canadaavailableavailableavailableavailableavailableavailableavailable
CH Switzerlandavailableavailableavailableavailableavailableavailableavailable
CZ Czechiaavailableavailableavailableavailableavailableavailableavailable
DE Germanyavailableavailableavailableavailableavailableavailableavailable
DK Denmarkavailableavailableavailableavailableavailableavailableavailable
EE Estoniaavailableavailableavailableavailableavailableavailableavailable
ES Spainavailableavailableavailableavailableavailableavailableavailable
FI Finlandavailableavailableavailableavailableavailableavailableavailable
FR Franceavailableavailableavailableavailableavailableavailableavailable
GB United Kingdomavailableavailableavailableavailableavailableavailableavailable
GR Greeceavailableavailableavailableavailableavailableavailableavailable
HR Croatiaavailableavailableavailableavailableavailableavailableavailable
HU Hungaryavailableavailableavailableavailableavailableavailableavailable
IE Irelandavailableavailableavailableavailableavailableavailableavailable
IN Indiaavailableavailableavailableavailableavailableavailableavailable
IT Italyavailableavailableavailableavailableavailableavailableavailable
JP Japanavailableavailableavailableavailableavailableavailableavailable
KR South Koreaavailableavailableavailableavailableavailableavailableavailable
LT Lithuaniaavailableavailableavailableavailableavailableavailableavailable
LV Latviaavailableavailableavailableavailableavailableavailableavailable
MX Mexicoavailableavailableavailableavailableavailableavailableavailable
NL Netherlandsavailableavailableavailableavailableavailableavailableavailable
NO Norwayavailableavailableavailableavailableavailableavailableavailable
NZ New Zealandavailableavailableavailableavailableavailableavailableavailable
PL Polandavailableavailableavailableavailableavailableavailableavailable
PT Portugalavailableavailableavailableavailableavailableavailableavailable
RO Romaniaavailableavailableavailableavailableavailableavailableavailable
SE Swedenavailableavailableavailableavailableavailableavailableavailable
SI Sloveniaavailableavailableavailableavailableavailableavailableavailable
SK Slovakiaavailableavailableavailableavailableavailableavailableavailable
TR Türkiyeavailableavailableavailableavailableavailableavailableavailable
UA Ukraineavailableavailableavailableavailableavailableavailableavailable
US United Statesavailableavailableavailableavailableavailableavailableavailable
ZA South Africaavailableavailableavailableavailableavailableavailableavailable
  • available
  • partial
  • not available

Every country and every exception, on the methodology page →

Any country where the engines officially answer. The exceptions are listed, not hidden.

2 Sep 2026 · methodology v1.0

Runs per query, and how the rates are computed

1 run is 1 query on 1 engine in 1 market, once. Rates are computed over the window's runs, and the run count is printed beside every rate.

1 run is 1 query × 1 engine × 1 market, once. Starter runs every query once a week per engine; Growth and Pro run every query twice a week per engine and the daily core once a day (20 queries on Growth, 100 on Pro). Rates are computed over all runs in the window shown; the run count is printed beside every rate.

How noisy the engines are

pending · not published yet
Table M3 · Run-to-run variation per engine · AnswerGeo's own 14-day replicate study · DE
EngineShare of cited domains kept between 2 runs on the same dayStandard deviation of the weekly mention rate at 1 run per weekat 2 runs per weekat daily runs
ChatGPTpendingpendingpendingpending
Geminipendingpendingpendingpending
Perplexitypendingpendingpendingpending
Claudependingpendingpendingpending
Google AI Overviewspendingpendingpendingpending
Google AI Modependingpendingpendingpending
Table M3 · Run-to-run variation per engine · AnswerGeo's own 14-day replicate study · DE

Around every rate we print a 95 % Wilson interval. It covers run-to-run randomness only: not model updates, not personalisation, not queries we do not track. On Starter (150 answers a week: 50 queries × 3 engines × 1 run) a country score carries about ±8 points; on Growth (500 answers) about ±4; a single query carries about ±35, so we never call 1 query up or down on 1 week of data.

The country bands assume runs of different queries are independent. If outcomes cluster by query, extra runs narrow the country band less than the arithmetic suggests; we estimate that effect from the replicate study and print it here; the pilot has not run, so there is no figure yet.

The randomness comes from the engines themselves: a large language model can return different answers to the same prompt even at temperature 0, because the arithmetic inside the model depends on how many other requests share the batch (Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference", 10 Sep 2025 ↗). That is why counts of runs are printed, never 1 run as a percentage.

When we say "up" or "down"

A weekly verdict says Up or Down only when the change between 2 windows exceeds the detection threshold computed from both run counts; otherwise it says Broadly stable and prints the threshold. Example (the figures are illustrative): "Broadly stable: −3 pts, within this week's detection threshold of ±6.5 pts." A single query's verdict is computed on a 4-week window.

2 Sep 2026 · methodology v1.0

What we count, and the raw answer behind every count

Mention, citation, cited first, competitor only: each has 1 definition, and the raw answer behind every count is 1 click away.

  • Mention: the brand name or one of its aliases appears in the answer text. Aliases are the ones you set in the project; matching is case-insensitive and whole-word.
  • Citation: one of the brand's domains appears in the engine's source list for that answer. Subdomains count; a mention of the domain inside the text without a source entry does not.
  • Cited first: the brand's domain is in position 1 of the source list.
  • Competitor only: a named competitor domain is cited and the brand is absent.
  • Sub-query: a search the engine exposed as run during the answer. Counted once per run; the chip shows the number of runs in which it appeared.
  • Sentiment: a word (positive, neutral, negative) assigned by a model to the sentence in which the brand appears, always shown beside the sentence and the one-line reason. Sentiment never takes a colour; you can overrule it per sample.
  • Open a raw stored answer from this week's showcase run ↗

How the Playbook score is computed

Score = severity × impact × ease. Severity: what the finding blocks (an AI bot blocked in robots.txt is critical; a missing page for a sub-query is high; a missing mention on a third-party page is medium). Impact: the number of queries × engines the finding touches. Ease: our estimate of the work, S, M or L. A finding closes when the next run no longer shows it; a finding you mark done gets a before → after table captioned with the weeks of data, and a verdict of Verified, Not yet or Regressed.

2 Sep 2026 · methodology v1.0

What changed, and when

3 logs: every run, every change to this method, every model change. The measurement log lists runs, not weeks.

3 logs. The measurement log lists every weekly run with its date, engines, markets, query count and anomalies. The change log lists every change to this method, dated and versioned; entries tagged measurement also appear on the public changelog. Model changes appear in both, and as footnotes on every trend chart that crosses them.

Measurement log

DE · deanswergeo.com
run #1284 · 5 Oct 2026
DateEnginesMarketsQueriesAnomalies
Mon 5 Oct 20265DE50No anomalies
Thu 1 Oct 20265DE50No anomalies
Every weekly run of AnswerGeo's own workspace, newest first, with its engines, market, query count and anomalies. · our own index, live · answergeo.com · DE · de · run #1284 · 5 Oct 2026 — not a sample · Open the raw answers ↗

2 Sep 2026 · methodology v1.0

Export everything

CSV on every paid plan, raw JSON from Growth, a REST API on Pro. The same fields as the console, from the same run rows.

Every paid plan exports CSV of the query × engine matrix with states, counts and deltas. Growth and above export the raw JSON of every sample: answer text, cited pages in order, sub-queries, brands with the quoted sentence and sentiment word, model id, timestamp, run id. Pro adds a REST API with the same fields. Public read-only report links and the PDF carry the same numbers as the console, from the same run rows.

What we cannot tell you

No assistant publishes prompt volumes or citation logs, so we show sampled frequency with its 95 % band, never a rank. 1 run is a count, not a rate; a change needs 2 runs to be judged. A change and a gain 2 runs later is correlation, not proof: engines change their models without notice, and we label the date and restart the baseline. Whether anyone clicked through is not in the answer, so it is not on this page.

What we refuse to do

  • Automate a consumer chat interface with a script
  • Rewrite a query's text to add a country name and fake a location
  • Run logged-in account farms or proxies
  • Print a single score to a decimal without a run count
  • Compare across a model change without a label
  • Attach the word "illegal" or any legal judgement to an answer
  • Restate another vendor's variance figures as our own

What AnswerGeo does and does not do

  • Measures who is cited, per query, engine and country, every week.Does not · Promise a citation or a ranking.
  • Generates queries in the market language; you edit them or import from Search Console.Does not · Write your content. The Playbook names the page and the sub-query; the writing is yours.
  • Uses engine APIs with a location, a SERP provider for Google's surfaces, and measures the gap to a human panel per market.Does not · Automate consumer chat interfaces. Copilot is measured by the panel only.
  • Reports rates with run counts and a 95 % range.Does not · Print a single score to 1 decimal place.
  • Tells you the date when an engine changes its model.Does not · Compare across a model change without labelling it.

Until the weekly worker writes its first run, every figure on this site marked "sample data" is drawn from a checked-in sample: the road-freight ledger for Germany and AnswerGeo's own workspace. The numbers are real in shape and fictional in value, and they are replaced, with their dates, the week the worker runs.

Written by Vitalii Kenyiz. Questions to ceo@answergeo.com.

Methodology v1.0 · not yet published · change log at the end

10 queries, 1 country, no account.

The first engine answers within minutes. No account until you save the result. Then 50 queries a week, from €99.

queries in English

10 queries · 1 country · 3 of the 5 engines · no account, no card · first answers within minutes

Already convinced? Start 14 days of Growth, no card →