We use Google Analytics cookies to understand which pages and tools are useful and improve the site. Privacy policy.

GEO testing

How to test whether an AI system can extract your answer

By bumpit Editorial2026-07-267 min read

Use a fixed question set, save complete outputs and citations, then compare the extracted claim with the page's entity, scope, conditions, date, and source.

Written for an editor or SEO lead testing retrieval and answer quality for important public pages.

Key facts

  • A test needs an expected answer before the team sees the model output.
  • Retrieval success and answer faithfulness require separate labels.
  • One prompt run cannot establish stable visibility across products or users.

A useful rule: make each important claim understandable and verifiable without requiring the reader to reconstruct your meaning from the rest of the page.

The direct answer

Choose a representative question and write the correct scoped answer from your source ledger before running the test. Query the named product in a documented state, save the full response and citations, and locate the likely source passage. Compare the output with the page across entity, conclusion, conditions, time, and source support. Repeat on a schedule and report observed behaviour for that product and date rather than a universal AI ranking. The finished work should let a reader or reviewer identify the subject, the intended result, the evidence behind the recommendation, and the next action without reconstructing your reasoning. Keep material conditions in the same passage as the claim they limit. Use the canonical public page as the source of truth, since search engines and answer systems retrieve pages rather than private briefs. Google describes useful, reliable, people-first content and ordinary search eligibility as the foundation for both search results and its AI features. No heading pattern or schema type can compensate for a page that gives a vague answer, hides its evidence, or serves a different intent from its title. Complete the task for one named reader first, then check how the result appears to crawlers and extraction tools.

  • A test needs an expected answer before the team sees the model output.
  • Retrieval success and answer faithfulness require separate labels.
  • One prompt run cannot establish stable visibility across products or users.

Sources: 1, 2, 3

Prepare the page and evidence before editing

Select questions that represent real customer tasks and pages with maintained evidence. Create a reference answer from the page's claims and sources. Record product, account state, location, date, and any search or model mode you can observe. Save a baseline before you change anything: the public URL, response status, canonical, visible title, main heading, opening answer, source links, and the date you checked them. Record the target question in the reader's words and write one sentence describing the decision the page supports. This baseline prevents a common measurement error where several edits ship together and nobody can tell which one improved the result. It also gives editors a compact source ledger. A reviewer can compare each material statement with the cited page, its jurisdiction or product version, and its checked date. If the task affects a generated template, inspect several representative URLs rather than assuming one record proves the template works for every content shape.

  • Freeze the prompt wording for the core test set.
  • Define supported, partially supported, unsupported, and no-citation labels.
  • Save the current page version and source-ledger state.

Sources: 1, 2, 3

Complete the process in five controlled steps

Work through the five steps in order and keep one output from each step. The order protects you from polishing copy while a crawl, canonical, intent, or evidence problem still blocks the page. Each output should be small enough for another person to verify from the public URL. Use plain labels and stable entity names throughout the page. When a changing fact controls the answer, cite the primary source beside that fact and include the relevant date or version. After each step, compare the output with the primary question. Remove any section that serves a different reader decision, and link to a separate guide when the adjacent task deserves its own page. This creates a focused answer instead of a broad page assembled from loosely related keywords.

  • 1. Run the prompt: Submit the exact question and save the full answer, links, and visible citations. Evidence of completion: The observation has enough context for later review.
  • 2. Find the passage: Open the cited page and locate text that could support the answer. Evidence of completion: The audit identifies evidence rather than URL relevance.
  • 3. Compare meaning: Check subject, conclusion, condition, jurisdiction or version, and date. Evidence of completion: The answer receives a documented support label.
  • 4. Inspect usability: Follow the citation and test whether the landing section answers the question and offers the right next action. Evidence of completion: The source works for a reader after the click.
  • 5. Repeat carefully: Run the fixed set at planned intervals and keep all results, including missing citations. Evidence of completion: The log can show variation without hiding unfavourable outputs.

Sources: 1, 2, 3

A worked example

The expected answer says a page must be indexed and snippet-eligible for consideration as a Google AI feature link. A tested answer says indexing guarantees inclusion and cites the correct Google documentation. The source supports eligibility as a prerequisite but does not support a guarantee. The audit labels the result partially supported, identifies the lost selection condition, and checks whether the page wording keeps that condition close to the claim. Retrieval occurred; faithful extraction did not. Treat the example as a model of the reasoning, not as a universal benchmark. The useful part is the chain from question to evidence to action. Preserve the exact entity names, scope, and conditions that a reader would need if an answer engine quoted the passage outside the page. If a number comes from a report, state the reporting window. If a result comes from a test, state the URL type, device or crawler, and date. A compact example earns its space when it helps the reader make the same decision on another page. Remove invented precision, anonymous authority, and conclusions that reach beyond the recorded evidence.

  • A correct URL can accompany an overbroad conclusion.
  • Predefined labels reduce pressure to call every citation a success.
  • The page can improve condition placement even when the answer system caused the overreach.

Sources: 1, 2, 3

Avoid the mistakes that weaken the result

GEO tests lose value when teams cherry-pick good outputs, change prompts between runs, or judge support from shared words instead of meaning. Fix the first mistake that changes eligibility or meaning before editing smaller presentation details. Keep source boundaries visible: one citation should support the nearby claim, while a separate claim should receive its own source. Do not repeat the target phrase to manufacture relevance. Search systems can use titles, headings, visible text, links, structured data, and other signals, so those elements should agree on the subject without copying one sentence across the page. Check the public result after deployment because a correct content record can still produce the wrong page through caching, layout inheritance, JavaScript failure, or a stale build.

  • Writing the expected answer after seeing output biases the audit. Correction: Set the reference answer from primary evidence first.
  • Counting only cited runs conceals retrieval failures. Correction: Store every result from the fixed set.
  • Reporting a position implies stable ranking data the test does not provide. Correction: Report product, prompt, date, citation, and support observations.

Sources: 1, 2, 3

Verify the result and choose the next action

Summarize citation rate, canonical-page rate, full-support rate, missing-condition rate, referral availability, and observed changes by test cycle. Keep small samples visible and avoid confidence claims they cannot support. Use a fixed observation window and compare like with like. Record the query set, country, device, page version, and publication or change date. Search impressions can show discovery and query matching; clicks and useful sessions show whether the result attracted the intended reader. Observed AI citations add a separate retrieval signal, but a citation count does not prove traffic or revenue. Review the cited passage when you can and check whether the answer preserved its subject, scope, conditions, and source. Keep the page stable long enough to collect evidence unless you find a factual error, broken route, security problem, or misleading claim. The next edit should respond to the strongest observed failure instead of a generic scoring recommendation.

  • Have a second reviewer score a sample without seeing the first label.
  • Link each finding to the exact answer, source passage, and page version.
  • Prioritize factual and entity errors before formatting experiments.

Sources: 1, 2, 3

Put it to work

Find the highest-impact fix on your site.

Use a fixed prompt, claim comparison, citation log, and condition-preservation review.

Test answer extraction

Sources

  1. 1.Bing Webmaster Tools: AI PerformanceChecked 2026-07-26
  2. 2.Google Search Central: AI features and your websiteChecked 2026-07-26
  3. 3.Liu et al.: Evaluating verifiability in generative search enginesChecked 2026-07-26
Published 2026-07-26 · Last reviewed 2026-07-26 · Review due 2026-10-27Search systems and content quality