← Back to Lab

Evidence Standard

A citation-lift claim without a denominator is marketing. Here's the evidence card every claim on this site must carry — and you should demand from anyone selling AEO.

Why anonymized case studies fail

"We increased AI citations by 47% for a B2B SaaS client in 30 days."

You've seen claims like this. They sound impressive. But what do they actually mean?

  • 47% of what? One intent? Five? A hundred? Did they start from zero, making the percentage meaningless?
  • Which assistant? ChatGPT? Perplexity? A specialized model? All three?
  • What else changed? Did they launch a major feature? Get press coverage? Run paid ads?
  • Can you verify it? Where are the raw probe results? The dated snapshots? The intervention log?

The truth is: you can't. Anonymized case studies are unfalsifiable. They might be real. They might be exaggerated. They might be entirely fabricated. You have no way to know.

This is a problem for buyers — and it's a problem for the entire AEO category. If every vendor's proof is unfalsifiable, the category becomes indistinguishable from snake oil.

The Evidence Card

Every outcome claim on this site — in the Lab, in reports, in case studies — must carry an evidence card with these seven components:

1. Numerator & Denominator

Not "citations increased 32%." Instead: "8 of 25 intents cited us (32%), up from 5 of 25 (20%)." The denominator matters. A 100% lift from 1 to 2 intents is not the same as 10 to 20.

2. Date Window

When was the baseline? When was the measurement? How long between them? "Baseline: July 1, 2026. Measurement: July 30, 2026. 30-day window."

3. Intervention

What specifically changed? Not "we optimized the site." Instead: "We added FAQ schema to 5 pages, published 3 answer-capsule rewrites, and submitted an llms.txt file."

4. Unchanged Variables

What stayed constant? "No press coverage, no paid ads, no new backlinks, no product launches, no pricing changes during the window." If something major changed, the lift might not be from your work.

5. Raw Artifacts

Where can someone verify this? Link the dated probe results, the git commits, the before/after snapshots. If you can't show receipts, don't make the claim.

6. Limitations

What can't this prove? "Single-sample snapshots, not statistical aggregates. Can't isolate this intervention from organic crawl timing. Correlation, not causation."

7. Control Group (when possible)

Did you change half your pages and leave half unchanged? Show the delta for both groups. "Changed pages: 3/10 → 7/10. Unchanged: 2/10 → 2/10." Control groups aren't always possible, but when they are, they're the gold standard.

Why seven components? Because this is what it takes to make a claim falsifiable. If any of these are missing, the claim becomes marketing instead of evidence.

Example: How we'd card a real outcome

Here's what an evidence card looks like for a real intervention on one of our properties:

Example Card

FAQ Schema Addition — File Conversion Tool

Claim

Adding FAQ schema to 8 format-specific pages increased citations for how-to intents from 2/8 to 5/8 in 30 days.

Numerator & Denominator

Baseline: 2/8 intents cited (25%). After: 5/8 (62.5%). 8 intents total, all "how to convert X to Y" queries.

Date Window

Baseline: June 15, 2026. Schema deployed: June 16, 2026. Measurement: July 15, 2026. 30-day window.

Intervention

Added FAQ schema blocks to 8 pages (PDF↔Word, HEIC→JPG, etc.). Each block contained 3-5 Q&A pairs extracted from existing page copy. No new content written.

Unchanged Variables

No backlinks added, no press, no paid ads, no product changes, no new blog posts. Traffic stayed flat (±5%).

Raw Artifacts

Probe results: ~/logs/probes/2026-06-15.jsonl (baseline), ~/logs/probes/2026-07-15.jsonl (after). Git commit: a3f8b21. Schema diffs available in repo.

Limitations

Single-sample snapshots (one probe per date). Can't isolate schema change from organic crawler timing. No control group (all 8 pages changed simultaneously). Correlation, not proven causation.

Why this is falsifiable: You could request the raw probe JSONs, replay the queries yourself today, check the git commit, and verify whether the schema blocks actually exist. The claim survives scrutiny because it's grounded in artifacts.

How to audit an AEO vendor's claims

Before you sign a contract with any AEO vendor (including us), ask for evidence cards on their case studies. If they can't provide them, the claims are unfalsifiable.

Ask: "Can I see the raw probe results?"

If they show you a dashboard with a line going up, ask for the underlying data. If they can't show you the actual assistant responses (with timestamps, model IDs, and source lists), the chart might be invented.

Ask: "What's the denominator?"

"We increased citations 300%" is meaningless without knowing the baseline. 1 to 4 is 300%. So is 10 to 40. The latter is a real signal. The former might be noise.

Ask: "What else changed during the window?"

If the client launched a viral feature, got TechCrunch coverage, or ran a major ad campaign during the measurement window, the lift might have nothing to do with the vendor's work.

Ask: "Can you show me the intervention log?"

What specifically did they change? If they can't show you git commits, CMS edit logs, or dated snapshots, you can't verify the work actually happened.

Ask: "What are the limitations?"

If they claim perfect results with no caveats, they're either lying or don't understand their own data. Real evidence comes with honest limitations.

Our commitment

Every outcome claim on this site — past, present, and future — will carry an evidence card or it won't be published. When we can't card it (because raw data doesn't exist or can't be shared), we'll say "we can't prove this" instead of making the claim.

This applies to:

  • The Agent Run Ledger (weekly probe results + intervention log)
  • Any "intervention reports" we publish from our portfolio
  • Design-partner case studies (when Autopilot launches)
  • Research reports about AI citation behavior

If you catch us making an unfalsifiable claim, email [email protected] and we'll either provide the evidence card or retract the claim.

Every $79 audit ships with evidence, not adjectives

Your deep audit includes dated probe results, raw source lists, and the exact intervention log Autopilot would execute. No vibes. No percentages without denominators.

Bot Traffic by AttractOS