Methodology · transparent by design

How we score AI visibility.

Every score AEO Owl shows you is computed from a published formula. No hidden weighting, no LLM-generated numbers, no industry benchmarks fabricated from a sample too small to matter. This page documents exactly how our two top-line scores are calculated, what each component means, and the things we deliberately don't do.

Methodology v2.0.1
Last updated 2026-06-30
Status Observed-denominator scoring — refined as customer data accumulates

Why two scores instead of one?

AEO Owl reports two top-line scores per audit. They measure different things and are never averaged into a single composite. Forcing two different questions into one number would obscure the diagnosis a customer most needs — which side of the equation to invest in.

Outcome
AI Visibility Score
Are AI engines actually citing and recommending your brand right now? This is the headline metric — the number your CMO references in a board update. Built from what we observe across 8 engines × your audit questions × 3 repetitions.
0 – 10 · hero metric
Foundation
AEO Readiness Score
How prepared is your website and brand to earn AI citations? This is the diagnostic companion. It tells you why your visibility is where it is — and where to invest if you want it to grow.
0 – 10 · diagnostic
The leading/lagging story. Readiness changes faster than Visibility. When you fix Technical structure or add Authority signals, AI engines need time to re-crawl, re-index, and update their training and retrieval. Expect Readiness to move first; Visibility tends to follow over 1–3 audit cycles. The Trend Report shows both as separate lines so you can see this story play out in your own data.

How is the AI Visibility Score calculated?

Computed from observed AI engine behavior on your audit. Four components contribute additively to a 0–10 scale. Each component is bounded; the total is bounded by construction.

AI Visibility = 4.0 × Citation Rate → max 4.0 + 3.0 × Share of Voice → max 3.0 + 2.0 × Engine Coverage → max 2.0 + 1.0 × Position Score → max 1.0 = max 10.0
ComponentWhat it measuresWeight
Citation RateMost direct signal — are you in the conversation at all? % of the observed (question × engine) pairs where your brand was cited at least once across 3 repetitions. Pairs an engine couldn't answer (timeout / outage / a retired model) are excluded — never scored zero — and a Google query where no AI Overview appeared counts as observed-but-not-cited. 40%
Share of VoiceCompetitive position — how much of the conversation is yours? % of total brand mentions (you + the competitors you specified) that are your brand. 30%
Engine CoverageResilience — how many engines know you exist at all? % of the engines we observed this audit (typically 8) that cited your brand in at least one response. An engine that was fully unavailable for the whole run drops out of the denominator rather than counting against you. 20%
Position ScoreProminence — when cited, how early do you appear? Average position of your brand mention within citing responses, normalized to 0–1. Lower (closer to 1st) is better. 10%
Why Citation Rate weighs heaviest. The single most direct signal that an AI engine knows your brand exists. Higher Citation Rate is the foundation that every other component builds on.

Why Position Score weighs lightest. Position measurement quality varies by engine type — search-integrated AI (Google AIO) has clear position data, but conversational engines (ChatGPT, Claude, Gemini chat) approximate from order of appearance. We weight this component modestly to acknowledge the measurement uncertainty.

How is the AEO Readiness Score calculated?

Computed from a deterministic crawl of your website + the citation source data from your audit. Three pillars contribute to a 0–10 scale, each itself broken down into three sub-components.

AEO Readiness = 4.0 × Authority Signals (pillar / 10) → max 4.0 + 3.5 × Content Clarity (pillar / 10) → max 3.5 + 2.5 × Technical Structure (pillar / 10) → max 2.5 = max 10.0
PillarSub-componentsWeight
Authority SignalsTrust signals AI engines reward most Source category diversity (review/social/news/competitor) + third-party citation ratio + category-leader citations (G2, Capterra, TrustRadius, Gartner, Forrester). 40%
Content ClarityRetrievable, AI-friendly content structure Question-shaped headings rate + direct-answer rate (short paragraphs after questions) + Q+A structured data (FAQPage / HowTo / Article schema) presence. Measured across up to 10 sampled pages. 35%
Technical StructureNecessary baseline for crawl access AI crawler access (GPTBot, ClaudeBot, PerplexityBot, etc. allowed in robots.txt) + schema markup presence (Organization, WebSite) + render method (SSR, hybrid, or CSR). 25%
Why Authority weights highest. In observed AI engine behavior across 2024–2026, the single strongest predictor of being cited is third-party recognition. AI engines actively prefer trusted brands and citation-worthy sources. Technical perfection without authority earns far fewer citations than moderate technical with strong authority. We weight accordingly. Authority is scored only from sources that actually cited your brand, and it ramps with citation volume so a single mention can't overstate it.

How are recommendations and projected impact calculated?

Every recommendation in your action plan ties to a specific audit finding and projects a measurable lift on one of the components above. Here's how those projections are computed — and the rules they must follow.

The LLM writes the words. The numbers come from a calibrated library. AEO Owl uses an AI model to draft the diagnosis and prescription text — that's a language task. But the projected impact number is NEVER generated by the LLM. It comes from a versioned tactic library where every action has a calibrated impact range (minimum, expected, maximum) per pillar. The LLM picks the tactic; the library provides the math.
RuleWhat it enforces
Headroom bounded A recommendation cannot project lift greater than your pillar has room to grow. If you're already at 2.4/2.5 Technical, a recommendation targeting Technical clamps to your remaining 0.1 headroom — not the library's full range.
Pillar-points, not score-points Impact is shown in the unit it actually affects (Technical pillar-points, Citation Rate component-points) AND translated to the headline-score impact alongside. The math is visible.
Secondary capped at 50% of primary A tactic that affects two pillars declares one as primary; the secondary impact range is capped at 50% of the primary range. Prevents recommendations from inflating their projection by spreading across pillars.
No additive sum The action plan never sums per-action projections into a single "projected total." Multiple actions targeting the same pillar overlap; we don't pretend they don't.
Category fit modifier v1.1 The same tactic doesn't move the score the same amount for every brand. "Claim your G2 profile" lifts a B2B SaaS brand meaningfully but a restaurant negligibly. Each tactic in the library declares per-category multipliers (bounded between ×0.2 and ×1.6); the worker applies the modifier before headroom clamping. When the modifier moves the lift up or down materially, a small "Strong fit" or "Lower fit" badge is shown on the recommendation card.

03b · The eight canonical categories

Category modifiers reference one of eight canonical brand categories. Your brand's freeform category is normalized to whichever of these is the closest fit; if no match is found, the modifier defaults to ×1.0 (no adjustment) so the recommendation still produces a usable lift estimate.

Canonical categoryTypical strong-fit tacticsTypical lower-fit tactics
B2B SaaSG2, Capterra/TrustRadius, integration partnerships, LinkedIn thought leadershipReddit founder presence (less impact than for consumer)
Enterprise ITConference speaking, G2, integration partnerships, entity graphReddit, Quora
Professional servicesLinkedIn thought leadership, original research, HARO/QwotedG2, Capterra, integration partnerships
Consumer SaaSReddit founder presence, Quora answersG2, Capterra
Ecommerce / DTCReddit, Quora, industry roundupsG2, Capterra, integration partnerships, LinkedIn
Media / PublisherPR pitches, HARO/Qwoted, entity graph (Wikipedia)Comparison pages, integration partnerships
Nonprofit / EducationPR pitches, entity graph (Wikipedia)G2, Capterra, integration partnerships
Local services(Mostly technical fixes — schema, robots, server-render — neutral 1.0)G2, Capterra, conference speaking, entity graph, integration partnerships

Modifiers are bounded ×0.2 to ×1.6. Tactics that are universally applicable (technical schema, robots.txt, server-side rendering, FAQ refreshes, content basics) carry no category modifier — they default to ×1.0 for every category. A category-modifier badge only appears on the recommendation card when the modifier is materially above or below 1.0; neutral lifts show no badge to keep the report quiet.

How is AI Mention Quality (AMQ) calculated?

AI Visibility measures how often you're cited. AMQ measures how well: every individual mention is scored 0–100 on how favorably, prominently, and credibly an AI engine portrays your brand. Your brand AMQ is the average across all your mentions in the audit — and because every competitor mention is scored on the same scale, AMQ is directly comparable head-to-head.

AMQ = 40 × Favorability (tone) → max 40 + 30 × Visibility (how central) → max 30 + 20 × Positioning (vs. competitors) → max 20 + 10 × Evidence (citation support) → max 10 = max 100
VectorWhat it measuresWeight
FavorabilityTone — is the portrayal positive? positive (praised / recommended / highlighted as strong) · neutral (informational, or a balanced "X vs Y" comparison) · negative (explicitly criticized, or a rival clearly preferred). A factual capability description is neutral, never negative. 40%
VisibilityProminence — how central are you? primary (the answer is about you, or you're named first) · secondary (one option among several) · tertiary (a passing / parenthetical mention). 30%
PositioningCompetitive stance vs. a named rival preferred · equal · alternative · discouraged — used only when a rival is named in the same mention · solo when no competitor is compared. 20%
EvidenceCredibility — is the mention backed by sources? strong (authority sources — G2, Gartner, TrustRadius — or ≥3 independent domains) · moderate (1–2) · weak (only your own site) · none. Derived deterministically from the citations in the response. 10%

Each mention's 0–100 score maps to a grade:

80–100 · Dominant 65–79 · Strong 50–64 · Established 35–49 · Under-indexed <35 · Deficient
Why AMQ is deterministic — the part that matters. The four vector labels come from a single language-model read of the mention at temperature 0, constrained to the fixed enums above. The score is then computed by the hardcoded formula — no model ever invents a number. Given the same frozen audit evidence, AMQ always recomputes to the same value, and every point is traceable to a label. The model + methodology version is pinned in amq_version, so a model upgrade is a new, dated measurement — never a silent rewrite of history.
The competitive layer — the AI Recommendation Landscape. Because every mention (yours and each competitor's) is scored on the same scale, AMQ powers three head-to-head views: the distribution of each brand's mentions across the five grades (the shape, not just the average — two brands can share a 50 while one is split all-100/all-0); the Recommendation Win Rate — in prompts where you and a rival both appear, how often your mention outscores theirs; and Recommendation Consistency — how steady your AMQ is across engines. Small head-to-head samples are always labeled with their sample size (n).

What we deliberately don't do

These are choices made for integrity. Each removes a feature a competitor might ship. Each is the right call.

We don't combine the two scores into one "Overall AEO Score."
Averaging outcome (Visibility) and input (Readiness) into a single composite obscures the very diagnosis customers most need. We show both separately. Forever.
We don't show "Industry Average" or "Percentile Rank" benchmarks yet.
Showing benchmarks computed off a small customer sample would itself be a fabrication. Benchmarks ship only when we have ≥50 customers per category — and the methodology for computing them will be published before they appear.
We don't let the LLM generate impact numbers.
Recommendation impact is sourced from a versioned, expert-calibrated tactic library. The LLM writes copy; the library provides math. No hallucinated projections.
We don't show projected lift that exceeds remaining pillar headroom.
If your Technical pillar is at 2.4/2.5, a recommendation can't project more than +0.1 Technical lift. The math is enforced at render time, not implicitly trusted from the LLM.
We don't change the methodology silently.
Every change to scoring formulas, pillar weights, or tactic library calibration bumps the version on this page and is logged in our internal methodology document. Customers receive an in-product notification when scores may shift due to a methodology change.
We don't hedge our projections behind "may improve" language.
Projected impact ranges are honest commitments — calibrated against AEO best practices. If you ship the action and observe a different result, that's calibration signal we use to refine, not noise we hide behind disclaimer language.
Calibration v1.1

Methodology version 2.0 (2026-06-21) moved Citation Rate and Engine Coverage to an observed denominator: a question×engine pair an engine couldn't answer (timeout, outage, a retired model) is excluded from the math rather than scored zero, and a Google query where no AI Overview appeared counts as observed-but-not-cited. The engine portfolio also moved from 9 to 8 (Microsoft Copilot retired). The tactic library and pillar weights are unchanged — calibration stays at v1.1.

Calibration version 1.1 (2026-06-15) added per-tactic category modifiers (eight canonical categories, multipliers bounded ×0.2 to ×1.6), widened the LinkedIn thought-leadership range to reflect higher observed variance, lowered the G2 profile activation expected range, bumped the integration partnerships range, and added a new entity-graph presence tactic for Wikipedia / Wikidata / Crunchbase.

Calibration version 1.0 (2026-06-14) means the inter-pillar weights and tactic library impact ranges are expert-estimated for MVP — based on published AEO research, documented AI engine citation behavior, and internal hypotheses about which signals retrieval systems weight most heavily. These ranges have NOT yet been calibrated against observed customer outcome data because AEO Owl is new. As real customers ship recommendations and observe results, the calibration will be refined and the version bumped. Every change is recorded.

Engineering reference. The canonical internal source-of-truth for every formula on this page is blueprints/methodology.md in the AEO Owl repository. Code and documentation must agree. If you spot a discrepancy between this page and what the product shows, please email admin@aeoowl.com — we treat methodology misalignment as a P0 bug.

Last updated: 2026-06-21 · methodology v2.0 · calibration v1.1