Methodology

How the index is measured

AI assistants generate answers. Those answers vary between runs, and no one can guarantee a brand a place in a generated answer. A visibility number is only meaningful when the method behind it is published in full. This page is the method behind Mira Index: what we ask, what we count, how the score is computed, and where the method stops.

Last updated 15 September 2026

01

Principles

Fixed question set

The score comes from a versioned question set we control and publish the shape of. A customer cannot add, remove or tune a scored question.

Licensed universe

Brands are resolved against the market's official licence register, with dated, human-approved status records. Not a hand-kept global list.

Repeated measurement

A single answer is a sample of one. Every question is asked 5 times per model, and every score carries a confidence band.

Published method

The formula, the weighting mechanism, the cadence and the limits are on this page. A score that cannot be explained cannot be audited.

No optimisation services

We measure visibility. We do not sell the work that moves the number. When the same party measures and optimises, the measurement has a stake in its own result; we removed that conflict of interest by not selling the second half.

02

What we measure: Core, Edge, Lens

Everything a customer sees belongs to one of three layers. Only one of them produces a score.

LayerWhat it isWho controls itEnters the score?
CoreThe fixed, versioned question set per market, demand-weighted, with every licensed brand measured on identical questionsThe market pack, versionedYes. This is the score.
EdgeCustom questions and watchlist brands a customer adds: answers, mentions and positions are shownThe customer, capped per planNever. No score, no rank, no trend.
LensHow a customer views the same data: their own row, chosen rivals, filters, alertsThe customer, freelyNo. A lens filters; it never recomputes.

Edge never enters a score because a score a customer can tune is not a benchmark. What customers ask through Edge is still signal: aggregated Edge demand is reviewed at scheduled version reviews, where it can shape the next Core version. It never changes the current one.

A Core version change is a comparability boundary. It is dated and disclosed, and no trend, diff or alert crosses it.

03

The question set

A market's question set is built from real player questions with commercial and reputational intent, written in the market's own languages, weighted by real search demand, and categorised. Each question maps to one theme of the engine's fixed taxonomy (Trust & safety, Bonus, Payout speed, Product & app, Live casino and Sportsbook) or is held out as monitoring-only, which is asked and archived but never scored.

Demand weighting

A market pack supplies each question's monthly search volume, never a weight. The engine derives every weight with one formula, freezes it when the version is seeded, and refuses a pack that tries to supply its own:

cap = 2 × median( monthly volume of the scored questions ) weight_q = min( volume_q , cap ) / cap a value in (0, 1]

Monitoring-only questions weigh zero everywhere. A market with no demand data weighs every question equally, and we say so rather than inventing precision. The per-market figures, including how many questions a set holds and its theme split, are published in each market's own pack disclosure.

Versioning

New questions mean a new version, never an edit in place. Stored answers and scores are append-only and preserved across versions. Trends, diffs and change alerts refuse to compare two scans of different versions: the comparison is declined, not fudged.

Questions never name a brand

No question text may contain any alias of any brand in the market's dictionary; the engine refuses to load a pack that breaks this. Because no question names a brand, one market scan measures every brand on identical input, and no brand can be handed a friendlier question.

Why the texts are unpublished

We publish the shape of the set and not the texts. A published question set is a gameable one: content could be written to the exact wording, and the benchmark would decay for everyone measured on it. This is the same logic that keeps Edge out of the score. The texts are fixed, versioned, and identical for every brand in the market.

04

The model panel

The active panel is 6 models: ChatGPT, Claude, Gemini, Perplexity, DeepSeek and Grok. Each is pinned to a dated release id, never a floating "latest" alias, so the instrument does not drift under the measurement. All 6 are scored.

ModelPinned idAnswers fromCitation track
ChatGPTopenai/gpt-4oTraining-time knowledge on the score trackYes. Re-asked with web grounding enabled.
Claudeanthropic/claude-sonnet-4-6Training-time knowledge on the score trackYes. Re-asked with web grounding enabled.
Geminigoogle/gemini-2.5-flashTraining-time knowledge on the score trackYes. Re-asked with web grounding enabled.
Perplexityperplexity/sonarLive web retrieval on every call, by designYes. Searches natively on every run.
DeepSeekdeepseek/deepseek-chatTraining-time knowledge onlyNo. Never web-grounded; contributes no citations.
Grokx-ai/grok-4.6Training-time knowledge onlyNo. Never web-grounded; contributes no citations.

The two columns matter for reading any result. A memory-only model answers from training-time knowledge: its answers reflect a slow consensus and move slowly, so a brand absent there has an awareness problem, not a news problem. A web-grounded run reflects what is retrievable today and can move week to week. Averaging the two without labels would be misleading, which is why every per-model view is labelled. On the citation track, ChatGPT, Claude, Gemini and Perplexity run web-grounded; DeepSeek and Grok are never web-grounded and contribute no citations.

The Perplexity asymmetry

Perplexity retrieves live web sources on every call by design, including on the score track, while the other panel members answer the score track from training-time knowledge. We accept this asymmetry rather than removing the model, because players use it exactly as it ships, and a panel without it would be tidier and less true. Its cells are labelled like every other model's, so a reader always knows which kind of answer produced a number.

Panel changes and retired models

The panel is capped at 6 models, and it changes only at a question-set version boundary, so no version ever mixes panels. A retired model stops being asked; its stored answers keep rendering in every historical scan. Retired so far: Llama. Every page reads the panel snapshot stored on the scan itself, never the current code list, so a panel change can never rewrite a past scan.

Request parameters are fixed: temperature 0.3, up to 500 tokens per answer. The question text is byte-identical on both tracks; web grounding is enabled through request parameters only.

05

Markets, licences, and why iGaming is different

A gambling brand does not exist globally. It exists per licence, per market, per domain: the same name can be a licensed operator in one market and an unlicensed one next door. A global visibility score averages those two facts into noise and hides the only thing that matters to a licensed operator: what AI models say inside the market where its licence applies. Mira Index measures nothing globally. Every number on every page belongs to exactly one market.

What a market pack contains

  • A licensed-brand dictionary: every brand on the market's official licence register, with its aliases and domains, so detection reads answers the way a player writes brand names.
  • Dated licence-status records: who is licensed, from when to when, approved by a named person.
  • A demand-weighted question set in the market's own languages, versioned as described above.
  • A watchlist of brands commonly named in that market without holding a licence there, plus an engine-level list of offshore brands licensed nowhere we measure.

Licence status is a dated lookup

When a scan is parsed, each detected brand resolves against the status record whose validity period contains that scan's start date. Licences lapse, renew, get revoked and get acquired. A product that silently rewrites its own past when a licence changes is not a measurement product. Here, a status change affects only scans from its effective date forward: every older scan keeps the status that was true on the day it ran.

Three mechanics back this up. Every scan records a hash of the exact dictionary it was parsed with. Brand and status records are append-only: a change closes the old record and adds a successor, and a brand is never deleted. And every recorded mention carries the licence status as resolved on that scan's date. Together they guarantee one thing: any stored result can always be read against the exact dictionary and statuses that produced it. They do not re-verify old answers, and they do not make the dictionary infallible; they make it accountable.

Unlicensed observation

When a model recommends a brand that holds no licence in the market, we record it, show it, and alert on it. It is never scored and never ranked. For a licensed operator this is recommendation share lost to operators outside the rules, on the exact questions players ask. For a regulator it is dated, quotable evidence of where AI answers send players. We record and report these observations; we claim nothing beyond that.

On the register itself, plainly: licence statuses are founder-curated. They are entered by hand, dated, and approved with a named approver, verified against the regulator's public register on a stated date. Automated register synchronisation is built but not yet live; when it runs, it will produce proposals for a human to approve, never automatic edits.

The engine is market-agnostic

The engine contains no market-specific logic. A market is data: a register, a dictionary, a question set. Adding a market adds no engine code, and every market runs the identical formula, cadence and rules published on this page.

Coverage today

MarketStatusWhat that means
the NetherlandsmonitoredWeekly index scans, full methodology on this page
the United Kingdomfree-scan instrumentedApproved licence register; free scans classify against it; no weekly index yet
SwedenannouncedOn the roadmap by demand; nothing is scanned yet

06

The scoring formula

The whole formula fits in a few lines. With the stored answers, the market's dictionary and this section, an analyst can recompute any score we publish.

position weight 1st mention = 1 2nd = 0.7 3rd or later = 0.5 absent = 0 cell point one question × one model: Σ position weight over its samples ÷ n successful samples score 100 × Σ( w_q × cell point ) ÷ Σ( w_q ) over every scored cell, where w_q is the question's demand weight
  • A mention is a case-insensitive match of a brand's dictionary aliases in an answer, at word boundaries. We store the brand, its first-appearance position, and the sentence containing it.
  • First-appearance position is counted among all detected brands, licensed and unlicensed alike. An unlicensed brand named first pushes every licensed brand down one position and costs it weight. A brand in neither list is invisible and pushes nothing.
  • A sample is one asking of one question to one model; n is the number of successful samples in the cell. Failed calls never enter a denominator.
  • The score ranges 0 to 100 and is reported to one decimal.

Every score carries a band: a 95% Wilson confidence interval on the appearance proportion, computed per cell and rolled up with the same demand weighting as the score. Fewer samples make the band wider. The band is always shown, never hidden.

The score deliberately excludes: unlicensed brands (recorded, never scored), monitoring-only questions, web-grounded runs, customer custom questions and watchlists, free lite scans, failed calls, and any model a scan's stored configuration marks as measured-but-not-scored.

We refuse to call a week-over-week change real unless the confidence bands do not overlap and the score moved at least 5 points. Below that bar the system stays silent, because a false alert costs more than a missed one.

07

Execution and repeatability

  • Cadence. Every monitored market is scanned weekly, on the deployed schedule 0 6 * * 1 (Mondays 06:00 UTC). A sweep on */5 * * * * (every five minutes) advances running scans, and a scan completes the same morning.
  • Sampling. Each question is asked 5 times per model on the score track, and 3 times per web-grounded model on the citation track.
  • Failures and retries. A failing call is retried up to 3 times in-call with backoff, then archived as a failed attempt. A failed slot is retried on later sweeps, at most 5 attempts per slot, and only within 48 hours of the scan's start. Failed attempts are kept forever; readers use the latest successful attempt.
  • Finalisation. A scan is terminal when every slot has a successful attempt, has exhausted its attempts, or the window has closed. It finalises with whatever succeeded.
  • The reliability floor. A scan whose share of successful calls falls below 70% is marked partial: it is quarantined from comparisons, never a baseline for a change alert, never the source of a derived view, and a human is emailed the failure. Silence means success.
  • Observed dates. The date shown next to any result is the date the answers were collected, never the date processing finished.
  • Idempotency. Every call is keyed by scan, question, model, sample, grounding and attempt. An interrupted scan resumes exactly where it stopped and can never duplicate a call.
  • Verbatim storage. Every raw answer is stored verbatim. Mentions, scores, citations and classifications are derived tables that can be rebuilt from the raw answers under a versioned parser; corrections are new rows, and old rows stay. Every rendered cell is a stored snapshot of what was actually returned, not a re-derivation.

The free lite scan is a different measurement

The free public scan runs the same formula at lower sampling: 5 questions × 3 models (ChatGPT, Claude and Gemini) × 3 samples. Its confidence band is wider, the report shows that band rather than hiding it, and its results are stored apart: a lite result never enters the index score, a leaderboard or a trend, and is not comparable to the index score on this page.

08

Citations and sources

A citation is a URL a web-grounded response itself attaches to its answer. We capture the citation payload verbatim from the response, store it, and parse it into queryable rows under a versioned parser. We record what the model cited, not what we think it should have cited. Citations come from web-grounded runs only; memory-only models contribute none.

Every cited domain is classified into one of 6 source types: operator, affiliate, media, regulator, forum and other. Classification has three layers of authority. Deterministic rules come first: a brand-owned domain is an operator, known regulator and forum domains classify by list. A manual override by us is final until replaced. Everything the rules cannot decide goes through one model-assisted pass whose decisions are logged append-only and can be overridden. The mapping is not fully deterministic, and we do not pretend it is; it is rule-based where possible, reviewed where not, and every decision is on the record. Unclassified domains still count.

We refuse to rank sources on thin data: a domain with fewer than 4 scans of history or fewer than 30 recorded citations is listed as early data, without a rank.

For an operator, the source mix matters more than the score itself. The affiliates, review sites, regulators and forums that models retrieve from are where the answers are actually formed. The score tells you where you stand; the sources tell you which specific domains decided it.

09

What we do not measure

  • AI surfaces outside the panel. Google AI Overviews, Google AI Mode and Microsoft Copilot are not measured. If it is not in the panel table above, we do not measure it.
  • AI referral traffic, registrations and deposits. We measure visibility, not the click and not the conversion. The referral layer, the utm_source=chatgpt.com kind, lives in your own analytics, and that is where to look for it.
  • Factual accuracy of what models say about a brand. Whether a model describes your licence status, payment options or availability correctly is not measured. Probe questions for this are drafted and not live; nothing on any dashboard today measures answer accuracy.
  • Personalised and logged-in answers. Every call is a fresh API request with no user history, so we measure the neutral answer, not the answer a signed-in player with a chat history sees.
  • Citations from memory-only models. DeepSeek and Grok are never web-grounded, so no citation data exists for them.
  • Retrieval by location. All calls run from one infrastructure location. Grounded answers can differ with the searcher's location; we do not measure that variation.

This list exists because a tool that only publishes its strengths cannot be audited.

10

Known sources of variance

  • Model updates under us. Providers change models. We pin dated release ids so the panel never drifts silently, and a deliberate panel change only happens at a dated version boundary. A provider can still adjust behaviour behind a pinned id; stored snapshots and the weekly cadence make such a shift visible as a dated break rather than an invisible one.
  • Run-to-run variation. The same model gives different answers to the same question. We ask every question 5 times at a fixed temperature and publish the confidence band that variation produces.
  • Prompt-order effects. Every question is an independent, single-turn call. No conversation state carries from one question to the next, so answers cannot be steered by what was asked before.
  • Retrieval differences. Web-grounded answers depend on what search returns that day, and search varies by time and place. Grounded runs are sampled on a fixed weekly rhythm from a fixed location, and each scan stores exactly what was retrieved and cited.

Repetition, a fixed question set, stored snapshots and dated version boundaries limit each of these. None of them is eliminated, and a vendor claiming otherwise is describing a different universe.

11

The buyer's checklist

Questions worth asking any AI-visibility vendor, answered for this one.

QuestionOur answer
Which AI engines do you monitor?The 6 models in the panel table, each pinned to a dated release id. Google AI Overviews, Google AI Mode and Microsoft Copilot are not covered.
How many times is each question run?5 times per model per weekly scan on the score track, and 3 times per grounded model on the citation track.
Can you segment by country?Measurement is per market by construction. A score is computed inside one market and is never aggregated across markets.
Can you monitor in local languages?Yes. A market's questions are written in the languages its market pack declares, not translated from a global template.
How is share of voice calculated?We do not report a share-of-voice percentage. The score is the position-weighted, demand-weighted formula published above, with a confidence band. A raw mention share hides position and demand, which is most of the signal.
Do you track citations as well as mentions?Yes. Web-grounded runs record every URL the model itself cites, and cited domains are classified into 6 source types. Memory-only models contribute no citations, and we say so.
Which iGaming brands are in the universe and how were they chosen?Every brand on the market's official licence register, with dated, human-approved status records. Brands commonly named without a licence are observed on a watchlist: recorded, never scored.
Can a customer influence the score?No. Customer questions and watchlist brands never enter a score, a rank or a trend, and the question texts are unpublished. Every brand in a market is measured on identical questions.
What is measurement and what is optimisation?Everything on this page is measurement. We sell no optimisation services, so the party producing the number has no stake in moving it.
How do you handle a model being retired or added?Only at a question-set version boundary, so no version mixes panels. A retired model stops being asked, but its stored answers keep rendering, and every page reads the panel snapshot stored on the scan itself.

12

Questions we get

Can you guarantee our brand appears in ChatGPT?

No. Nobody can. AI answers are generated fresh on every request and vary between runs. That is why we ask every question 5 times per model and report a confidence band instead of promising a placement.

Is this GEO or AEO?

The measurement side of it. The industry calls the discipline generative engine optimisation (GEO) or answer engine optimisation (AEO); we measure AI visibility and sell no optimisation. One caution on the acronym: in gambling, GEO usually means a geographic market. On this site a market is always the licensed jurisdiction, and when we mean the optimisation discipline we spell it out.

We already use a generic AI visibility tool. Do we need this?

A generic tool measures a brand as a global name. In regulated gambling, visibility only counts inside the market where the licence applies. Mira Index resolves every brand against the market's official licence register with dated statuses, asks its questions in the market's own languages, and records when models recommend operators that hold no licence there. If your current tool does those three things, you do not need us.

How often does the score update?

Weekly. Every monitored market is scanned once a week on a fixed schedule, and every question is asked 5 times per model in that scan.

What happens when the question set changes?

A change is a new version with a dated boundary, never an edit in place. Old scans keep their scores. Trends and change alerts refuse to compare across the boundary, so a version change can never masquerade as a market movement.

Why not publish the question texts?

A published question set is a gameable one: content could be written to the exact wording and the benchmark would decay. We publish the shape instead: the themes, the weighting formula, the versioning rules and the cadence. The texts stay fixed, versioned and identical for every brand.

Who owns the data?

Measurement data, meaning the model answers we collect and the scores computed from them, is our published measurement. Your account and personal data is yours, with GDPR access, correction and deletion via hello@miraindex.ai. Your own measurement data is exportable at any time.

13

Change log

  • 15 Sep 2026First published.

Last updated 15 September 2026. Seen something on this page that does not match what we do? Email hello@miraindex.ai. We publish corrections in the change log above rather than silently editing this page.