Skip to main content
Visibility tells you how often the AI names your brand. Intent tells you in which kinds of question it names you, what the AI did with that question, and how your brand sat inside the answer — because “named” covers everything from “my top pick” to seventh in a list of nine, and a mention in an explanation is not the same finding as a mention in an answer that steers the reader to a shop. Bluemoon measures this at three levels, and each level is read from evidence the pipeline already stores before any AI is asked to judge anything. The first two share the same three labels on purpose. It is the taxonomy every marketer already knows from keyword tools, and the symmetry is what makes drift legible: “commercial question → informational answer” reads without a legend. The Intent Map draws exactly that drift.

How to find it

Intent Map

Intent Map in the left sidebar (/intent). The question-intent → answer-treatment flow, the per-engine matrix, and the share of answers that recommend you. Described in its own guide.

Visibility → Question type

The Question type tab on the Visibility page, alongside Topic, Persona and Platform: your mention rate per question shape.

Every AI answer

Open any answer, anywhere in the product. The viewer shows the question’s intent, the answer’s treatment with the evidence that earned it, the role of every tracked brand named in it, and the funnel stage the two derive to.
For a long time the Intent screen was reachable only by URL: it restated the question-side score, and that is a dimension, not a destination. The map changed that — question intent × answer treatment × your presence × engine is a reading you come back to — so /intent is in the sidebar today.

The evidence-first rule

Every level follows the same discipline, and it is worth stating once because it is what makes the numbers checkable:
1

Rule from stored evidence first

A deterministic rule reads facts the pipeline already recorded — cited retailers, prices in the text, brands the mention detector found, the brand’s sentiment, the sentence around its name. This verdict is final wherever the evidence speaks outright. Two retailer citations is a shop-steering answer whatever a model thinks.
2

Ask an AI only where the evidence stops short

A cheap model (gpt-4o-mini) reads the text only for the ambiguous band: some signal, below every threshold. It never overrides an evidence verdict, because it is never asked about one.
3

Store both, once, with provenance

The verdict shown carries its source (evidence or llm), and the evidence verdict is stored beside the AI one so their agreement stays measurable. A verdict is never re-asked. When the model’s reply cannot be parsed, nothing is stored and the answer is retried on a later run — a guessed label is never written.

Level 1 — Question intent

Every tracked question carries a commercial intent score from 0 to 100: how close the asker sits to a purchase. On the distribution card it is shown in four bands (0–24, 25–49, 50–74, 75–100), because the underlying signal is coarse and four buckets is roughly the resolution it can defend. For the map, the same score is cut into thirds so that the question axis mirrors the answer axis: A question with no score yet is unscored: it is excluded from the map and counted in a note under it, never folded into “informational”.
The commercial score is a weighting and prioritisation signal, not a predictor. Measured on production data, it correlates with visibility overall — and that correlation collapses to near zero once the topic tier is held constant. Use it to decide which questions are worth your attention first. Do not read it as a forecast of where the AI will name you; question shape is the axis that survives the control.

Where the number came from

Two scores of 62 can be worth very different amounts. Every score carries a source, shown next to it everywhere it appears and never mixed into a blended average.

Ad prices

Strongest. Anchored on Google Ads competition data for the question’s priceable keyword, percentile-ranked against your own keyword population within a single country and language. This is what advertisers actually pay to reach the asker.

Read from the question

Middle. Derived from how the question is phrased — a coarse 0–5 reading of buying proximity, blended at half weight with the topic estimate. No market data exists for this keyword yet.

Topic estimate

Weakest. A fallback from the topic tier plus buying language in the wording. Free, deterministic, and available for every question from day one — but a prior, not a measurement.
The anchor is percentile-ranked within your own keyword population, inside a single (country, language) market, because raw ad competition is not comparable across industries. A keyword Google returned no competition data for stays null and the score falls back to the wording — “unknown competition” is never rendered as “no competition”. Classification runs on a schedule and works through the backlog. Rescoring is triggered by an input change — the question text moved, or the classifier version moved — and never by a timer: a question’s intent is a stationary property of the question, and re-deriving it on a schedule would move the weighting under every rate each sweep.

Question shapes

Alongside the score, every question is classified into exactly one of ten shapes — Definition, How to, Comparison, Best-of list, Alternatives, Pricing, Review, Recommendation, Troubleshooting, Availability — with a fixed precedence order so the same question always lands in the same bucket. The Question type tab on Visibility reports your mention rate per shape, with a confidence interval and the number of answers behind it; a question needs at least 3 answers of its own before it counts toward its shape. The Core topics column restricts the same rate to core-tier topics, which is the control that separates “the shape matters” from “the tier is speaking through the shape”. A shape is called a blind spot only when the optimistic end of its interval is still below 5% and at least 3 tracked questions carry that shape. When nothing clears the bar, the callout renders nothing rather than filling the slot.

Level 2 — Answer treatment

The treatment is what the AI did with the question. It is read from four facts stored for every answer: The rules, in order of commitment: An answer can carry more than one treatment: one that cites two shops and weighs four brands is both transactional and commercial. The most committed label is the primary — the one the map draws — and the answer viewer shows the full list. “Informational” is only ever alone: it is the label for “nothing else fired”, not a quality an answer has in addition to steering or comparing. Every classification can therefore answer “why”. The viewer prints it: Evidence: 3 retail sources cited · 4 brands named · 0 comparison pages · prices in the answer.

The AI refinement

The evidence has a middle: an answer that names one brand, or quotes one price, or cites one comparison page, has some signal but sits below every threshold. There, framing decides, and framing is in the prose the rules never read. Those answers — and only those — are sent to an AI read of the question and the answer together, since a price in an answer to “how much does X cost” is the expected shape, not a sales push. The model returns every applicable treatment and the primary one.
  • The verdict is stored once, with the evidence verdict beside it, and marked refined by an AI read of the answer wherever it is shown.
  • An answer with no signal at all — no brands, no prices, no comparison sources — is confidently informational and is not sent to a model. Paying to confirm thousands of those a month is how a refinement pass becomes a bill.
  • The map reports how many answers in the window were refined this way.
Bluemoon measures how often the AI read corrects the evidence verdict in that band. The figure is being re-measured on the current corpus and is not quoted here until it is.

Level 3 — Brand role

“Named” says the brand was in the room; the role says how. The role is a property of one brand inside one answer, so every answer that names at least one tracked brand gets one role per brand — yours and each competitor’s. The evidence rules read, per brand: its rank among the tracked brands the answer names; the brand’s sentiment from the sentiment pass; a cue phrase in the sentence around its first mention (English, French and German forms, within a short window of the name — a cue two paragraphs away does not count); and how many tracked brands the answer names. Precedence, most defensible claim first:
  1. Negative sentiment attributed to the brand → Caveat. A measured, brand-attributed signal outranks any word in a short window.
  2. A recommendation phrase near the brand, or rank 1 with a list marker (“1.”, “#1”) right before it → Recommended.
  3. A caveat phrase near the brand → Caveat.
  4. No cue and three or more brands named → Listed.
  5. No cue and one or two brands → Mentioned.
The AI read is asked only where the evidence hints and stops short: a recommendation cue beside a caveat cue, or no cue at all with one or two brands (where “mentioned” is a default, not a finding). One call reads every named brand in the answer at once, because roles are relative — “listed” only means something next to the others. A brand the model skips keeps its evidence verdict; an unusable reply stores no rows at all for that answer, so it is retried rather than half-stored.
“Brands named” counts tracked brands only. An answer listing eight outdoor brands of which you track two is, to the rules, a two-brand answer — which is precisely the kind of case the AI read exists for, and why its verdict is stored with its source.

The funnel stage

From treatment × roles, each answer derives a stage. It is a rule, not a third classifier — a finer stage than the evidence can carry would be false precision:

”Not measured” is never zero

Every surface on these three levels follows one rule: a value that was not measured is shown as a dash or as “not measured yet”, never as 0 or 0%.
  • An answer whose roles have not been read yet shows Roles not read yet, and its funnel stage is left blank rather than computed without them — a stage derived without the roles would silently under-report “decision”.
  • “Recommended in X of N answers that named you” is only printed when at least one such answer has been role-read. “Recommended in 0 of 41” and “not measured” are two different facts, and only the first one is a finding.
  • A question with no intent score is excluded from the map and counted, not hidden.
  • A measured rate that rounds to 0.0% is shown as <0.1%: a handful of mentions is not the same finding as none.
The reason is the same everywhere in Bluemoon: “we asked and you were not named” and “we have not looked yet” must never look alike, because collapsing them is the fastest way to make every other number untrustworthy.

Cadence and windows

The three levels are not computed the instant a scan finishes. So after a scan, the evidence verdicts are visible immediately; the AI refinements and the roles land about two minutes behind each answer, and until then the corresponding surfaces say so. On a 90-day view, answers older than 30 days carry their evidence treatment but no AI refinement and no roles — they read as not measured on the role surfaces, not as zero.

Limits

The transactional signal leans on knowing which cited sites are retailers. Early in an account’s life that table is thin and the signal under-fires; the map says how many retail sources have been identified so far rather than letting a quiet number read as a strong one.
Rank 1 means first among the brands you track, not first in the answer — which is why rank alone is never a recommendation cue and needs a list marker beside it.
Recommendation and caveat cues are whole-word phrases in English, French and German. A recommendation phrased in some other way is not caught by the rules — it falls to the AI read only if the answer is in the ambiguous band, otherwise it stays at the evidence verdict.
Once an answer has an AI verdict — for its treatment or for its roles — it keeps it. Improving the rules changes future answers, not stored ones.
A low share of recommendations tells you the AI did not pick you when it named you. It does not tell you why, and Bluemoon does not claim to know. Open the answers and read them — that is the evidence, and it is one click away.