Skip to main content
Every AEO tool will tell you what to do. Almost none will tell you whether doing it changed anything. The Impact page closes that loop: for each action you have marked Done, it reports how the mention rate moved on the questions that action targets, measured against the rest of your domain over the same period, with a confidence interval and a significance test attached.

How to find it

Navigate to app.trybluemoon.com and select Actions → Impact from the left sidebar. The page has two parts: Suggestion quality at the top (how good Bluemoon’s own recommendations have been) and Measured impact below it (one card per completed action). It lists the 50 most recently completed actions for the selected domain.
The measured-impact list stays empty until you mark at least one action Done in the Earned or Owned catalogue. Marking Done is what starts the measurement — it is the event Bluemoon anchors the before-and-after to. The suggestion-quality panel above it renders straight away.

What happens when you mark an action Done

1

Bluemoon photographs the before

At the moment you mark the action Done, it counts, over the trailing 21 days: how many recorded answers mentioned your brand on the questions this action targets, and how many mentioned it everywhere else on the domain. A recorded answer is one tracked question asked on one AI platform, with the answer stored. Both direct and implicit brand mentions count.
2

The target population is fixed at that moment

Most actions target a topic. Some target a question shape — every “how-to” question, for instance — and those span topics by construction, so the baseline records the exact list of questions instead. Whichever it is, it is written down at capture and replayed later. A before and an after computed over different sets of questions would produce a confident number that means nothing, which is worse than reporting nothing.
3

The control cohort is captured alongside it

The same two counts are taken for everything the action does not target — the rest of your domain. This is the control.
4

The same picture is taken again on demand

Every time you open the Impact page, the after-window is recomputed live: the identical query, over the identical target population, across the trailing 21 days ending now.
5

The lift is reported as the difference of the differences

Bluemoon reports how much the target moved, minus how much the control moved, in percentage points.
The baseline is captured once and never overwritten. Re-completing an action does not move the goalposts.

Why the control cohort is the whole point

Suppose an outdoor retailer publishes a comparison page for its “waterproof jackets” topic and the mention rate on that topic goes from 20% to 32% over the next three weeks. That looks like a 12-point win. It might not be. AI answer engines shift underneath everyone. A model update, a fresh index, a competitor going quiet — any of these can lift or sink your whole domain at once, and none of them has anything to do with the page you published. So Bluemoon asks a second question at the same time: what did the rest of your domain do? If every other topic also rose, from 15% to 19%, then 4 of those 12 points were happening to you anyway. That subtraction is the method. It is called a difference-in-differences, and it is the reason the number on the card is smaller than the raw improvement you could screenshot yourself. The control absorbs whatever moved your whole domain, so what is left is the part that is plausibly attributable to the action rather than to the weather.
The control is “everything else on this domain”. If you complete several actions at once, each one sits inside the others’ control cohort and they mute each other’s measured lift. Stagger completions when you want a clean read on a specific action.

Why the number comes with an interval

A mention rate is a proportion measured on a sample. “35%” from 20 recorded answers and “35%” from 400 recorded answers are the same number and completely different claims — the first would look entirely different if you had asked on a different day. So each card shows before and after not as two dots but as two bands: a 95% Wilson confidence interval around each rate, drawn to scale, with the observed rate marked inside it. Wide bands mean few answers behind the figure. If the two bands overlap heavily, you are looking at noise wearing the costume of a result. Alongside the bands, a two-proportion test asks whether the target’s own before-to-after change is bigger than sampling noise would comfortably explain. At p < 0.05 the card reads significant lift; otherwise it reads no significant change.
Read the two together, and note what each covers. The net lift figure is the one that nets out domain-wide drift. The significance test is run on the target’s own before-to-after change, not on the netted figure. So “significant lift” means the target genuinely moved; the net lift number tells you how much of that move survives once the control is subtracted.
The card also prints n= the number of recorded answers behind the after figure. Small n is not a defect to hide — it is the reason the band is wide, and it tells you what to do next (track more questions on that topic, or wait for more polls).

The three states a card can be in

Measuring

The measurement window is still open. Shown for the first 14 days after completion, with the day count on the chip. The numbers are live but provisional.

Measured

At least 14 days have passed and both windows carry enough recorded answers. This is the reading to act on.

Not enough data

Fewer than 8 recorded answers on the targeted questions in either the before or the after window. Bluemoon will not compute a verdict from that.
“Not enough data” is a refusal, not a failure. With a handful of answers behind it, a diff-in-differences will happily produce a large, confident-looking lift that is pure sampling accident. Bluemoon would rather tell you it cannot say. Treat the status chip as the finding: when it reads “not enough data”, the figure printed beside it is not a result. The calibration tile in the panel above holds itself to the same standard. Where the correlation cannot be computed yet it says so in words or shows a dash, rather than printing a zero — below 5 completed actions the calibration tile states that there are not enough of them to judge, and a correlation that cannot be computed reads , never “no correlation”. “We did not measure this” and “we measured this and it was zero” are different claims, and collapsing them would let an empty dataset masquerade as a bad outcome. The adoption and grounding tiles beside it do not yet make that distinction: with nothing decided, both print 0%. Read them as empty rather than as a verdict until you have judged a few suggestions.

Leading signals

Two fast, directional flags appear on a card before the statistics have anything to say. They are indicators, not proof — the numbers above are the measurement.
  • Your page is now cited — one of the owned URLs recorded with the action was found in a citation captured after the completion timestamp. Up to 10 URLs are watched per action, matched as literal substrings against cited URLs. This is a coincidence in time, not a causal proof.
  • First mention on [platform] — a platform that had zero recorded brand mentions for this domain before the completion timestamp has at least one after. This flag is computed across the whole domain, not just the action’s target, so it is domain news that happens to have surfaced on this card. Do not credit it to the action on its own.

Suggestion quality

The upper panel grades Bluemoon’s own recommendations, independently of any realised lift. It is unusual for a tool to publish this: the panel exists so you can tell whether the queue is worth working through, which necessarily means it can tell you that it is not.

Adoption rate

Of the suggestions you actually decided on, the share you kept (moved to To do or Done) rather than declined. Suggestions you have never touched count in neither half, so the rate is about decisions, not about backlog. Broken out per format so you can see which kinds of recommendation you keep and which you consistently throw away.

Grounding

The share of suggestions generated for this domain that carry real evidence and are not aimed at a question your brand already wins. A recommendation to go win a question you are already mentioned in is noise, and it is counted as such here.

Prediction calibration

Bluemoon ranks its suggestions with a signal score. Calibration is the rank correlation (Spearman) between that score and the change actually measured afterwards on each completed action’s target. It answers one question: does a higher-ranked suggestion really tend to move more?

Withheld until it can be said

Calibration is not reported at all below 5 completed actions with baselines, and is only labelled reliable at 8 or more with a rank correlation above 0.4 in the right direction. Until then the panel says so plainly rather than printing a number.
What to do with it. A high adoption rate with high grounding means the queue is worth working top-down. A low adoption rate on one format is a signal about that format, not about you — decline them and the number will say so. Calibration is the coarsest of the three: while it reads “not yet reliable”, treat the ordering of the queue as a reasonable heuristic and nothing more, and use each action’s own measured impact as the verdict.
The signal score is a prioritisation signal. It orders the queue; it does not forecast your visibility. Calibration is exactly the check on whether that ordering is earning its keep, and it is allowed to come back negative.

What this measurement cannot tell you

It does not verify that you shipped anything

Marking an action Done is a declaration. Bluemoon does not crawl your site to confirm the page exists — it takes the timestamp and starts measuring. Mark Done when the work is actually live, or the before-window is anchored to the wrong moment.

It measures mentions, never causes

Everything here is the mention rate in AI answers on the questions you track. A lift is evidence that the target moved more than the control did over that window. It is not proof that your action caused it, and Bluemoon does not claim otherwise.

Low-traffic targets stay unmeasurable

A topic or question shape with almost no recorded answers will not become measurable by waiting — the window slides, so it stays roughly as empty in a month as it is today. What moves it is tracking more questions on that target, or polling more platforms. Waiting alone does not.

The window is trailing, not anchored

Both counts use a trailing 21 days. Because a card graduates to “measured” at day 14, about a third of its after-window still predates the completion. The reading gets cleaner as the window clears; at day 21 and beyond it contains only post-completion answers.
Two further limits worth knowing. Calibration compares each completed action’s signal score against the raw change on its target topic, with no control subtraction — so unlike the per-action lift, it carries domain-wide drift inside it, and it is a coarse check on the ranking rather than a verdict on any single action. And because the whole feature reads from recorded answers, a platform that has gone quiet for your domain contributes nothing to either window; see Agent Analytics for the other half of the measurement picture.