Impact

Every applied fix is measured on the KPI it was meant to move, as a difference-in-differences against your own store trend — with honest verdicts, including "no change".

The question every CRO tool avoids is the only one that matters six weeks later: did that fix actually do anything?

Impact answers it. When a fix is applied, DynoWeb captures a baseline. After the lag appropriate to that kind of fix, it measures the same metric again and reports the before → after. Not a score, not a projection — a measured delta with its measurement window stated.

Impact table with money at top and ranked fixes below.

The right KPI per fix

Measuring every fix on revenue would be dishonest, because most fixes do not move revenue directly and the ones that do take a long time to show it. So each fix resolves to its own outcome metric through a registry — which metric, at which scope, which direction counts as better, and how long to wait.

Kind of fixMeasured onBetter means
Frustration fix on an elementRage / dead / error clicks on that selectordown
Hidden or weak CTAClick-through rate on that elementup
Content below the foldScroll reach on that pageup
Form frictionForm completion rate, field abandon rateup / down
SEO fixOrganic sessions to the fixed pageup
GEO / AI-search fixAI referral sessions, or verified-liveup / binary
Speed fixField LCP p75 for that pagedown
Conversion or pricing changeConversion rate, revenue, channel revenueup

The scope is as specific as the fix. An element fix is measured on that selector, not on the page. A page fix is measured on that page, not store-wide. Measuring a button change against store-wide conversion rate is how tools manufacture both false wins and false failures.

Difference-in-differences

Why a plain before/after is not enough

Suppose you fix a CTA and its click-through rate rises 12%. If your whole store's engagement rose 11% that fortnight — a sale, a press mention, seasonality — your fix did almost nothing, and a naive before/after would have credited it with all twelve points.

Impact measures the change in the target and the change in the store-wide control over the same window, and reports the difference. What survives is the movement that is specific to what you changed.

This is not a randomised experiment and Impact does not claim it is. There is no synthetic control group and no counterfactual store. It is the strongest honest comparison available when the intervention has already shipped to everyone — which, for a theme fix or an SEO field, it necessarily has.

Windows and thresholds

SettingValueWhy
Baseline window28 days before the fixLong enough for a stable baseline, short enough not to reach back into a different merchandising regime
Minimum sample50Below this, Impact reports "not enough data" rather than a number. A 3-sample lift is not a lift
LagPer KPIA rage-click rate is readable in days; organic sessions are not. Each KPI carries its own settled-read delay
Provisional readSome KPIs onlyWhere the metric comes from same-day telemetry, an interim read is taken early and clearly labelled provisional. Metrics with lagged external sources get one read, at the end

A row is measuring until its lag elapses. Measuring rows are not blank — they show the captured baseline, so you can see the "before" number DynoWeb is going to compare against, on the day you shipped.

Verdicts

Impact reports what it found, including when what it found is nothing:

Improved

The target moved in the intended direction, beyond the store-wide trend, on enough sample to say so.

No change

The delta did not clear the bar. This is a real, common and useful result — it stops you scaling something that did nothing.

Regressed

It moved the wrong way. Shown as prominently as a win, because this is the one that saves you money.

Not enough data

Below the sample threshold. Honest rather than decorative.

Every figure is labelled an estimate. There are no guarantees anywhere in this surface, and that is deliberate.

Money, carefully

A $/mo revenue estimate appears only where there is a clean path from the metric to money — conversion rate, revenue, channel revenue, or organic sessions (traffic that can be valued at your own conversion rate and order value).

For a click-through-rate win or a rage-click reduction, no dollar figure is produced. It would require inventing a conversion assumption for the improvement, and that invented number would then be the headline. Instead those fixes report their real movement — "−42% rage clicks", "+18% CTR" — as secondary metric chips, alongside the primary verdict.

The monthly total at the top of Impact combines measured fix revenue with Popups incremental revenue — the holdout-normalised figure, not gross totals — so it reconciles exactly with the dashboard's earned-this-month headline and the Popups attribution view. Three surfaces, one number.

What enters Impact

  • A suggestion you marked done — clickable straight back to the suggestion that proposed it
  • A batch applied by SEO Autopilot
  • Fixes verified by read-back, where "verified" means DynoWeb checked the value is live

Limits

These are properties of the method. Stating them is cheaper than arguing about them later.

  • No counterfactual. You are compared against your own past plus a store-wide control, not a held-out group. External events inside the window — an algorithm update, a competitor's sale, a supply problem — are not controlled for.
  • Seasonal stores are harder. A swimwear store fixing something in April and measuring in May is comparing across a demand ramp. The store-wide control absorbs some of that, not all of it.
  • Concurrent changes share credit. Ship three fixes to the same page in the same week and the attribution between them is not separable. Space out changes you actually want to learn from.
  • Small stores get "not enough data" often. The 50-sample floor is a real constraint. Saying nothing is the correct output when the data cannot support a claim.
  • Proxies are labelled as proxies. Organic sessions stand in for search clicks; they are correlated, not identical. The registry records which KPIs are hard measures, which are proxies, and which are binary applied/not-applied.
  • "No change" is sometimes measuring the truth about the apply, not the fix. If a change was marked done but never actually reached the live storefront, the metric will correctly report nothing happened.

Common questions

How long until a result is trustworthy?

It depends entirely on your traffic and the size of the effect. Impact will not declare a result until the affected traffic supports it, which on a low-volume store can mean weeks rather than days.

Is this the same as an A/B test?

Not quite. A/B testing splits traffic simultaneously; Impact compares comparable periods before and after. For nudges, DynoWeb does run true holdout splits — see Popups. Period comparison is the pragmatic option for theme and copy changes where a split is impractical.

What if I change several things at once?

Impact will tell you the combined effect and will not pretend to isolate individual contributions. If attribution per change matters to you, stagger them — the tool will say so when overlap makes a result unattributable.

Does it account for seasonality?

Comparison windows are chosen to be like-for-like where possible, and known distortions such as sale periods are flagged on the result rather than silently averaged in.