Skip to main content

Prediction performance by change type

Prediction performance

SN21 models predict what each Google Ads change will do before the outcome is known. When it settles, every prediction is scored — this shows how reliable they are, one row per kind of change.

How to read this page

What we measure
The change's effect on cost, conversions and the goal metricCPA (cost per lead) for lead-generation accounts, ROAS (return on ad spend) for ecommerce.
At what level
The campaign the change belongs to — not the whole account. An ad, keyword or ad-group change is measured on its campaign's results against a 60-day baseline.
Over what window
The next 7, 14, or 28 days after the change — switch the window on the table below.
The rating Reliable lean on it Roughly right direction right, numbers less so Not reliable don't lean on it yet Too early not enough has settled yet
A rating is prediction accuracy — how close the models were — not a probability the change will succeed, and not advice to make it.

Changes settled & scored

953

served changes settled since 20 August 2026

Kinds of change

9

one line each, below

Models compared

126

every model predicts every change

Best model today

UID 160

leaderboard #1

Where the number comes from

16 days, 2026-08-20 to 2026-09-04
  1. 1. Captured248,719raw changes per day, before any rule
  2. 2. Measurable episodes15,191since 20 August 2026, after the published rules and 72-hour clustering
  3. 3. Served to models3,162of 6,423 eligible; 1,264 held out for validation, 53 on our own accounts
  4. 4. Settled & scored953this page — served changes old enough to have settled
  5. 5. Predictions scored105,556every model predicts every change

Each stage is deliberate. The page counts the fourth: a change reaches it about 15 days after it happened, once its first outcome window has closed and settled.

Budget change — when a daily budget was raised or lowered 278 Reliable UID 75

Across 278 such changes (31,542 model predictions), a prediction here is reliable. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Broken down by kind:

down 91 Roughly right UID 75
down large 56 Reliable UID 75
flat 26 Reliable UID 204
unknown 2 Too early
up 40 Roughly right UID 124
up large 63 Roughly right UID 189
Bid strategy switch — when the bid strategy was switched from one type to another 1 Too early

Across 1 such changes (119 model predictions), a prediction here is too early. What that's made of:

  • Roughly right Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Reliable Range — how often the real result fell inside the predicted range
  • Roughly right Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Bid target change — when a bid-strategy target value changed (e.g. Target CPA $60 → $50) 32 Roughly right UID 75

Across 32 such changes (3,677 model predictions), a prediction here is roughly right. What that's made of:

  • Roughly right Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Roughly right.

Broken down by kind:

ad-group bid change 4 Too early
maximize conversion value · down 1 Too early
maximize conversion value · flat 3 Too early
switched to Maximize Conversion Value 1 Too early
maximize conversion value · up 1 Too early
maximize conversion value · up large 1 Too early
maximize conversions · down 4 Too early
maximize conversions · down large 3 Too early
switched to Maximize Conversions 7 Roughly right UID 244
maximize conversions · up 4 Too early
maximize conversions · up large 3 Too early
Negative keyword add — when negative keywords were added 135 Reliable

Across 135 such changes (14,898 model predictions), a prediction here is reliable. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Targeting change — when keywords, audiences, placements or other targeting were changed 158 Reliable

Across 158 such changes (17,934 model predictions), a prediction here is reliable. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Broken down by kind:

audience change 2 Too early
bid on a keyword / audience / placement 31 Reliable UID 171
targeting criteria change 37 Reliable
criterion enable 5 Reliable UID 186
criterion pause 24 Reliable UID 124
criterion remove 2 Too early
criterion url change 1 Too early
geo change 2 Too early
keyword add 37 Roughly right UID 225
keyword remove 13 Reliable
negative keyword remove 3 Too early
schedule change 1 Too early
Campaign and ad-group pause / enable — when an existing campaign or ad group was paused or turned back on 27 Reliable UID 225

Across 27 such changes (3,101 model predictions), a prediction here is reliable. What that's made of:

  • Not reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Roughly right Range — how often the real result fell inside the predicted range
  • Reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Broken down by kind:

adgroup pause 3 Too early
campaign pause 24 Reliable UID 225
Ads & assets — when an ad or asset was created, edited, paused or removed inside an existing campaign 18 Reliable UID 230

Across 18 such changes (1,998 model predictions), a prediction here is reliable. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Broken down by kind:

ad create 6 Reliable UID 214
ad pause 9 Reliable UID 254
asset change 3 Too early
Combined changes — when several kinds of change happened together 300 Reliable

Across 300 such changes (31,825 model predictions), a prediction here is reliable. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Not reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

Broken down by kind:

adgroup change + 2 more 3 Too early
adgroup enable + 1 more 1 Too early
adgroup pause + 1 more 3 Too early
adgroup pause + 2 more 5 Roughly right UID 13
adgroup pause + 3 more 1 Too early
adgroup pause + 4 more 2 Too early
adgroup pause + 5 more 3 Too early
adgroup pause + 6 more 3 Too early
adgroup pause + 8 more 4 Too early
adgroup target change + 1 more 1 Too early
ad create + 1 more 4 Too early
ad create + 2 more 9 Roughly right UID 186
ad create + 3 more 6 Reliable UID 69
ad create + 4 more 2 Too early
ad create + 5 more 1 Too early
ad create + 7 more 1 Too early
ad enable + 1 more 2 Too early
ad enable + 3 more 1 Too early
audience change + 1 more 2 Too early
bid strategy change + 9 more 1 Too early
budget change + 1 more 70 Roughly right UID 211
budget change + 2 more 19 Reliable UID 230
budget change + 3 more 6 Reliable UID 225
budget change + 4 more 8 Reliable
budget change + 5 more 5 Too early UID 13
budget change + 6 more 2 Too early
budget change + 8 more 2 Too early
campaign pause + 1 more 7 Reliable UID 239
campaign pause + 3 more 2 Too early
criterion bid change + 1 more 6 Reliable UID 194
criterion change + 1 more 3 Too early
criterion enable + 1 more 2 Too early
keyword add + 1 more 34 Reliable
keyword add + 2 more 13 Reliable UID 1
keyword add + 3 more 22 Reliable
keyword add + 4 more 4 Too early
keyword add + 5 more 3 Too early
keyword remove + 1 more 1 Too early
keyword remove + 2 more 2 Too early
negative keyword add + 1 more 17 Reliable
negative keyword add + 2 more 7 Roughly right UID 34
negative keyword add + 3 more 1 Too early
target value change + 1 more 8 Roughly right UID 42
target value change + 2 more 1 Too early
Other changes — when other changes 4 Too early

Across 4 such changes (462 model predictions), a prediction here is too early. What that's made of:

  • Reliable Direction — called the change up or down correctly on the goal metric
  • Reliable Size — how close the predicted size of the effect was
  • Not reliable Range — how often the real result fell inside the predicted range
  • Reliable Goal — how close on the goal number itself (CPA or ROAS)

Today's leaderboard #1 (UID 160) on this change type: Reliable.

A row reads Too early until at least 500 of its changes have settled. Open any row for the score breakdown and the best model.

Counted: changes to things that already exist, once their outcome has settled.

Not counted: brand-new campaigns & ads (no "before" to score), and recent changes still settling.

The winning model over time

The same winning model, week by week — is the best model on the subnet getting better? Every point is the mean of per-entry scores already published in the daily receipts, so anyone can recompute it.

No scored days have published yet. The first 7-day outcomes settle on 18 August 2026, and the first point appears the week after.

Each point is the mean score of the leading model's predictions that settled in that week, at that horizon. Every score is published per entry in the day's signed receipt, so any point here can be recomputed from the same documents used to verify a day — this chart reports those numbers, it does not compute a separate one. The leading model is shown without identity. A horizon with no settled entries in a week has no point rather than a zero: the 28-day line cannot begin before 8 September 2026.

Where these numbers come from

Every score on this page is the settle-flow score in the daily receipts, published per prediction at /v1/daily — the same numbers that feed the leaderboard. This page and the leaderboard can never disagree, because there is nothing here to disagree with: it is the same data, added up by change type.