History Frozen at publish · superseded for credibility
An early check on made-up example data
2026-05-29 · 5 questions · made-up example accounts
Scope
- Date
- 2026-05-29
- Scope
- 5 questions (one per category) · made-up example accounts
- Setups
- bare Claude (plain and strict) · Google Ads MCP · AdTAO MCP
- Model
- Claude Sonnet 4.6
- Type
- First check across categories, small and on made-up data, for direction not headline numbers
- Status
- Frozen at publish · kept as history · superseded for credibility
What it covered
Five questions a manager would ask
The five questions each stood for a job you do every week, with the same prompt sent to every setup:
- Diagnose a drop: is it CVR, wasted spend in the search terms report, or a change in the change history?
- Make a change: adjust a bid strategy or budget, mindful of the learning period.
- Compare performance across accounts on CPA and ROAS.
- Judge a CVR against a benchmark for a niche business.
- Run a safety check: add negative keywords at the right match type without cutting impression share on terms that convert.
The fair, durable result
Carried forward to the comparison overview
It would not invent a benchmark
One question asked whether a conversion rate was good for a niche business, where no reliable benchmark exists. It is the one question here that is fair to every setup, because it does not depend on reading a specific account.
Bare Claude invented a conversion-rate benchmark:
"Verdict: Good… service businesses typically 2–5%, you're in the upper range."
No such benchmark exists.
AdTAO refused honestly:
"Corpus too thin (fewer than 50 accounts) to give a confident percentile,"
Instead of inventing a CVR figure, it pointed to the real checks a manager can run: the CVR trend in your own conversion tracking, wasted spend in the search terms report, and impression share for the segment. A made-up benchmark is the kind of confident wrong answer that moves a bid strategy the wrong way.
This contrast is carried forward as the evergreen example on the comparison overview.
Honest limits
Honest limits of this run
-
Made-up example accounts can't fairly test tools that read your real account.
AdTAO and Google's MCP both work against live Google Ads accounts; handed a made-up one, AdTAO correctly asks which account rather than inventing an analysis of a search terms report it cannot see. So the account-specific questions here under-represent both, and only the benchmark question is a clean comparison.
-
Google MCP leg ran without Google credentials.
Its response times and step counts reflect a degraded setup. Treat them as indicative only.
-
5 questions is a quick check, not a credible suite.
The complete programme is much larger. This run was published for direction, not as a headline benchmark.