Blog

Every AI PPC Copilot Is Reading a Redacted Report. Here Is How to Benchmark One Yourself.

AI PPC tools benchmark 2026: a Google Ads search terms report with part of the query column redacted, next to a copilot recommendation panel.

Every AI copilot pointed at a Google Ads account reads the same privacy filtered search terms report that you and I read, and not one of them prints how much of your spend that report leaves out. I am not accusing anybody of lying. They summarise a document with a hole in the middle of it, the hole is a different size in every account, and because it varies that much no published figure will be true of yours. So the only benchmark worth running is one you run yourself.

What can an AI PPC copilot actually see in your account?

It sees the rows Google chose to disclose and nothing at all about the rows Google withheld. The search terms report only lists queries clearing a volume threshold, described in the help page posted by Pallavi Naresh on September 9, 2021 as a way to "ensure user anonymity by only reporting on terms that have seen sufficient search volume across all Google searches" (Google Ads Help). Google has never published the number that counts as sufficient, so you cannot calculate what your account ought to be missing, which is why everything below is a measurement procedure, not a lookup table.

The spend attached to those unnamed queries does not leave your account. It sits in the campaign cost totals, it fed whatever the bidding system learned that month, and it carries no query label, so it cannot be n-grammed, cannot be negated, and cannot be reasoned about by anything whose input is that report. A copilot's input is that report.

Seer measured this once in 2020, so we went and measured our own

Seer Interactive put search term visibility by cost at 98.7% on August 31, 2020 and 71.0% on September 1 (Seer Interactive, September 2, 2020). Their sample was "all of our clients" with no account count given, and six years on nobody has replicated it, which is most of why we went and pulled our own numbers instead of quoting theirs. Seer also reported mining 5.1 million data points across 30+ businesses and finding 15% of budgets going to mostly hidden low volume queries that were not converting, and that 15% is still the only published estimate of the hidden tail's size, from 2020, unreplicated.

So: three accounts SEMoptimize manages, window fixed at May 1 to July 31, 2026 before a single figure was read, same method on all three. Visible share of Search spend came out at 87.7%, at 68.4%, and at 15.9%, pooled at 77.5%. Three accounts from one agency's book over one quarter is exactly the criticism I just made of Seer, which is why I am giving you the spread rather than the average. One account could see nearly everything it was buying and another about a sixth of it, under the same threshold, in the same quarter, managed by the same team.

We ran the ratio on clicks too, as a check that the cost figure was not a few expensive hidden terms distorting things, and the two agreed closely on the two accounts where we had it. On the third we closed the view before capturing the Search click total, which happens when you click a reporting UI at speed.

The account with the least query volume was the most redacted by a wide margin, and the mechanism is straightforward: an absolute threshold gets harder to clear the less volume you have, so a small account's queries individually fail it while a large account's clear it comfortably. That was the hypothesis when the accounts were picked and it held, on three accounts, which makes it an observation rather than a law. Three months is also not long enough to rule out something seasonal behind the 15.9%. The numerator is Google's own Total: Search terms subtotal, so it inherits whatever Google does to that row.

The machines win the volume problem outright

They win it and it is not close. Adalysis aggregates one, two and three word n-grams at campaign and account level, and its auto resolve setting, which the docs call your search terms autopilot, sits at Audit > Audit settings > Prebuilt > Search terms (Adalysis docs). Opteo's N-Gram Finder runs the same analysis across Search, Performance Max and Shopping at open an account > Toolkit page > N-Gram Finder, and previews which keywords and search terms a proposed negative would catch before you add it (Opteo changelog). No human reads a million search terms.

The best evidence against my argument comes from a vendor. Optmyzr looked at 17,380 accounts running at least 90 days and found 95% do not accept Google's suggestions, and that the 333 accounts which both accepted them and held a 90 to 100 Optimization Score had the best performance in the study, with the 90 to 100 band beating the sub 70 band on ROAS by 186% (Optmyzr, August 23, 2024). My reading is that the causal arrow runs through active management rather than the recommendations, because an account somebody is working on gets both a high score and good performance for the same underlying reason, and that reading is mine and their data cannot prove it, so take it as the hypothesis it is. Optmyzr's own conclusion is that you should not dismiss Google's suggestions out of hand, and on their numbers that is fair.

Why do your tool's search term numbers differ from Google's?

Because at least one vendor filters the report a second time before you see it, and says so on the page. Adalysis is explicit: "The Google Ads report will show you search terms even when they've been blocked by a negative keyword. The Adalysis report is different: we only show you actionable search terms. This is why your search term numbers will be different in Google Ads versus Adalysis." A blocked search terms toggle puts the filtered set back. I have no complaint about the design, and documenting it puts Adalysis ahead of Google, which has never documented the size of its own filter. The consequence is two subtractions between the auction and the recommendation, so compute your visibility ratio inside a tool and you get a different denominator with no way to tell which layer moved it.

The Opteo page describing N-Gram Finder carries no publication date anywhere on it, which is why I opened an account and confirmed the menu path existed rather than writing from a changelog alone.

Google's own documentation marks where the reporting stops

Google documents one of its own blind spots in writing, on the AI Max reporting page: "When filtering for match type = 'AI Max' the search terms report may show lower numbers for AI Max. This filter doesn't consider 'Other search terms'" (Google Ads Help). The same page gives the keywords report two aggregate rows at the bottom, Total: AI Max expanded matches and Total: AI Max landing page matches, the second defined as "Total traffic from search queries that matched because of your landing pages or assets, outside of your keywords." An aggregate row cannot be negated, so it is unusable as the input to a negative keyword decision, and the same page tells you to use negative keywords sparingly. Starting February 2027, campaigns using Dynamic Search Ads get automatically upgraded to AI Max, confirmed by Search Engine Land on June 23, 2026, and a stray search result I hit saying September 2026 is wrong.

Performance Max search terms did start appearing in the standard report on a gradual rollout first spotted by Hana Kobzová in March 2025. That article says the terms appear. It does not say the reporting is complete, and no share of hidden PMax queries has ever been published, so PMax is somewhere I would run the ratio rather than assume parity.

Google's Ask Advisor, which SEL describes as keeping advertisers in control of campaign decisions, and Microsoft's Copilot page, dated 16-07-2026, both document diagnostics and performance analysis, and neither claims query level search term analysis, so I am not going to tell you they do one.

Why does Ad Strength keep turning up in the recommendations you receive?

Because Ad Strength measures structural completeness rather than results, which is Optmyzr's own conclusion after looking at roughly 20,000 active accounts: it "reflects structural completeness ... not how well your ad will actually perform" (Optmyzr, April 6, 2026). Ads rated Average came in at a $12.43 CPA. Ads rated Excellent came in at a $28.68 CPA and a 4.97% conversion rate, the worst pair in the set. Ads rated Poor produced the best ROAS at 327.65%. Optmyzr says plainly that this is observational rather than controlled, that their customers skew toward experienced practitioners, that some subsets are tiny with the headline casing cut down to 31 accounts, and that results are directional and not population wide absolutes, and taking that seriously is why I use the study to stop reporting Ad Strength as a KPI in client decks rather than as a licence to pin everything.

Optimization Score has its own problem. Jyll Saskin Gales, six years at Google, writes that "Dismissing a recommendation gives you the exact same score uplift as applying it", and that Recommendations "actually started as an internal sales tool for Google Ads sales representatives ... Now that Recommendations surface automatically in every account, that human filter is gone" (Search Engine Land, December 10, 2025). That is commentary from a named practitioner rather than a study with a sample size.

Google's remove redundant keywords recommendation auto applied and removed Optmyzr's own brand keyword on January 12, 2023, eight days after Google widened that recommendation's scope to all match types, and the keyword went from 7 conversions to 0 while CPC rose by more than 130% (Optmyzr, February 20, 2023). One incident, on the vendor's own account, three years ago, is not a failure rate and I will not turn it into one, and it is still why I check those boxes on accounts we inherit.

The protocol: benchmark one copilot against your own search terms

Six steps, in Google Ads, read only. Nothing here changes a setting.

1. Fix the window before you look at anything, and write it down. Three complete calendar months is stable enough to trust and recent enough to describe your current account. Do not adjust it once you have seen a result, because a window tuned after the fact will show you what you wanted.

2. Get the visible spend. Search terms report, your window, the Total: Search terms row, Cost column. That row sums only the disclosed terms. It is your numerator.

3. Get the total Search spend, and this is where the arithmetic goes wrong. Campaigns report, same window, filtered to Campaign type = Search, Cost column. Both sides of the division have to be restricted to Search, because a search term can only ever attach to a Search campaign, so leaving Shopping, Performance Max, Display and Video spend in the denominator manufactures a redaction rate that is really just a picture of your campaign mix. Adding channel=1 to the URL applies the filter reliably, which matters because the UI filter is easy to lose between the two reports.

4. Divide, then cross check on clicks. Numerator over denominator is your visible share, and the remainder is spend Google took against queries it will not name. Run the same ratio on the Clicks column. If cost and clicks disagree badly, your hidden spend sits in a few expensive terms rather than a long tail, which changes what you do about it.

5. Ask your copilot for its search term work on the same window and sort every recommendation into three piles. Confirmed means you can point at a disclosed term that supports it. Contradicted means a disclosed term argues against it. The third pile holds recommendations that neither the disclosed terms nor the aggregate rows can speak to, and it should come out roughly the size of your hidden share. Much smaller and the tool is being more confident than its input allows.

6. Write your visible share at the top of the report, and re-run it quarterly. A recommendation arriving without that number beside it is missing its own margin of error. Re-run after enabling AI Max, after a campaign type shift, and before February 2027 if you still have Dynamic Search Ads running.

What a passing score looks like

A copilot passes when it tells you its own coverage before it tells you what to do, and I have not yet seen one that does. Short of that, a tool passes on your account when its unverifiable pile is about the size of your hidden share and the other two piles survive a look at the underlying terms. A dozen crisp recommendations on an account where you can see 15.9% of Search spend tells you about the tool's defaults.

This measurement is why we built our queryDNA approach around query level structure rather than platform scores, and the same reasoning runs through what we have written on over-reliance on Google Ads automation and through our case studies. Smart Bidding, PMax and AI Max are where the volume is and refusing them is not a strategy. Knowing the size of the gap you and your tools are both working across, per account, with a date on it, is.

FAQ

What percentage of Google Ads search terms are hidden in 2026? Nobody can tell you, because it depends on the account and no current public measurement exists. The three accounts we measured from May 1 to July 31, 2026 came in at 87.7%, 68.4% and 15.9% visible by Search spend, pooled at 77.5%, and the honest use of that range is as evidence the number varies enough to be worth measuring yourself. Figures like "roughly 40% hidden" trace back to aggregator blogs rather than any primary measurement.

Does the Google Ads API return more search term data than the interface? I do not know, and I looked. The search_term_view reference documentation was too large for me to retrieve in full and I will not assert an equivalence I did not verify, so treat any tool's claim of extra coverage as something to test with steps 2 through 4.

Can I trust AI PPC copilot recommendations? Trust them the way you would a well briefed analyst who has read part of the file, which means finding out which part. Their query level reasoning inherits your redaction rate and they will not print it, so once you know your own number you can size the blind spot without disputing any single recommendation.

Do Performance Max search terms appear in the search terms report now? Yes, since a gradual rollout in March 2025, and you can add negatives directly from the report. Whether that reporting is complete has never been documented by Google, so run the ratio on PMax separately.

Should I turn off auto apply recommendations? I turn it off on every account we take over and review by hand instead, and the reasoning is that dismissing gives the same Optimization Score uplift as applying, so the score costs you nothing either way. The boxes are at Recommendations tab > All Campaigns view > Auto-Apply Settings. The only documented case of harm I can cite is one 2023 incident on a vendor's account, which makes this a cheap precaution and not a response to a measured failure rate.

Justin Schroder is CTO and co-owner at SEMoptimize, a Denver PPC agency.

All insights and articles

Reading about it is the easy part

Stop guessing. Start knowing your growth potential.

Get a free, no obligation audit of your paid search account, real math from your own query data, not a sales pitch.

The Optimal Path to Conversion

With precision bidding strategy, delivering the right Keyword/Match Type, to the right Ad Message, to the right Landing Page.

Wasteful Themes Pruned - Negatives
See how queryDNA works →