Blog

ChatGPT and Perplexity Are Citing PPC Advice. We Checked Whether It Is Right.

Geometric vector illustration of a citation chain of linked nodes, some lit and some faded, on a light field

AI search PPC advice is accurate on mechanics and unreliable on anything carrying a date. The published tests that scored marketing answers put the error rate between 13 and 20 percent, and NP Digital's 600 prompt test found the outright wrong answers clustering on multi part questions, niche domain questions, and anything about a recent change. That last category is where PPC lives most of the time.

One thing up front, because this piece is about accuracy and ours has to hold. We did not run our own lab test of ChatGPT and Perplexity. What follows is an audit of the published record and of the ad platforms' own documentation, with every AI answer I characterize taken from a dated, named test that wrote it down. No invented transcripts, no numbers from an audit that does not exist.

How accurate is AI search PPC advice?

An answer engine is a search tool, such as ChatGPT search or Perplexity, that replies with a synthesized answer and citations instead of a list of links. Three published measurements are worth anything here.

Susie Marino at WordStream put 45 PPC questions to five tools, 225 answers total, and scored them. Her headline is that about 20 percent were wrong, with AI Overviews at 26 percent, ChatGPT 22, Meta AI 20, Perplexity 13, and Gemini 6. Worth saying, since accuracy is the subject: those five rates come to 40 wrong answers of 225, which is 17.8 percent rather than 20. Carry it as "about one in five, on WordStream's own scoring." She ran the same exercise on 50 SEO questions and got 13 percent wrong or misleading, then wrote the comparison herself: PPC scored materially worse than SEO in the same author's hands.

NP Digital published the bigger one in February 2026, a survey of 565 US digital marketers alongside a 600 prompt accuracy test across six models, graded by humans as fully correct, partially correct, or incorrect. ChatGPT came out highest at 59.7 percent fully correct with 7.6 percent outright wrong, Claude had the lowest error rate at 6.2, and Grok was worst on both ends at 39.6 and 21.8. Those fully correct numbers look alarming until you notice they do not sum with the error rates, because the remainder is partially correct, which is what a colleague gives you most days.

Outside marketing, the EBU and the BBC had journalists grade more than 3,000 AI answers about news, and found a fifth carrying major accuracy issues they describe as hallucinated or outdated information.

What the engines get right, and it matters that they do

Gemini got 42 of 45 PPC questions right in WordStream's test. That is a good score, and it does not travel: the EBU and BBC graders put Gemini last on news, with 76 percent of its responses carrying significant issues. Two named studies, opposite rankings, which is the finding rather than a contradiction somebody needs to resolve. Meta AI was the only tool of the five to get every average Facebook Ads cost and performance question correct, a nice illustration that the accuracy domain follows platform ownership. Claude's misses in the NP Digital test were mostly omissions rather than inventions, the error type you can catch on review.

And when Marino asked for a Google Ads script that pauses high CPC keywords, ChatGPT, Perplexity and Meta AI all produced one. Gemini declined on security grounds, and AI Overviews returned a sentence explaining that Google Ads scripts use JavaScript, which is true and also not a script. Google's two tools were the two that would not help.

The fundamentals hold up. Ask an engine what exact match actually matches in 2026 and you will get a usable answer.

Google changed its own headline number, and the corpus has not caught up

On May 6, 2025, Brian Burdick, Senior Director of Product Management at Google Ads, published the AI Max launch post. The claim: advertisers who activate AI Max in Search campaigns "will typically see 14% more conversions or conversion value at a similar CPA/ROAS," rising to 27 percent for campaigns still mostly on exact and phrase keywords. The footnote almost never travels with the number, and it is load bearing: Google internal data, 2025, based on campaigns with more than 70 percent of conversions from exact or phrase match keywords, non Retail.

Google's current published figure is different. On its own Accelerate announcement page, published May 13, 2026, the sentence reads: "AI Max for Search campaigns already using search term matching see an average of 7% more conversions or conversion value at a similar CPA/ROAS by enabling text customization and final URL expansion." Footnote: Google internal data, Global, 2026, non Retail.

Those two numbers are not a correction of each other. The 2025 figure compares AI Max on against off, and the 2026 figure compares the full feature suite against search term matching alone. Both are Google's, and one has had sixteen months of blog posts and agency explainers built on top of it while the other sits on a page almost nobody has opened.

The independent read belongs next to both. Mike Ryan at Smarter Ecommerce analyzed 250 plus Search campaigns, reported by Anu Adegbola in Search Engine Land in March 2026: median revenue up 13 percent, median CPA up 16 percent, a ROAS range running from plus 42 to minus 35, and up to 63 percent of the time broad match expansion was recycling coverage the account already had rather than finding new queries. Ryan's verdict: "Turning on AI Max is essentially a coin toss: you may see a lift, but efficiency likely won't follow."

The ground is still moving. Google emailed advertisers on August 5, quoted by Search Engine Roundtable and reported by Search Engine Land and PPC Land, that from September 1, 2026 campaigns using automatically created assets or the campaign level broad match setting are upgraded to AI Max automatically. No Google blog post announces it, which is why I am naming the reporting chain.

AI answers are averages. Your account is not.

Query level data is performance measured at the individual search query, underneath the keyword and campaign averages the ad platforms report by default. Most bad AI PPC advice is a platform average answer to a query level question, and Marino's test caught it twice.

She fed a poor performing campaign to the tools: 13 clicks, 1.19 percent CTR, $48.66 average CPC, 1,094 impressions, zero conversions, $632.59 in two weeks. Meta AI called the CTR "slightly below the average CTR for Google Ads (around 1.5-2%), but it's not terrible." WordStream's own benchmark put average Google Ads CTR at 7.52 percent. That range was true in a prior era and it survived into the answer with the confidence intact.

Then the CPL question. Correct answer, per WordStream's Facebook benchmark report, $21.98. Perplexity and AI Overviews both answered $10 to $50, and ChatGPT said $5 to $30. Both ranges contain the right number and neither will help you set a budget. ChatGPT separately put CPC in expensive industries at up to $300, which Marino flags as way too high.

Why a point estimate is the wrong shape of answer shows up in WordStream's own 2026 benchmark report, built on over 13,000 search campaigns across 23 industries running April 2025 through March 2026. Average CPC $5.42, ranging from $1.63 in Arts and Entertainment up to $9.87 in Attorneys and Legal Services. Average CPL $66.69, from $26.84 up to $131.63 in those same two verticals. A universal CPC figure is wrong for almost every account that asks, and no amount of model improvement fixes that, because the question has no single right answer.

I should declare the obvious. SEMoptimize sells query level optimization, so we have a commercial reason to argue the query is the unit of truth. WordStream's test concludes you should rely on third party marketing partners, and WordStream sells being one. NP Digital sells AI assisted marketing services. Discount all three of us by the same amount.

Google documents the visibility limit itself, which beats any vendor's estimate of it. Its search terms report help page says the report shows terms "used by a significant number of people," and that terms without enough query activity are omitted for privacy and aggregated as "other queries." No number is published for what counts as significant.

Why do AI engines get PPC advice wrong?

NP Digital named three prompt types that broke every model in the test: multi part prompts, recently updated or real time topics, and niche domain specific questions. A normal PPC question is usually all three at once. "Should I turn on AI Max for a legal services account running tROAS" is multi part, recent, and niche, and it wants a number attached.

The four error types they catalogued are fabrication, omission, outdated information, and misclassification. In enterprise paid search the expensive one is outdated information, because it arrives in the same register as a correct answer. A 2024 benchmark presented as current looks completely normal.

Should you trust ChatGPT for PPC recommendations?

Yes, for orientation, vocabulary, first draft structure, and for explaining a setting you have not used before. The rule that holds: never act on an account specific claim generated without account data, including anything an engine tells you about your own visibility in AI answers.

The five minute pressure test is four questions. Ask for the date and source of any number, and check the source exists. Ask whether the figure is for your vertical, because a $5.42 average CPC means nothing to a legal account. Pull your search terms report and see whether the recommendation survives the queries you actually received. And ask a second engine, since NP Digital measured a 20 point spread in fully correct rates across six models on the same prompts, 59.7 at the top and 39.6 at the bottom.

Questions people ask about AI PPC advice

Is PPC advice from ChatGPT accurate? Mostly. WordStream scored ChatGPT wrong on 22 percent of 45 PPC questions in 2025, and NP Digital scored it the most accurate of six models on a broader 600 prompt set later that year, at 59.7 percent fully correct and 7.6 percent flatly wrong.

What is ai search ppc advice accuracy, and how do you measure it? It is the share of PPC recommendations from an answer engine that survive checking against a known correct source. Everybody measuring it so far has graded against published benchmarks or platform documentation, a reasonable proxy and not the same as grading against outcomes in a live account.

Which is more reliable for marketing advice, ChatGPT or Perplexity? The two published tests disagree, which is the useful finding. WordStream had Perplexity ahead on PPC questions, 13 percent wrong against ChatGPT's 22. NP Digital had it the other way, ChatGPT at 59.7 percent fully correct against Perplexity's 49.3. Rankings flip by domain and by year, so a per model league table is not a decision input, and neither knows your account.

Why does AI give outdated Google Ads advice? Because the public corpus is full of the old number and the platform keeps changing the new one. Google's 14 percent AI Max uplift claim from May 2025 was written about for over a year, and its current 7 percent figure, against a different baseline, has been up since May 2026.

The engines won a round here, and it deserves saying plainly. On mechanics, definitions and scripts they are good, measurably, and better than the "AI is confidently wrong about everything" line suggests. Where they cost money is the number with a date on it, delivered in the same calm voice as everything else. If an AI answer just prompted a strategy question in your account, your search terms report settles it faster than another prompt will.

All insights and articles

Reading about it is the easy part

Stop guessing. Start knowing your growth potential.

Get a free, no obligation audit of your paid search account, real math from your own query data, not a sales pitch.

The Optimal Path to Conversion

With precision bidding strategy, delivering the right Keyword/Match Type, to the right Ad Message, to the right Landing Page.

Wasteful Themes Pruned - Negatives
See how queryDNA works →