B2B SaaS only • $10k/mo minimum See if you're a fit →

← All articles

AI in the ad account

AI Google Ads Management: What It Gets Wrong in 2026

By James Gregg · August 18, 2026

AI Google Ads Management: What It Gets Wrong in 2026

An AI model recently told me my cold outreach was disqualifying me with the exact buyers I most want to reach. Then it explained why — and the explanation was backwards.

“SpyFu cannot see search terms. It sees keywords and estimates.”

AI can analyze a Google Ads account. It can’t manage one. Not for lack of reasoning power — because it reasons about your account from outside it. No access to your CRM, your closed-won deals, or your search terms report. It works from public documentation and an industry corpus that is itself confused about the one distinction that decides whether a B2B SaaS account works: keywords versus search terms. The model didn’t just get that wrong. It got it exactly inverted.

Here’s what it built on top of that premise: telling a Director of Demand Gen “here are some bad search terms at the top of your account” supposedly reads as this person doesn’t know the difference between a bid keyword and a query. Credibility problem, not a conversion problem. Confident strategic advice, resting on a broken model of the plumbing.

Worth unpacking — because the same confusion is probably costing you money right now. I’ll also get to what AI is genuinely good for in an ad account. I use it every week.


What’s the difference between a keyword and a search term?

Two definitions, and the gap between them is where B2B SaaS budgets go to die.

  • A keyword is what you bid on. It lives inside your account. You control its match type, its bid, whether it’s on or off.
  • A search term is what a person actually typed. It’s the touch point between you and a buyer.

You pick the keyword. Google picks the search terms it’ll show you for. The looser your match type, the more of that decision belongs to Google.

Now back to the claim. SpyFu works by running queries against Google and recording which advertisers got served. It cannot see your keywords — nobody outside your account can. Bid keywords are first-party data. They sit behind your login. Every competitive intelligence tool on the market is looking at the query side of the equation, because the query side is the only side that’s visible from the outside.

So the correction was inverted in the one way that matters: it named the single category of data the tool structurally has no access to.

And there’s a second thing the model missed. When SpyFu’s crawler runs a query, Google’s auction doesn’t know it’s a crawler. It renders a real SERP and makes a real ad-serving decision. What you’re looking at is Google’s actual matching behavior — who it chose to serve, for which query. That’s why the diagnostic works and why it never depended on a human having typed the query. If a tool can search “AI” two hundred times and your ad appears on a hundred and fifty of them, your net is too wide. It doesn’t matter who ran the search.

Is SpyFu accurate? Not really. Cost, volume, and click figures are estimates, SpyFu says so itself, accuracy degrades on long-tail terms, and the database refreshes monthly. All of that is fair criticism, and I’d made the same criticism myself before anyone raised it.

But it’s criticism of the estimates, and my method never leaned on the estimates. It leans on the pattern — which queries a company keeps showing up for. That’s the tell. The accuracy critique attacked something I’d already conceded, then smuggled a false category claim in underneath it. Watch for that move. It’s the most convincing shape a wrong answer can take.

One thing the model was right to poke at: a third-party tool’s ranking isn’t literally “the top of your account” — your account is represented by your own search terms report. So the honest phrasing describes what’s visible from outside. The message I actually send says “[Company] shows up most for these terms.” That’s a claim about ad-serving behavior, which is exactly what the data supports.

If you want the operator’s version of this distinction, I wrote a whole search term evaluation matrix for scoring the terms once you’re looking at them.


Why did an AI get such a basic distinction backwards?

Because it wasn’t a hallucination. That’s the part worth sitting with.

Nothing was invented. SpyFu does publish estimates. Google’s glossary does distinguish keywords from search terms. Every input fact was true and citable. What went wrong was the assembly — two correct facts stacked into a conclusion that inverted reality. A category error built from accurate parts. That kind of error survives a fact-check, which is precisely why it’s more dangerous than a made-up statistic.

But here’s the more useful read, and it’s the reason this is a blog post instead of a gripe:

The model didn’t invent the confusion. It averaged it.

If the industry corpus treats keywords as “targeting” and search terms as “a report you check sometimes,” then a model trained on that corpus will hand that back to you in a confident voice. The AI error isn’t the story. It’s evidence that the misinformation is everywhere — and this was the first time I’d seen it stated cleanly enough to screenshot.

Which brings up the structural problem, and this is the actual argument of this post. An LLM reasoning about your Google Ads account is in the same epistemic position as a SERP scraper: outside, looking in. It can’t see your closed-won values. It can’t see which of last quarter’s demos turned into a customer. It can’t see your search terms report. It can reason beautifully about the documentation and still be wrong about your account, because your account is made of first-party data it has never touched.

And the stakes here weren’t academic. The advice riding on that wrong premise was don’t send this — it will cost you credibility. Strategy handed down with total assurance, from a system that had the mechanism upside down. If I’d taken it at face value, I’d have killed an angle that works.


How does Google blur the line between keywords and search terms?

Four mechanisms, and none of them require anyone at Google to be acting in bad faith. Incentives are enough.

1. Match type gets sold as opportunity, not as surrendered control.

Target “construction management software” in phrase, then in broad, and the framing is always that you’re expanding the relevant searches you can appear for. More reach. More volume. What doesn’t get said is that each step outward hands more of the targeting decision to Google. The pitch wants you thinking about the upside and not thinking too hard about who’s now choosing your traffic.

2. Google physically moved search terms out of the targeting UI.

The search terms report used to live under the Keywords tab. It now sits under Insights and reports. That’s not cosmetic — it reclassifies your search terms from a facet of targeting into a report you go read. And search terms are the single most direct evidence you have of who Google is actually putting you in front of.

3. You only see part of the data.

On September 1, 2020, Google limited the search terms report to terms “a significant number of users searched for,” citing privacy. Seer Interactive measured the damage at the time: search-term visibility of cost fell from 98.7% to 71.0%, and of clicks from 98.3% to 77.9% (Seer Interactive).

Google did walk some of it back. From February 1, 2021, advertisers would see “on average 6.5x more queries” — while being explicit that this was not a reversal of the 2020 change. In February 2022, historical data from before the cut that didn’t meet the current threshold was deleted outright.

Now look at the unit in that restore metric. 6.5x more queries. Not impressions. Not spend.

Those are very different things. Add a long tail of one-impression search terms and your unique-query count explodes while the share of impressions you can account for barely moves. The metric Google chose to report is the tell — and it explains something I see constantly: you can scroll a report full of single-impression terms and still not be able to account for most of your traffic.

In my accounts, post-restore, I still typically see somewhere between 20% and 60% of impressions represented in the search terms report. I’ve seen it as low as 10%. That’s my own observation across client accounts, not a published figure — which is exactly why you should go measure your own.

The five-minute test: Open your search terms report. Sum the impressions the listed terms account for. Compare that to the total impressions for the ad group or campaign you’re looking at. The gap is what you’re buying blind.

4. The keyword view has knobs and totals. The search terms view has neither.

At the keyword level you get control — match type, bids, pause, activate — and tidy math: cost per conversion, clicks, conversion breakdowns, summed right there in the column. The search terms view gives you none of that. You just have to trust what’s shown.

Attention follows controllability. The place where the truth lives has no knobs and no totals, so almost nobody looks.


Why is this worse for B2B SaaS than for e-commerce?

Because a loose match is a rounding error in B2C and a structural failure in B2B.

If someone searches “red shoes” and gets served an ad for a blue shirt, that’s not a catastrophe. They might like the shirt. They might buy it on a whim. They might land on the site and find red shoes anyway. Cross-selling off an imperfect match is a completely realistic B2C outcome — and B2C is what most Google Ads features are built for.

Now price it out for your account. Annual contract value in the tens or hundreds of thousands. Adoption by dozens of people. Sign-off from several more. Nobody in that buying process converts off a generic query. Ever.

I have seen the search term “AI” — just those two letters — sitting near the top of a procurement software campaign. That is not a buyer. Even if one in a hundred thousand people searching “AI” happens to be in the ICP, that person is not signing a six-figure contract off that click. No amount of trusting the algorithm changes the math.

The independent data on AI Max lines up with this. In a B2B SaaS project-management account, a same-account match-type comparison found AI Max converting at 0.76%, against 4.67% for exact match — roughly six times worse. The detail that should stop you: broad match beat AI Max on conversion rate and cost per click (1.23% at €0.72 vs 0.76% at €0.89). 68.6% of AI Max’s queries were entirely new, and skewed heavily toward informational searches from people nowhere near ready to buy (Search Engine Land).

That’s one account, and I’d treat it as one account. For scale, Smarter Ecommerce’s Mike Ryan analyzed 250+ campaigns and found a median revenue lift of +13% alongside a median CPA increase of +16% — with a ROAS range running from +42% to −35%, and only 22% of campaigns landing anywhere near their original targets (Search Engine Land). That set is retail, so don’t read it as a B2B verdict. Read it as a warning about variance: the median hides an enormous amount of downside.

And Google’s numbers deserve a hearing. Google reports 27% more conversions at similar CPA versus campaigns built mostly on exact and phrase match. ClickUp — a B2B SaaS company, not a retailer — reported a 15% higher conversion rate, 22% lower CPA, and a 20% lift in incremental conversions.

I believe those numbers. The variable nobody mentions is volume. ClickUp is a category leader running enormous query volume through a mature account. Loose matching works when you have enough conversions to train on and enough budget to absorb the misses. At $5,000–$15,000 a month, you have neither. You get the misses without the training data.

Which raises the fair objection: exact match alone starves Smart Bidding. In a low-volume B2B niche you can’t feed the algorithm the conversion volume it wants, so some looseness is a deliberate trade, not sloppiness.

True — and the answer isn’t to loosen your match types. It’s to stop using the bid strategy that demands volume you don’t have. Max CPC still exists, and in a lot of these accounts it’s the right call. Keep your match types tight, control your bids manually, stop feeding a model that will never get enough data to be good at this.

Then, when you do need to expand: phrase match or AI Max, with audience layering over the top, and real discipline on keyword management. Every new query worth having gets harvested into its own ad group with its own landing page. Expansion becomes something you manage, not something that happens to you.

Match types are only one of the two filters you control, incidentally. The other is the ad itself, which does most of its work by disqualifying people before they ever cost you a click.

That’s also, incidentally, why so many accounts show a healthy CPL and a pipeline that never moves — a gap I’ve written about in why your CPL looks fine but pipeline stalls.


How do you know if your campaign only works because of brand?

Run this test, because the answer is usually uncomfortable.

An accounting practice-management SaaS account we work on had a phrase match campaign that looked fine. Phrase match, worth noting, is basically the new broad — it just sounds precise enough that people stop worrying about it. The ad group had good keywords in it, things like “accounting practice management software.” The campaign was producing demos at roughly $400 each.

Then I pulled the search terms for the previous 90 days and filtered for conversions. Every converting term was brand. Every single one.

We negated brand. That was the only change we made.

Over the next 20 days, that campaign spent $4,000 and produced zero conversions. At its established efficiency, $4,000 should have bought about ten demos.

The $400 cost per demo was never real. It was a blended average subsidized by brand traffic — people who already knew the company, searching for it by name, being counted as demand generation. Negating brand didn’t break that campaign. It revealed that the non-brand half had never worked at all.

Before someone raises it: negating brand changes delivery, and Smart Bidding can re-enter a learning period. Google puts that at up to 50 conversion events or three conversion cycles, and in a low-volume account it can drag on longer than that.

But a learning period explains unstable performance, or worse performance. It doesn’t explain zero. Twenty days and $4,000 of traffic that produced nothing isn’t a model recalibrating — it’s traffic that was never going to convert. There was nothing to learn from.

Here’s the version you can run this week:

  1. Open the search terms report on a broad, phrase, or AI Max campaign.
  2. Filter to terms with conversions. Count how many are brand.
  3. Negate brand in that campaign. There’s no reason to be buying your own name inside a campaign that isn’t a brand campaign.
  4. Watch for two to three weeks.

To be clear, brand isn’t the villain here. Competitors bid on your name and you often need to defend it. Some of our clients run a deliberate brand budget purely to learn what a good buyer looks like, then feed that signal back into targeting elsewhere. Brand campaigns are fine. Brand hiding inside a demand-gen campaign is not, because it makes a failing campaign look successful — and it does that by siphoning credit from the channels that actually created the demand.


So what should you actually use AI for in a Google Ads account?

Plenty. I use it constantly. The line I draw is this: AI is excellent at applying a rubric I wrote, at a scale I can’t match by hand. It is not good at strategy, and I don’t let it exercise judgment.

The clearest example is negative keyword work. I built a matrix for evaluating search terms, and each campaign type has its own grading system across those dimensions. I feed search terms in small batches and the model returns verdicts against my criteria. It surfaces the negatives worth adding, fast — and that’s become more valuable, not less, as Google keeps loosening match types. If you’re expanding into phrase or AI Max because you need volume, this is how you keep it reined in.

Note the shape of that arrangement. I wrote the matrix. The AI applies it. Reverse those roles and you’ve got a very confident intern with no access to your CRM.

There’s a second problem, and it’s the one that made me write this post. Push back on an AI — “really, are you sure?” — and watch how fast it folds. That’s not a good partner in a PPC strategy. The same system that stated the SpyFu claim with total confidence will abandon a correct position the moment you sound skeptical. Confidently wrong when unchallenged, instantly agreeable when pushed. Neither of those is judgment.

And don’t use Google’s AI to grade Google’s traffic. Whatever tooling you use for this, it should be independent. Google’s models carry Google’s incentives, and asking the platform to tell you whether the platform is spending your money well is asking the referee to bet on the game.

The genuinely hard part isn’t the AI, by the way. It’s getting your data into a shape it can reason about — concise, framed, with the right context attached. That’s most of the work.

None of this replaces the operator. It’s leverage on the tedious half of the job, which is exactly where you want leverage. The channel-level decisions still need a human with a P&L in mind — the argument I made in our guide to the B2B SaaS PPC channel mix.


Isn’t Google hiding search terms for user privacy?

It’s a fair challenge and I’ll give it a fair hearing, because I can’t rule it out.

Google’s threshold counts how many people searched a term across all of Google — not how many impressions it got in your account. So a term with a single impression in your campaign can still clear the bar easily if the wider world searches it constantly. If Google is genuinely stripping out queries containing names, addresses, or personal details, I’m all for it. I have no idea how I’d draw that line myself.

But privacy explains why specific rare queries get withheld. It doesn’t explain the aggregate gap.

For most of an ad group’s impressions to be privacy-withheld, your budget would have to be flowing overwhelmingly to queries that almost nobody else on Google is searching. So one of two things is true: either the hidden slice is small — or Google is spending your money on the near-unique long tail. Both can’t be right, and Google doesn’t tell you which one you’re living in.

Worth noting what you can still see: location-based searches, income-related searches, industry searches. The genuinely privacy-adjacent material isn’t what’s missing.

I don’t think this is a conspiracy. I think it’s a line someone drew, and the line happens to sit exactly where it also conceals how low-quality the matching can get. Maybe that’s data storage costs. Maybe it’s a real privacy program that went a little further than it needed to. Either way, the effect on you is identical: you’re accountable for spend you can’t fully inspect.


The AI wasn’t the problem. It was the mirror.

It told me, in a confident and articulate voice, exactly what our industry believes — including the part our industry has backwards. That’s the thing worth taking from this. The corpus is wrong, so the model is wrong, and the reason the corpus is wrong is that the platform benefits from the confusion.

Meanwhile the distinction itself is doing enormous work in your account. It’s the difference between a campaign that generates pipeline and one that quietly resells you traffic you already had.

I still use AI every week. I just don’t let it decide anything I haven’t already taught it how to judge.

If you want a second pair of eyes on what your account is actually showing up for, that’s what our free PPC audit is for.

Frequently asked questions

Can AI manage a Google Ads account?

Not on its own. AI can analyze data, apply scoring rubrics, and draft copy at scale. It can't see your CRM, closed-won revenue, or your search terms report, so it reasons about your account from the outside. Use it for execution against criteria you define — not for strategy or judgment calls.

How do you manage Google Ads with an AI agent?

Give it a human-authored rubric and a narrow job. The most reliable use is search term review: feed terms in small batches against a scoring matrix, and let it return verdicts on what to negate. Keep the agent independent of Google's own AI, which carries Google's incentives.

What's the difference between a keyword and a search term in Google Ads?

A keyword is what you bid on — first-party, inside your account, with controls attached. A search term is what a person actually typed to trigger your ad. You choose keywords; Google chooses which search terms to match them to. The looser your match type, the more of that choice is Google's.

Does SpyFu show search terms or keywords?

It can't show keywords — bid keywords are first-party data that no external tool can see. What SpyFu observes is queries your ads were served on, which is real Google matching behavior. Treat the estimates for cost and volume with skepticism, but the pattern of queries is diagnostic.

Why doesn't Google show all of my search terms?

Since September 2020, Google has limited the report to terms a significant number of users searched, citing privacy. It restored some coverage in February 2021, claiming 6.5x more queries — but that metric counts queries, not impressions. Measure your own gap by comparing search-term impressions to campaign totals.

Is AI Max effective for B2B SaaS?

Usually not, at the budgets most B2B SaaS teams run. In one B2B SaaS account, AI Max converted at 0.76% versus 4.67% for exact match. Google reports 27% more conversions on average, and ClickUp saw real gains — but high-volume accounts can absorb loose matching that a $5–15k/month budget cannot.

Should I negate my brand term in a non-brand campaign?

Yes. Brand traffic inside a demand-gen campaign inflates its apparent performance and hides whether non-brand targeting works at all. Run brand as its own campaign if you need to defend it from competitors — just keep it separate so you can see each one's real contribution.

Want to know what's actually holding PPC back?

Get a free audit with real fixes—not a pitch deck.

If we're not the right fit, we'll tell you—and you'll still leave with useful next steps.