Welcome to the world championship in acronyms and trendy words.
That’s what digital marketing feels like at the moment. Marketers are expected to have a view on LLMs, MCPs, agents, and whatever gets announced next Tuesday, while the advice on offer is either “AI will replace you” or “AI is a bubble.” Neither helps much when you have a budget to allocate on Monday.
The problem isn’t the jargon. It’s that all of it gets sold as one thing. AI. As if a chat window, a bidding algorithm, and a forecasting model were the same product with different logos. They aren’t; they fail in different ways, and picking the wrong one costs real budget.
Which kind of AI does the job?
Never use a complex tool when a simple one is enough. That’s one of the rules of the universe, and there’s no tool more complicated than AI.
For ads, the test is short: does your input affect the output? It does. Add a location, edit an audience, change creatives, and the result changes. Fifty keywords behave nothing like 20,000. So ads need optimizing, and it keeps getting harder. Google is becoming two channels in one: traditional search and AI Mode, with different behavior and different levers. Soon, we’ll optimize for agents as well as humans.
So, with what?
Not with an LLM, for most of it
The hint is in the name. Large language models work on probabilities. They predict the next token, and a token can be a word, part of a word, or a number. They do it roughly the way a weatherman predicts tomorrow’s weather, and better than most weathermen, at least the Swedish ones. But it’s still a prediction, not a calculation. One plus one doesn’t come out the same every time.
For language that’s a feature. Ask Claude for two headlines and you get two good ones, neither right nor wrong. For a budget it’s a problem. Confusing 100 euros with 1 euro is unlikely. Not impossible, but unlikely. Mistaking a 2.39 ROAS for 2.3 is quite likely, and it gets likelier the longer the chat or the project runs. That’s the context window at work.
Two failure modes are worth knowing by name. Both get worse as the context window fills.
Chunking without aggregation
When there’s too much data, the system evaluates in chunks and summarizes each separately. Ask for return on ad spend across Google and Meta, and you can get a 5 on Google, a 1.5 on Meta, and a tidy 3.25 average. If you spend ten times more on Meta than on Google, that number is fiction. You need total cost and total revenue, not the average of two averages.
Subword tokenization
This one is sneakier. As the context window fills, the system stops reading full words and reads the first part, assuming the rest matches. Fine for prose. Less fine for SKUs or keywords in long lists. Read the first six characters of an SKU, ignore the last two, and a lot of different products quietly become the same product.
Neither of these is a bug. It’s what happens when you ask a language tool to do accounting.
Connecting it to your ad accounts doesn’t fix it
Someone usually mentions MCPs here, so: an MCP is a way for an LLM to talk to another platform, Google, Meta, or us. That’s it. No optimization layer inside it, no intelligence of its own. Instead of a recommendation you take yourself, the model performs the action. Useful. Not the same as knowing which action is correct.
Campaign structure, headlines, descriptions, a written voiceover for a table of numbers: yes, that’s the right use. Letting it decide budgets, or judge which assets, audiences, or keywords are performing: a hard no.
And there’s a second half nobody talks about
Most of this conversation stops at the math. Every system you send instructions to has rules. Character limits are the obvious ones. Less obvious is how the action affects that platform’s own optimization: what resets learning, what a pause does versus a delete, how a lead gets attributed in your CRM. Being right is not enough. It has to be correct according to the system receiving it.
An example. We don’t set keyword-level bids. Since automated bid strategies became standard, bidding is Google’s job, so we work on budget allocation instead. For Performance Max, we manage the budget, build the asset groups, and set audience signals and search themes from what Search data actually shows, then let Google’s AI handle its part. Knowing where to stop is part of the design.
On Meta, creative testing looks like this: five ads per test ad set, 80% of the budget to what’s scaling and 20% to testing, ads checked daily for fatigue and paused rather than deleted. None of that is a judgment call made in a chat window.
Where we do use Claude: the voiceover for performance reports, ad copy generated from a crawl of the customer’s site, and a check on whether a keyword matches the landing page before it’s added. All three are language tasks, all isolated, so nothing depends on a context window staying clean.
So use several tools, not one
Which brings me back to where I started. AI is not only LLMs. Machine learning, reinforcement learning, and sequential decision optimization are all AI, all self-learning, and all better than an LLM at the jobs they were built for. Working out which fits where is most of the work.
Claude has become an umbrella tool people reach for by default, and I understand why. Prompting is intuitive, and nothing comes close, so use the LLM as the interface. Let the systems underneath do the calculating and executing.
If an AI test of yours has failed, that’s usually what it was telling you. Not that AI doesn’t work. That you pointed the wrong kind at the job.
There’s a reason Gemini isn’t the thing optimizing Google Ads.
Amanda AI is a deterministic model: the same inputs always produce the same output, so the answer is calculated rather than predicted. It uses a mathematical decision process, published and peer-reviewed, to allocate budget across campaigns and channels, updating continuously as results come in. Inside the Amanda AI platform, Generative AI handles only the language work: ad copy, keyword relevance, and the written summaries in reports.
Want to try us? Book a meeting or create an account at amandaai.com.