What happens when AI gives SEO strategy advice?

by Mihajlo Bunčić ⏐ July 8 2026

Every trend the SEO industry has hyped is sitting in an AI model’s training data. We measured AI’s tendency to recommend trending SEO strategies over advice that more accurately fits specific business situations. The results: AI is not immune to hype.

Earlier this year, a marketing director copied three months of falling organic traffic data into ChatGPT and asked whether to pivot her budget toward AI search. The answer came back in seconds: “Yes, pivot now. AI search is eating traditional clicks, and companies that wait will fall behind.” It even sketched a migration plan.

AI was wrong. The traffic decline affected only informational pages that had never generated any revenue. Her money pages were steady. Conversions were up. 

The crisis ChatGPT diagnosed was a rounding error dressed up as a trend, and the pivot it recommended would have moved budget away from the content that was actually working well.

ChatGPT wasn’t wrong because it was dumb. It was wrong because the marketer’s data never said which pages were losing clicks — and rather than ask, AI filled that gap with the loudest trend in the industry. And it did so with total confidence.

That gap between confident advice and correct advice is what we set out to measure. 

We run an SEO agency, and our clients’ marketing teams ask AI strategy questions every day, so we tested it on our own field against 4 models, 6 real decisions, more than 3,400 measured AI responses.

Our Methodology

We picked 6 SEO strategy decisions based on real-world scenarios. For each one, a trendy strategy competes with a less popular strategy:

  • Refresh old content, or publish new content
  • Optimize for AI search, or traditional Google search
  • Build a Reddit presence, or invest in your own site
  • Improve “AI visibility score,” or referral traffic
  • Generate content with AI, or keep humans in the loop
  • Add schema markup, or fix the fundamentals

For each strategy decision, we prompted AI with realistic situations brands encounter everyday. But with a catch. In most situations, the trendy strategy was the wrong call based on the background described in the scenario.

Every scenario needed a verified right answer before we could test the models against it. So two senior SEO strategists on our team reviewed each one independently. They worked separately to determine the best strategy decision for each scenario. They agreed on a correct approach for 23 of the 25 scenarios. 

One scenario was discarded outright because the strategists couldn’t agree on the right answer. A second scenario was flagged by both reviewers as unclear, not wrong, and got rewritten rather than dropped. More on that one below.

Each model got a single prompt asking it to assign confidence scores to both strategies, splitting 100 points between them instead of picking just one. This approach made the models’ biases more visible even when a model chose the right answer. 

The scores showed how confident it was in the trendy option, not just which way it leaned. And because a single AI response can vary from run to run, we ran every situation through each model 20 times and averaged the results, so we were measuring a real tendency, not a fluke. In total: more than 3,400 measured answers.

Without context, AI defaults to hype — mostly

First, we prompted each model to pick an SEO strategy with zero client context (no industry, no business goals, no background).

Simple prompt: “Should we refresh old content or publish new content?”

The trend bias showed up immediately. With nothing to reason from, every model leaned toward refreshing old content. Asked a similar no-context question about AI-generated content, one model put 88% of its effort behind “keep humans in the loop”. The safe, popular answer, even with no evidence that this strategy was actually the right call.

The bias wasn’t universal. Schema markup may be the most hyped tactic in AEO circles right now. But the models defaulted against it. Hype in the market doesn’t always translate into bias in the model.

Figure 1. The trendy default is real — but uneven

Give the models real facts, and the bias mostly disappears

Then we added the situations, and the picture changed. On five of the six decisions, the trend bias collapsed once the models saw real facts. On AI content, the default of 79% dropped to 21% when generating at scale was the right call. On content refreshes, 59% fell to 21%.

This is good news, and it refines the original finding rather than contradicting it. That study gave models thin context: a company’s size and industry. We gave them the details a consultant would demand. With enough facts on the table, most models stopped chasing trends and started reasoning.

Figure 2. Give the models real facts and the bias mostly disappears

One decision refused to collapse: GEO versus classic SEO. Its trendy pull barely moved. When we dug in, the entire gap traced to a single scenario — and that scenario became the most interesting result in the study.

The theory that failed

We had a hunch about what drives the remaining trend-chasing: fear. SEO content is soaked in panic right now — “AI is stealing your clicks.” Maybe fearful wording pushes models toward the trendy fix.

We tested it properly. We took the same situations and wrote each one three ways: neutral, threatening, and opportunity-framed. Every fact stayed identical, verified byte for byte. Only the emotional wording changed.

The result was flat. Neutral wording scored 39%. Threatening wording scored 42%. Opportunity framing scored 45%. On the key scenario, adding fear to identical facts moved the answer by less than one point. Our theory was wrong, and we dropped it.

Figure 3. Fear didn’t move the needle

What actually triggers it — and the model that would not move

The strategist review pointed us to the real answer. Both reviewers, working blind to each other, flagged the same GEO scenario as unclear. It described a scary-looking drop in click-through rate, but never said where the drop was happening — on pages that make money, or pages that don’t.

So we rewrote the scenario with that gap closed. The facts now said plainly: the drop sits on low-value pages, the money pages are fine, and the business is not declining. Then we ran it again.

Three of the four models corrected. Llama fell from 60% on the trendy option to 20%. Mistral fell from 61% to 23%. Qwen dropped 18 points. Their “trendslop” was never blind trend-chasing. It was a model filling a gap in the data with the trend.

One model did not budge. GPT-4o put 69% of its effort toward the trendy option on the unclear version — the same answer in all 20 runs, without a single wobble. After we made the facts unambiguous, it still averaged 71%. The other models needed clarity and used it. GPT-4o did not.

Figure 4. Clear facts fixed three models. GPT-4o didn’t move.

What this means for your team

Four rules fall straight out of the data.

Never ask a strategy question without context. The no-context answer is the trendslop zone. Every model showed its strongest bias when given nothing to reason about. If your prompt could apply to any company, expect the trendy answer.

Ambiguous requests are the danger zone. Models fill gaps with the trend. Before you ask, bring the numbers a consultant would demand: where the drop is, which pages convert, what the trend has actually cost you. In our test, closing one gap in the facts cut the trend-chasing by more than half for three of the four models.

Confidence is not a signal. The most wrong answer in our study was also the most consistent one. GPT-4o repeated its answer 20 times out of 20 without variation. Consistency measures conviction, not correctness.

Get a second opinion. The models disagreed most on exactly the scenarios where judgment mattered. Run the same question through a second model, or better, past a human who knows your business. Where the answers split, you have found a judgment call, not a fact.

One more reassuring note. On the decision closest to a vanity metric — chasing an “AI visibility score” versus real referral traffic — the models mostly got it right. Given real data, they picked traffic and conversions over the score. The machines are not the only ones tempted by shiny numbers; on this one, they resisted better than many teams do.

The marketer who almost pivoted her budget

Go back to the marketing director I mentioned at the beginning of this article. Her prompt never said which pages on her website lost traffic. ChatGPT, faced with that gap, recommended our industry’s most hyped SEO strategies rather than asking clarifying questions.

We saw the same pattern across all 3 AI models and all 3,400 responses in our study. The AI models aren’t broken. They’re guessing the most relevant strategy based on the information provided. 

Give them very little background on the situation, and they recommend the most hyped SEO strategy. Give them the full picture, and they almost always agreed with our experts on the correct strategy.

The unsettling part isn’t that AI was often wrong. It’s that AI was confidently wrong. GPT-4o recommended the trendy strategy at 71% confidence even after we clarified the facts and made the correct answer obvious. Twenty runs out of twenty. The other models corrected. GPT-4o didn’t budge. That’s the real risk for teams leaning on AI for SEO strategy: not that it will get things wrong, but that it will get things wrong with confidence.

The good news buried in these numbers: give AI the specifics a consultant would demand, and on most SEO strategy recommendations, it stops chasing trends and starts reasoning. The AI tools we tested work well for strategy only when they have the full picture.

About this study

We tested four models from four makers: Qwen 3.6 (Alibaba), Llama 3.3 70B (Meta), Mistral Small (Mistral), and GPT-4o (OpenAI). These are the free-tier-accessible versions available at test time, not necessarily each maker’s newest release. Every scenario’s “right answer” was reviewed independently by two senior SEO strategists. One scenario was excluded after the reviewers disagreed; one was re-tested with clarified wording, as described above. Answers used a 100-point split of effort, with 20 repetitions per model per scenario and identical settings across models. The full methodology, every prompt, and all raw data are available on request.