Skip to main content
Skip to content
Back to Blog
Guide

Why Internal Linking Matters More in AI Search

Mathias Decourt
Mathias Decourt·Co-Founder - CEO
July 23, 20268 min read

In May 2026, Google published its first official guide on optimizing for AI features. Its position on the acronyms the industry spent two years building: optimizing for generative AI search is optimizing for search, and that is still SEO.

Most people read this as bad news.

It isn't. If AI features run on the same ranking systems as normal search, then every proven SEO lever already works for AI search. No new playbook needed.

But not all of SEO carries over. Some of it has now been tested, and two of the loudest tactics failed.

Two tactics sold hard as AI-specific have failed controlled tests. Internal linking is not one of them. What carries over is what controls access to the index and how precisely a page answers a question. What doesn't carry over is what only ever signalled site reputation.

Take schema markup, the code that describes your page to machines.

An analysis of 6 million URLs found that AI-cited pages were three times more likely to use it. That number circulated for a year as proof. So Ahrefs tested it properly, tracking 1,885 pages that added schema against 4,000 matched control pages:

  • +2.4% on Google AI Mode, indistinguishable from noise
  • +2.2% on ChatGPT, same
  • −4.6% on AI Overviews, small, and the authors won't blame schema for it

Their explanation for the original correlation matters more than the result. Schema tends to sit on well-maintained sites. Those sites also write better content and rank better. The markup was a symptom, not a cause.

Google says the same thing from the other side. Its guide states that Search doesn't use llms.txt files, and that structured data isn't required for AI search.

One limit to that study almost nobody mentions. Every page in it already had 100+ AI Overview citations before the test began. So it measures whether schema helps pages that are already visible. It says nothing about helping a page become visible.

Two different questions. The difference runs through the rest of this article.

Here is what the evidence supports:

SignalCarries over to AI search?Why
Discoverability through HTML linksYesDecides if the page enters the index AI pulls from
Retrieval positionYesStrongest measured predictor of citation
Page-to-question matchYesStrongest measured content signal
Backlinks and domain authorityIndirectlyNo direct link to citations, but they still drive rank, and rank drives citation
Schema markupNoNo effect in causal testing

The backlinks row needs care. In the AirOps data, backlinks and domain authority showed no positive correlation with citation.

That is not "backlinks are dead." They still build the rankings that put a page in front of the AI. What died is the straight line from backlink to citation.

What Actually Decides Whether You Get Cited

Two signals dominate: where your page lands in the AI's results, and how precisely it answers the question. Covering more ground does not help. That has a direct consequence for how you structure content, and it contradicts a decade of advice.

The largest public study of what happens between a query and a citation looked at 16,851 queries and 353,799 pages inside ChatGPT's retrieval process. The results, summarized by SparkToro:

  • 58.4% citation rate for the first result retrieved, dropping to 14.2% by position 10
  • 26–50% coverage of sub-questions beats 100% coverage, once relevance is held constant
  • Two sub-questions is what most prompts generate, not the five or six people assume

The authors put it bluntly: a page that nails one question beats a page that adequately covers five.

This is not permission to split content into one page per sub-question. Google's guide names that directly, warning that creating separate content for every search variation falls under its scaled content abuse policy. The test is whether a real reader has that question, not whether a tool listed it.

Why precision beats breadth

There is a technical reason for this that hasn't reached the SEO world.

Researchers at Michigan State and MIT traced what happens inside a language model while it answers questions from retrieved documents, then compared correct answers to wrong ones. Their May 2026 paper found a consistent pattern: correct answers come from deep, well-connected reasoning. Wrong answers come from shallow, fragmented reasoning.

They name the failure mode. When the evidence only matches the question on the surface, the model drifts toward whatever the retrieved text happens to say.

This is lab work on question answering, not on crawled web pages, so treat it as an analogy rather than proof. But it makes the retrieval data less strange. When a system rewards pages whose heading closely matches the question, that isn't an arbitrary formatting preference. Surface-level matching is a known failure inside the model.

The practical version: build pages that answer one question completely, not pages that mention ten questions adequately.

Which creates a new problem.

Focused Pages Are Fragile. Linking Is What Makes Them Viable

Splitting content into focused pages gives you more deep pages. Each one is a better citation candidate on its own, and structurally weaker: further from the homepage, thinner in authority, easier to strand. Contextual linking doesn't create citations. It keeps the focused structure alive.

Start with the hard constraint. As of June 2026, no major AI crawler runs JavaScript. Not GPTBot, not ClaudeBot, not PerplexityBot. An analysis of over 500 million GPTBot fetches found zero JavaScript execution.

Google dropped its own JavaScript SEO warning in March 2026. That applies to Googlebot only.

So a focused page reaches an AI system one way: an HTML link that something followed to get there. A page with no inbound links isn't a weak candidate.

It isn't a candidate.

That's the access argument, and nobody disputes it. The second argument is about what the link says.

Bolted on: "We cover this in more detail. Read our guide."

Woven in: "Covering more than half the sub-questions correlates with lower citation rates once relevance is controlled."

Same destination. The second anchor says what the target page proves, and the sentence around it carries the claim. The first says nothing.

That surrounding sentence matters more than it looks. Patent analysis of AI Mode suggests the system may pull up to five chunks before and after a relevant one for context. Retrieval isn't purely atomic: the text next to a passage can travel with it. The sentence carrying your link is part of what gets retrieved.

The experiment that limits this argument

An honest version of this case includes the result that cuts against it.

SE Ranking built a fake brand in March 2026 and tracked 825 prompts across five AI engines. One test was structural: a hub page linked to 10 supporting articles, all indexed, all properly linked.

It produced zero AI citations.

Meanwhile, 30 thin pages on a test domain produced more than 1,800 citations between them.

The authors' reading is the right one: internal linking may help a search engine understand a site, but AI systems still need a reason to cite a specific page for a specific answer.

That's the boundary. Internal linking doesn't create demand. It controls access and clarity once the demand exists. A perfectly linked cluster on a topic nobody asks about stays invisible. Linking was never the missing piece there.

Which is why it belongs with the foundations, not the growth levers. Foundations don't produce results by themselves. Their absence stops everything else from producing results.

What to Do About It

Four actions, ordered by what they cost you when ignored.

1. Check what survives without JavaScript

Turn off JavaScript and load your key pages, or run curl and read the raw HTML. Anything that disappears is invisible to every AI crawler, however good it looks in a browser.

A page nothing links to can't be reached by a crawler that only follows HTML. Screaming Frog or a sitemap-versus-crawl comparison finds them in an afternoon. Every deep page you care about should get contextual links from related content.

3. Split exhaustive guides, without inventing satellite pages

A guide covering eight distinct questions competes against eight focused pages and loses on each. Split it where the reader's question actually changes, then connect the pieces.

Don't generate a page per keyword variation. That's the pattern Google flagged as scaled content abuse.

4. Rewrite generic anchors

"Learn more" and "click here" say nothing about the target page. Neither do two-word keyword anchors dropped into a sentence that didn't need them.

The anchor should describe what the linked page proves, inside a sentence that would still work without the link.

Then measure it. Since June 2026, Search Console includes a generative AI performance report showing impressions from AI Overviews and AI Mode, page by page. No click or query data yet, and history only goes back to May 2026.

Page-level impressions is exactly the granularity this question needs. For the first time you can watch whether the deep pages you just connected start showing up in AI answers, using your own data instead of a third-party estimate.

Frequently Asked Questions

Mathias Decourt

Written by

Mathias Decourt

Co-Founder - CEO

Website performance specialist helping businesses identify the few actions that truly move the needle. Turning complex data into clear, actionable insights that drive growth.