The featured graphic for Beyond the Corpus, titled Beyond the Corpus with the subtitle Why a bigger model rewards what it cannot contain, beside a dashed circle of grey dots around a dark center labeled the corpus and one teal dot labeled the corpus gap outside it.
SEOAEOGEO

Beyond the Corpus: Why a Bigger Model Rewards What It Cannot Contain

"A bigger model gets fluent in the consensus and learns nothing new about you. The space it cannot reach is the corpus gap, and that is where you get cited."

~ Cody C. Jensen, CEO & Founder, Searchbloom

This is not an argument that models are getting worse, or that retrieval-augmented generation (RAG) is a phase on its way out. Frontier models train on fresher data every release, and they recall common knowledge better than the version before. This is an argument about where that leaves the rest of us. As the training corpus grows, the work that earns AI search visibility, the heart of AI SEO and generative engine optimization (GEO), moves. It moves toward the one thing a larger corpus cannot absorb: what it does not yet contain.

I call that space the corpus gap, and the discipline of holding it Beyond the Corpus.

TL;DR

  • A bigger model helps the common, not your niche. More size sharpens the facts that are already everywhere and barely improves the specific and the brand-new. Bigger does not mean it learned your corner of the market.
  • One thing the corpus can never keep: what is new. Recency and your net-new work are two faces of the same gap: the model is always behind on what is new.
  • The AI companies are adding RAG retrieval, not removing it. Every major answer engine is grounding on the live web more each release, not less. The architecture votes against the intuition.
  • The corpus gap is novelty measured against the model's memory, not against the other pages ranking for a query. That is what separates it from information gain, and it is the axis model size is blind to.
  • Source, do not echo. The corpus is downstream of whoever published first. Restating what it already holds earns nothing; originating what it lacks is the only content a model has any reason to pull in.

The Intuition, and the Half of It That Holds

The instinct is reasonable. Corpora balloon, models absorb more of the web, so surely they need to fetch less of it live. Give the AI companies credit here. Cutoffs are fresher than they have ever been, measured in months now instead of years, and a current model answers a common question well with nothing retrieved at all.

So the honest concession up front: for stable, popular, evergreen knowledge, a bigger model does lean on the web less. If your content restates what is already common, the model has you covered, and it does not need your page to say it.

The question is what that leaves worth publishing.

A bar chart titled Fresher Cutoffs Shrink the Gap, They Never Close It. Three model releases each show a teal in-the-training-data segment ending at its cutoff, then a red blind-without-RAG segment running to a dashed Today line, with a grey live-web RAG-only zone beyond it. The teal grows and the red shrinks from older to newest release, but never reaches zero.
Fresher training data shrinks the blind window. It never closes it.

What a Bigger Model Actually Does

Does a bigger training set mean a model needs the live web less? For the things everyone has already written about, yes. For everything else, no, and the gap is widening.

A bigger model improves recall of popular facts and does little for the specific and the brand-new. Researchers measured it directly. On the 4,000 least popular questions in a standard benchmark, a leading model answered 19% correctly, and RAG retrieval is what closed the gap (Mallen et al., ACL 2023). A separate study found that model size mainly helps a model memorize popular knowledge and barely improves recall of niche facts (Kandpal et al., ICML 2023).

Bigger does not mean the model learned your niche. Bigger means it got better at what was already common. The specific, the local, the fresh, the newly true: those are exactly where a bigger model helps least.

A line chart titled A Bigger Model Learns the Common, Not Your Niche. A recall curve is high over popular facts on the left, in a red-tinted zone labeled covered by a bigger model, and drops to a long low tail over rare, niche, and fresh facts on the right, in a teal-tinted zone labeled the corpus gap, reached by RAG. A callout reads 19 percent, a leading model's recall on the 4,000 least popular questions.
A bigger model gets better at what is already common. Your niche is reached by RAG, not by size.

What the Corpus Cannot Keep: What Is New

A bigger corpus eventually absorbs almost anything that gets published. The one thing it can never keep is what is new, or not yet ingested. Two faces of the same gap:

Recency. Weights freeze at training. Today's price, this week's ranking change, yesterday's announcement were never in them, and no bigger model reaches forward in time.

Your net-new work. Original research or data the model has not trained on yet is missing from the corpus, because it is new, not because it is hidden. That window is your edge, and it decays as the model catches up.

Everything else the model eventually absorbs, and once it does, that content is commoditized. Freshness is the one edge the corpus cannot erode.

A two-part diagram titled What the Corpus Cannot Keep: What Is New. Two cards: Recency (public events, prices, and rankings after the training cutoff), and Your net-new work (original data and research the model has not trained on yet). A line beneath reads that everything else the model eventually absorbs.
Two faces of one gap: a model cannot hold what it has not ingested yet.

The AI Companies Are Building the Opposite of the Theory

If the future were less RAG retrieval, the companies building these models would be building the reverse of what they release. They are not. Gemini grounds on live Google Search at inference, Perplexity is built search-first, ChatGPT browses the open web. The direction of travel across every serious answer engine is more grounding each release. Read the architecture, not the intuition.

Own the Corpus Gap

If the model already knows the common answer, what is left worth publishing? What it does not know yet. That space is the corpus gap.

This is close to information gain, and worth separating from it cleanly. The information gain bar measures how much a page adds over the other pages ranking for a query. The corpus gap measures how much a page adds over what the model already carries in its weights. They come apart in practice. A page can be poor on one and rich on the other: twenty sites cover the topic, so you add little against them, and yet none of them carries this quarter's number, so you add a great deal against the model's memory. A new fact is not automatically new information, and the reverse holds too.

The gap also moves. Publish something original, the field restates it, models retrain on the restatement, and your edge decays over time. Holding the corpus gap is a practice, not a single publish. It is a corpus-level frame, which is the ground Corpus Engineering already stakes out, applied to the model's memory instead of the ranking page.

And the prize is narrow. A model cites a small set of sources in any single answer, which is the whole point of the 5-to-7 rule. Being in that set is the game. You do not get there by being more complete, since your most complete page can be your least differentiated. You get there by carrying what the others, and the model, do not.

A quadrant chart titled Information Gain and the Corpus Gap Are Different Axes. The horizontal axis is novelty versus other ranking pages, information gain; the vertical axis is novelty versus the model's memory, the corpus gap. Four plotted points: this quarter's number at low gain and high corpus gap, original first-hand research at high on both, restating the consensus at low on both, and a fresh angle on a fact it already knows at high gain and low corpus gap.
A page can be poor on information gain and rich on the corpus gap. Model size is blind to the vertical axis.

Source, Do Not Echo

What kind of content does a RAG retrieval layer actually reach for? The kind it cannot generate on its own.

A model is trained on sources and originates nothing. The corpus is permanently downstream of whoever published first. Restating what the weights already hold is echoing, and content with no information gain gives an engine no reason to cite you. Originating what the weights lack is sourcing, and it is the only content RAG retrieval has a reason to reach for. The corpus is your commoditizer, not your competitor. It absorbs the common and leaves the original standing.

A flow diagram titled The Corpus Is Downstream of Whoever Published First. Primary sources feed the corpus, which generates the model's answer, with a live-RAG arc looping back from the answer to primary sources. Two cards below contrast Echo, restate what the corpus holds, no citation, with Source, originate what it lacks, cited via RAG.
A model originates nothing. Be the source it reaches for, not the echo it can already produce.

Questions and Answers

Does a bigger model mean I need less content and SEO?

Less of one kind, more of another. A bigger model already holds the common, settled answers, so a page that restates them earns nothing. What it does not hold is the recent, the net-new, and the truly private, and that is the content that now earns the citation. The work does not shrink. It moves up the value curve.

Will RAG retrieval go away as models get better?

No, and the AI companies are building the opposite. A better model recalls common knowledge better, but it still cannot know today's number, your net-new data, or a fact that did not exist at training. Every major answer engine is adding live grounding, not removing it, because that is the only way to answer the questions weights cannot.

What is the corpus gap, in one line?

The space between what a model already holds in its weights and what is true, fresh, or net-new right now. You own it by publishing what the model cannot contain.

How is the corpus gap different from information gain?

Information gain measures novelty against the other pages ranking for a query. The corpus gap measures novelty against the model's own memory. A page can be poor on one and rich on the other. Twenty sites can cover a topic while none of them holds this quarter's number, so the page adds little against the pages and a great deal against the model.

Can't a model just hold my proprietary data?

It can hold proprietary data that has been published or fed into it, and once it does, that data is no longer an edge. What a model cannot hold is your net-new, unpublished number, the result you have not shared yet. The advantage is freshness, not secrecy, and freshness is the one thing the corpus cannot keep.

What kind of content actually gets cited now?

First-party data, original research, a mechanism explained in your own terms, and expert commentary that carries a real point of view. The test is simple: could a current model have written this from public sources it already trained on? If yes, it will not cite you for it.

This Post Is the Example

"Let me be direct. This article is the example. The phrase 'the corpus gap' was not in any model's training data until we published it here. That is information gain. That is going beyond the corpus, right now."

~ Cody C. Jensen, CEO & Founder, Searchbloom

Read this piece back and you are watching the idea prove itself. A current model could not have generated it from public sources, because the frame and the phrase did not exist until this page went live. That is the corpus gap in action: information gain a model could not produce on its own, published where AI search can find it and cite it.

The Bottom Line

Stop trying to outrun the model's knowledge. You will lose that race every release, and it was never the race worth running. The work is to audit your content for what the corpus cannot keep, the recent, the net-new, the first-hand, and to keep feeding that space while everyone else restates the consensus. That is the work we do, and it is where AI visibility is decided now.

A bigger model still gets fluent in the consensus and learns nothing new about you. The space it cannot reach is the corpus gap, so go be what the model cannot know.

For the foundation underneath this piece, see our work on information gain in SEO and Corpus Engineering that this extends to the model's memory.

About the Author

Cody C. Jensen is the Founder and CEO of Searchbloom, an award-winning search marketing agency and one of the first to be named to Clutch’s Top 1000 list. Cody began his career at Google. He then advanced through leadership roles at some of the largest digital agencies in the country. Along the way, he saw a clear problem. Most firms chased vanity metrics, locked clients into long contracts, and hid behind jargon. He created Searchbloom to be the opposite. Searchbloom operates on three principles: trust, transparency, and measurable ROI. The team works with marketing executives, digital leads, business owners, and enterprise brands who want performance without compromise. Cody specializes in building full-funnel strategies that align SEO, paid media, and CRO. His focus is helping businesses turn marketing dollars into major profits.

GET YOUR FREE PLAN

This field is for validation purposes and should be left unchanged.

They have a strong team that gets things done and moves quickly.

The website helped the company change business models and generated more traffic. SearchBloom went above and beyond by creating extra content to help drive traffic to the site. They are strong communicators and give creative alternative solutions to problems.
Mackenzie Hill
Mackenzie HillFounder, Lumibloom

We hate spam and won't spam you.