Grounding Page: The Fact Sheet AI Systems Quote From
A grounding page bundles the verifiable facts about a company for AI systems. What belongs on it, how to build it and why it does not carry on its own.
August 14, 2026 · 14 min read
SEO Galaxy is now GetMentioned. The platform stays the same. Want a demo? Book your slot here →
Keyword density, minimum word counts and the fear of duplicate content date back to a time when search engines counted words. Seven common rules checked against Google's own documentation, plus a grid for signing a text off.
David Hahn · August 19, 2026 · 27 min read
The short version
An SEO text is a text written for a known search query. Its job is settled before the first sentence.
The rules such texts are still written by come from a time when search engines counted words. Keyword density, minimum word counts and the fear of duplicate content are arithmetic for a system that no longer works that way.
This article checks seven widespread rules against Google's own documentation. One of them survives.
After that comes a grid of six criteria. It lets you sign off a text, a commissioned one as readily as one written in house.
The term SEO text belongs to the industry. It does not appear in Google's documentation.
It still describes a real way of working, in which a text is produced for a search query.
An SEO text is a text for a website that is meant to answer a known search query in full.
Three things are fixed before writing starts. The topic, the query and the page the text will live on.
An article written without those three is still an article. It simply serves no search.
In practice the label covers very different things. Category copy in a shop, guide articles on a blog and landing pages for an offer all travel under it.
The three have different jobs. Category copy frames a range of products. A guide article answers a question. A landing page leads to a decision.
Anyone buying a text should name the type. A price per word says nothing about what arrives.
The umbrella term also hides a decision. A guide article rarely serves a query with buying intent, and the reverse holds too.
Early search engines sorted pages by word matching. A page that used a term more often counted as a closer match.
The familiar numbers grew out of that mechanic. Keyword density, minimum length and placement in the opening paragraph all describe the same model.
The model has changed in several steps. Google now names RankBrain, neural matching, BERT and a passage ranking system among its own systems.
The numbers stayed anyway. They are easy to measure and easy to put in a brief.
In our experience that is exactly why they persist. A buyer can check a keyword density. Whether a text hits the question behind the search has to be read.
The seven sentences below appear in almost every article on the subject. Each is checked here against the place it would have to come from.
The checking was done in Google's documentation and in the instructions for its external raters.
Each rule ends with what holds instead, plus wording you can lift into a brief.
The figure appears in no Google document. The term keyword density is absent from the helpful content guidance and from the SEO starter guide.
The 182 pages of rater instructions contain the word keyword three times. Two deal with spam.
The third is a rating example, and it turns the rule around. For a query about tooth loss in a five-year-old child, Google shows a result about tooth loss in pike fish that happens to carry the words "five years old".
The instruction there says the result fails to meet the user intent because it has keyword matches only. It sits on page 137 of the Search Quality Rater Guidelines dated 11 September 2025.
A density of 0.7 percent would not have saved that result. The rating hangs on the subject.
The simpler version holds. A text names the thing it is about as often as the subject demands. Nobody counts.
Not this: "Place the focus keyword 'invoice software' at a density of 0.8 percent, spread evenly across all paragraphs."
Better, ready to paste into a brief: "The text treats invoice software as its subject throughout. Synonyms and related terms are free to use. There is no target frequency."
Google contradicts this rule outright. Its list of warning signs for search engine-first content contains a question Google answers itself.
The helpful content guidance puts it this way: "Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"
The question sits in a list Google heads with content to avoid. Writing to a target word count is itself the warning sign.
The rater instructions go further. They carry a section on filler, described as low-effort content that adds little value and does not support the purpose of the page.
Filler can artificially inflate content, the wording says, creating a page that appears rich while lacking anything visitors find valuable. A page carrying a lot of it should be rated Low.
Google's own quality standard therefore points the other way. Length without substance lowers the rating.
The question takes the place of the number. Length follows from what has to be answered. A definition can be finished after 200 words, a decision guide not after 3,000.
Not this: "For rankings the page needs at least 350 words, ideally 800."
Better, ready to paste into a brief: "The text answers the question in full and then stops. If 400 words do the job, it is 400 words. There is no minimum length and no paragraph that exists to add bulk."
This rule is wrong at its core. Google describes duplicates within one site as normal.
The page on canonicalization says it plainly: "Some duplicate content on a site is normal and it's not a violation of Google's spam policies."
What Google describes is a selection. Where several similar pages exist, Google picks one canonical version and shows that one.
Google names a system type for it. In the ranking systems guide they are called deduplication systems.
Your own canonical tag stays a suggestion in that process. Google writes that indicating a canonical preference is a hint and not a rule.
So the consequence of a duplicate is not a deduction. The consequence is that one of the two pages does not appear.
The case the rule aims at sits elsewhere. Google's spam policies cover copying other sites under the term scraped content.
That means taking content from other sites, often automatically, and hosting it to manipulate rankings. It is a violation, and the consequence can reach removal from the index.
Two cases with two different consequences remain. Your own duplicates cost visibility for one of the pages. Copying other people's content breaks the policies.
Not this: "Every page needs 100 percent unique content or Google will hand out a duplicate content penalty."
Better, ready to paste into a brief: "The text exists in this form on no other website. Passages lifted from third-party sources are excluded. Where similar pages exist on our own domain, one canonical version is chosen."
No Google source supports this rule. The SEO starter guide explains headings without stating a keyword requirement.
On the number of headings Google writes that there is no magical, ideal amount of headings a given page should have.
On order it goes further. Semantic order is called excellent for screen readers, while from Google Search's perspective a different order does not matter.
The rule still makes sense, for another reason. A heading that names the subject tells a reader in a result list what they are getting.
In our experience that only shows once a heading stands alone. In a table of contents, a shared preview, or the single line quoted in an AI answer.
The rule holds on different grounds. The H1 names the subject of the page because people read it. As a ranking instruction it is undocumented.
Not this: "The H1 must contain the focus keyword or Google cannot classify the page."
Better, ready to paste into a brief: "The H1 names the subject so that it still makes sense without the text below it. Every subheading describes its own section instead of announcing a type of passage."
This figure has no Google source either. Neither the SEO starter guide nor the helpful content guidance names a position in the text.
The rater instructions do carry a statement about order, without keywords. They ask that the most helpful content sit near the top.
Behind it sits an observation anyone recognises. Ask a question, get three paragraphs of throat-clearing, and you move on.
So the usable core of the rule concerns the answer, not the word. A page that has said nothing after 100 words loses readers.
The rule survives in a different form. The answer to the query stands in the first section. Whether the searched term falls inside it usually takes care of itself.
Not this: "Place the focus keyword within the first 100 words, ideally in the opening sentence."
Better, ready to paste into a brief: "The opening section answers the question in the heading in two or three sentences. Passages about why the topic matters are cut."
This rule holds. It is the only one of the seven that survives unchanged.
Google covers it in the spam policies. Keyword stuffing there means filling a page with keywords or numbers to manipulate rankings in Google Search results.
The examples are lists of phone numbers without added value, blocks of text listing cities and regions a page wants to rank for, and words repeated until they sound unnatural.
On consequences Google writes that sites violating the policies may rank lower in results or not appear in results at all.
One detail of the definition is rarely quoted with it. The violation hangs on the intent to manipulate rankings.
Google names no percentage anywhere. There is no value below which a text is safe and none above which it automatically offends.
What remains to be judged is how the text reads. A text that reads unnaturally falls under this policy. Arithmetic does not replace reading here.
Not this: "Stay under 2.5 percent keyword density and keyword stuffing does not apply."
Better, ready to paste into a brief: "No sentence carries a word that is there for the search engine alone. Lists of cities, regions or product variants without a statement of their own are excluded."
This rule asserts a chain of cause. Unique text is supposed to raise dwell time, and dwell time is supposed to improve rankings.
The second link is missing from Google's documentation. The ranking systems guide lists seventeen systems, from BERT through the freshness systems to the spam detection systems.
Dwell time does not appear among them. Neither does bounce rate or click-through rate.
The list is Google's own account of what it names publicly. It is not a list of every signal. What is documented is only that Google does not name this one.
The first link fails to measure what it claims. Analytics measures time on your own page. It says nothing about how long somebody stayed before returning to the result list.
For the effect of originality Google does carry an entry. Its original content systems show original reporting prominently ahead of sites that merely cite it.
The original share still works, along a route other than the one claimed. Google names a system for original content and no metric for dwell time.
Not this: "Unique content raises dwell time and therefore rankings."
Better, ready to paste into a report: "The article carries two figures that appear on no other page. Whether dwell time moves as a result is observed and not claimed as a cause."
| Rule | Status | What the source says |
|---|---|---|
| Keyword density 0.5 to 1 percent | no source | the term does not appear at Google |
| At least 350 words | contradicted | Google states it has no preferred word count |
| Duplicate content is penalised | wrong at the core | duplicates within one site are normal |
| Main keyword in the H1 | no source | no ideal number, order does not matter |
| Keyword in the first 100 words | no source | Google names no position in the text |
| Keyword stuffing is penalised | holds | stated in the spam policies |
| Unique content raises dwell time | chain unproven | dwell time is in no system list |
One rule in seven survives unchanged. Two carry a sound core and a false reason. Four appear in no Google document at all.
I have never seen a keyword density rescue a text. I have often seen a single figure somebody could look up make a text impossible to replace.
The rules just checked share one advantage. They can be counted.
The criteria that hold have to be read. A fixed grid helps, so that sign-off does not hang on taste.
Six criteria are enough. They apply to a commissioned text exactly as they do to one written in house.
Google's raters judge results on a five-step scale. Every step is defined through user intent.
The top step applies to queries with one clear intent and the one matching result. The bottom step describes a result that completely fails to meet the needs of all or almost all users.
None of the five steps names a text length. None names a keyword density.
For sign-off that produces one question. Does the text answer the question somebody asks with this query, or a neighbouring one?
The most common failure is a text pitched at the wrong stage of a decision. Somebody wants a comparison and gets a definition.
An article on "invoice software pricing" that spends 2,000 words explaining what invoicing is misses this step.
Better suits an article that opens with the pricing models in the market, breaks down what each tier includes, and closes with contract terms.
The second criterion asks for evidence. Every figure in the text needs the place it came from.
A source counts as named when a name, a date and a link are present. A bracket without a link is a claim wearing a name badge.
Two checks cover it. Does the link reach a live page, and does the statement actually appear there?
In practice we see the second fail regularly. The link points at a vendor's overview page while the figure came from a third-party article.
Not this: "Studies show that longer content ranks better (source: various SEO studies)."
Better, ready to lift: "Google's helpful content guidance states that it has no preferred word count, checked 19 August 2026." The link then points at that exact page.
Where many such claims about one brand collect, they deserve a page of their own. How that page is built sits in the article on the grounding page.
The third criterion is the hard one. It asks what stands in this text and in no other.
Google's rater instructions name three measures. Effort is the extent to which a human being actively worked to create satisfying content. Originality is the share not available on other websites. Talent or skill is the craft that went into it.
The instructions name the opposite case too. Pages whose main content is almost entirely copied, paraphrased, embedded, auto or AI generated, or reposted with little effort, little originality and little added value should be rated Lowest.
Licensed and syndicated content is expressly excluded from the word copied in the same instruction.
For sign-off one question is enough. Which figure, which case or which sentence in this text could be found nowhere else?
A section that summarises three competing guides and turns them into a list fails on this question.
Better, ready to lift with your own numbers: "We compared eighteen quotes from content providers in the first half of 2026. Twelve of them named no research scope at all."
Anyone without a data point of their own can put a judgement with a name behind it. That counts as an original share too.
The fourth criterion is about structure. Every section has to hold when read alone.
The reason lies in how content is retrieved today. Google names a passage ranking system that assesses single sections of a page.
Use in AI answers adds to it. What appears there is usually one paragraph or one sentence.
That gives sign-off two checks. Does each section carry one main statement? And is the subheading understandable without the paragraph before it?
A subheading reading "An example" above a paragraph that makes no sense without the previous two pages fails here.
Better, ready to lift: "Category copy for an online shop as an example" as the heading, with a paragraph underneath that explains the case in two sentences of its own.
The fifth criterion is the only one that checks quickly. It concerns the number of claims per sentence.
A sentence with three claims makes the reader sort them. Three sentences with one claim each do that work for them.
The test takes a minute. Read any paragraph from the middle out loud and count how often you draw breath.
Not this: "We have been in the market for a number of years and mainly look after mid-sized clients across the German-speaking region, with a focus on offerings that need explaining."
Better, ready to lift: "GetMentioned has been in the market since 2016. It is operated by SEO Galaxy GmbH, based in Germany. The portfolio holds more than 100,000 vetted publishers."
The sixth criterion asks how current the figures are. Every figure in the text needs a year.
Without a year every reader takes a number as current. A market figure from 2021 becomes a false statement about 2026.
Google warns about a related practice in the same list as the word count question. It names changing page dates to seem fresh when the content has not meaningfully changed.
Sign-off gets two questions from that. Does every figure carry a year? And does the page date match the last substantive edit?
Not this: "Around 60 percent of companies invest in SEO content."
Better, ready to lift: "In the most recent available survey on this, the figure was 60 percent. It dates from 2021 and nobody has measured since."
Six questions, six answers. A text that fails more than two goes back.
In our experience the third question fails most often. It is also the only one a rewrite does nothing for. It needs new information.
Google's quality standard is public. It sits in the instructions for the external raters who evaluate results.
The document is called General Guidelines, runs to 182 pages and carries the date 11 September 2025.
A caveat belongs with it. These ratings change no individual ranking. They measure whether a change to the system improved the results.
The document is still the most precise description of what Google means by quality.
The section on main content quality names four measures. Effort, originality, talent or skill, and accuracy.
Effort is drawn broadly. It covers the writing itself and also the building of functionality on the page.
As a counter-example Google names the automatic creation of thousands of pages. The case described is freely available content run through translation software without oversight or curation.
Originality is measured through a plain question. Where other sites carry similar content, the rater considers whether this page is the original source.
Accuracy is named for informational pages. For topics touching health, finances or safety, consistency with well-established expert consensus is added.
A section of its own deals with filler. It is headed "Filler as a Poor User Experience".
It describes low-effort content that adds little value and does not support the purpose of the page. Google adds that filler can artificially inflate content.
Two cases lead to a Low rating. A page with a large amount of unhelpful filler, and a page where prominent filler makes the helpful part hard to find.
That leaves the minimum word count doubly exposed. Google names no preferred length and names padding as a reason for a lower rating.
The four measures translate into one sign-off question each.
These four questions take longer than a count. In exchange they hit the measure Google describes in the instruction.
Another section covers the most common borderline case. It concerns pages that repackage content from higher quality sites, with low effort, low originality and low added value.
The examples are social media reposts with little added comment, pages of embedded third-party content without curation, and "best" lists built from existing reviews.
In our experience most commissioned texts land exactly here. They are not wrong. They are interchangeable.
A text now serves two audiences. The reader in the result list and the system assembling an answer. What that shift changes across a whole site sits in the article on AI search optimization.
The same grid applies to the second, for a different reason. A system building an answer from several sources needs sentences that stand alone and claims it can verify.
The Pew Research Center published an analysis of AI summaries in Google Search in July 2025.
The method is open. It used the browsing behaviour of 900 US adults who installed a tracking app. The period was March 2025.
The dataset covers 68,879 unique Google searches. An AI summary appeared for 12,593 of them.
Pew reports the outcome in two numbers. With an AI summary present, users clicked a traditional search result on 8 percent of visits. Without one it was 15 percent.
A click on a source inside the summary happened on 1 percent of visits.
Two caveats belong with it. The sample is US adults. Google publicly disputed the analysis and called the set of queries unrepresentative.
No new type of text follows from those numbers. What follows is a shift in purpose.
Part of visibility now happens without a visit. Being named becomes the outcome rather than the click.
For the text that means three things. Statements have to be quotable on their own, claims have to be verifiable, and the brand behind them has to be identifiable.
How answer engines assemble sources sits in the article on answer engine optimization.
In our experience this adds almost no work. The same six criteria that make a text hold up for readers make it quotable.
The question of the good text ends with who writes it. There are three routes.
All three produce usable results. They differ in internal effort and in where the subject knowledge comes from.
The first route keeps the text inside the company. The subject knowledge sits where it is created.
That makes the third criterion the easiest to meet. Numbers from daily operations, cases from client work and a view of your own are all in the building.
The bottleneck sits elsewhere. Time and regularity disappear the moment operations get busy.
In our experience this route rarely fails on the quality of the first text. It fails on the fifth, which never gets written.
It suits topics where nobody outside the company knows the facts.
The second route buys the writing in. The market runs two models for it.
Content marketplaces price per word and quality tier. That works for standardised product and category copy at volume.
Freelance writers work per assignment and research alongside. The price then hangs on research effort rather than word count.
One job stays in house either way. The subject knowledge has to reach the brief, or the original share is missing from the result.
In our experience that produces the most common disappointment. The delivered text is cleanly written and contains nothing that was not already online.
The route suits volume work on topics that need little specialist knowledge.
The third route hands over research, writing and checking, and keeps sign-off in house.
We fall under this answer ourselves, so the boundary belongs here. GetMentioned is a link building agency with its own placement marketplace. Full service in link building, not full-service SEO.
We do not work on your own website. Technology, site structure and CMS upkeep stay with your team or your SEO agency.
We do write and deliver SEO texts, including texts for your own pages. Publishing them there is your side of the line. The difference is delivery rather than intervention.
Ours is the intent research, the outline for approval, the text, the editing and the checking of every figure. Yours are the topic, the input and the sign-off.
Where an article is to appear with a publisher, coordination with the desk comes on top. A different standard applies there, namely what that publication's audience expects.
| In house | Freelancer or marketplace | Written and delivered | |
|---|---|---|---|
| Subject knowledge from | internal | the brief | the brief and research |
| Internal effort | high | medium | low |
| Original share in the text | easiest | depends on the brief | depends on the input |
| Scales across many texts | poorly | well | well |
| Fact checking | internal | rarely included | before handover |
| Who publishes | you | you | you |
The last row reads the same across all three. Publishing on your own domain stays with whoever runs the website.
Some services sound adjacent and are not. So here is the boundary as a list.
A good text meets one condition among several. It makes a page capable of answering.
Whether the page gets found also depends on how much weight the domain carries in its subject.
Google lists link analysis systems and PageRank alongside BERT and the freshness systems. Both layers run at once.
An excellent text on a domain without weight ranks for queries nobody contests. For contested ones it rarely reaches the top.
The reverse holds as well. A strong domain with an interchangeable text takes positions and loses them again as soon as somebody answers the question better.
Which signals from other websites count here and how they come about sits in the article on link building.
The diagnosis comes down to three cases.
In our experience the second case is handled wrongly most often. A complete text gets a fourth rewrite while the gap in referring domains stays put.
Checking it takes ten minutes. Open the ten results for the query, hold your own answer against them, then compare referring domains.
Of seven widespread rules for SEO texts, one holds. Keyword stuffing sits in Google's spam policies, with no percentage attached and with intent at its centre.
The other six are either unsourced or wrong at the core. They persist because they can be counted.
The criteria that hold demand a reading. Coverage of the search intent, dated evidence, an original share, sections that stand alone, one claim per sentence and a visible year on every figure.
Those six questions apply to any text, one written in house as readily as one delivered. They need no tool.
That leaves the other half. A text makes a page capable of answering, and weight comes from outside it.
An SEO text is a text for a website meant to answer a known search query in full. Topic, query and target page are fixed before writing starts. The term comes from the industry and does not appear in Google's documentation.
There is no documented minimum. Google's helpful content guidance states outright that it has no preferred word count and lists writing to a target word count as a warning sign. Length follows from the question the text answers.
Google names no value. The term keyword density appears neither in the SEO starter guide nor in the instructions for Google's raters. What does bind is the spam policy on keyword stuffing, and that turns on the intent to manipulate rankings.
Not in the form the rule is usually quoted in. Google writes that some duplicate content on a site is normal and not a violation of its spam policies. Where pages are similar, Google picks a canonical version. Copying other sites is a separate matter and sits in the spam policies as scraped content.
Google requires that nowhere. The SEO starter guide names no ideal number of headings and states that heading order does not matter from Search's perspective. An H1 that names the subject still makes sense, because people read it in result lists and tables of contents.
Take three samples. Does the opening section answer the question in the heading? Does the first figure carry a source with a date and a link? And is there any claim in it you could not find on another page? If one of the three fails, run the full grid.
Google judges the substance rather than the origin. Its documentation states that using automation, including AI generation, primarily to manipulate search rankings is a violation of the spam policies. Google's rater instructions put pages whose content is almost entirely copied or generated with little effort, originality and added value on the lowest step. What decides is the editorial stage afterwards.
About the author
Managing Director, GetMentioned
David has been building link acquisition and digital PR processes since 2016, first as an agency under SEO Galaxy, today as a platform with GetMentioned. He has scaled his own projects from zero to seven-figure monthly traffic and delivered thousands of campaigns for clients. Here he writes about what works in practice, and about what only costs budget.
Compare publishers, SEO data and prices in one place and book the placements that fit.
A grounding page bundles the verifiable facts about a company for AI systems. What belongs on it, how to build it and why it does not carry on its own.
August 14, 2026 · 14 min read
Every AI system runs its own crawlers, and each has a different job. Why the widely repeated advice about Google-Extended does not work, and what Google and Microsoft actually ask for.
August 16, 2026 · 11 min read
AEO works on the individual section, not the whole page. What Google says about the term, which levers hold up and how to build a section an answer engine can use.
August 15, 2026 · 18 min read