GEO & AI search

AI-generated content and Google: what the guidance actually says

What Google has actually published about AI-generated content, where the spam policies draw the line, what models do well and badly, and how to audit the content shipped on your domain.

10 min read

What Google has actually published

The panic and the hype about AI content share one feature: neither cites the source. Google's guidance on AI-generated content has been public since early 2023 and it is remarkably short. The core sentence is that Google focuses on the quality of content rather than on how it is produced. Automation, including AI, is not against the rules in itself. What breaks the rules is automation used primarily to manipulate search rankings. That is essentially the whole documented position, and almost every confident claim you will hear from a vendor is an interpretation layered on top of it.

A second documented signal is a wording change. The helpful content guidance once asked whether material was written by people, for people. Google revised that phrasing to created for people, and stated plainly that using AI confers no special advantage: it is just content, rewarded if it is helpful, original and shows what Google calls E-E-A-T, meaning experience, expertise, authoritativeness and trustworthiness. The revision is small but deliberate. It moved the test from authorship to outcome, which is exactly the shift most buyers, and a fair number of agencies, have still not internalised.

In practice that means nobody can honestly tell you your blog was demoted for being AI-written, because Google publishes no such signal and has never claimed to identify production method reliably. Rankings move for reasons the documentation does list: unoriginality, thin coverage, missing first-hand experience, weak site-level quality. AI makes those failures far cheaper to mass-produce, which is why they now show up together. The same underlying logic governs answer engines, and the relationship between the two disciplines is unpacked in our comparison of SEO and GEO.

Where the line is actually drawn: scale, not tooling

The sharpest documented line arrived in March 2024, when Google updated its spam policies and renamed spammy automatically generated content to scaled content abuse. The published definition rewards a literal reading: creating many pages for the primary purpose of manipulating search rankings rather than helping users, typically large amounts of unoriginal content that provides little or no value, no matter how it is created. That last clause is the important one. Hand-written spam and generated spam are treated identically. The policy targets a pattern of production, not a category of tool.

That makes scale the real risk variable. Ten AI-assisted articles reviewed by someone who genuinely knows the subject sit nowhere near the policy. Four hundred pages a month, one per keyword variation, assembled from an identical template, sit uncomfortably close to it whether a person or a script produced them. The same March 2024 update also named site reputation abuse and expired domain abuse, two other tactics that only pay off at volume. Read together, the three describe one consistent worry: industrial output aimed at the index rather than at a reader.

Beyond that, be honest about where the documentation stops. Google has published no page-count threshold, no tolerated percentage of AI assistance and no detection method. Anyone quoting one is guessing. What seems reasonable inference, and we flag it as inference rather than policy, is that programmatic pages built on a genuine proprietary dataset behave differently from programmatic pages built on paraphrase, because the first adds something the index does not already hold. The parallel with manipulative link acquisition is close, and the same volume logic applies there.

AI-assisted versus AI-abandoned content

The useful distinction is not human versus machine, it is assisted versus abandoned. Assisted content has a person who decides what the piece argues, supplies the parts a model cannot know, checks every claim that could be wrong, and puts their name to the result. Abandoned content has a prompt, an output and a publish button. Both may run on the same model. They produce different artefacts, they carry very different risk, and after two paragraphs an experienced reader can usually tell which of the two they are holding.

Abandonment is rarely a philosophical choice. It is a margin decision. A retainer priced for four researched articles a month is easier to deliver as twelve generated ones, and the client rarely reads them all. Which is why the question of AI content is really a question about how your content marketing engagement is priced and staffed. If the fee assumes an hour of expert time per thousand words and the observable output implies six minutes, something is absorbing the difference, and it is almost always the reader.

One practical test cuts through most of the argument. Pick any published page and ask who would be embarrassed if a sentence in it turned out to be false. If a specific, named person would, the process behind it is probably sound. If the honest answer is nobody, you are looking at abandoned content regardless of how it was drafted. That question mirrors how Google frames its own self-assessment, which asks who created the content, how it was created, and why it exists at all.

Where models genuinely help

Used inside a real process, current models are genuinely good at a narrow set of tasks, and pretending otherwise costs money. They compress research: twenty sources condensed into a comparison a human then verifies against the originals removes hours of work. They are strong at structure, turning a messy pile of notes into an outline that exposes the gaps in your thinking before anyone writes a paragraph. And they are dependable at mechanical transformation, which covers schema markup, alt text, metadata drafts and consistent formatting across a large library.

Translation is the clearest win and the most commonly botched. A model produces a fluent first draft in French, German or Spanish far faster than a human translator could. What it does not know is that search demand differs by market, that legal wording differs, that the local competitor set is entirely different. Teams working with agencies in France or specialists in Germany use models for the draft and a native practitioner for terminology, intent mapping and everything else that decides whether the page ranks in that country.

The last honest strength is pattern spotting across your own site. Models are good at proposing internal links between pages a human forgot existed, at flagging two articles that now cover the same intent, and at drafting a first-pass keyword cluster from a crawl. Treat each suggestion as a hypothesis rather than an instruction, because the model cannot see your traffic. Most serious SEO and GEO tools now ship these features, which makes them a commodity rather than a reason to pick one vendor over another.

  • Research synthesis: many sources condensed into something a human then checks against the originals
  • Outlines and structure: exposing the gaps in an argument before anyone writes a paragraph
  • Translation drafts: a fluent first pass in a second language, finished by a native practitioner
  • Structured data: schema markup, metadata and alt text at volume, verified on the rendered page
  • Internal link suggestions: candidate connections a human approves against real traffic data
  • Editing passes: catching inconsistent terminology, broken logic and repetition across a long library

Where they reliably fail

The failures are just as predictable, and they map almost exactly onto what makes content worth reading. A model has no first-hand experience. It has never migrated your platform at two in the morning, never watched a client's traffic collapse after a template change, never negotiated with a supplier in your category. It can imitate the register of experience convincingly, which is worse than failing openly, because the imitation reads well enough to survive a quick review by somebody who was not in the room.

It also has no data of its own. Every number in a generated draft is borrowed, approximated or invented, and the invented ones arrive in exactly the same confident phrasing as the real ones. Currency is the related failure: models are trained to a cutoff, while pricing, product features and regulation keep moving. Fabricated citations are the most dangerous case of all, because a plausible reference to a study that does not exist is far harder to catch than an obviously wrong sentence.

Finally, models do not make judgement calls, they average them. Asked what a company should do next, a model returns the consensus of everything written on the subject, which is precisely the advice your competitors are already following. The value of an expert is the defensible minority position and the willingness to be accountable for it. That is also, awkwardly for the generated-content business model, what answer engines tend to quote, as the mechanics of AI citation make clear.

  • First-hand experience: what actually happened on a real project, with the details only participants know
  • Original data: your own numbers, benchmarks and results, which no model can produce for you
  • Current facts: anything that moved after the training cutoff, including prices, features and regulation
  • Judgement: the defensible minority view rather than the average of everything ever written
  • Verifiable citations: real sources, checked by opening them, not plausible-looking references

The editorial process that makes AI assistance defensible

What makes AI-assisted content defensible is boring and procedural. A named human owns each piece and is accountable for its claims. Every factual assertion has a source that somebody actually opened. Anything describing experience comes from an interview with a person who has that experience, not from a model imagining what it might feel like. Nothing publishes without review by someone who could have written the piece unaided, which is the step that quietly disappears first when volume rises and margins tighten.

Expertise attribution matters more than it used to. Google's quality guidance asks that readers be able to see who stands behind a page, and an empty byline or a generic team credit removes the only signal a reader has. Attribute a piece to the person whose expertise it draws on, even when a model produced the prose and a writer did the assembly. Reviewer credits work well for technical material. The person named should be able to defend every sentence in a meeting.

On disclosure, the documented advice is narrower than the debate around it. Google suggests AI or automation disclosures where a reader would reasonably wonder how something was made, and treats them as unnecessary where the automation is obvious or trivial. It has not asked for a label on every assisted paragraph. Separately, the European AI Act adds transparency obligations for some published AI-generated text from August 2026, with an exemption where a human reviewed the content and a person holds editorial responsibility. Check the detail with counsel, not with your agency.

  • A named human owner per piece, accountable for every claim it contains
  • Sources opened and checked, never cited on the strength of how plausible they look
  • Experience sourced from an interview with somebody who actually has it
  • Review by a person who could have written the piece without any assistance
  • Disclosure where a reader would reasonably wonder how the page was produced
  • A publishing cadence that matches the review capacity you actually have

How to audit what is being published on your domain

You do not need a detector to audit an agency, and detectors are unreliable enough that leaning on one will mislead you in both directions. Use process evidence instead. Ask for the brief, the source list and the reviewer name for three recent articles. A team running a real process produces all three within the hour. Ask who was interviewed for the pieces that describe experience. Ask what the model was used for, specifically, and treat a refusal to answer as the answer.

Then read the output as a buyer rather than as a proofreader. Uniform section counts, identical paragraph rhythm, conclusions that merely restate the introduction and a complete absence of anything specific to your company are the reliable tells, far more so than any stylistic tic. Cross-check against performance data: a content cohort that gained impressions but almost no clicks is usually being indexed and ignored. Directories that rank suppliers on verified Search Console performance are useful for the same reason.

Fix the rest in the contract rather than in an awkward conversation nine months later. Specify who reviews, who is named as author, what disclosure applies, and what happens to pages that underperform, and keep ownership of drafts and accounts. The list in our agency selection questions covers most of it, and the usual contract red flags cover the rest. If the answers do not hold up, rebuilding a shortlist from the agency directory or from independent consultants is not hard: compare candidates on evidence, not on how confidently they discuss AI.

Frequently asked questions

Does Google penalise AI-generated content?

Not for being AI-generated. Google's published guidance states that it focuses on the quality of content rather than how it was produced, and that using AI confers no special advantage or disadvantage. What its spam policies do target is automation used primarily to manipulate search rankings, in particular large volumes of unoriginal pages that offer little value to a reader.

What is scaled content abuse in Google's spam policies?

Scaled content abuse is the name Google gave in March 2024 to a spam policy that previously covered spammy automatically generated content. It describes creating many pages for the primary purpose of manipulating search rankings rather than helping users, typically large amounts of unoriginal content with little value, explicitly no matter how that content was created. The policy therefore applies equally to human-written and machine-written spam.

Do I need to disclose that an article was written with AI?

Google's guidance suggests an AI or automation disclosure where a reader would reasonably wonder how the content was made, and considers it unnecessary where the automation is obvious or trivial. It has never required a label on every assisted paragraph. Separately, European transparency rules that apply from August 2026 cover some categories of published AI-generated text, with an exemption where a human reviewed it and holds editorial responsibility, so check your specific case with a lawyer.

Can Google detect AI-generated content?

Google has never published a detection method for AI-written text, and no public tool identifies it reliably enough to base a decision on. Commercial detectors produce both false positives on careful human writing and false negatives on edited machine output. What Google does measure, and describes in its documentation, is originality, usefulness and overall site quality, which is what actually moves rankings.

How can I tell if my SEO agency is using AI to write my content?

Ask for the brief, the source list and the name of the reviewer for three recent articles, plus the name of whoever was interviewed for any passage describing experience. A team with a real editorial process supplies all of that within an hour. In the published pages themselves, the reliable signals are uniform structure, no company-specific detail and conclusions that only restate the introduction.

Should I delete the AI-generated articles already published on my site?

Not automatically. Sort them by whether they are useful and original rather than by how they were drafted. Pages that earn impressions but almost no clicks, duplicate an intent already covered elsewhere, or contain claims nobody can source are the candidates for rewriting, merging or removal. Pages that answer a real question accurately can stay, ideally after a human review and a proper byline.

Read next