How we actually use AI in content production (and where it stops)
Blu Mint CEO, David Bailey leads our team of writers who use AI in their work
A few months ago, SEO analyst Lily Ray published something worth taking seriously: she'd spent months tracking more than 220 websites that had scaled their content production using AI tools, many of them publicly cited as customer success stories by the vendors selling those tools.
Her finding, in her own words, was that it works, until it doesn't.
Across those 220-plus sites:
54% lost 30% or more of their peak organic traffic.
39% lost half or more.
22% lost three-quarters or more.
The shape of the pattern was consistent: rapid growth in published pages over six to twelve months, an organic traffic peak a few months after that, then a steep decline that erased most of the gain, often within a year, and often below where the site started.
Ray is careful about what this does and doesn't prove.
Her data is third-party traffic estimates, not first-party analytics, and she's clear that correlation isn't causation.
However, the pattern held across industries, cybersecurity, travel, SaaS, healthcare, B2B services, and it lines up closely with two Google actions on record: the September 2023 Helpful Content Update, and the March 2024 Core Update, which introduced a spam policy called Scaled Content Abuse, explicitly targeting pages generated at volume to manipulate rankings, regardless of whether a human or an AI wrote them.
What actually gets flagged
Reading through Ray's analysis, the sites that lost the most traffic tended to share a specific pattern: content published to fill a template, at volume, without anyone reviewing whether each individual page actually needed to exist.
Comparison pages built for every possible product pairing in a category. Glossary entries for every term a competitor's page might also target.
Self-promotional listicles where the company names itself the best option, without evidence anyone actually tested the alternatives.
FAQ pages built one question per URL, structured purely for AI extraction rather than for a person who typed that exact question into a search bar.
Location pages multiplied across every city or country a template could reach, regardless of whether the company had anything real to say about that specific place.
None of these page types is inherently broken.
Ray makes the point directly: they work, which is exactly why so many companies scaled them.
The problem shows up specifically at scale, when a template gets multiplied across thousands of pages faster than any human is reviewing what's actually going out.
She frames the test worth applying before publishing anything at that pace plainly: would a real customer need this specific page, or does it exist because a search engine or an AI model might cite it?
Could a competitor generate a near-identical version tomorrow with the same prompt? That second question alone rules out most of what gets published this way.
That distinction, AI-assisted versus AI-unsupervised, turns out to matter more than whether AI touched the content at all.
We've written before about Ahrefs' research into AI Overview citations, which found that 87.8% of the pages Google cites are a mix of AI and human writing, and that the correlation between a page's AI-generated percentage and its citation position is effectively zero.
Two entirely separate studies, run by different people looking at different data, land on the same conclusion from opposite directions; it was never really about whether AI was involved. It's about whether a person was still making the calls.
What the AI-human split actually looks like for us
Every piece we publish goes through the same shape of process, and it's worth being specific about it rather than gesturing at ‘balance.’
A human sets the strategy, the argument and the structure before anything gets drafted
AI accelerates the research and produces a first pass against that structure, which is genuinely faster than starting from a blank page
A human editor takes it from there: checking every fact, adding real experience and judgement, and rewriting until it actually sounds like someone who knows the subject wrote it
A human makes the final call on whether it's good enough to publish under a real name
That last step is the one the 220+ declining sites in Ray's dataset seem to have skipped. Her own recommendation, closing the piece, is that AI tools are genuinely useful for research, briefs and accelerating a workflow ‘where a human expert is still in the loop,’ and that companies using them should be transparent about it, ‘which is recommended by Google.’
We already do that.
Our own AEO piece discloses, with a source link to Anthropic's own documentation, that Claude embeds a watermark in AI-generated text and that Anthropic has committed to that under the EU AI Act's Article 50(2) Code of Practice.
That's not a claim about our process. It's the actual mechanism sitting inside the tool we use, disclosed because Ray's advice, and Google's, is that this is what responsible use looks like.
What AI still can't supply
Google's Helpful Content system rewards a signal AI cannot manufacture on its own: genuine first-hand experience, the "Experience" component added to E-E-A-T at the end of 2022. We've written about why that shows up as a permanent ranking signal now, not a passing update, and it's the same reason our own process keeps a human in the loop for the parts that actually carry that signal: real product knowledge, a real customer's actual words, a real judgement call about what's worth saying.
That's live in our work right now, not just something we did once and wrote up afterwards.
We're currently drafting content for JALG, a Scandinavian design maker of handmade, Red Dot Award-winning minimalist TV stands based here in Tallinn, selling through dealers from Denmark to Japan.
JALG's previous content wasn't badly intentioned.
Like a lot of companies over the past couple of years, they'd read the GEO and AEO coverage that's been everywhere, took the advice seriously, and used AI to write and illustrate articles aimed at getting cited by AI systems.
What got missed wasn't effort; it was Google's own E-E-A-T guidance, the same "Experience" signal covered above, and the research into what actually earns AI visibility we've written about separately.
The old articles had AI-generated hero images with garbled, illegible text rendered onto the TV screens, a generic "Key Takeaways" bullet format bolted onto the top of every post, and a materials list- wood, glass, metal, particle board, sintered stone, acrylic- that read like a template rather than a description of what JALG actually sells.
No named author anywhere.
The rebuild starts with AI, structured prompts through ChatGPT and Claude to get a working first pass fast, then moves entirely into human territory. Every piece now carries a real byline, JALG co-founder Veiko Kallas, and links through to the actual products- oak stands, birch stands, the wooden TV stands collection- rather than a generic materials list.
One new article splits a care-and-cleaning table into JALG's own four materials against common alternatives for a reader who bought their stand elsewhere, a structural decision as much as a proofreading one. The specificity a template can't fake shows up in the detail: the birch finish is sealed rather than open-grained, so wax sits on top instead of soaking in, while oak needs re-oiling rather than polishing, distinctions that only exist because someone who actually makes these stands was interviewed and quoted directly.
The new imagery follows the same logic, designed by Kairi from our web design and development team, replacing decorative AI renders with genuinely functional diagrams, an annotated room photo showing TV width, height, eye level and viewing distance, a side-by-side comparison of the Regular (42–55") and XL (55–77") stand sizes, and step-by-step guidance on measuring a TV and reading the room it'll actually sit in, right down to marking the stand's footprint on the floor with masking tape before committing to a size.
Every image now carries real alt text describing what it actually shows, not a generic file name, which matters for accessibility and for search in a way a decorative AI image never could.
Run that same distinction against Ray's own checklist, and it holds up cleanly: a real customer's question gets a real answer, informed by a real product and a real person's words, not a template a competitor could reproduce with the same prompt tomorrow.
Where human content writers (and web designers) fit in
This is the ‘next on the list’ piece we flagged in the article on who we build for: the honest version of how content actually gets made here, rather than a paragraph promising to explain it later.
If your team is trying to work out whether to scale content production with AI, the answer isn't whether to use it. Every credible study on this now says the tools themselves aren't the risk. The risk is removing the human who decides what's actually worth publishing. A genuine content writing agency does that work deliberately, not as an afterthought bolted onto a faster drafting process.
Get in touch if you'd rather build that in from the start than find out the hard way, the way 220 other companies already have.
-
Not on its own. A study tracking over 220 sites that scaled content with AI found the sites losing traffic shared a specific pattern, content published at volume without anyone reviewing whether each page needed to exist. Separate Ahrefs research found the correlation between a page's AI-generated percentage and its AI citation position is effectively zero. It was never really about whether AI was involved, it is about whether a person was still making the calls.
-
Two questions. Would a real customer need this specific page, or does it exist because a search engine or AI model might cite it. Could a competitor generate a near-identical version tomorrow with the same prompt. That second question alone rules out most of what gets published purely at scale.
-
Yes, directly. Our own AEO piece discloses that Claude embeds a watermark in AI-generated text, and that Anthropic has committed to that under the EU AI Act's Article 50(2) Code of Practice. That is not a claim about our process, it is the actual mechanism sitting inside the tool we use.
-
A human sets the strategy, argument and structure before anything gets drafted. AI accelerates research and produces a first pass against that structure. A human editor checks every fact, adds real experience and judgement, and rewrites until it sounds like someone who knows the subject wrote it. A human makes the final call on whether it is good enough to publish under a real name.