How we actually use AI in content production (and where it stops)
Blu Mint CEO, David Bailey leads our team of writers who use AI in their work
A few months ago, SEO analyst Lily Ray published something worth taking seriously: she'd spent months tracking more than 220 websites that had scaled their content production using AI tools, many of them publicly cited as customer success stories by the vendors selling those tools.
Her finding, in her own words, was that it works, until it doesn't.
Across those 220-plus sites:
54% lost 30% or more of their peak organic traffic.
39% lost half or more.
22% lost three-quarters or more.
The shape of the pattern was consistent: rapid growth in published pages over six to twelve months, an organic traffic peak a few months after that, then a steep decline that erased most of the gain, often within a year, and often below where the site started.
Ray is careful about what this does and doesn't prove.
Her data is third-party traffic estimates, not first-party analytics, and she's clear that correlation isn't causation.
However, the pattern held across industries, cybersecurity, travel, SaaS, healthcare, B2B services, and it lines up closely with two Google actions on record: the September 2023 Helpful Content Update, and the March 2024 Core Update, which introduced a spam policy called Scaled Content Abuse, explicitly targeting pages generated at volume to manipulate rankings, regardless of whether a human or an AI wrote them.
What actually gets flagged
Reading through Ray's analysis, the sites that lost the most traffic tended to share a specific pattern: content published to fill a template, at volume, without anyone reviewing whether each individual page actually needed to exist.
Comparison pages built for every possible product pairing in a category. Glossary entries for every term a competitor's page might also target.
Self-promotional listicles where the company names itself the best option, without evidence anyone actually tested the alternatives.
FAQ pages built one question per URL, structured purely for AI extraction rather than for a person who typed that exact question into a search bar.
Location pages multiplied across every city or country a template could reach, regardless of whether the company had anything real to say about that specific place.
None of these page types is inherently broken.
Ray makes the point directly: they work, which is exactly why so many companies scaled them.
The problem shows up specifically at scale, when a template gets multiplied across thousands of pages faster than any human is reviewing what's actually going out.
She frames the test worth applying before publishing anything at that pace plainly: would a real customer need this specific page, or does it exist because a search engine or an AI model might cite it?
Could a competitor generate a near-identical version tomorrow with the same prompt? That second question alone rules out most of what gets published this way.
That distinction, AI-assisted versus AI-unsupervised, turns out to matter more than whether AI touched the content at all.
We've written before about Ahrefs' research into AI Overview citations, which found that 87.8% of the pages Google cites are a mix of AI and human writing, and that the correlation between a page's AI-generated percentage and its citation position is effectively zero.
Two entirely separate studies, run by different people looking at different data, land on the same conclusion from opposite directions; it was never really about whether AI was involved. It's about whether a person was still making the calls.
What the AI-human split actually looks like for us
Every piece we publish goes through the same shape of process, and it's worth being specific about it rather than gesturing at ‘balance.’
A human sets the strategy, the argument and the structure before anything gets drafted
AI accelerates the research and produces a first pass against that structure, which is genuinely faster than starting from a blank page
A human editor takes it from there: checking every fact, adding real experience and judgement, and rewriting until it actually sounds like someone who knows the subject wrote it
A human makes the final call on whether it's good enough to publish under a real name
That last step is the one the 220+ declining sites in Ray's dataset seem to have skipped. Her own recommendation, closing the piece, is that AI tools are genuinely useful for research, briefs and accelerating a workflow ‘where a human expert is still in the loop,’ and that companies using them should be transparent about it, ‘which is recommended by Google.’
We already do that.
Our own AEO piece discloses, with a source link to Anthropic's own documentation, that Claude embeds a watermark in AI-generated text and that Anthropic has committed to that under the EU AI Act's Article 50(2) Code of Practice.
That's not a claim about our process. It's the actual mechanism sitting inside the tool we use, disclosed because Ray's advice, and Google's, is that this is what responsible use looks like.
What AI still can't supply
Google's Helpful Content system rewards a signal AI cannot manufacture on its own: genuine first-hand experience, the "Experience" component added to E-E-A-T at the end of 2022.
We've written about why that shows up as a permanent ranking signal now, not a passing update, and it's the same reason our own process keeps a human in the loop for the parts that actually carry that signal: real product knowledge, a real customer's actual words, a real judgement call about what's worth saying.
That's live in our work right now, not just something we did once and wrote up afterwards.
We're currently drafting content for JALG, a Scandinavian design maker of handmade, Red Dot Award-winning minimalist TV stands based here in Tallinn, selling through dealers from Denmark to Japan.
The drafting starts with AI, structured prompts through ChatGPT and Claude to get a working first pass fast.
Every piece then goes through human proofreading against JALG's actual products and real quotes from their creative directors- the exact detail no AI model can supply on its own, because it doesn't exist anywhere for a model to have learned it from.
The AI gets us to a draft quickly.
The specificity that makes the draft worth publishing comes from a person who actually knows the product.
Run that same distinction against Ray's own checklist, and it holds up cleanly: a real customer's question gets a real answer, informed by a real product and a real person's words, not a template a competitor could reproduce with the same prompt tomorrow.
Where human content writers fit in
This is the ‘next on the list’ piece we flagged in the article on who we build for: the honest version of how content actually gets made here, rather than a paragraph promising to explain it later.
If your team is trying to work out whether to scale content production with AI, the answer isn't whether to use it. Every credible study on this now says the tools themselves aren't the risk. The risk is removing the human who decides what's actually worth publishing. A genuine content writing agency does that work deliberately, not as an afterthought bolted onto a faster drafting process.
Get in touch if you'd rather build that in from the start than find out the hard way, the way 220 other companies already have.