I’ve talked a lot about good practices to make AI content. But I have to admit I’ve made a lot of mistakes over time. When I read what I did a year ago, I cringe. So I recently spent a long time analyzing what I did wrong and how to fix it.
Single prompt generation
Ok I’ve never done this one. I knew from the beginning I needed more context than just asking “write me a newsletter about common mistakes when generating AI content”.
But I prefer to be explicit about it: don’t do that!
Tropes are not a styleguide
I thought my list of AI tells was enough. For Claude, it almost is. Claude has a fairly neutral default and it follows examples. Tell it what not to do and it usually stops.
The others do not. GPT wander. Grok will follow the list, then still write like a model unless you also tell it what to do.
Tropes.md is the don’t list. A styleguide is the do list: contractions, first person, mix of long and short sentences, no slogan lines, headings that name the problem. Asking to be “friendly but professional” doesn’t do much. Rules plus examples produce something you won’t have to paraphrase and can easily edit.
No references
You will not get your voice from instructions alone.
My process for a new format: pick the type of post, find the best examples I can, write down what they do well and what they do badly, generate a draft from these examples and my styleguide, then edit until it is something I would send. That edited piece becomes the reference for the next ones.
Raw model output as a “reference” is how you make slop. Generated-then-hand-edited is the way to go.
Dirty references rot everything after them
I messed this up last week. It is why I sat down and reviewed the whole setup.
My agent made a few mistakes in a new post. I did not fix them all. I used that post as the reference for the next batch. Everything new came out with the same tells.
Your references have to be clean. If you skip the edit, you teach the model to copy the slop.
Using the wrong model
A few months ago only Claude was actually good at writing. Now the gap has closed but I made the mistake of trusting the others too much.
Claude still has the least AI-sounding first draft. It’s neutral, follows references, works even when the styleguide is thin.
GPT 5.6 Sol and Astra are inconsistent on this job. Sometimes they match Claude, sometimes they ignore tropes.md and don’t follow references.
Grok 4.6 is the compromise I like most. The native style is not as good as Claude but it follows references and it is consistent. With a real styleguide it gets close enough that I stopped paying for a standalone Claude sub. I still have Opus and Fable in Cursor when I lose patience but it’s rare.
People keep saying Gemini 3.8 and GLM 5.3 are as good or better now. That is not my experience. GLM is getting close. Gemini is still not better than a research intern: great at analyzing sources, but outputs a corporate memo.
High effort makes copy worse
Ultra thinking helps agents write code but on copy it overthinks the sentence. Low or medium is enough. Plus it saves tokens!
Hermes is a mediocre writing harness
I like my Hermes agent because I can just queue a lot of research and get my dose of vitamin D in the meantime. But when it comes to generating content, I’m still on the fence. Like GPT models, it seems to take some liberties with my references. It might be a configuration issue but for now I usually do the generation with Cursor or another harness locally.
Skills will override your voice
Hermes will write its own skills. And they might not match how you want the copy to sound. If you leave them unmaintained, they inject AI tells back into every run.
Open-source copywriting skills are mixed. A repo can have one good skill next to three bad ones. You need enough copy judgment to know which skills to use.
Skills also outrank your styleguide. If a skill says something that goes against your voice, the skill wins. Either fix the skill or turn it off. Don’t wait or you’ll spend a lot of time fixing everything later.
“Don’t add claims” and pick another model
When I write about health, I bring sources that do not always match mainstream science. ChatGPT treats that as a fact-check assignment. You can push back, but it is more work than it is worth.
Telling the model not to add claims helps. Switching to a model that won’t fight you helps more.
It still won’t necessarily pass the Pangram test
Even by doing all of that, you’ll have to do some manual editing.
Because removing all AI tells is almost impossible and no model will get your tone of voice right. Even when you think you caught all the AI tells, you notice new ones you were not before since there were so many tells.
But it has cut something like 80% of my editing time.

