I wrote the headlines by hand. That mattered to me. The rest of the report, my piece on why my code agent starts over from zero every Monday, was model-generated, sure, but the headings I wanted for myself, so I sat with them and filed away until they landed. Strong verbs, one image per line, nothing sanded flat. The kind of language where you can hear that somebody meant it. One of them took me ages. “Eine Schicht trägt nicht. Zwei halten. Drei werden Verbund.” (One layer won't hold. Two will. Three become a composite.)

Then, half out of curiosity and half out of vanity, I pushed the report through Gemini and asked which parts sounded like AI.

Gemini was quite sure of itself. The machine-written text? Passed as human. And the headlines, the one stretch in the entire document that came out of my own head from start to finish, flagged as AI.

I read that twice. No close call, no “hard to say”. The model had gone in exactly the wrong direction, and it went there with conviction. My first reflex was to defend myself, because yes, that was me, I still have the drafts. Then came the question I haven't been able to shake since. How does a detector get it that wrong and stay that confident?

The detector isn't measuring what you think

The answer starts with a letdown. Gemini didn't measure anything here. It guessed, and they all guess, the commercial ones included, the expensive ones included. No detector owns an instrument that physically separates machine-written text from human text. What they have is a sense of how AI text usually looks, dense, rhythmic, deliberately built, and my hand-tuned headlines happened to look precisely like that. Which is also what happens when a person really tries.

That would be a nice anecdote about a confused algorithm if there weren't a clean empirical finding underneath it that flips the whole thing around. If the detector isn't catching “machine”, what is it catching?

A paper from early 2026 answered the question precisely, and the answer is more uncomfortable than it sounds. The researchers ran base models and their instruction-tuned siblings through GPTZero and Pangram, the two big commercial detectors. Llama-3-8B as a raw base model: rated human 96.7 and 98.8 percent of the time. Then the same model, same architecture, only trimmed afterwards to follow instructions, and suddenly it counts as human just 30.3 and 17.1 percent of the time.1

Sit with that for a second. The machine stays exactly the same. The one thing that changes is a post-processing step. And that step decides whether the detector says “human” or “AI”. So it isn't tracking machineness at all. It's tracking the fingerprint of alignment training. What it takes to be the essence of AI text is really a scar the polishing step leaves behind, and my hand-filed writing happened to carry a few of the same marks.

Machines aren't the only ones who get fooled here. We do too. In the cleanest Turing test study so far, a frontier model was taken for the human 73 percent of the time, more often than the real human in the same test.2 What I believe about that number isn't “people are blind”, which strikes me as a stretch. Which text I like better is a very different question from which one a machine wrote, and the two get confused all the time.3 Someone who works with the technology daily spots it differently than someone whose last brush with AI was in 2023. Detection is a matter of practice, not a fixed human defect. But most casual readers genuinely don't notice, and that gap is where the trouble sits. “Sounds like AI” and “is AI” are two separate things. They correlate, granted, but the detector only sees the first one and claims the second.

The smoothness is scar tissue

On the left a real fingerprint with irregular ridges, on the right the same print as a perfectly even set of concentric rings.
On the left, a print that exists only once. On the right, what survives post-training.

Which leaves the question of where that ironed-flat sound comes from in the first place. The base-versus-instruct numbers already hold the answer. The base model writes like a human because it was trained on human text and knows nothing else. Then RLHF arrives, the alignment step, and grinds the humanity off.

Technically, two things happen. The first is mode collapse. The model could draw from an enormous range of phrasings, and after training it no longer does. It collapses onto a narrow corridor of phrases the reward liked. Any single text reads fine, but across a thousand texts the model keeps saying the same thing the same way. The second is sycophancy. The model learns to please, and pleasing language comes out soft, reassuring, allergic to conflict. It won't risk a sharp edge, a hard sentence, anything with a downside.

Together the two produce what people call GPT-isms. “delve”, the everlasting groups of three, the “es ist nicht X, sondern Y” (it's not X, it's Y), the dutiful three-part list at every opportunity. Text that reads as though no person wrote it, but rather the statistical average of all persons. Call that a style if you like. It's scar tissue from the training run. The process that makes a model polite and helpful and safe is the same process that makes it recognizable.

I barely own any writing of my own

A receding archive shelf: the binders at the front are half transparent, at the very back stands a single solid binder covered in cobwebs.
The last folder I'd put my hand in the fire for sits all the way at the back.

I could have filed all of this away as an interesting detector problem. What ruined the little smile about Gemini's misjudgement came later, when I turned the thing around and tested myself.

I wanted a counter-example. A longer piece of writing where I can say with total certainty: not one line of this has ever been seen by a model. No draft pass, no “mach das mal runder” (make that read smoother), no quick reshuffle of a paragraph. Pure human.

Almost nothing came to mind.

My emails run through suggestions. Commit messages, half autocompleted. Blog posts, reports, documentation, all of it now happens in dialogue with a model. The last texts whose authorship I could guarantee at a hundred percent are years back. Handwritten diary pages maybe, a few old letters.

And it isn't that I handed myself over to the AI. The thoughts are mine, no question about that. What's in the texts, what they argue towards, which examples show up, that comes out of my head. What's never quite mine are the words. The phrasing belongs to the model, and honestly, most of the time I don't like it. I iterate, push back, ask for another version, then another, and still land almost every time at the point where the result annoys me. I'd probably be faster typing the paragraph myself. I still don't. I've gotten lazy, and that's the unpleasant half of the confession.

If somebody asks me tomorrow whether a text is mine, the honest answer for nearly all of it is: sort of. The rest gets settled by a detector that's guessing.

Why smooth language isn't harmless

You could treat all of this as an aesthetic problem, a question of taste, smooth against rough. The moment somebody stops noticing they're talking to a machine, it turns into a problem with consequences. Because whoever misses that gets handled differently, and swayed harder.

The measurement already exists. An AI with access to a few personal details out-persuades a human conversation partner 64.4 percent of the time.4 As long as recognition holds, you raise a defensive reflex against a stranger who's trying to convince you. When recognition fails, the reflex goes with it, and you're defending yourself against something you take for one of your own.

The cultural stake is becoming visible right now. In April 2025 it was reported that OpenAI is working on its own X-like social network, an internal prototype built around a feed of AI-generated content, on which Sam Altman was gathering outside feedback.5 What's reported is an early prototype, not a finished product. The direction is hard to misread though, and my reading of it is bleak: a platform where the machine ends up supplying the content and the human only supplies the idea. Once the statistical average of all people becomes the dominant voice in the feed, we forget what a single real person sounds like. It goes further than culture, too. What effortless, frictionless AI does to our thinking is a barrel of its own, big enough that I won't open it until the end of the series.6

The bandage and the wound

A boat with a hole in the hull: water gushes in while a man bails it back out with a bucket.
The humanizer bails out the water. The hole stays where it is.

There's a whole market for making AI text “human” again. Humanizers, they call themselves. You push your generated text in, they roll dice on synonyms, break up the sentence lengths, sprinkle in a few irregularities, and at the end the text slips past the detector. An arms race that treats the symptom. You take the smooth, alignment-polished output and rub sandpaper over it afterwards so it looks rougher.

That never convinced me, and after the Gemini moment I understood why a little better. The humanizer works at the end of the chain, on the finished text. The smoothness happens far earlier, inside the model, during training, long before the first word appears. A humanizer covers the scar. The question that's stuck with me since: what if you prevented it from forming at all?

So I'm building it myself

Work exists in that direction, and it's good. Sam Paech published a method called FTPO, Final Token Preference Optimization, and the idea is surgical instead of blunt. You first profile which tokens and phrases a given model overproduces, build a preference dataset out of that automatically, then finetune against exactly those patterns at the single position where the model would swerve onto the slop track. No overhaul of the whole model, just a damping of the worn paths, careful enough that the rest stays intact. The results are unusually clean for an intervention like this: roughly 90 percent less slop, while the math and reasoning benchmarks hold steady or even tick up a bit.7 The surgery repairs the language without damaging the capability. And the gap it's fighting is grotesque. Some of these patterns show up over 1,000 times more often in model output than in human writing. Nothing about that ratio is subtle drift. It's a statistical scream. The full pipeline is open, MIT licensed.

That's a root-cause fix rather than a bandage. You don't humanize the text, you repair the model that produces it.

Which leaves the catch that makes the whole thing personal for me. All of it exists in English. The slop profiles, the banlists, the datasets, the finished anti-slop adapters, all trained in English and measured against English corpora. For German there's none of it. German AI text has its own fingerprint, different stock phrases, different reflexes, and nobody has measured it systematically so far, let alone ground it off.

So that's what I'm building. A German slop profile from real model output, a preference dataset that captures the specifically German tells, and at the end a finetuned model that writes German without every second paragraph closing on a mini-conclusion and every list running to exactly three points. No humanizer bolted on the back. Never let the smoothness form.

Except it turned out that before you can train slop out of a model, you have to know exactly what slop is. And that's where it got properly interesting. Part 2 is about that.