Your Name Is on the Cover. Did You Do the Thinking? 

100% AI? That May Be the Wrong Question. History Has Entered the Age of Cheap Certainty

An AI detector can tell you that writing resembles machine-generated text. It cannot tell you who did the research, who developed the argument or whether the claims are actually true. In the age of generative AI, perhaps we are testing the wrong thing.

Someone recently sent me screenshots of an AI-detection report on a piece of published writing.

I am deliberately not going to identify the work, the writer or even the subject. None of that is necessary for what follows.

What caught my attention was the confidence of the result.

There was an enormous percentage on the screen. The language sounded almost forensic. If you saw the screenshots without knowing anything else, the temptation would be obvious.

Caught.

That was more or less my first reaction too.

Then I noticed the small print.

The detector itself warned that its result reflected a likelihood and was not definitive.

Ah.

That changes things.

So I started looking into AI detectors themselves.

Slightly annoying, actually, because the original story would have been much easier to write.

Someone publishes something. Detector gives spectacular result. Writer gets exposed. Everyone argues on Facebook.

Done.

Except the evidence wouldn’t let me write that story.

AI detectors are simply not reliable enough to establish authorship on their own. OpenAI withdrew its own experimental classifier in 2023 because of its low accuracy. Turnitin, which continues to offer AI-writing detection, warns that false positives can occur and that its AI score should not be used by itself to take adverse action against someone. Independent research has also found considerable variation in how well different detectors perform.

So I cannot tell you that the writer whose work was sent to me used AI.

Maybe the detector is right.

Maybe it isn’t.

I don’t know.

And where the evidence ends, so should the accusation.

Fine.

Forget the detector for a moment.

Because once the easy story fell apart, I realised there was a much more interesting question sitting underneath it.

Who actually did the thinking?

What Does “AI-Generated” Even Mean?

Before going further, there is some terminology worth clearing up because these detector reports can look more scientific than they really are.

Different systems use different categories, so these are not universal scientific definitions. But in ordinary language, detectors are often trying to distinguish roughly between three things.

AI-generated basically means: this text looks to our system like a machine probably produced the wording.

Imagine asking an AI:

“Write me 1,000 words explaining the rise and decline of an ancient trading port.”

It produces the passage. You copy most of it into your manuscript.

That is the straightforward example.

But the detector did not watch you do it.

It did not see your screen. It did not examine your notes. It did not interview your editor.

It looked at the finished words and classified them.

Then there is AI-assisted or paraphrased.

This is the messy middle.

Perhaps I wrote a paragraph myself and told AI:

“This is clumsy. Make it clearer without changing my argument.”

The research and argument are mine, but the machine substantially changes the wording.

Or perhaps AI produced the first draft and I rewrote it heavily.

Maybe it reorganised my sentences. Maybe it translated them. Maybe it paraphrased something I had already written.

Human and machine contributions can become tangled surprisingly quickly.

Finally, likely human means the detector thinks the writing resembles human-produced prose more closely.

Again, that doesn’t mean the software somehow witnessed a human writing it.

It is still a classification.

And here is the important part:

These labels describe what the finished text resembles. They do not reconstruct what actually happened at the writer’s desk.

A human can write something a detector thinks looks artificial.

An AI can generate something a detector thinks looks human.

And something can pass through both human and machine hands so many times that asking who “wrote” the final sentence becomes surprisingly difficult.

Which is why I think we need a better test.

Actually, three.

My Three Checks

The first is the obvious one.

Check One: How was this probably written?

This is where the AI detector belongs.

It may give me a reason to become curious. Perhaps even suspicious.

But it is the weakest of the three checks because it is trying to reconstruct a writing process from linguistic patterns in the finished product.

Useful clue.

Not proof.

Then I move to Check Two: Is what this person is saying actually true?

Now we are getting somewhere.

Forget whether the prose came from ChatGPT.

Take an important factual statement and check it.

If a writer says an event happened in the 14th century, check the chronology.

If archaeological evidence supposedly establishes something, find out what was actually excavated.

If one civilisation supposedly influenced another, what evidence establishes the connection?

If a sweeping claim is made about the origins of a community, culture or identity, slow down.

When?

Where?

According to whom?

What exactly do those words mean in the period being discussed?

This is where beautiful prose sometimes starts getting into trouble.

Consider a sentence like:

“Communities developed through centuries of complex cultural interaction and exchange.”

Sounds intelligent.

I could probably put it into a museum panel tomorrow.

But what did it actually tell me?

Which communities?

What interaction?

What was exchanged?

When?

What changed?

How do we know?

Generative AI is very good at this sort of writing. Give it enough context and it can produce paragraphs that sound thoughtful, balanced and historically literate.

Everything flows.

Nothing jars.

The civilisation flourishes. Trade expands. Identities evolve. Cultural practices develop through complex interactions over many centuries.

Lovely.

Now show me what happened.

Humans have been writing like this for a very long time, of course. AI did not invent vague historical prose.

It merely became extraordinarily efficient at producing it.

And that brings me to Check Three.

The one I care about most.

Show me the evidence.

Where did the claim come from?

Show me the footnote.

Show me the document.

Show me the archaeological report.

Show me the scholar whose work you are relying upon.

And then, because merely having a footnote does not impress me very much, show me that the source actually supports what you wrote.

That last part matters.

A source can be genuine and still be irrelevant.

A quotation can be accurate and still be stripped of context.

A famous scholar can appear in a footnote without ever having argued what the author says he argued.

Ten books repeating the same claim do not necessarily give us ten independent pieces of evidence. Sometimes all ten lead back to the same unsupported assertion.

This is where the three checks become useful.

Check One asks how the words may have been produced.

Check Two asks whether the claim is accurate.

Check Three asks whether the evidence actually supports it.

The first can make me curious.

The second can make me suspicious.

The third determines whether I trust the work.

Now Something Interesting Happens

Imagine a researcher spends years studying a subject.

She reads the primary sources. She knows the scholarship. She develops her own interpretation, understands the disagreements and can defend every major claim.

But her English isn’t particularly elegant.

So she uses AI heavily to turn rough prose into something readable.

An AI detector might have a fit.

Now imagine another writer.

Every word is proudly human.

No ChatGPT.

No Claude.

No Gemini.

Pure human craftsmanship.

Wonderful.

Except the chronology is dubious, the citations are weak and the grand claims collapse the moment somebody bothers to check them.

Which writer worries me more?

Faced with those two choices, I would trust the AI-assisted researcher who can show me the evidence before I trusted the proudly human writer who cannot.

That doesn’t mean authorship no longer matters.

Quite the opposite.

It forces us to ask what authorship actually means.

Typing Is Not the Same as Thinking

I use AI myself.

So I have no interest in pretending technological purity.

If I research an argument, develop the idea, write the rough prose and then use AI to improve the grammar, am I still the author?

I think so.

Now move the machine further upstream.

I provide my research notes and AI drafts the chapter.

Hmm.

Move it further.

AI gathers the material, synthesises it, constructs the argument, decides which interpretation sounds convincing and generates most of the chapter. I read it, make some changes and put my name on the cover.

Now the question becomes uncomfortable.

What exactly did I author?

There is no magical percentage at which authorship changes hands.

I don’t think counting machine-generated words will solve it either.

I am more interested in where the intellectual work happened.

Who selected the evidence?

Who understood it?

Who noticed when two sources contradicted each other?

Who decided one interpretation was stronger?

Who recognised that a confident paragraph was actually nonsense?

Who knew when the evidence was insufficient?

Who was prepared to leave something unresolved rather than manufacture certainty?

That is authorship too.

Perhaps it is the more important part.

I don’t particularly care who physically typed the sentence.

I care who can answer for it.

This Gets More Uncomfortable With History

This question bothers me particularly when we move into historical writing because that is where I spend much of my own time.

History is unusually vulnerable to beautiful nonsense.

A sentence can be elegant, confident, perfectly readable and historically weak.

Sometimes the problem is obvious.

Often it isn’t.

Take subjects involving ancient societies, archaeology, migration, trade, cultural origins or identity. A few broad facts can be assembled into an enormously persuasive narrative.

A happened.

B happened.

C happened.

Therefore A caused B, which produced C.

Except perhaps there are 300 years between them.

Perhaps the archaeological evidence does not establish the claimed connection.

Perhaps the word being used for a community in the 21st century did not mean the same thing six centuries ago.

Perhaps scholars disagree.

Perhaps we simply don’t know.

AI can smooth all those awkward gaps beautifully.

So can a talented human writer.

Which is precisely why “sounds convincing” has never been a good historical method.

AI Has Made Certainty Cheap

This is where I eventually arrived after thinking about those screenshots.

I went into this expecting the detector score to be the interesting part.

It wasn’t.

The bigger change brought by generative AI may not be that machines can write false sentences.

Humans already had that market covered.

The change is scale.

Previously, producing 80,000 words of plausible pseudo-history required considerable effort.

You still had to sit there and write the bloody thing.

Now a machine can generate coherent, authoritative-sounding historical prose almost instantly.

It can synthesise a narrative.

Produce an explanation.

Suggest citations.

Smooth contradictions.

Fill gaps with plausible language.

And keep going.

That changes the economics of misinformation.

The cost of producing something that looks like knowledge is collapsing.

The cost of actually knowing something has not.

That is the asymmetry that concerns me.

AI can produce ten historical claims in seconds.

Checking those ten claims might take me an afternoon.

One obscure claim might take days if I have to chase the original source.

An archaeological interpretation might require finding an excavation report, understanding its dating and comparing it with later scholarship.

A claim about cultural identity may require us to first decide whether the category being used even makes sense for that historical period.

Generating certainty has become cheap.

Verifying it remains expensive.

And that, I think, is a much bigger problem than whether a detector thinks somebody’s sentence looks robotic.

So Should Writers Disclose AI?

I think there is a reasonable line here, although I don’t pretend it is universally settled.

Nobody needs to confess because software fixed three commas.

Grammar correction does not bother me.

Translation assistance does not automatically bother me.

Using AI to challenge an argument, organise research or make difficult prose clearer can be genuinely useful.

I do some of these things myself.

But the closer AI moves towards doing the intellectual work, the stronger the case for telling the reader.

If it generated substantial passages, that matters.

If it synthesised the literature for you, that matters more.

If it constructed the argument and you mainly approved what appeared on screen, I think readers are entitled to know.

Because when I buy serious non-fiction with someone’s name on the cover, I assume that name represents something beyond ownership of the Word document.

I am buying judgement.

Someone has read.

Someone has thought.

Someone has made choices.

Someone is prepared to defend those choices.

Otherwise, what exactly did I pay for?

Humans Were Quite Capable of Bad History Before AI

There is another reason I don’t want this to become an anti-AI argument.

People were producing bad history long before ChatGPT arrived.

We already had pseudoarchaeology.

Grand civilisational claims.

Romantic stories about ancestors.

Unsupported heritage narratives.

Claims repeated so many times that repetition eventually became mistaken for evidence.

Humans did all of that perfectly well by ourselves.

We hallucinate too.

We misunderstand sources.

We cherry-pick.

We become emotionally attached to beautiful theories.

Sometimes we decide what happened first and then go shopping for evidence afterwards.

There is nothing uniquely artificial about intellectual laziness.

AI simply gives it industrial capacity.

That is why obsessing over whether a sentence is “human” misses the point.

Human does not mean true.

AI does not mean false.

The standard has to be somewhere else.

Maybe the Best AI Detector Is Still a Human Being

Perhaps I don’t need another percentage.

Sit the author down.

Ask questions.

Why this source?

Where did that claim come from?

What evidence contradicts you?

Why this chronology?

Who disagrees?

What remains uncertain?

What would make you change your conclusion?

And my favourite:

Show me the evidence.

That is a brutal detector.

It works beautifully on AI-generated writing.

Unfortunately for bad human writers, it works on them too.

And unlike a software probability, the answers actually tell me something useful.

Perhaps that is what bothered me about those screenshots in the end.

Not that they might have caught somebody using AI. I still don’t know whether they did.

What bothered me was how easily the number itself could become the evidence.

We risk doing exactly what we worry AI will encourage us to do: accepting a confident answer before asking how that answer was produced.

So I don’t particularly want the percentage anymore.

Show me the claim.

Show me where it came from.

Show me the evidence.

And if your name is on the cover, be prepared to explain how you got from one to the other.

I have said for some time:

The internet rewards attention. History rewards evidence.

AI has made the first part extraordinarily cheap.

The second remains stubbornly expensive.

Good.

It should be.

Because a machine can generate the words.

The author still has to answer for them.

CONCISE SOURCE / REFERENCE NOTE

AI-detection results should not be treated as forensic proof of authorship. OpenAI withdrew its experimental AI-text classifier in 2023 because of low accuracy, while Turnitin’s guidance acknowledges false positives and cautions against using its AI-writing score as the sole basis for adverse action. Independent research has likewise found that detector performance varies across systems, models and types of text. Detector labels such as “AI-generated”, “AI-assisted/paraphrased” and “likely human” are best understood as classifications of linguistic patterns in submitted text rather than direct evidence of what occurred during writing. The three checks proposed in this essay deliberately move beyond that limitation: how was the text probably produced; are its claims accurate; and do reliable sources actually support them? The first provides a clue. The latter two test the work itself.

error: Content is protected !!