An AI detector made me doubt my own authorship. A second gave the opposite answer and made me realise I had trusted the detector more than I trusted myself. What began as a test of accuracy became an investigation into authorship, accessibility and recognition.

I started with what I thought was a fairly straightforward question about the accuracy of AI detectors. By now, you have probably realised that Substack had introduced Pangram, a tool making very strong claims about its ability to distinguish human writing from AI-generated content.
I was curious, but also dubious.
My experience of genAI detection in higher education has not given me much confidence in these draconian systems. That lack of confidence is supported by a growing body of research questioning the reliability and fairness of AI detection in higher education (Liang et al., 2023; Perkins et al., 2024; Hadra et al., 2026). We already have evidence that multilingual students can be disproportionately flagged as having used genAI and, worse, falsely accused of misconduct. Research has also shown that AI detectors struggle to distinguish hybrid human, AI writing from work that is entirely human or entirely AI-generated.
I had not planned to spend the day testing multiple AI detectors. I certainly had not intended to write another essay. But the question did not sit patiently in my head. One test led to another, the results became increasingly contradictory, and before long I was documenting different versions of my writing, comparing scores and trying to understand what any of those scores could reasonably tell me.
Let me take you through the experiments.
The result that made me trust the detector
Experiment 1
I began with a short section I had written that morning about AI detectors themselves. The original draft was entirely mine, straight from my head onto the page in my usual imperfect, messy style. It contained the familiar traces of my first-draft process: spelling mistakes, grammatical slips, compressed reasoning and several jumps that made sense inside my head but would inevitably need to be made more visible to the reader.
I entered the draft into ChatGPT and used a tightly constrained prompt. I asked it to keep as much of my voice and wording as possible, improve the flow and point out any jumps that a reader might struggle with.
The AI identified several places where the reader needed more information. I then:
researched the missing information and read more widely around it;
clarified what I meant by some of my language; and
made the central point of the essay explicit when the AI tried to take me off on a different tangent.
I then took that AI-supported version back into Word and edited it again before deciding it was finished. I ran all three versions of the passage through Pangram: the original human draft, the first AI-supported revision produced under the tight prompt, and the final version after I had edited it again myself.
I expected the results to show some kind of progression. Perhaps the untouched draft would appear human, the AI-supported version would show some assistance, and the final human-edited version would sit somewhere between the two.
They did not.
All three versions were deemed human. I was relieved. Pangram appeared to agree that my use of AI was completely valid and not in any way cheating. It understood the difference between AI-generated and AI-supported writing… phew!
I then wondered whether Pangram had worked some kind of magic and improved AI detection far beyond anything we had previously had access to in higher education. To test this, I ran the same three passages through several other detectors.
The results showed huge inconsistencies across the different models. Three of the six detectors classified all three versions as 100 per cent human.
The mean human score fell across the three versions, but that average masks the real story. The detectors were not converging on a shared judgment. Some did not detect the AI intervention at all. One appeared to think the writing was predominantly AI-generated before AI had even touched it. Another became even more convinced that the passage was AI-generated after further human revision.
All in all, the results did not show a reliable relationship between what had actually happened and what the detectors said had happened.
But Substack was using Pangram. It all looked good, right? Nothing to see here, right?
My optimism was short-lived.
The result that made me trust myself less
Experiment two
I put the opening of my latest essay, Sit Properly!, through Pangram. It had taken me around six weeks and probably 20–30 hours of active work to research, write and hone, not including all the time I spent thinking about it before and between drafts.
Pangram classified it as 100 per cent AI-generated.
I then tested a section from the middle and another from the end of the same essay, using the same number of words each time to make the comparison as fair as possible. Every section came back as 100 per cent AI-generated.
For those of you who have not read the essay, you can find it here. But first, let me give you some background about how it came into being, because it feels important to be transparent about how I use genAI.
The trigger: The essay began to form in my head late one evening after a long meeting with someone whose company I genuinely enjoy. It was one of those conversations where ideas seemed to generate more ideas; we were bouncing them around like ping-pong balls. At some point, without really thinking about it, I slipped off my shoes, brought both feet onto the seat and folded my legs in front of me. I felt relaxed, comfortable and completely immersed in the conversation. My body had settled into the position that felt most natural. Later that evening, I began replaying the meeting. Had I made the other person uncomfortable? Had I looked unprofessional? Had I crossed an invisible line without realising it?
The essay: That experience became the starting point for a much larger question. Why had I absorbed the idea that professional posture was not expected to be comfortable? Why are office chairs designed as though there is only one correct way for a human being to sit? And what does something as ordinary as a chair reveal about the hidden rules we place on bodies, behaviour and belonging?
The experience was mine. The later anxiety was mine. The question grew from my own experience of unmasking at work. The research, the interpretation and the movement from a specific bodily experience into a wider systems argument had taken months to develop. The essay itself shows that progression clearly: the meeting, the bodily comfort, the retrospective social analysis and then the broader question about design and conformity.
Yet Pangram, the detector that had reassured me so completely with my earlier A+ grade in Human, had now classified every section of the essay as 100 per cent AI-generated.
The result unsettled me more than I expected.
I let an AI detector make me doubt my own authorship
I knew how the essay had been written. I knew where the experience came from, how the question had developed and how long I had spent working out what I wanted to say. I also knew that I had used AI as part of the editing process. I had used it to help with structure, flow and the places where my thinking moved faster than the reader could reasonably be expected to follow.
Still, seeing the label made me question myself. Had I used too much support? Had I accepted more changes than I realised? Had the writing somehow stopped being mine? Had I crossed a boundary without noticing it?
For a moment (well, a few hours) the detector’s score seemed to carry more authority than my own knowledge of the writing process.
Then I ran the same passage through another detector.
It classified the text as 100 per cent human.
The remaining tools placed the passage at almost every point between those two extremes: 70% human, 24% human, 15% human, 3% human.
The same piece of writing was apparently entirely human, entirely artificial and almost everything in between.
That was the moment the question I was asking changed.
I was no longer asking whether the first detector had caught me doing something wrong.
I was asking why I had trusted it more than I trusted myself.
And I knew I was not alone. I had seen other writers on Substack sharing similar experiences: inconsistencies that did not make sense or follow any clear pattern, alongside calls to keep writing despite Pangram’s looming presence.
Everyone knows AI detectors are bullshit, right?
Yet even with everything I knew about my own writing process, and everything I had seen of AI detection in higher education, that F grade for being 0% human made me spin.
If I was questioning my own writing, I could only assume that somebody else would question it too.
That made me sad, because what I was trying to say in that essay, and, in fact, in all my essays, matters. I try to reveal blind spots, bring hidden ableism to the surface and challenge the normative culture we have learned to accept without even seeing it.
The detector had not changed how the essay was made.
It had changed how I imagined other people might see it.
15% human and 85% neurodivergent?

One of the results stayed with me for another reason. One detector listed some of the features it had noticed: specific narrative detail, personal reflection, uneven sentence lengths and idiosyncratic punctuation… are you kidding me?! Apparently, even my commas have failed the humanity test.
My immediate, flippant response was: so, 15% human and 85% neurodivergent?
I do not mean that the features listed were diagnostic signs of autism or ADHD. They are not, and neurodivergent people do not share one writing style. But the description felt unexpectedly close to aspects of my natural communication style.
I notice details. I replay social interactions. I analyse the invisible rule underneath what happened. I often move between longer, associative passages and shorter statements that hold the central point. My punctuation follows the rhythm of my thinking, particularly in an early draft. My writing often begins with a concrete experience before widening into a question about the system that produced it.
That is recognisably mine.
I cannot prove from one small, home-grown experiment that a detector misclassified the passage because I am autistic. I cannot conclude that a particular system is biased against disabled writers. But I also cannot ignore that some of the features it appeared to treat as suspicious overlap with the way I naturally think and communicate.
This becomes more concerning when placed alongside evidence that people writing in English as an additional language are already more likely to be falsely accused of using AI. If these systems learn a statistical model of what human writing usually looks like, whose writing sits closest to the centre of that model? Whose writing becomes less recognisable as human because it follows a different rhythm, structure or pattern?
The problem may not be that a system has been explicitly taught to discriminate.
It may be that the assumed norm is too narrow.
That should make us cautious about assuming these systems perform equally well for other groups whose language may also fall outside the statistical norm, including neurodivergent writers. If a detector learns what “human writing” looks like from large collections of texts I want to know:
· Which humans are represented in those collections?
· What happens when a multilingual writer uses English differently?
· When an autistic person organises ideas differently?
· When an ADHD writer follows an associative route that makes perfect sense to them but not to a system trained to recognise more conventional patterns?
Measuring the wrong thing
The more I thought about my results, the more I realised I had been, yet again, asking the wrong question.
I wasn’t really trying to find out whether Pangram could detect AI.
I was trying to understand what it was actually detecting.
An AI detector can only examine the finished piece of writing. It cannot observe the process that produced it. It cannot see the months spent reading papers, the lived experience that shaped an idea, the conversations that refine our thinking, or the many discarded drafts and whole essays that never make it the Substack feed. It can only analyse patterns in language and infer how those patterns were produced.
This raises a deeper question about authorship.
If I develop an idea over months or years, write the first draft myself, then use AI to help me organise, clarify or edit my language, what exactly is the detector measuring? Is it detecting who did the thinking, or is it detecting that the finished text contains linguistic features associated with AI assistance?
This is not only a problem of false detection. It is a problem of recognition. When a system learns what human writing looks like, whose humanity does it learn to recognise?
This is not a new problem
We have long used visible outputs as proxies for things we cannot directly observe: an exam for learning, an essay for understanding, and polished language for the quality of the thinking beneath it. AI detection extends that same shortcut to authorship..
In her essay Jagged Intelligence, Melanie Mitchell explores the profoundly uneven capabilities of current AI systems. They can appear remarkably intelligent in one context, then fail unpredictably when the task changes only slightly. Try asking one to count the words in a passage and see how confidently it gets the answer wrong.
Yet their fluent language encourages us to assume a depth of understanding that their performance does not consistently support. This is why we repeatedly tell students not to outsource their thinking. It is also the concern that tools such as Pangram are intended to address. No one wants to read an essay produced without meaningful human thought behind it.
Interestingly, Mitchell also examines whether language alone can ever produce humanlike intelligence.
We can take that question and flip it.
An AI detector analyses language as a proxy for authorship because language is all it can see. But the finished text cannot reveal the experience, thought, intention and judgement beneath it. If language alone cannot fully capture intelligence, why should we trust it to determine authorship?
AI as a translation partner
There is another reason this matters to me.
I know from this Substack community, and from conversations with other neurodivergent writers, that many of us are not using AI because we have nothing to say. We are using it because we have too much to say and cannot always get it out in a form that other people can follow.
Some people use it to reduce the anxiety of beginning. Some use it to organise thoughts that arrive all at once. Some use it to check whether the tone of a message might be misunderstood. Others use it to translate the internal shape of an idea into language that can travel between minds.
I am autistic and ADHD, and communication has always been difficult for me. Trying to say exactly what I mean can be really challenging. I can hold complex and connected ideas in my head, but turning them into a linear explanation is much harder. I often know what I mean and still cannot find the words that will make it visible to somebody else. The thought exists, fully formed in one sense, but not yet in a form I can easily articulate.
AI has helped me bridge that gap.
It has not given me a voice.
The voice was always mine.
It has helped me access it.
For the first time, I feel able to communicate some of the complex and nuanced ideas that would previously have remained stuck in my head, or confined to conversations on a run or over coffee with friends.
I can put my thoughts onto the page and share them in places like Substack because I can ask AI to identify where I have jumped too quickly, where the reader needs more context, or where the structure I can see so clearly in my head has not yet been made visible on the page. I can use the response to understand what is missing and then decide how to express it.
That is not the same as asking a system to think for me.
It is closer to translation: from inside to outside, from simultaneous thought into sequential language, from something I understand intuitively into something another person can enter.
Jaime Hoerricks, PhD said it beautifully:
Technology helps me reach language. It does not replace the mind reaching for it.
For me, that process has reduced anxiety rather than replaced effort. It has made writing more possible, not less mine.
This is why the language of “AI-generated” can feel so blunt. It collapses the difference between outsourcing thought and gaining access to your own. It can turn an accessibility tool into evidence against the person using it.
And this is where the question of bias becomes even more complicated.
What the detector could not see
The question of whether some writers are more likely than others to be misread by these systems deserves much closer examination. There is emerging evidence that autistic writing can be disproportionately classified as AI-generated (Chambers & Kelley, 2025), just as earlier research identified serious concerns for people writing in English as an additional language. That is a much bigger question than I can do justice to at the end of this essay, so I will write a separate piece on this.
For now, these two experiments left me with three concerns.
1) I began the day worried that AI detectors might not be accurate enough. I still think that concern is justified, but the problem is more precise than simply saying the tools get things wrong. In my small experiment, the detectors did not agree with one another. They did not consistently identify when AI entered the writing process, and some became more certain that a text was AI-generated after further human editing. The same passage could be classified as entirely human by one system and entirely AI-generated by another.
Take home: That is not a stable basis for making confident claims about authorship.
2) My second concern is nuance. The categories offered by these tools may appear to distinguish between human, assisted and generated writing, but my results did not reliably diagnose the extent of AI involvement. Some detectors found no difference between my original draft and the AI-supported version. Others treated relatively limited support as near-total generation.
Take home: A percentage cannot tell us who did the thinking.
3) My final concern is what happens after the score appears. Readers do not see an “AI-generated” label and think only about statistical probability. They may see dishonesty, cheating. They may assume the experience was invented, the argument a jumble of newspaper clippings or that the writer contributed little beyond the prompt.
Take home: The score does not simply classify the text. It changes how the writer is seen.
The most important result of the past few days has not been my percentage score. It is moment I realised that I had allowed a detector to make me doubt my own authorship.
A person who knows exactly how a piece was made can still find themselves defending it against a system that does not know how it was made at all.
I am not arguing that every use of AI is the same, or that transparent use does not matter. It does. There are legitimate questions about where assistance becomes generation.
But those questions cannot be answered responsibly by looking only at the finished words and pretending they reveal the whole writing process.
An AI detector sees the surface.
It does not see the months of thought beneath it. The hours and days writing and re-writing.
The danger is not only that the technology may be wrong. It is that we are asking a pattern-recognition system to answer a human question it was never capable of seeing:
Who did the thinking?
If you would like a copy of the prompts I used. Or if you want to replicate the experiments I did with your own writing just message me.
References
Chambers, S., Kelley, M.C. (2025). The Misclassification of Autistic Writing as AI-Generated. In: Cristea, A.I., Walker, E., Lu, Y., Santos, O.C., Isotani, S. (eds) Artificial Intelligence in Education. AIED 2025. Lecture Notes in Computer Science(), vol 15879. Springer, Cham. https://doi.org/10.1007/978-3-031-98420-4_7
Liang W, Yuksekgonul M, Mao Y, Wu E, Zou J. GPT detectors are biased against non-native English writers. Patterns (N Y). 2023 Jul 10;4(7):100779. doi: 10.1016/j.patter.2023.100779. PMID: 37521038; PMCID: PMC10382961. GPT detectors are biased against non-native English writers - PubMed
Perkins et al. Int J Educ Technol High Educ (2024) 21:53 https://doi.org/10.1186/s41239-024-00487-w
Hadra et al. International Journal for Educational Integrity https://doi.org/10.1007/s40979-026-00213-1




This made me think that we need to place AI in a longer history of human beings acquiring new ways to extend perception, memory, and expression.
The issue is not confined to the Modern Age. From the period after the medieval world, art increasingly developed alongside technologies that changed what artists could see and do: optical devices, new pigments, printmaking, photography, and later many other technical means. The much-debated possibility that Vermeer used optical aids in Girl with a Pearl Earring would not make it less his painting. It would raise a more interesting question: what did the artist do with the possibilities the tool created?
The same is true much more broadly. Writing itself was once regarded as a troubling technology. In Plato’s Phaedrus, Socrates worries that writing will weaken memory and produce the appearance of wisdom without its reality. That was not an absurd concern, but it did not follow that writing should be rejected. Writing became one of the conditions under which human thought could develop in new ways.
AI may create real dangers, including the outsourcing of thought and the erosion of learning. But it may also enable forms of reflection, communication, and creative development that we have not yet understood. The crucial question cannot simply be whether a person used AI. It is what relationship they have formed with it: did it replace their thought, or did it help them develop, express, test, and share thought that remained their own?
That is why it seems especially perverse to hand the judgment over to AI detectors. In trying to protect human creativity, we risk allowing AI to decide which forms of human expression it is prepared to recognise as human.
Hi again Jayne, this Pangram thingummy has caused a furore. Did you see Sher's (the cognitive ecologist's) post about this? Also @sam illingworth at Slow AI about the new Salem witch trials. Both brilliant in different ways.
I started to worry as I write in a similar way to you...throwing down tangents and thoughts and associations as I go. A room would be filled with sticky notes if I still used them or A3 sheets with illegible scribble all over.
I checked a few posts and worked out pretty much what sends it into its ridiculous judgements. I could waste time trying to make alterations but what for?
My attitude is...I couldn't care less if someone uses AI. I enjoy reading the thing or I dont. Over time I would know if a writer wasnt writing theor own thoughts...through interactions with them. Not through AI tells but I do tend to see these sometimes.
But I do care, very much, that neurodivergent and ESL writers are being pushed into disadvantage and castigation with this attitude.
I think it takes a lot to understand the issues. I am following the one year course with Slow AI to educate myself, for one thing.
In many places of work I have been required to complete regular safeguarding and data protection training and equality and diversity training. I am wondering if we all should be required to properly educate ourselves on, at least the use of LLMs and image and voice AI tools.
Thank you for this comprehensive article.