← Back to Blog

AI Essay Writers vs Human Writing: What Professors Can (and Can’t) Detect

2026-08-05

Author: Tian Yi | Published: August 2026 | Reading time: 14 minutes

Abstract: As AI writing tools become ubiquitous on university campuses, the cat-and-mouse game between detection software and student workflows has intensified. But the conversation has been dominated by a misleading question: “Can Turnitin catch ChatGPT?” This article moves beyond surface-level paranoia to examine what actually happens when a professor sits down to evaluate a paper. We explore the mathematical logic behind AI detection metrics like perplexity and burstiness, dissect the linguistic signals that make prose feel authentically human, and map the widening gap between what automated detectors flag and what experienced professors actually notice. The answer is less about detection scores and more about the nature of academic writing itself — a craft that demands critical thinking, not just coherent sentences.

The 2 AM Panic

It is 2:13 AM. Marcus stares at a blank Google Doc, his political science paper on Latin American populism due in nine hours. ChatGPT sits open in the next tab. Whether he submits the raw output, lightly edits it, or uses it only as scaffolding determines not just his grade but his academic standing. And the professor reading his submission? She is looking for something deeper than a probability score.

What AI Detectors Actually Measure: The Mathematics of Suspicion

Before we can understand what professors detect, we must understand the two core mathematical concepts behind AI detection: perplexity and burstiness. Neither is a fingerprint of AI authorship — both are statistical proxies with significant limitations.

Perplexity: How “Surprised” a Language Model Gets

Perplexity measures how predictable text is according to a language model. Low perplexity means every word is roughly what the model expected; high perplexity means the text is surprising. The counterintuitive truth: AI-generated text tends to have lower perplexity than human writing. LLMs gravitate toward the most statistically probable next token, while humans make idiosyncratic choices no probability model would predict. Detectors essentially ask: “Would a language model have been surprised by this sentence?” If not — if the text is smooth and predictable — the detector raises a flag.

Burstiness: The Rhythm of Human Variation

Burstiness measures how sentence length, structure, and complexity vary within a passage. Human writing is naturally bursty — a short, punchy sentence followed by a longer, meandering one with subordinate clauses and semicolons, because that is how human thought flows. Then another short one. AI-generated text exhibits low burstiness: sentences are similar in length and structure, paragraphs follow predictable patterns, the rhythm never changes.

Table 1: AI Detection Metrics — What They Measure and What They Miss

Metric

What It Measures

What AI Text Looks Like

What It Misses

Perplexity

Statistical predictability of word choices

Low perplexity (predictable, safe word choices)

Human writing that is formulaic or technical

Burstiness

Variation in sentence length and structure

Low burstiness (uniform rhythm, repetitive patterns)

Human writing with consistent style (e.g., scientific prose)

Token probability distribution

How evenly probability is spread across possible next words

Concentrated on top-1 or top-5 tokens

Edited AI text where humans have introduced variation

Repetition analysis

Frequency of recycled phrases and structures

Higher n-gram repetition within and across paragraphs

Deliberate repetition for rhetorical effect

This reveals why detection is fundamentally a probability game, not a certainty game. Low perplexity and burstiness could signal AI generation — or a meticulously edited human draft, a non-native speaker’s careful prose, or a technical paper where formulaic writing is the norm. The metrics measure deviation from an imagined “average human,” and actual humans deviate from that average constantly.

The Signals That Make Writing “Sound Human”

If detectors hunt for low perplexity and burstiness, what positive signals make writing register as authentically human? Understanding these is essential for students using AI tools ethically — the goal is genuine intellectual engagement, not tricking detectors.

Sentence Length Variation: The Invisible Signature

Human writers vary sentence length instinctively. A paragraph of human prose breathes: long sentences build momentum, short ones land like punctuation. This variation serves rhetorical purpose.

Consider two versions of the same argument:

Version A (low burstiness): Populist movements in Latin America have historically emerged during periods of economic instability. These movements tend to promise redistribution of wealth. They also appeal to nationalist sentiment. Their leaders typically position themselves against established political elites. This pattern has repeated across multiple countries.

Version B (high burstiness): Populism in Latin America does not emerge from a vacuum. It rises, almost without exception, during moments of economic rupture — when inflation devours wages, when the gap between the governed and the governing becomes unbridgeable, when people look at their political class and see, rightly or wrongly, an enemy. Then come the promises. Redistribution. National pride. Revenge against the elites. The pattern repeats because the conditions repeat.

Both convey the same information, but Version B feels human because its rhythm breathes — long sentences stretch, short fragments punctuate, the final sentence resolves. This is what burstiness indirectly measures and what professors unconsciously register as authorial presence.

Vocabulary Diversity and the Long Tail

Human vocabulary follows a Zipfian distribution, with individual lexical fingerprints: pet words, unusual adjectives, occasionally misused words in revealing ways. AI-generated text operates in the safe middle, avoiding the long tail of rare words. This lack of lexical ambition signals disengaged writing to professors in humanities and social sciences.

Structural Unpredictability: The Argument That Surprises

Human arguments loop back, anticipate objections, build toward points not obvious from the first paragraph. This structural unpredictability is hard for AI to replicate — it requires understanding not just what argument is being made, but what the reader expects. The best student essays contain “the turn” — a moment where the argument pivots, complicates itself, or acknowledges a limitation. AI essays march linearly without ever surprising the reader.

Table 2: Human Writing Signals — How They Manifest and Why They Matter

Signal

Human Writing Pattern

AI-Generated Pattern

Why Professors Notice

Sentence length variation

Wide range; short sentences for emphasis, long ones for development

Uniform medium-length sentences; limited rhythmic contrast

Creates the “feel” of engaged thinking vs. automated generation

Vocabulary diversity

Zipfian distribution with personal word preferences

Concentrated in high-frequency vocabulary; avoids rare words

Lexical ambition signals genuine engagement with the material

Structural unpredictability

Arguments that pivot, loop, or complicate themselves

Linear progression from thesis to supporting points to conclusion

Surprise is a marker of original thought, not just coherent arrangement

Idiomatic fluency

Natural use of colloquial transitions, hedging, and emphasis

Stiff or overly formal transitions; limited use of hedging

Reveals whether the writer “thinks” in academic English or translates into it


Beyond the Detection Score: What Professors Actually Look For

Here is a fact that complicates the entire detection debate: most professors do not rely primarily on AI detection scores. The score from Turnitin, GPTZero, or any other tool is, at best, a starting point for inquiry. Professors evaluate papers holistically, and the signals they rely on overlap only partially with what algorithms measure.

Citation Quality: The Tell Nobody Talks About

AI language models can generate citations. What they cannot reliably do is use citations well. A student who has engaged deeply with sources integrates them at the sentence level — distinguishing what a source explicitly claims, what it implies, and what it leaves unaddressed. AI-generated citations cluster at paragraph ends, are rarely integrated mid-sentence, and often reference sources that do not quite say what the text claims. Some AI tools even hallucinate citations entirely.

This is where Sodpen takes a fundamentally different approach. Rather than generating citations from a language model’s memory, Sodpen’s citation engine draws from authoritative, vetted academic sources with standardized citation formatting — ensuring every reference maps to a real, verifiable source. This distinction — between hallucinated citations and curated ones — often separates a paper that triggers suspicion from one that survives scrutiny.

Argument Coherence: The Thread That Connects Everything

A strong academic paper is a chain of reasoning where each link supports the next. AI-generated essays often exhibit “surface coherence” — the text flows smoothly, but the underlying argument does not hold together. Claims are stated but not defended. Transitions are decorative. The paper reads like someone summarized the topic rather than argued about it.

Voice Consistency: The Ghost in the Machine

Every human writer has a voice — patterns and quirks that persist across assignments. A sudden shift in voice between papers, or even between sections of the same paper, is often more revealing than any detection score. The laziest form of AI cheating — pasting raw output unchanged — is also the easiest to catch: the voice is not the student’s but a language model’s, and that voice, however polished, is recognizable.

Table 3: AI Detector Tools vs. Human Professor Evaluation — What Each Catches and Misses

Evaluation Dimension

Turnitin AI Detection

GPTZero

Human Professor

Perplexity-based flags

✅ Scores text predictability

✅ Core detection mechanism

❌ Does not consciously evaluate this

Burstiness patterns

⚠️ Limited; more reliance on perplexity

✅ Explicit burstiness scoring

✅ Intuitively senses rhythmic monotony

Citation authenticity

❌ Not evaluated (separate plagiarism check)

❌ Not evaluated

✅ Checks sources against actual publications

Citation integration quality

❌ Not evaluated

❌ Not evaluated

✅ Notices shallow vs. deep source engagement

Argument coherence

❌ Not evaluated

❌ Not evaluated

✅ Primary evaluation criterion

Voice consistency

❌ Not evaluated

❌ Not evaluated

✅ Notices shifts across assignments

Factual accuracy

❌ Not evaluated

❌ Not evaluated

✅ Subject-matter expertise catches errors

Template detection

⚠️ Indirectly through structural patterns

⚠️ Indirectly through burstiness

✅ Recognizes formulaic argument structures

False positive risk

⚠️ Well-documented; flags ESL and formulaic writing

⚠️ Similar limitations

✅ Contextual judgment reduces false positives

False negative risk

⚠️ Edited or paraphrased AI text often passes

⚠️ Heavy editing defeats detection

✅ Can still detect through argument depth

The asymmetry is clear: automated detectors are good at what they measure and blind to everything else. Professors are the inverse — not optimized for detecting statistical word-choice patterns, but exceptionally good at detecting intellectual engagement, or its absence. A paper that passes Turnitin with a 0% score can still fail because the argument is hollow, the citations are decorative, and the voice sounds like nobody in particular.


The Great Irony: Bad AI Papers vs. “Robotic” Human Papers

The irony: an AI-generated paper with poor argumentation often scores lower on AI detectors than a carefully written human paper in a technical field.

Why? Because the metrics work against engagement. The best human writing in technical disciplines — engineering reports, lab write-ups, quantitative social science — is low in burstiness and perplexity by design. These fields reward clarity, consistency, and convention. A well-written methods section is supposed to be predictable. Run that exemplary human writing through an AI detector, and the detector panics.

Meanwhile, a student who prompts ChatGPT for a philosophy essay might receive output that — while substantively weak — varies its sentence structure enough to avoid the lowest burstiness thresholds. The essay is philosophically empty but statistically diverse enough to register as “possibly human.”

Table 4: Four Quadrants of Detection — Quality vs. Detectability


High AI Detection Score

Low AI Detection Score

Strong Academic Quality

Carefully written humanities paper with idiosyncratic style that nonetheless triggers perplexity flags due to polished prose

Strong technical/scientific paper with formulaic structure that reads as “predictable” to detectors; strong argument despite low burstiness

Weak Academic Quality

Raw AI output submitted without editing — detectable and substantively poor

Heavily edited/paraphrased AI output, or poor human writing that happens to be statistically diverse — undetectable but still earns a bad grade

This quadrant reveals the fundamental misalignment: AI detectors and academic quality measure fundamentally different things. Optimizing for “not getting caught” may still produce a failing paper. A student who never touches AI may produce work detectors flag as suspicious. The only stable strategy is genuine academic quality — and professors are already grading for that.

How AI Writing Tools Should Be Used Ethically

The binary narrative — “AI is cheating” versus “AI is the future” — is a false choice. The question is not whether to use AI but how and at what stage.

The Research and Ideation Phase: Where AI Excels

AI tools are genuinely useful as research accelerators and brainstorming partners. They can:

· Summarize complex readings to help identify key arguments before diving into source texts

· Generate outlines providing structural starting points — what Sodpen’s unlimited free essay outlining feature is designed for

· Suggest counterarguments that widen the scope of analysis

· Identify gaps in preliminary research by mapping existing scholarship

The ethical line: AI should help you find the territory, not write the map. Every claim must be verified against primary sources; every outline reshaped by your own analytical priorities.

The Drafting Phase: Scaffolding, Not Substitution

Using AI to generate a complete first draft that you lightly edit violates academic integrity at most institutions. But using AI to generate paragraphs you substantially rewrite, provide alternative phrasings for sentences you drafted, or check logical flow across your own writing sits in a gray zone most policies have not fully addressed. The guiding principle: if the intellectual work of argument construction, evidence evaluation, and rhetorical choice was performed by you, AI played a legitimate supporting role. If those cognitive tasks were outsourced, AI wrote your paper.

This is precisely the philosophy behind Sodpen’s design: unlimited free essay outlining, authoritative and vetted academic sources with standardized citation formatting, and AI-powered draft generation — not to replace the writer, but to transform blank-page paralysis into a structured starting point the student builds upon with their own critical thinking.

The Revision Phase: The Safest Integration Point

Using AI to review a draft you have written — to check grammar, suggest clarity improvements, or identify weak transitions — is generally uncontroversial. Grammarly has done this for years without triggering academic integrity debates. The key distinction: AI as editor vs. AI as author.

The Academic Integrity Gray Zone: Where Does “AI Assisted” End?

As of 2026, most institutions distinguish between “unauthorized AI use” (outright generation and submission) and “authorized AI assistance” (instructor-approved tools). But the boundary is a gradient, not a bright line. Consider:

· Using AI to generate an outline, then writing every word yourself. Most professors accept this.

· Writing a rough draft, then using AI to “improve academic tone” throughout. Substantial AI-generated language — crosses the line at many institutions.

· Using AI for topic sentences only, writing evidence and analysis yourself. Genuinely ambiguous, case-by-case.

· Writing the entire paper, then using AI only for grammar and citation formatting. Widely accepted, functionally equivalent to Grammarly.

The safest — and most educationally sound — approach: use AI for pre-composition (research, outlining, source gathering, structural planning) and write the paper yourself. This aligns efficiency with integrity: you learn more, produce better work, and never wonder whether you crossed a line.

The Counterargument: What If AI Detection Is Fundamentally Flawed?

There is a strong argument that AI detection is so unreliable it should not be used for academic integrity decisions at all.

OpenAI itself discontinued its AI text classifier in July 2023, citing “low accuracy.” Multiple studies demonstrate that AI detectors disproportionately flag writing by non-native English speakers — a deeply concerning bias. Sophisticated paraphrasing tools can reduce detection scores to near zero. And as language models improve, the statistical gap narrows — GPT-4 produces higher-burstiness text than GPT-3.5, and future models will close the gap further.

This does not mean academic integrity is doomed. It means the question was never “can software catch this?” The question was always “does this paper demonstrate genuine learning?” — the one professors have been asking since long before LLMs existed.

The Verdict: Detection Is a Distraction — Quality Is the Signal

After examining the mathematics of detection, the linguistics of human writing, how professors evaluate work, and the ethical landscape of AI use, one conclusion emerges:

AI detection scores are a weak signal. Academic quality is a strong signal. And the two are poorly correlated.

The student who submits an AI-generated essay with no substantive engagement will likely receive a poor grade — whether or not a detector flags it — because the essay lacks argument depth, citation integration, and intellectual voice. The student who uses AI thoughtfully as a research and outlining assistant, then writes a paper reflecting genuine critical thinking, will likely earn a strong grade — and need not worry about detection because there is nothing to detect.

Professors describe a consistent experience: suspicions arise not from Turnitin scores but from reading a paper and sensing that nobody was home. The argument is present but not alive. Citations are formatted but not engaged. Sentences are grammatical but not thoughtful.

This is, paradoxically, good news for students doing the work. It means the system is not rigged by unreliable algorithms. It means the path to a good grade remains: read deeply, think carefully, write honestly. And for students paralyzed by fear of a false positive, the evidence is reassuring: professors are educators evaluating whether learning occurred — and learning leaves traces no language model can simulate.

Ready to move from a blank page to a structured, well-sourced draft — without crossing academic integrity lines? Sodpen provides unlimited free essay outlining, standardized citation formatting from authoritative academic sources, and AI-powered draft generation designed to support your writing process, not replace it. Write better. Grade higher. Start building your paper at sodpen.com.