02 August 2026

AIn't Necessarily So, Part 2


AI robots writing essays like this one
Part II — A Rigorous Approach to Identifying AI‑Generated Writing

Last time, we focused on fuzzy ‘soft’ clues in determining AI fingerprints in text, especially identifying common phrases and constructs. Computer people refer to fuzzy methods as heuristics.

We noted triplets.
“He was tall, dark, and handsome.”
We noted parallel negatives.
“It wasn’t black. It wasn’t white. It was charcoal grey.”
We noted common phraseology.
“It’s not just warm, it’s burning.”

Today we slide away from soft dissection to hard facts and more scientific methods. Assertions have been made that combining hard and soft techniques can result in 80-90% accuracy. Google claims it can identify close to 98%.

Note that I used Google Gemini and Microsoft Copilot to research and organize this, part 2 of this article. Thus throughout, you’ll notice artifacts as unintended examples.

Objective Linguistic Signals, Statistical Profiles, and Computational Methods

While intuitive reading can reveal stylistic quirks that feel artificial, a scientific approach to AI‑text detection relies on measurable linguistic properties. These properties emerge from how large language models generate text: token by token, guided by probability distributions learned from massive corpora. By analyzing those distributions, we can identify patterns that differ from human writing in consistent, quantifiable ways.

This section outlines the major pillars of a hard, objective approach — the kind that complements your Part I by grounding intuition in data.

➀ Lexical Analysis: Measuring Word‑Level Patterns

Lexical analysis examines the words themselves — their frequency, diversity, and distribution. AI‑generated text tends to exhibit:

  • Lower lexical diversity — measured via type–token ratio; AI favors mid‑probability vocabulary.
  • Function‑word overuse — glue words like however, moreover, indeed, additionally.
  • Uniform vocabulary patterns — humans spike unpredictably into slang, rare words, or idiosyncratic phrasing.
  • Rare‑word avoidance — models avoid low‑probability tokens unless prompted.

These signals arise because AI models optimize for clarity and coherence, which pushes them toward “safe,” middle‑probability word choices.

➁ Syntax Analysis: Sentence Structure and Rhythm

Syntax analysis examines how sentences are built — clause structure, punctuation, and rhythm. AI text often shows:

  • Consistent clause length — sentences follow similar patterns and pacing.
  • Over‑regular grammar — few deviations, few fragments, few stylistic breaks.
  • Predictable transitionsMoreover, In addition, Ultimately, However.
  • Low syntactic entropy — humans vary complexity more dramatically.
  • This uniformity reflects the model’s goal: maximize coherence and minimize confusion.

    ➂ Punctuation

    Perhaps I give less credence to punctuation than I should because the Macintosh option key makes typographical life easy. Instead of two hyphens to simulate an n-dash, option-hyphen plugs in a real dash. Likewise, option-colon drops in a true ellipsis instead of typing three dots.

    Even without a Mac, word processors like Microsoft Word performs conversions such as transforming flat quotation marks to curly quotes, thus my reluctance to attach too much attention to perfected punctuation.

    That said, AIs sprinkle in more m-dashes than a breathless self-published Mary Sue story. Current AIs throw in a lot of dashes, the really wide ones.

    Perfected Punctuation
    character human proper
    ellipsis...
    n-dash--
    m-dash--
    single quote'‘’
    double quote"“”

    ➃ Semantic Analysis: Meaning, Depth, and Conceptual Structure

    Semantic analysis looks at how ideas are organized and expressed. AI text tends to exhibit:

  • High semantic consistency — few contradictions or digressions.
  • Topic over‑coverage — AI exhaustively lists subtopics to “cover the space.”
  • Shallow originality — limited conceptual leaps unless prompted.
  • Generic framing — broad, universal statements anchoring paragraphs.
  • Humans, by contrast, often wander, contradict themselves, or introduce unexpected angles.

    ➄ Perplexity: Statistical Predictability of Text

    Perplexity measures how surprising a passage is to a language model.

  • Low perplexity → predictable → typical of AI
  • High perplexity → surprising → typical of humans
  • AI‑generated text tends to have very low perplexity, because it is produced by the same statistical engine used to measure it. Tools that use perplexity include:

  • GPTZero
  • OpenAI’s classifier
  • Cross‑entropy scoring
  • Edited AI text, however, can raise perplexity — making detection harder.

    ➅ Burstiness: Variation in Sentence Length and Complexity

    Burstiness measures how much sentence structure varies.

    • Humans show:
      • Long sentences
      • Short fragments
      • Abrupt shifts
      • Irregular rhythm
    • AI shows:
      • Low burstiness
      • Smooth, even pacing

    Burstiness is one of the strongest statistical indicators of human authorship, especially in long‑form writing.

    ➆ Stylometric Fingerprinting: Author Identity Through Writing Style

    Stylometry analyzes the “fingerprint” of an author’s writing style — their idiolect, quirks, and habits. AI text typically lacks:

  • Idiolect — no personal quirks or signature phrasing.
  • Stylistic variability — humans shift tone depending on mood, audience, or genre.
  • Natural digressions — AI rarely meanders.
  • Stylometric tools include:

  • JStylo
  • Writeprints
  • Signature stylometric analysis
  • These methods are especially powerful given long samples.

    ➇ Structural and Metadata Clues

    Even when prose looks human, structural patterns can reveal AI origin. Common signals include:

  • Perfect paragraph symmetry
  • Overuse of enumerated lists — AIs love lists
  • Overuse of bulleted lists — Consider this an example
  • Predictable section ordering
  • Lack of temporal markers — humans reference time, place, personal context.
  • These clues often appear in polished AI essays and reports.

    ➈ Hybrid Detection Tools: Combining Multiple Signals

    Modern detectors combine perplexity, burstiness, stylometry, and semantic analysis. Examples of tools include:

    • GPTZero
    • DetectGPT
    • Turnitin AI
    • Copyleaks
    • QuillBot AI
    • Grammarly
    • HuggingFace detectors

    These tools go beyond analyzing basics, they compare text against known AI‑generation patterns.

    Conclusion: Hard Detection Complements Soft Deduction

    Part I focused on intuition — tone, phrasing, emotional cadence, suspiciously neutral voice. Part II provides the science — statistical regularities, lexical smoothness, syntactic uniformity, semantic predictability.

    Together, they form a dual‑system framework:

  • Soft heuristic detection: “This feels like AI.”
  • Hard science-based detection: “Here’s measurable evidence.”
  • This combination is far more reliable than either method alone.

Note: Wikipedia is exceptionally vulnerable to AI contamination. They operate a project to identify and mitigate AI effects. Their article, Signs of AI Writing, is well worth a read.

AI robots competing for creativity

No comments:

Post a Comment

Welcome. Please feel free to comment.

Our corporate secretary is notoriously lax when it comes to comments trapped in the spam folder. It may take Velma a few days to notice, usually after digging in a bottom drawer for a packet of seamed hose, a .38, her flask, or a cigarette.

She’s also sarcastically flip-lipped, but where else can a P.I. find a gal who can wield a candlestick phone, a typewriter, and a gat all at the same time? So bear with us, we value your comment. Once she finishes her Fatima Long Gold.

You can format HTML codes of <b>bold</b>, <i>italics</i>, and links: <a href="https://about.me/SleuthSayers">SleuthSayers</a>