Key Points
- The Economist compared 1.2m words of its own prose with text from four leading chatbots.
- Only Anthropic's Claude used more em-dashes than human writers; ChatGPT used markedly fewer.
- The study says AI tells are now long words, long sentences and thin punctuation.
The latest:
Em-dashes are no longer a reliable sign of machine-written text, according to a study by The Economist that compared its own journalism with articles rewritten by OpenAI’s ChatGPT, Anthropic’s Claude, Google’s Gemini and xAI’s Grok. The comparison covered 55,940 sentences and 1.2m words. The publication said the models’ giveaways had shifted with successive software updates.
Details:
- The method: The Economist asked the four models to rewrite its articles without consulting the web, using AI-generated article summaries as prompts. The resulting corpus was checked against journalism from CNN, the New York Times and the Washington Post, plus excerpts from hit novels published between 1950 and 2022. The study’s full methodology was not detailed in the piece.
- On em-dashes: The widely held belief that chatbots stuff prose with em-dashes no longer holds after the most recent updates, the study found. Only Claude used more of them than human writers, while ChatGPT used markedly fewer than any other writer examined. No figures for the size of those gaps were published.
- The new tells: Models now favour polysyllables such as significant, increasingly and consequences, rarer words like interdependence and reindustrialisation, scientific terms including parameter and methodology, and nominalisations. All four models did this, Gemini and Claude most of all. The study did not rank the models numerically.
- Punctuation: A better marker is text with little punctuation at all, according to the study: the models used fewer commas and semicolons than humans and hardly any parentheses, partly because they write longer sentences and partly because they do not quote experts. And was their most overused word.
- Sentence shape: Chatbot sentences tend to be long, with paragraphs rarely broken by short statements, the study said. For liveliness the models reached for rhetorical devices including not X but Y, not only but also, and the rule of three. ChatGPT and Claude used these more per 1,000 sentences than other models and humans.
- Expert caution: Karolina Rudnicka, a linguist at the University of Gdansk, said there is no single style of AI writing, just as there is no single style of human writing, noting that writers have idiosyncrasies and bots may too. The article did not say whether she reviewed The Economist’s data.
- Detection tools: Pangram, described as a leading detection firm, claims 99.98% accuracy and has partnered with the blogging platform Substack on such a tool. The Economist said detectors are black-box algorithms that can produce false positives and give no reasons for their conclusions. Pangram’s response to that criticism was not reported.
- A disputed prize: Some alleged that AI-generated prose won this year’s Commonwealth Short Story prize, with judges praising its quiet authority. The Commonwealth Foundation denied the claim. The report named neither the accusers nor the entry involved, and gave no outcome to the dispute.
- Scale of use: AI text drafts more than a third of new websites by one count cited in the report, and is helping students write essays and probably scientists write papers. The count’s source and date were not identified, and no figure was given for academic use.
Background:
Earlier attempts to detect machine text relied on scouring for suspicious words or comparing papers written before and after chatbots became publicly available. The Economist said those methods struggle to separate AI quirks from wider language trends.
Between the lines:
The study’s own framing points to a closing window: it reports that with every update AI writing grows more similar to human prose, and that Pangram, while successful now, may struggle in future. Tommie Juzek of Florida State University noted the models are trained on human writing and human feedback, adopting what people find impressive and dropping what they do not. That mechanism implies any published tell invites its own removal.
What’s next
Whether the next model updates erase the remaining markers the study identified — Latinate vocabulary, long sentences and sparse punctuation — and whether detection firms such as Pangram publish accuracy results against those newer versions.