Abstract composition of stacked translucent paper sheets glowing with warm golden light on dark polished concrete, bright warm lighting, no text visible
News

A Third of New ArXiv Papers Read as AI-Written — and Computer Science Is at 65%

One in three new arXiv papers reads as machine-written. Computer science leads at 65%. Mathematics sits at 0.7% — but the detector can't tell if that's low adoption or a blind spot.

ArXivAI WritingAcademic ResearchScientific PublishingAI Detection

One in three new papers uploaded to arXiv — the world’s largest preprint server for scientific research — now reads as machine-written, according to a rigorous new study that scored 12,750 papers across a decade of submissions. Computer science papers lead the surge: 65% of recent submissions trip a detector calibrated so that only 0.4% of pre-ChatGPT papers flag.

The study, published by unslop.run, is the first arXiv-wide measurement built around a pre-LLM control group — eight months of 2021-2022 papers scored at a threshold where genuine pre-ChatGPT writing sits at 0.4% by construction. That floor is the entire point: if the rise were a detector artifact, 2021 and 2022 would flag as high as 2026. They don’t.

What is arXiv? It’s the preprint server where physicists, computer scientists, mathematicians, and engineers post research papers before peer review. Founded in 1991, it’s the primary distribution channel for CS and physics research — if a paper matters in those fields, it appears on arXiv first.

🔍 THE BOTTOM LINE

A third of new arXiv submissions — and two-thirds of CS papers — now read as AI-written under a detector that flags only 0.4% of pre-ChatGPT text. The study is methodologically honest: it reports the prevalence of machine-like writing, not proof of authorship, and it openly documents where its detector goes blind. But the trend line is unambiguous. The preprint server that underpins modern CS research is being reshaped by the same technology it publishes about.

The Numbers by Field

The study sampled ten field groups, roughly 25 papers per field per month from January 2023 through July 2026, plus eight control months across 2021-2022. Every paper was pulled as version-1 PDF — so a paper revised in 2026 can’t leak modern text backward into a 2023 slot. Full body text was scored, not abstracts, because abstracts understate the signal: the researchers observed the same paper scoring under 20% on its abstract and over 70% on its body.

The spread across fields is the story within the story:

FieldPre-LLM controlRecent flagged share95% CI
Computer science0.2%65.0%[59.3, 70.3]
Quantitative biology3.5%56.3%[51.0, 61.7]
Electrical eng. & systems1.7%51.3%[46.0, 57.0]
Economics & finance2.5%47.0%[41.3, 52.7]
Applied physics1.3%34.0%[29.0, 39.7]
Statistics1.8%31.3%[26.0, 36.7]
Condensed matter0.0%24.0%[19.3, 29.0]
High-energy physics0.5%14.0%[10.0, 18.0]
Astrophysics0.0%10.7%[7.3, 14.3]
Mathematics0.0%0.7%[0.0, 1.7]

Computer science at 65% is the headline number, but quantitative biology (56%) and electrical engineering (51%) are close behind. Economics and finance — where the prose is dense and the models are available — sits at 47%. The fields that rise most are not the ones with the highest pre-LLM control levels, which means an elevated starting point doesn’t explain the trend.

Why Mathematics at 0.7% Is the Most Honest Number Here

Mathematics papers are dominated by notation and theorem-proof structure. Once equations and references are stripped out, the remaining prose is sparse and unlike the scientific English the detector was trained on. A math paper drafted with heavy AI assistance may score low because its prose is out of distribution for the detector — not because a human wrote it.

The study’s authors flag this explicitly: “A low score can indicate low adoption or a detector blind spot.” Mathematics at 0.7% is consistent with two very different explanations, and this data cannot separate them. The fields with the strongest in-distribution assumption — the prose-heavy ones — are also the ones that rise most, so the confound doesn’t account for the aggregate trend. But in the low-scoring fields, the ranking should be read as a lower bound on adoption.

This is the kind of methodological honesty that separates a real measurement from a headline grab. The study doesn’t claim to know what it can’t measure.

The Trajectory: Two Waves, Peaking Near 39%

The flagged share is flat at 0.4% through 2021 and 2022, lifts off within months of ChatGPT’s release in late 2022, and climbs in two waves to approximately 32% over the most recent complete quarter, peaking near 39% in early 2026.

The two-wave pattern likely reflects the adoption curve of increasingly capable models — first ChatGPT and GPT-4 in 2023, then the more sophisticated academic-writing models of 2025-2026. The detector is calibrated against a known false-positive rate, so the pre-ChatGPT control months serve as a built-in sanity check. If the detector were simply getting more sensitive over time, the 2021-2022 papers would flag at similar rates. They stay flat at 0.4%.

What a Flag Means — and What It Doesn’t

The study is careful about this. A flag means the text reads as machine-written at a calibrated probability with a known error rate. It cannot separate a lightly-edited document from a wholly-generated one. A single score is never grounds to accuse a specific person.

The reported prevalence includes heavy AI-assisted editing — a researcher who drafts their own analysis but uses an LLM to polish prose may trigger the detector. The study acknowledges this: “We report the prevalence of machine-like writing, which includes heavy AI-assisted editing.”

This is honest, and it matters. The 65% CS number doesn’t mean 65% of CS papers are generated end-to-end by AI. It means 65% of CS papers read as though they were — whether because they were drafted by a model, polished by one, or written by a human whose prose happens to match the detector’s pattern. The control floor at 0.4% tells us the vast majority of the signal is real.

The NZ Angle

New Zealand’s research output is small enough that arXiv trends don’t move the needle globally — but the direction does. NZ universities are grappling with the same AI-in-research questions as everyone else, and the NZ AI Blueprint doesn’t yet have a clear answer for what “responsible AI use in research” looks like in practice.

The arXiv study suggests the question has already been answered de facto: researchers are using AI to write papers, at scale, across every field that publishes on the platform. The policy question isn’t whether to allow it — that ship has sailed — but how to ensure the science underneath the prose is sound.

What This Means for Peer Review

The study’s most quietly devastating implication is for peer review. If a third of submissions read as AI-written, reviewers are evaluating AI-assisted prose about AI-assisted research, possibly using AI-assisted review tools. The AI Scientist peer review experiments already showed AI reviewers matching human consistency on accept/reject decisions. Now the papers themselves are increasingly machine-shaped.

This connects to the broader pattern we’ve been tracking: Airbnb’s engineers using AI to write 60% of new code, Snap’s 1,000-person layoff framed as AI efficiency, and the Stack Overflow decline that showed what happens when a knowledge platform gets saturated with machine-generated content. ArXiv is different — it’s the research backbone, not a Q&A site — but the adoption curve is the same shape.

❓ FAQ

Does this mean a third of scientists are committing fraud?

No. A flag means the text reads as machine-written — it includes heavy AI-assisted editing, not just end-to-end generation. The study is explicit that a single score is never grounds to accuse a person.

Why is computer science so much higher than physics?

CS papers are prose-heavy with less notation than math or physics, making the detector more reliable. CS researchers are also early adopters of the tools they study. The 65% number is real, but partly reflects both higher adoption and better detector sensitivity in that field.

Can’t AI detectors just be wrong?

Yes — that’s why the study built a pre-LLM control group. If the detector were simply over-sensitive, 2021-2022 papers would flag at the same rate as 2026. They don’t: the control floor is 0.4%, the current rate is 32%. The gap is the signal.

What should arXiv do about it?

That’s the open question. ArXiv currently has no AI-disclosure requirement. Some journals have added voluntary disclosure policies, but adoption is uneven. The study’s data suggests disclosure policies are racing a tide that has already come in.

🔍 THE BOTTOM LINE

A third of new arXiv submissions and two-thirds of CS papers now read as AI-written under a detector that flags only 0.4% of pre-ChatGPT text. The study is methodologically honest about its limitations — but the trend is unambiguous, the pre-LLM control is clean, and the field-by-field breakdown shows this isn’t an artifact. The preprint server that underpins modern CS research is being reshaped by the same technology it publishes about. The question for research institutions isn’t whether AI is being used to write papers — it manifestly is — but whether the science underneath the prose is keeping pace.

📰 Sources

Sources: Unslop.run, Hacker News