Frozen at GPT-4 It Caught 0.6%. Hexagon Took the Proofs

An October arXiv paper shows a GPT-4-frozen detector catching at most 0.6% of GPT-5-era rewrites. arXiv capped submissions on October 1. Hexagon is archiving proofs no author will claim.

Printed detector-score table on a desk next to a laptop showing an arXiv abstract, desk lamp, no logos or readable secrets

A new screening study on arXiv as 2610.11599 retrains an AI-writing detector at each model release and then freezes older ones. Frozen at GPT-4, it catches at most 0.6 percent of the rewritten abstracts from any GPT-5-generation version they tested. Frozen at Qwen2.5, it still catches 86 percent of Qwen3.8. That split is the paper. It is not a vendor ranking.

TechTimes on October 8 described the other half of the same week: arXiv’s October 1 caps (two submissions per calendar month, three active at a time) after 40,363 submissions in September, and a nonprofit archive called Hexagon for mathematical results that no human will sign. Detectors that do not travel across generations and a preprint server that will not take unclaimed LLM proofs are the same operational problem. The text keeps moving. The policy is a rate limit.

What they actually measured

The authors distinguish AI-generated text from AI-rewritten text. The experiment here is the second: take a human abstract, rewrite it with a given LLM, and see whether a detector trained on earlier rewrites still flags the new ones at a fixed false-positive rate on human abstracts.

They retrain at each release, using versions up to that point. They also freeze a detector on one generation and run it forward. The freeze is the result people will quote. A GPT-4 freeze catches at most 0.6 percent of GPT-5-and-later rewrites. A detector trained on GPT-4 catches 0.2 percent of GPT-5.4 rewrites in their pairwise grid.

This is not the same study as “detectors fail on unseen models” in the abstract. They already knew that. The shape they add is that the loss concentrates at a few generational boundaries rather than accumulating smoothly with age.

A word-frequency detector with no neural network shows the same collapse at 5 percent and 0.5 percent false-positive operating points, and with two alternative rewrite prompts. If your procurement slide says you will beat this with a better transformer classifier, the n-gram control is the slide you should keep.

Two collapses, and several that did not happen

When they train on everything before version n and test on n, 18 of 20 versions with prior data stay above 85 percent true positive rate. Median among those 18 is 95.4 percent. The lowest survivor is 85.5 percent on GPT-5.1.

Two transitions fall off the table:

  • GPT-5, first of that generation: 26.6 percent
  • Muse-Glimmer, first open-weight model of Meta’s Muse family: 3.8 percent

Qwen’s move from Qwen2.5 into Qwen3 still detected at 94.7 percent. Meta’s move into Llama 4 mixture-of-experts still detected at 99.2 percent. Architecture change is not a reliable predictor of a screening outage. Generation labels sometimes are.

Cross-vendor contemporaneous models can transfer. A detector trained on GPT-4o (November 2024) catches 99.5 percent of Llama 3.3 rewrites in their grid. Same-vendor later generations can fail. Plan around generations, not logos.

Muse-Glimmer is also the weak spot for their combined screens: union of per-version detectors at 83.6 percent true positive versus 66.2 percent for a pooled detector. If you only remember one open-weight name from the paper, remember that one.

Union versus pooled is a 1 percent false-positive argument

They calibrate a union of per-version detectors against a single pooled detector. At the operating point they chose, the union flags 1.1 percent ± 0.4 percent of held-out human abstracts; the pooled detector flags 0.8 percent ± 0.3 percent. Median true positive rate is 99.3 percent for the union and 97.2 percent for pooled across the 23 versions in the table.

That is a small FPR gap for a large TPR gap on the ugly versions. Journals that “just want one model” are choosing the pooled column. Labs that can afford to keep a detector per generation are choosing the union column. Neither column is a human reader.

We already had a different failure mode in insecure reporters hiding the negative result. That paper was about models omitting the finding. This one is about screens that do not see the rewrite after a generation jump. Do not merge them in a slack summary.

arXiv’s October 1 cap is a queue, not a detector

TechTimes reports that on October 1, 2026, arXiv announced a cap of two submissions per calendar month and three total active submissions, after calling the volume a watershed. September 2026: 40,363 submissions, more than four times the same month a decade earlier. Moderators are volunteers. Entirely LLM-generated content is already prohibited; the cap is what you do when prohibition does not slow the firehose.

A cap does not tell you whether an abstract was rewritten by GPT-5. It tells you a lab can no longer spray preprint numbers at the wall. If your group had a habit of slicing one project into five arXiv notes, that habit is now a calendar.

The paper’s 0.6 percent freeze result is why a cap is not a substitute for screening. You can rate-limit garbage and still fail to flag the rewritten abstract that came in under the cap.

For open-weight serving costs we already separated tokens from the bill. Screening cost is a third column: retraining when the generation changes, or accepting that a frozen GPT-4 detector is theater.

Hexagon is for proofs nobody will sign

The TechTimes piece describes Hexagon as a nonprofit archive standing up because arXiv will not host fully AI-generated mathematical results, even when those results might be real and checkable. The failure mode they sketch is specific: frontier models surface adjacent lemmas in algebraic geometry, combinatorics, or theoretical computer science while solving something else. No author wants to claim them. The old Ginsparg rule, human intellectual contribution, has no slot for an unclaimed lemma.

Whether Hexagon becomes a serious venue is a 2027 question. The 2026 fact is that a math community is building a side door instead of pretending arXiv will relax the human-contribution rule.

If you run a lab, do not treat Hexagon as a place to dump generated papers that arXiv rejected for being slop. The stated job is unclaimed, possibly verifiable math with no human claimant. That is a narrower object than “the model wrote a PDF.”

Hospital annotation work we covered at temperature zero had a human in the loop on purpose. Hexagon is the opposite experiment: keep the artifact when the human steps away. Different risk. Do not copy the workflow.

What to log this month if you screen manuscripts

Write the detector’s training cutoff next to the score. “AI probability 0.92” with no generation tag is how you get the 0.6 percent freeze and still believe the dashboard.

When a major generation ships, retrain or add a detector. The paper’s sequential results say most in-generation steps stay above 85 percent TPR if you do that. The GPT-5 and Muse-Glimmer rows say you do not get a grace period at the boundary.

Keep a word-frequency baseline. If it also drops, you are not looking at a transformer bug.

Do not use arXiv’s two-per-month cap as evidence your screening works. The cap is a moderator survival tool.

If a colleague forwards a Hexagon link, ask whether the object is an unclaimed proof with a machine-checkable artifact, or a blog post with theorems. Only the first is the story TechTimes actually reported.

If you run a journal, put the false-positive number in the author-facing policy. Union at 1.1 percent ± 0.4 percent of human abstracts is still a handful of innocent people in a thousand. Pooled at 0.8 percent ± 0.3 percent is slightly kinder to humans and worse on Muse-Glimmer. Pick one, publish it, and say which generation the detector last saw.

Do not hide the 23-version table behind a vendor name. GPT-4o (May 2024), GPT-4o (August 2024), and GPT-4o (November 2024) are separate rows in their grid. Llama 3.1, 3.3, and 4 Maverick are separate rows. Qwen2.5, Qwen3, and Qwen3.8 are not interchangeable for a freeze. If your internal wiki says “we detect GPT and Llama,” you have already flattened the paper.

A reasonable lab log for October:

  • detector id and training cutoff date
  • operating FPR
  • whether the screen is pooled or union
  • the model family you think the submission used, if you even ask
  • the action when the score is high: human review, not auto-reject

Auto-reject on a frozen GPT-4 screen is how you reject nothing from GPT-5 and still congratulate yourself.

Hexagon does not belong in that log unless the object is an unclaimed, checkable proof. Mixing “we might post this lemma” with “we screened this abstract” is how both processes get sloppy.

The 40,363 September submissions are the volume number that made arXiv reach for a monthly cap. Four times a decade-ago September is a staffing problem. Volunteer moderators cannot be the detector. The paper’s freeze result is why adding more volunteer eyes also fails: they are not reading 40,000 PDFs for GPT-5 rewrite tells.

If your lab hit the two-per-month cap, merge notes. Do not open a second account. The cap is per the policy TechTimes reported, not a suggestion.

Qwen3 surviving at 94.7 percent after Qwen2.5 is the counterexample to “every new version breaks the screen.” GPT-5 at 26.6 percent and Muse-Glimmer at 3.8 percent are the examples that do. Write both on the same wiki page so nobody ships a single slogan.

GPT-5.1 recovering to 85.5 percent in the sequential setup is why “we retrain” is the actual mitigation, not “we bought a better classifier.” The freeze at GPT-4 never recovers. That is the 0.6 percent line.

Llama 4 Maverick staying detectable after a mixture-of-experts jump (99.2 percent in their sequential test) is the other counterexample. Do not use MoE as a synonym for undetectable.

Muse-Glimmer at 83.6 percent even on the union screen is the row that should scare a journal that only tests on GPT-4o rewrites. 66.2 percent pooled is worse. If your vendor demo used last year’s OpenAI models, ask for Muse-Glimmer and for GPT-5, not only for GPT-4o May 2024. Ask for the training cutoff too.

The practical sentence for an editor: a screen trained last spring is a GPT-4 freeze with extra UI. Budget the retrain, or stop telling authors you can tell.