Authorship verification works better as similarity, not classification
The arXiv paper “Contrastive Learning for Authorship Verification” shows why pairing two texts and learning their stylistic distance can beat a direct classifier, with useful caveats for real-world attribution systems.
TL;DR: Treat authorship verification as a similarity problem, not a label-picking problem, but do not confuse a high benchmark score with proof of identity in the wild.
What did the paper actually show?
The primary source is the arXiv paper “Contrastive Learning for Authorship Verification,” listed under both cs.CL and cs.LG. The paper reports that contrastive learning outperformed a classification-based approach for authorship verification under the tested settings.
That framing matters.
A classifier asks, roughly, “same author or different author?” and learns to output that decision directly. A contrastive setup trains the model to place writing samples from the same author closer together, and samples from different authors farther apart. That sounds like a small modeling choice, but it changes the job from memorizing decision boundaries to learning a usable style space.
The strongest reported result is a ModernBERT Bi-Encoder that reaches 98.4% accuracy on the PAN21 authorship verification task. That is a serious number for that benchmark. It is also not a universal authorship machine.
PAN21 is a task with defined data, splits, and evaluation rules. A model that does well there may still struggle when the writing is short, edited by someone else, translated, ghostwritten, AI-assisted, or intentionally disguised. The paper’s own details point in that direction: performance depends on loss function, batch size, training duration, the pre-trained model, input context length, and random text span data augmentation.
That is not a footnote. It is the story.
Why does contrastive learning fit authorship better?
Authorship verification is naturally comparative. Most real questions look like this: “Did the same person write these two things?” Not, “Which one of 10,000 named people wrote this?”
Contrastive learning fits that shape. It can compare two samples without requiring a closed list of candidate authors. That makes it more useful for practical workflows where the candidate pool is incomplete or changing.

The bi-encoder choice also matters. A bi-encoder processes each text separately, then compares their representations. That can be faster and easier to scale than approaches that require every pair to be processed jointly from scratch. For an operator, that opens the door to indexing writing samples and comparing new material against a reference set.
But speed is not truth. Style is a signal, not a fingerprint.
People shift tone by audience. Editors flatten style. Teams write under one brand voice. LLMs can imitate style, and humans can prompt them to do it. Short posts may not contain enough signal. Long context helps, and the paper identifies input context length as one of the important factors, but length alone does not solve the attribution problem.
The interesting practical pattern is not “AI can identify writers.” It is narrower: if you have enough comparable text, and if the task resembles the training and benchmark setup, contrastive models can produce a strong similarity signal.
Where would this actually be useful?
I would use this kind of model as a triage layer, not a judge.
For platforms, it could flag suspicious account takeovers when a user’s writing style changes sharply. For publishers, it could compare submitted work against known contributor samples. For security teams, it could help cluster scam messages or impersonation attempts by stylistic similarity. For academic or legal settings, it might support investigation, but it should not stand alone.
The catch is calibration. A 98.4% benchmark accuracy does not tell you your false positive rate on your own corpus. You need local evaluation: same document lengths, same genres, same editing patterns, same incentives to evade detection. You also need a policy decision for what the score means. Is it a quiet flag? A human review trigger? A block? Those are different systems.
Practitioner’s Take: If you are building authorship verification, start with a contrastive bi-encoder baseline instead of a direct classifier. Build a small internal eval set from your real writing pairs, include edited and AI-assisted samples, and measure false positives before you wire the score into any action. The missed catch is that authorship models are most useful when they reduce a review queue, not when they pretend to prove who wrote something.