Based on Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization. Code: PersonalBench.

I gave a digital twin five topics derived from my own writing and asked it to write about each one. Then I compared its answers with writing I had kept out of the profile.

The result was 0.616.

On its own, that number says almost nothing. Three reference points, shown here alongside the twin, gave it meaning:

Comparison LUAR similarity
Same person, held-out writing 0.938
Digital twin 0.616
95th percentile of unrelated authors 0.512
Mean of unrelated authors 0.423

The twin was closer to my writing than every author in the 50-person calibration set. It was also nowhere near the similarity between two groups of my own writing. That is a much more useful result than “it sounds like me.” It locates the twin between unrelated writers and the real person it is trying to represent.

A horizontal LUAR scale showing unrelated authors centered at 0.423, their 95th percentile at 0.512, the digital twin at 0.616, and held-out writing by Yash at 0.938.

The real PersonalBench run used in this post. Higher means closer to the profile writing. The cross-author values are computed against the same profile, so they are reference points for this run rather than universal thresholds.

Why “sounds like me” is hard to measure

A twin can repeat your favorite phrases, discuss your interests, or follow a style instruction. Each can look like personalization. None establishes that its writing carries your authorship signal.

The obvious evaluation also leaks the answer. Give a model one of your essays and ask it to continue the essay, then the source itself supplies vocabulary, punctuation, sentence rhythm, and tone. Even an unpersonalized model can borrow those features. In the paper, using the first sentence of the target post as the prompt inflated the unpersonalized same-author rate by 28 percentage points.

PersonalBench separates the topic from the writing used as ground truth. It also gives the final similarity score three reference points:

  1. How similar is the twin to the person’s profile?
  2. How similar is genuine held-out writing by that person to the same profile?
  3. How similar are unrelated people to that profile?

An optional fourth comparison runs a non-personalized model on the same prompts. This measures how much the twin improves over its underlying model.

How the evaluation works

Start with two disjoint sets of writing. The profile represents the person. The reference set stays held out and supplies both neutral topics and genuine-human answers.

PersonalBench flow: five profile documents form one LUAR embedding; five held-out references are reduced to topic-only prompts; a twin answers those prompts; LUAR compares the profile with twin answers, held-out human writing, and unrelated-author embeddings.

The twin receives only the five topic prompts. Profile writing and held-out answers stay inside the evaluation bundle.

The bundle freezes the documents, prompts, selected profile IDs, random seed, model identifiers, and file hashes. That matters when two twin versions are compared months apart. They should face the same test.

Step 1: turn five documents into an authorship representation

LUAR stands for Learning Universal Authorship Representations. It was trained for authorship verification: given writing samples, produce an embedding that helps distinguish authors. PersonalBench uses the LUAR-MUD checkpoint as its primary style metric.

For one evaluation group, define:

  • p₁ ... p₅: five profile documents
  • t₁ ... t₅: the twin’s answers
  • h₁ ... h₅: the corresponding held-out human documents

LUAR aggregates each set into one vector:

P = LUAR(p₁, ..., p₅)    profile embedding
T = LUAR(t₁, ..., t₅)    twin embedding
H = LUAR(h₁, ..., h₅)    held-out human embedding

The two headline scores are cosine similarities:

twin likeness = cosine(P, T)
human ceiling = cosine(P, H)

The name “ceiling” is shorthand for the same-person reference in this run. It is not a mathematical maximum. A different split of the same person’s writing can produce a different value.

Why five documents? Authorship is noisy at the level of one answer. Topic and content can dominate a short sample. Aggregating five documents gives LUAR repeated evidence about sentence construction, function words, punctuation, and other author-level regularities. The paper validated this choice on the Blog Authorship Corpus: authorship discrimination rose from AUC 0.76 for individual posts to 0.96 for five-post aggregates.

PersonalBench still computes a 1v5 LUAR diagnostic for every answer. The primary number is one 5v5 comparison, not the average of those five per-answer scores.

Step 2: calibrate the score

The same profile vector P is compared with 50 precomputed author vectors. Each calibration vector contains five documents from one unrelated author in the Blog Authorship Corpus.

sⱼ = cosine(P, Cⱼ)    for each of 50 calibration authors

This yields a distribution rather than a single “random human” number. PersonalBench reports its mean, median, standard deviation, 5th and 95th percentiles, and the twin’s percentile within it.

The human comparison answers one question: how close can two held-out groups from this person look? The cross-author distribution answers another: how close would unrelated human writing look? A twin score becomes interpretable only after both are visible.

If baseline answers are supplied, PersonalBench also calculates:

improvement = twin likeness - baseline likeness

             twin likeness - baseline likeness
gap closed = ----------------------------------
             human ceiling - baseline likeness

The fraction is left unclamped. A negative value means personalization hurt style likeness. A value above 100% is also possible. The calculation is unavailable when the denominator is zero or negative.

Reading my result

My profile contained ten posts. With seed 42, PersonalBench selected five of them and created five cases from five held-out posts. The twin wrote two paragraphs for each neutral topic.

The primary result was:

human ceiling       0.9379
twin likeness       0.6161
cross-author p95     0.5124
cross-author mean    0.4230

The twin landed at the 100th percentile of the packaged cross-author distribution. On this sample, its style signal is measurably closer to my profile than the unrelated-author reference set. The remaining gap to genuine held-out writing is 0.322.

The individual answers ranged from 0.488 to 0.587 on the 1v5 diagnostic. The lowest was the model-comparison case. Its twin answer began:

What I would want to understand here is not just whether the RL-trained model performs better, but what actually changed inside the network to produce that behavior.

My held-out version began with the concrete event first:

Reinforcement learning made this model better at math. Same 1B network, one RL step on verifiable rewards, and afterwards it works through word problems it used to fumble.

The difference is visible. The twin starts by describing an analysis someone could perform. My version starts with what happened and makes the internal change the problem. That observation fits the lower per-case score, though it does not explain the score by itself. LUAR is a learned embedding, and a single pair of openings is not a feature attribution method.

The report also includes function-word cosine, punctuation cosine, sentence-length difference, vocabulary-richness difference, and ROUGE. These are diagnostics. ROUGE measures overlapping content, so it should never be read as style likeness. In the paper, LUAR, an LLM trait judge, and function-word similarity had pairwise correlations below 0.07. Combining them into one convenient average would hide that they measure different things.

Run the same evaluation

Install the package and verify the offline fixture first:

git clone https://github.com/yashsawant22/personalbench.git
cd personalbench
python -m venv .venv
source .venv/bin/activate
pip install -e '.[eval]'
personalbench verify-install

Put at least five UTF-8 documents in each split. Keep the filename stems unique across both folders.

my-writing/
├── profile/
│   ├── email-001.txt
│   ├── essay-002.txt
│   └── ...
└── reference/
    ├── journal-101.txt
    ├── post-102.txt
    └── ...

Prompt preparation uses meta-llama/Llama-3.2-3B-Instruct. Accept its Hugging Face access terms, authenticate locally, then prepare a frozen bundle and export its neutral prompts:

hf auth login
personalbench prepare my-writing --output my.bundle
personalbench prompts my.bundle --output prompts.jsonl

Give prompts.jsonl to the twin. Save every response as one JSON object per line:

{"case_id":"case_0001","output":"The twin's response..."}

Then score it:

personalbench score my.bundle \
  --answers twin-answers.jsonl \
  --output runs/my-twin

If you also collect answers from the unpersonalized base model, add them to the same command:

personalbench score my.bundle \
  --answers twin-answers.jsonl \
  --baseline-answers base-answers.jsonl \
  --output runs/my-twin

The first real scoring run downloads the pinned LUAR files. The generated report.md contains the calibrated comparison and per-case diagnostics. The JSON artifacts retain the exact scores and reproducibility metadata for later analysis.

What the number cannot establish

My trial contains exactly five evaluation cases, which creates one complete 5v5 group. There is no bootstrap confidence interval with one group. A larger reference set would show whether 0.616 is stable across topics instead of being specific to these five technical posts.

The topics are also close to writing already present in my profile: model internals, robotics, and world models. PersonalBench separates topic wording from style, but topic and style are never fully independent in natural writing. Testing across more domains would make the claim stronger.

The frozen bundle makes this completed evaluation reproducible. Regenerating it has a weaker edge: the prompt model revision in this run is recorded as main, so an upstream model update could alter future topic summaries. Pinning that model to an immutable commit would close the gap.

Finally, this is a writing-style evaluation. It does not establish that the twin remembers facts about me, shares my preferences, gives correct answers, or has avoided training-data leakage. A high score says something narrower: across these prompts, its authorship representation moved toward mine.

A measuring stick for the next version

The PersonalBench paper started from a disappointing result. Across 50 authors and 1,000 generations, four inference-time personalization methods scored between 0.484 and 0.508 on LUAR. The cross-author human floor was 0.626 and the same-author ceiling was 0.756. The personalized outputs differed from one another, yet remained in the model’s style space. Those paper values are not thresholds for my run. Each profile needs its own held-out and cross-author comparisons.

The practical tool turns that research protocol into a test for one person’s twin. Its main value is repeatability. Freeze one bundle, run every new twin version against it, and keep the human and cross-author references fixed. Then “this version sounds more like me” becomes a claim with a number that can move, regress, and eventually close a measurable gap.