Pev: A Calibrated Fast-Decision Model for Personal Agents
Summary
Personalized agents, assistants that keep a long-term memory of one user and act on that user's behalf across shopping, travel, email and calendar, are quickly becoming a product category. Their bottleneck is less open-ended generation than personal judgement: deciding, many times per request, what this particular user would want. Asked to “re-order the dog food”, an agent has to infer that the user switched brands after an allergy, that the new $54 bag now exceeds the user's $50 approval rule, that an address the user asked it to forget must not be reused, and whether the request is worth interrupting the user for. We cast these judgements as seven typed question families and train Pev-27B, a fast decision model that answers each with a calibrated probability from a single forward pass, and release Pev-Bench, built from real public behaviour with user-disjoint splits. On a public test set of 720 new users, Pev-27B reaches 0.915 family-macro accuracy, ahead of gpt-6-astra (0.873), Kev-27B (0.823), Jev (0.788) and its base model (0.762).
- Pev-27B0.915
- gpt-6-astra0.873
- Kev-27B0.823
- Jev0.788
- Qwen3.8-27B (base)0.762

Details
- Pev-27B is a rank-16 LoRA on Qwen/Qwen3.8-27B, supervised on a single option-label token and read out as a temperature-calibrated distribution over the options.
- Pev-Bench is generated from real public behaviour (Amazon Reviews 2023, UCSD Google Local) and real distractor text (Enron, OpenFlights), with labels fixed by a knowledge base before any text is rendered, exact label and answer-position balance, verbatim anchor checks and user-disjoint splits.
- Under a pre-registered protocol, the frozen model passed a one-shot hidden gate: family-macro accuracy rose from 0.754 to 0.905 (+15.1 points, 95% CI [+13.5, +16.8]), safety false negatives fell, and the share of questions it can answer alone within a 5% error budget rose from 0.54 to 0.93.
- The public test set covers 720 new users, half rendered by gpt-6-astra and half by Claude Opus 5.5. Pev-27B leads overall and on both halves (Holm-adjusted p = 5 × 10⁻⁴ for every comparison).
Scope and limits
Predicting which option a user will pick from real behaviour remains unsolved for every model, and gpt-6-astra keeps slightly fewer safety false negatives on two families.
Citation
@techreport{envloop_pev,
title = {Pev: A Calibrated Fast-Decision Model for Personal Agents},
author = {{EnvLoop Research}},
institution = {EnvLoop},
url = {https://github.com/EnvLoop/Pev}
}