EnvLoop.ai
EnvLoop Research

Researchmeasured, verified, released.

Technical reports, benchmarks and models from EnvLoop. Every entry ships with its data, evaluation protocol and stated limits.

Featured · Technical report · Model · Benchmark

Pev: A Calibrated Fast-Decision Model for Personal Agents

EnvLoop Research

Personalized agents, assistants that keep a long-term memory of one user and act on that user's behalf across shopping, travel, email and calendar, are quickly becoming a product category. Their bottleneck is less open-ended generation than personal judgement: deciding, many times per request, what this particular user would want. Asked to “re-order the dog food”, an agent has to infer that the user switched brands after an allergy, that the new $54 bag now exceeds the user's $50 approval rule, that an address the user asked it to forget must not be reused, and whether the request is worth interrupting the user for. We cast these judgements as seven typed question families and train Pev-27B, a fast decision model that answers each with a calibrated probability from a single forward pass, and release Pev-Bench, built from real public behaviour with user-disjoint splits. On a public test set of 720 new users, Pev-27B reaches 0.915 family-macro accuracy, ahead of gpt-6-astra (0.873), Kev-27B (0.823), Jev (0.788) and its base model (0.762).

Family-macro accuracyPublic TEST set, 720 new users
  1. Pev-27B0.915
  2. gpt-6-astra0.873
  3. Kev-27B0.823
  4. Jev0.788
  5. Qwen3.8-27B (base)0.762
Full summary and BibTeX
Bar chart of family-macro accuracy on the public TEST set and its two halves. Pev-27B is highest in every group, followed by gpt-6-astra, Kev-27B, Jev, the Qwen3.8-27B base model and Qwen3.5-4B.
Family-macro accuracy on the public TEST set, overall and on the gpt-6-astra-rendered and Claude Opus 5.5-rendered halves.
Other papers
Paper · Benchmark · Interactive demo

Recursive-Play: AI-Generated Interactive Challenges for Improving AI Agents

EnvLoop Research

AI-assisted task construction can supply executable challenges that evaluated AI agents do not yet complete, creating a concrete resource for subsequent improvement. Recursive-Play implements this idea with 150 interactive puzzle environments, ten mechanic families and six sequential levels per task. Author-only constructive witnesses and deterministic state verifiers establish reachability, while players receive pixels and neutral controls without the rules or solutions. Across eleven configurations and 1,650 independently eligible outcomes, 15 tasks remain unfinished by every configuration under the fixed 512-action / 720-second protocol, and two yield no completed level for any configuration.

Full summary and BibTeX
Ranked bar chart of complete-task rate and cleared-level rate for eleven model configurations on 150 tasks. The top configuration completes 135 of 150 tasks; three configurations complete none.
Complete tasks and cleared levels for the eleven evaluated configurations, same model order, two denominators.
Technical report · Code · Evidence explorer

CUA-RSIBench: Executable Data Research for Verifiable Computer Use

EnvLoop

Can a researcher agent turn browser-agent failures into better training data? CUA-RSIBench is a controlled pilot of data-centric research for verifiable computer use. Frontier researcher models write Python data factories, construct native task states in a real Kanboard application from public issue metadata, collect independently verified GUI demonstrations from a fixed gpt-5.6-sol teacher, and revise their datasets using selection feedback. Each valid candidate trains a fresh Qwen3.5-4B LoRA adapter through Tinker, and Harbor evaluates the student in E2B with an independent saved-state verifier. In the separate gpt-6-sol / gpt-6-luna extension, the student selected by gpt-6-sol scores 2/12 (16.7%) across two six-task repetitions, against 0/12 for the base student.

Full summary and BibTeX
Pipeline diagram: researcher code, native task states, GUI experience from a gpt-5.6-sol teacher, training data, Tinker SFT of a Qwen3.5-4B LoRA, E2B and Harbor trials, selection, and sealed final tests, with permitted feedback from selection back to the researcher.
The executable data-research loop. The researcher cannot access credentials, hidden answers, or final-test feedback.
Collaborate with EnvLoop Research

Questions, data requests or partnerships.