CUA-RSIBench: Executable Data Research for Verifiable Computer Use
Summary
Can a researcher agent turn browser-agent failures into better training data? CUA-RSIBench is a controlled pilot of data-centric research for verifiable computer use. Frontier researcher models write Python data factories, construct native task states in a real Kanboard application from public issue metadata, collect independently verified GUI demonstrations from a fixed gpt-5.6-sol teacher, and revise their datasets using selection feedback. Each valid candidate trains a fresh Qwen3.5-4B LoRA adapter through Tinker, and Harbor evaluates the student in E2B with an independent saved-state verifier. In the separate gpt-6-sol / gpt-6-luna extension, the student selected by gpt-6-sol scores 2/12 (16.7%) across two six-task repetitions, against 0/12 for the base student.

Details
- Two separate cohorts: the original cohort (gpt-6-astra, gpt-5.6-sol) completes 10 research rounds and 7 training candidates; the extension (gpt-6-sol, gpt-6-luna) completes 10 research rounds and 6 training candidates.
- Selected students are evaluated on each cohort's own sealed final instances. Scores, training resources, invalid submissions and infrastructure failures are reported separately by cohort and are not pooled into a four-model ranking.
- A separate v0.6 qualification note documents environment and verifier work for six real-software application cells; it reports qualification evidence, not full-study results.
Scope and limits
This is a small data-research pilot, not evidence of sustained recursive self-improvement or broad computer-use generalization. It uses DOM-assisted browser use on one application.
Citation
@techreport{envloop_cua_rsibench,
title = {CUA-RSIBench: Executable Data Research for Verifiable Computer Use},
author = {{EnvLoop}},
institution = {EnvLoop},
url = {https://github.com/EnvLoop/cua-rsibench}
}