# Computer Use Benchmark: Environment Validation Status

Updated: 4 October 2026. This snapshot assumes Tinker is unavailable. It records application controls and executable infrastructure, with zero full-study model outcomes. Completed controls below are environment and verifier checks.

## Evaluation idea

An agent reads evidence in real software, changes the requested business or document state, and preserves unrelated state. An independent verifier reopens the saved result. Each reference trio checks a correct result, a deliberately wrong result, and unchanged state after fresh initialization. Source review, actual native inputs, ownership, time limits and cleanup remain separate requirements.

The six cells cover Excel web, PowerPoint web, LibreOffice desktop, Odoo Community, GitLab CE and Magento. The full protocol retains four researcher configurations, 20 selection tasks and 100 final tasks per cell. The study follows the improvement-and-evaluation structure of RSIBench; task details are specified in the repository.

## Executable components

The latest non-training container CI passed build, network-isolated smoke and registry push. Root reopened all 1,727 files in its source archive and confirmed the continuation code matches the reviewed source. These checks made zero model or Tinker requests. Public source can be built independently; anonymous container-image pull remains unverified because organization policy keeps the package private.

| Cell | Current implementation | Verified controls | Remaining execution |
| --- | --- | --- | --- |
| Excel web | Original Office worker and SEC data factory | Offline original packages and independent formula checks | Current native cloud edit, save, download and reset |
| PowerPoint web | Original Office worker and WDI data factory | Earlier-source cloud controls remain preserved | Current upload has no server acknowledgement; native lifecycle remains incomplete |
| Desktop | Calc, Impress and Writer under Source52 | Three public TRAIN trios; first original selection trio saved 0/1/0, three distinct closed guests, 16 native decisions independently reopened | Other 19 selection tasks, original final cohort and all execution paths |
| Odoo Community | Native14 and Qualification9 | All 20 original selection trios with source review; first original final-candidate trio independently reaudited | Original 100-task qualification is running; complete saved and source review afterward |
| GitLab CE | Source146 and original V3 reference workflow | First original selection trio saved 0/1/0; isolated cold-start diagnostic passed | Original remaining 19 trios; separate continuation and complete source review |
| Magento | Native10 and original reference workflow | 12 original selection trios independently audited; 11 retained with original provenance and one fresh trio | Remaining eight trios; aggregate audit and complete source review |

The desktop repair queries the next required cell after formula commit. It retains the original position classifier, input guard, verifier, formulas and budgets. Previously failed attempts receive zero control credit. The Magento continuation preserves original evidence and rejects old or partial records relabeled as fresh. GitLab startup failure and unscored startup success remain distinct evidence; a successful boot does not establish a service fix.

## Training boundary and remaining study

Tinker supplies fine-tuning, checkpoint storage and checkpoint sampling when enabled. It is not required to build the environments or independent verifiers. Training and checkpoint comparison remain paused. Ordinary API baselines can run only after the corresponding native evaluation gates pass.

Official totals remain 0/600 admitted final identities, 0/24 completed researcher chains and 0/3,000 reconciled matched model outcomes. The running Odoo final-candidate controls do not change those totals. The final empirical paper requires completed, independently audited model comparisons; this PDF is an environment-validation report.

Office uses the existing account and owned test folder. No second account is required. Credentials, account identifiers, private task bodies, hidden allocations and gold state are excluded from public artifacts.

## References and evidence

RSIBench: https://rsibench.co/. Repository: https://github.com/EnvLoop/cua-rsibench.

Verified build: https://github.com/EnvLoop/cua-rsibench/actions/runs/37163903827. Public evidence: docs/evidence/non-training-progress-2026-10-04.json.
