Sign InOpen Brain
arXivPaperNeeds Review

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward finds that VLM judges often approve failed computer-use runs. Its benchmark and open reward models offer a more grounded way to evaluate trajectories without paying frontier-model costs.

arXiv · Jul 30, 2026
Open Source Open MarkdownOpen JSON
Source Summary

OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.

Practical Implication

Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.

Agent-Ready Context
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.

Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.

The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.
Connected Context · Feed7 Judgment

This adds a standardized, cross-platform test of reward models used to judge computer-use agents and identifies false approval as a shared failure mode. It strengthens the case for human-calibrated trajectory evaluation and explicit tracking of leniency, while narrowing the reported cost advantage: the supplied evidence does not establish transfer to a builder’s particular interfaces, platforms, or task distribution.

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?Desktop-Delta Bench offers a complementary diagnostic for false approvals by checking whether models understand the GUI state change produced by each action, not merely apparent completion.Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude InstituteHuman-verified trajectories and repeatable reward scoring fit Harbor’s loop of reproducible rollouts, outcome verification, and trajectory inspection.The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AIThe observed leniency of model judges supports retaining layered deterministic and human checks even when agent-based judging is added for variable trajectories.From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AIReplayable production traces provide the task-specific calibration set needed to test whether OSReward models generalize beyond their released benchmark distribution.
Context Map
benchmarkcoding#computer-use#agent-evals#agent-reliability
Uncertainty
The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.