OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward finds that VLM judges often approve failed computer-use runs. Its benchmark and open reward models offer a more grounded way to evaluate trajectories without paying frontier-model costs.
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed.
Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring.
OSReward evaluates vision-language judges on human-verified computer-use trajectories and adds **OSReward-Hard** and **OSReward-Multi**. The study finds a shared leniency bias: judges often classify failed runs as completed. Do not treat one model-judge verdict as ground truth for browser or desktop agents. Calibrate against human-labeled failures, track false approvals, and consider the released **OS-Shepherd 9B and 35B** models for repeatable scoring. The authors report commercial-level judging at **30–60% lower cost** than frontier models, but the supplied material gives no per-platform scores or evidence that these reward models generalize to a builder’s own UI and task mix.
This adds a standardized, cross-platform test of reward models used to judge computer-use agents and identifies false approval as a shared failure mode. It strengthens the case for human-calibrated trajectory evaluation and explicit tracking of leniency, while narrowing the reported cost advantage: the supplied evidence does not establish transfer to a builder’s particular interfaces, platforms, or task distribution.