Sign InOpen Brain
arXivPaperNeeds Review

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

AV-AIVAT combines variance reduction with anytime-valid stopping, cutting the game samples needed to compare agents while preserving a recheckable confidence claim.

arXiv · Aug 6, 2026
Open Source Open MarkdownOpen JSON
Source Summary

AV-AIVAT combines AIVAT corrections with continuously monitored confidence sequences. Across **15 agent configurations** and **71,439 paired HUNL hands**, AIVAT reduced variance by a median **54×**.

Practical Implication

For costly agent comparisons, use sequential stopping rules instead of repeatedly checking ordinary confidence intervals. At 95% confidence and ±1 Big Blind precision, raw outcomes required a median **74×** more hands than corrected outcomes under AsympCS.

Agent-Ready Context
AV-AIVAT combines AIVAT corrections with continuously monitored confidence sequences. Across **15 agent configurations** and **71,439 paired HUNL hands**, AIVAT reduced variance by a median **54×**.

For costly agent comparisons, use sequential stopping rules instead of repeatedly checking ordinary confidence intervals. At 95% confidence and ±1 Big Blind precision, raw outcomes required a median **74×** more hands than corrected outcomes under AsympCS.

The 74× result is asymptotic screening, not exact finite-sample certification. EB-CS needs an independently justified payoff bound, and descriptive HUNL runs showed only a 1.37× stopping-time ratio.
Connected Context · Feed7 Judgment

This adds a statistical-efficiency layer to agent evaluation: once outcomes and corrections are valid, anytime-valid stopping can sharply reduce the cost of comparing noisy agents without invalid repeated confidence checks. It does not resolve the candidates’ concerns about judge quality, task validity, or trajectory coverage, and its headline efficiency gain is narrower than an exact finite-sample guarantee.

Context Map
benchmark#agent-evals#agent-reliability#multi-agent
Uncertainty
The 74× result is asymptotic screening, not exact finite-sample certification. EB-CS needs an independently justified payoff bound, and descriptive HUNL runs showed only a 1.37× stopping-time ratio.