Sign InOpen Brain
arXivPaperNeeds Review

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

A structured multi-agent reviewer closed part of the gap on rule-heavy documents, suggesting explicit taxonomies, specialized skills, and verification beat a single generic review pass.

arXiv · Aug 6, 2026
Open Source Open MarkdownOpen JSON
Source Summary

GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts.

Practical Implication

For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**.

Agent-Ready Context
GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts.

For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**.

The benchmark targets Chinese national-standard documents and uses generated counterexamples, so results may not transfer directly to other regulated corpora. The remaining expert gap is still substantial.
Connected Context · Feed7 Judgment

GB/T-Bench supplies quantitative evidence that rule-intensive review remains far from expert performance, while showing that a staged, taxonomy-driven reviewer can materially narrow the gap. Against the prior candidates, it strengthens the case for specialized skills plus explicit diagnosis and verification, but also narrows broad claims about skill-based or multi-agent reliability because the evidence comes from generated errors in one regulated Chinese document domain.

The Regression Tax: Decomposing Why Skills Help and Hurt LLM AgentsIts staged verification supports the candidate’s warning that procedural skills need output checking, while the remaining expert gap reinforces that skill gains should not be treated as uniformly reliable.Claude Science, an AI workbench for scientists, is now availableGB/T-Reviewer provides benchmark evidence for a domain-specialist workflow resembling the workbench’s coordinator, specialist, and reviewer decomposition, though in a narrower document-review setting.What Does Done Even Mean? Agents and Paperclip's Liveness Model - Dotta, PaperclipTraceable error labels and a separate verification stage make review completion evidence-based, reinforcing the candidate’s distinction between agent-declared completion and verified acceptance.The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationThe separation of global inspection, targeted diagnosis, rule scanning, and verification aligns with the candidate’s claim that long-horizon performance needs structured transitions rather than a collection of atomic skills alone.
Context Map
agentdata#multi-agent#skills#agent-reliability
Uncertainty
The benchmark targets Chinese national-standard documents and uses generated counterexamples, so results may not transfer directly to other regulated corpora. The remaining expert gap is still substantial.