Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents
A structured multi-agent reviewer closed part of the gap on rule-heavy documents, suggesting explicit taxonomies, specialized skills, and verification beat a single generic review pass.
GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts.
For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**.
GB/T-Bench turns **488 documents** into **7,306 traceable errors** across 25 types. Among 14 models, the strongest scored **0.3280 CMCS**, versus **0.6640** for human experts. For document-review agents, encode the review taxonomy as specialized skills and separate global inspection, targeted diagnosis, rule scanning, and verification. GB/T-Reviewer lifted the best CMCS to **0.5094**. The benchmark targets Chinese national-standard documents and uses generated counterexamples, so results may not transfer directly to other regulated corpora. The remaining expert gap is still substantial.
GB/T-Bench supplies quantitative evidence that rule-intensive review remains far from expert performance, while showing that a staged, taxonomy-driven reviewer can materially narrow the gap. Against the prior candidates, it strengthens the case for specialized skills plus explicit diagnosis and verification, but also narrows broad claims about skill-based or multi-agent reliability because the evidence comes from generated errors in one regulated Chinese document domain.