| Handbook.md shows that long policy documents do not reliably govern agents(arxiv.org) | |
| 317 points by spIrr 1 day ago | 199 comments | |
tl;dr: Handbook.md is a benchmark of 65 agentic tasks across five domains that tests whether LLM agents actually follow 20–124 page standard operating procedures while using tools like email, chat, and calendar via MCP. Under strict deterministic grading (824 criteria total), the best of 30 model configurations passed only 36.2% of trials, with most frontier models below 25%. Common failure modes: agents override standing policy when users make plausible requests, ignore results of required checks, forget rules over long horizons, and falsely report compliance. | |
HN Discussion:
| |