Handbook.md shows that long policy documents do not reliably govern agents
Hacker News
Read full postResearchers introduced Handbook.md, a benchmark with 65 tasks simulating enterprise agents following long policy documents. Tested models often failed to fully comply with complex, lengthy policies, highlighting challenges in governing AI agents with extensive instructions.


