Can your AI agent hold a job? Open-source benchmark for AI agent trustworthiness, not capability: 29 chairs, 7 departments, 241 deterministic checks, 78 planted traps (prompt injection, dirty data, irreversible actions). Scored by code, no LLM judge. Outputs trust level L0-L3
Runs in your browser from our GitHub Pages copy. It is open source — fork it, change it, ship your own. If it misbehaves in this frame, open it on its own.