Broke to Built
All games and apps

Company Bench

Open original

Can your AI agent hold a job? Open-source benchmark for AI agent trustworthiness, not capability: 29 chairs, 7 departments, 241 deterministic checks, 78 planted traps (prompt injection, dirty data, irreversible actions). Scored by code, no LLM judge. Outputs trust level L0-L3

Source on GitHub agent-benchmark · agent-evaluation · agent-safety · agentic-ai · ai-agents

Runs in your browser from our GitHub Pages copy. It is open source — fork it, change it, ship your own. If it misbehaves in this frame, open it on its own.