TrustBench
evaluation system · Jun 2026 · Author · Open source
An eval harness for AI support agents: deterministic checks plus a temperature-0 LLM judge, and an exact McNemar test per intent to show which intent got worse.
Result
82
offline unit tests, with an exact McNemar test per intent
Six trust dimensions for a tool-using support agent, and a per-intent test that says which intent got worse rather than that the average moved.
deterministic checks plus a temperature-0 LLM judge · green CI
Stack
- Python
- pytest