TrustBench

evaluation system · Jun 2026 · Author · Open source

An eval harness for AI support agents: deterministic checks plus a temperature-0 LLM judge, and an exact McNemar test per intent to show which intent got worse.

Result

82

offline unit tests, with an exact McNemar test per intent

Six trust dimensions for a tool-using support agent, and a per-intent test that says which intent got worse rather than that the average moved.

deterministic checks plus a temperature-0 LLM judge · green CI

Stack

  • Python
  • pytest

Proves