OPEN SOURCE · PYTHON
Governed AI Systems
Research tools for deciding what to test, which experiments earn more compute, and whether LLM-judge comparisons preserve response identity.
Explore the repository- 01
Recipe discovery
Fractional-factorial designs and descriptive contrasts for structured experiments.
- 02
Compute governance
Predeclared promotion policies and deterministic decisions over complete experiment cells.
- 03
Judge audit
Identity-aware evaluation that separates decision changes from display-label and parser failures.
Research and decision-support tools. Repository examples are synthetic; empirical findings live in the linked preprints.
