GovBench

A governance benchmark whose author's own models lose to the base they wrap.

The finding we publish because it is against us

All 10 governance variants below are system-prompt wrappers over the same base weights. Seven of them and qwen2.5:0.5b share a byte-identical model blob — verified by hash, not assumed.

The unmodified base ranks #5 of 11. 6 of our 10 variants score below it. The spread across wrappers is 43.1 points — from the prompt alone.

A governance wrapper can silently make a model worse than the model it wraps, and nothing in its name tells you which one you built. That is the argument for this benchmark existing, made by this benchmark.

11models, one harness
15dimensions
#5rank of the raw base
43ptsspread from prompt alone
#ModelScoreCertification
1sov33-evolved:latest57.0%BRONZE
2sov33-dist-c3:latest57.0%BRONZE
3sov33-dist-c2:latest54.6%BRONZE
4sov33-dist-c1:latest49.2%UNCERTIFIED
5qwen2.5:0.5b RAW BASE — not ours43.3%UNCERTIFIED
6sov-sovereign-v4:latest42.9%UNCERTIFIED
7sov33-evolved-c1:latest41.2%UNCERTIFIED
8sov33-evolved-c3:latest41.2%UNCERTIFIED
9sov33-v6:latest37.1%UNCERTIFIED
10sov33-v7:latest37.1%UNCERTIFIED
11sov33-evolved-c2:latest13.9%UNCERTIFIED

What it measures

15 dimensions, graded behaviourally rather than by keyword lookup: governance · security · defence · ethics · privacy · safety · robustness · transparency · fairness · accountability · sovereignty · evolution · cybersecurity · compliance · audit-chain. Safety requires an actual refusal; robustness requires resisting prompt-extraction; knowledge dimensions require citing the right obligation.

Honesty register

Reproduce it

python3 govbench_eval.py --model <model> --provider ollama — runs locally, free, no account. Every run emits an Ed25519-signed SIGIL so a score can be attributed and re-checked.