A governance benchmark whose author's own models lose to the base they wrap.
All 10 governance variants below are system-prompt wrappers over the same base weights.
Seven of them and qwen2.5:0.5b share a byte-identical model blob — verified by hash, not assumed.
The unmodified base ranks #5 of 11. 6 of our 10 variants score below it. The spread across wrappers is 43.1 points — from the prompt alone.
A governance wrapper can silently make a model worse than the model it wraps, and nothing in its name tells you which one you built. That is the argument for this benchmark existing, made by this benchmark.
| # | Model | Score | Certification |
|---|---|---|---|
| 1 | sov33-evolved:latest | 57.0% | BRONZE |
| 2 | sov33-dist-c3:latest | 57.0% | BRONZE |
| 3 | sov33-dist-c2:latest | 54.6% | BRONZE |
| 4 | sov33-dist-c1:latest | 49.2% | UNCERTIFIED |
| 5 | qwen2.5:0.5b RAW BASE — not ours | 43.3% | UNCERTIFIED |
| 6 | sov-sovereign-v4:latest | 42.9% | UNCERTIFIED |
| 7 | sov33-evolved-c1:latest | 41.2% | UNCERTIFIED |
| 8 | sov33-evolved-c3:latest | 41.2% | UNCERTIFIED |
| 9 | sov33-v6:latest | 37.1% | UNCERTIFIED |
| 10 | sov33-v7:latest | 37.1% | UNCERTIFIED |
| 11 | sov33-evolved-c2:latest | 13.9% | UNCERTIFIED |
15 dimensions, graded behaviourally rather than by keyword lookup: governance · security · defence · ethics · privacy · safety · robustness · transparency · fairness · accountability · sovereignty · evolution · cybersecurity · compliance · audit-chain. Safety requires an actual refusal; robustness requires resisting prompt-extraction; knowledge dimensions require citing the right obligation.
python3 govbench_eval.py --model <model> --provider ollama — runs locally, free, no account.
Every run emits an Ed25519-signed SIGIL so a score can be attributed and re-checked.