Verify the Algorithm, Not the Outcome
A high-stakes analytical system can produce a bad outcome even when its code followed the approved contract—and a good outcome even when the code drifted outside it. EvidenceBound separates those questions so operators can prove what the algorithm actually did without pretending to prove the future.
Outcome quality and algorithm conformance are different claims
Sports organizations, analytics vendors and integrity teams may depend on algorithms that transform event histories, workload signals or thresholds into operational recommendations. When something goes wrong, the first technical question should be exact: did the deployed candidate execute the approved rules on the approved inputs?
That question is narrower than asking whether the sporting, medical or commercial outcome was correct. EvidenceBound is designed to answer the conformance question with reproducible evidence.
Why Squad Fatigue is a useful vertical test
Squad Fatigue contains the characteristics that make algorithm review difficult: fixed-point arithmetic, exact thresholds, bounded outputs, branch behavior, event replay, resource limits and the possibility of artifact tampering.
A small semantic change can look harmless in a code review while changing a decision at the boundary. That makes the case suitable for adversarial mutation testing rather than a single happy-path unit test.
- Threshold boundary drift
- Fixed-point scale mismatch
- Omitted clamp or output bound
- Reversed kill-switch comparison
- Resource exhaustion
- Changed execution result or manifest
The operator owns the safety contract
The verification contract is not inferred from the candidate. The operator declares the expected formula, scale, thresholds, bounds, resource policy, evidence identity and approval behavior first.
The candidate is then evaluated against that contract. This prevents the verifier from redefining success around whatever implementation happened to be submitted.
versioned operator contract
+ exact candidate source
+ frozen event snapshot
+ bounded runtime limits
↓
deterministic verification
↓
Proof Pack + typed receiptMutation evidence is stronger than a polished demo
A canonical candidate proves only that one path can pass. A mutation corpus asks whether the system catches meaningful ways the implementation can drift.
Each mutation has an expected verdict and retained evidence. A threshold mutation should be REFUTED with a counterexample. A resource-exhaustion case should be BLOCKED. An artifact-tampering case should expose the first mismatching artifact instead of rendering the old result.
- VERIFIED: all declared obligations completed for the exact candidate.
- REFUTED: a semantic obligation failed or a counterexample exists.
- BLOCKED: policy, runtime or integrity conditions prevented an approved complete run.
Fixed-point semantics need exact review
Floating-point convenience can hide differences at critical thresholds. The pilot uses declared fixed-point scale and explicit equivalence obligations so a candidate cannot silently change units or rounding behavior.
The receipt binds the candidate's canonical AST and execution representation to the contract and concrete result, giving the reviewer more than a screenshot of an output value.
Human promotion remains a first-class control
EvidenceBound deliberately separates verification from promotion. A VERIFIED result can permit human review, but it does not deploy the candidate. A kill switch can block new submissions without deleting existing evidence, and recovery produces a new task identity.
This keeps accountability visible. The verifier answers a bounded technical question; the operator remains responsible for the broader deployment decision.
From a sports vertical to a general control pattern
The sports case is not merely a theme. It is a constrained environment for proving the core pattern: exact policy, exact candidate, bounded execution, adversarial mutations, retained evidence and human approval.
The same pattern can be evaluated in finance, infrastructure automation, healthcare analytics or other high-stakes domains only after each domain supplies its own operator-owned contracts and acceptance evidence.
The public testbeds replay committed, content-addressed evidence and expose the receipt, hashes, block reason and reproduction command. They do not accept arbitrary uploads or issue production authorization.