research
An Evidence Ladder for Self-Improving Trading Systems
A self-modifying system needs stronger evidence than a one-off model because every accepted change can alter future decisions. This note outlines a staged evidence ladder for governing those changes.
Improvement claims need a reference point
A system cannot meaningfully say it improved unless the comparison identifies what changed, what stayed fixed, and which outcome is being measured. Comparing a new strategy with an old strategy on different data or under different costs mixes the effect of the change with the effect of the environment. A useful evaluation therefore freezes the target, metric definition, cohort, time boundary, and relevant execution assumptions before interpreting the result.
This sounds procedural, but it is central to learning. Without a stable reference point, a system can appear to improve simply because the evaluation moved in a favorable direction. Versioned experiments and immutable evaluation definitions make the comparison auditable and allow later reviewers to distinguish a genuine candidate improvement from a changed scoring rule.
Historical evidence should be cheap and skeptical
Historical evaluation is useful for filtering ideas quickly. It can reject candidates that fail basic performance, robustness, or cost tests without waiting for new market outcomes. Because the candidate was designed with historical information available, however, historical success should remain a weak form of evidence for promotion. The same data ecosystem that inspired the change can also reward it.
A strong historical gate therefore asks for robustness rather than one optimized score. Nearby parameters, alternate windows, conservative costs, and untouched validation periods reduce the chance that a candidate survives only because it matches one accidental path. Passing this gate earns a prospective test; it does not justify treating the candidate as demonstrated improvement.
Prospective evidence separates prediction from explanation
Prospective paper or shadow evaluation freezes the candidate before later outcomes are known. That ordering matters because it blocks a common source of accidental hindsight. The system records what it expected, waits for the outcome, and then evaluates the already-recorded decision. Restart-safe timestamps and immutable evidence identifiers help prove that the observation was not rewritten after the result arrived.
Prospective evidence should also compare the candidate with an appropriate baseline. A candidate that improves one metric while degrading another may not represent a useful capability gain. Direction accuracy, magnitude error, turnover, cost sensitivity, and risk can answer different questions, so the promotion rule should identify which tradeoffs are allowed before observing the result.
Promotion should be rarer than proposal
An adaptive system should be allowed to generate many hypotheses while requiring much stronger evidence to change the active policy. This asymmetry makes experimentation cheap and deployment conservative. Rejection is not a system failure; it is evidence that the evaluation process prevented an unsupported change from affecting later decisions.
The ladder can continue after promotion. Small live exposure, execution comparison, regime monitoring, and the possibility of rollback provide evidence that the improvement survives contact with real operations. The central principle is that learning should be represented as a sequence of increasingly realistic tests with preserved evidence, not as a single score that gives the system permission to rewrite itself.