Verdict
Reject, retest, observe, or advance: one of four, stated plainly.
Approach
These are the rules the machine runs by. They are adversarial by design: the work is to find the reasons a result should not be believed, before anyone commits capital or engineering to it. The same six apply to a candidate the system generated for itself and to a model somebody brings in from outside, because the question does not change with the author.
Method
The hypothesis, the universe, the horizon, the cost assumptions and the bar for success are written down first. A result graded against a bar that moved is not a result, and the only way to prove the bar did not move is to have published it.
On a k-day hold, daily observations overlap, so the naive t-statistic is inflated by roughly the square root of k. Significance is recomputed on non-overlapping blocks with a serial-correlation correction, and the Sharpe is deflated for the number of attempts actually made.
A candidate is compared against a shuffled twin and a plausible baseline, not against zero. Beating zero is easy. Beating a permutation of your own data, and a naive forecast that already works, is the question worth asking.
A result below the bar is not automatically a null. If the test could not have resolved an effect large enough to matter, the finding is 'underpowered', not 'no effect'. Absence of evidence and evidence of absence are recorded differently.
Rejected candidates are registered rather than discarded, with the reason and the scope of the rejection. A search that remembers only its successes will keep rediscovering the same dead ends and will overstate its own hit rate.
Surviving candidates are frozen and evaluated forward, point-in-time. A forward track cannot be refitted, which is precisely what makes it slow, unflattering, and worth more than any backtest.
One bounded deliverable
Bounded on purpose. It answers one question about one claim, and it attaches everything needed to disagree with it. This is the smallest useful thing the loop can produce, which is why it is the first thing a customer can buy. It is one workflow the platform performs, not the whole of what the platform is for.
Reject, retest, observe, or advance: one of four, stated plainly.
The tests run, the nulls used, and the numbers as computed, including the unflattering ones.
What the result does not establish: the universe it was not tested on, the regimes it did not see, the costs it assumed.
Where the claim breaks, and which conditions would change the verdict.
A report describes what the evidence supports about a specific claim under stated conditions. It is not investment advice, not a recommendation, and not a representation that any strategy will perform in future.
The same rules run continuously inside the system, on candidates nobody asked for. You can watch that happening in the observatory, or read how the pieces fit together on the platform page.
Bring a claim to test