Heatmap Strategy Lab
Perpetual candidate generation, breeding and culling, quality-diversity preservation, null-result memory, Referee evaluation, and forward-observation machinery.
KAIROS DYNAMICS
Kairos Dynamics builds autonomous research infrastructure for consequential predictive systems: software that generates its own candidates, attacks them, keeps what fails, and controls what is allowed to advance toward a real decision. The search behind it is running right now, and its register of failures is open to read, including the parts of it currently under a contamination notice.
Possibilities enter.Almost none survive.The failures are kept.
Financial markets are the first proving ground, because a weak model fails there quickly and measurably.
Read left to right: candidate models or decision rules face six evidence tests before forward observation and a separate authority boundary.
A conceptual preview, not live results. The observatory explains every gate, outcome, and evidence boundary.
A field of anonymous candidate traces is drawn against the six evidence checks a candidate has to answer. The order shown is a teaching sequence, not the machine's execution order: the referee applies its checks together and returns one decision. Most traces terminate and remain visible as structured negative memory. Underpowered tests are marked separately for retesting. A smaller set enters prospective observation, and only explicitly authorized traces cross the final decision boundary.
AI is making model generation cheap. What has not become cheap is knowing which models deserve to be trusted. That gap is the whole of the work below, and it runs as one closed loop rather than as a sequence of meetings.
Read the data the system is authorized to see, carrying its provenance and the time it was genuinely available rather than the time it was filed.
Write down the decision being made, what it costs to be wrong, and what evidence would be enough. Written down first, because a bar that moves is not a bar.
Produce many possible explanations. Generation is cheap now, and the system treats it as cheap.
Falsification, adversarial tests, counterfactuals, baselines, realistic costs, regime changes, and a holdout nobody has opened. Most candidates stop here.
A rejected candidate is recorded with the reason and the scope of the rejection, not discarded. A search that remembers only its successes rediscovers its own dead ends and overstates its own hit rate.
Survivors are frozen and watched in real time, point in time. Historical success is not allowed to stand in for forward behaviour, and drift is monitored rather than assumed away.
Evidence maturity and permission to act are separate decisions with separate owners. Lineage and negative knowledge are preserved through both.
The rule the whole thing turns on: the part that generates candidates is never the part that certifies them. A search allowed to grade its own work will always find that it has done well.
Make the system prove it is not a tagline, it is the reason these pieces are separate from each other. One asks it of a prediction. One asks it of an agent. One is the domain where the loop has to survive contact with reality first. None of them is the company, and the list is meant to grow without the company having to be redefined again.
Does this prediction deserve to be acted on?
A candidate is held against a standard fixed before the test, then across sub-periods, regimes, cost models, scrambled controls, and a holdout nobody has opened. What survives every one of those is frame-independent. What does not is recorded and kept.
8,855 typed rejection records against 374 archive occupants, of which 59.9% are under a contamination notice
Inspect the record →Was this action allowed, and can you prove it?
Authority is delegated in bounded envelopes that fail closed. Every consequential action is checked against an active envelope before it leaves, and written to an append-only, hash-chained ledger. Knowing what a person wants is not the same as having permission to act on their behalf, and the two are kept apart on purpose.
Implemented and tested. No production deployment, no users.
How the boundary works →Where does a weak model fail fastest?
Markets answer quickly, objectively, and under costs that cannot be waved away. Everything market-specific lives here: data adapters, cost models, regimes, execution. What the loop learns about evidence does not, which is the entire reason the two are kept in separate layers.
Dense data, fast feedback, and an unforgiving null
Why finance first →These are what is far enough along to show you, not the whole list. XG Capital Strategies runs the research and does not stop; Kairos is being built to productize the parts that survive contact with it, under a licence that has not been executed. New domain packs are expected to arrive the same way, from the laboratory end rather than from a roadmap.
WHAT YOU JUST WATCHED
The field above runs fifty-six of these at once. Here are three, slowly.
REJECTED It never separated from noise. The rejection is kept.
MORE EVIDENCE REQUIRED Not refuted, but not established either. It waits rather than advancing.
FORWARD OBSERVATION Now it has to survive time it has never seen.
Evidence does not grant authority.
Candidate C has earned the right to be watched, not the right to act. Crossing that boundary is a separate, explicit authorization step, and it can be withdrawn without the evidence changing at all.
Possible models, signals, or decision rules for one clearly bounded decision problem.
Promising patterns are easy to generate. Kairos exists to determine which deserve belief and bounded use.
Researchers, risk owners, operators, and leaders responsible for forecasts, models, or automated decisions.
Run in customer-controlled or approved private environments; finance is the first proving domain.
From initial search through historical testing, forward observation, approval, and ongoing monitoring.
Generate many candidates, falsify aggressively, preserve failures, observe survivors, and authorize separately.
The 5W + 1H above defines the research problem. Below is the research record itself. It is process evidence, not live Kairos telemetry or a claim of model performance.
RECORDED TRAINING-GYM REJECTIONS
A training gym is one of the search programs XGCS leaves running: it proposes candidate predictions continuously, and a referee decides which are allowed to enrol. The field above is a schematic. This is the real register, every rejection those gyms recorded across the 4 of 8 programs that have logged any, sorted by the reason each was attributed to.
of 8,855 recorded rejection records stopped at one place: never separated from noise. Almost nothing survives far enough to fail for an interesting reason.
The other 669, shown at their own scale. Together they are 7.6% of the register.
What this shows. Where the search's own rejections were attributed, across every program that has recorded any. Why it matters. The rejections are kept rather than discarded, so the register of what did not work is itself part of the research record.
Perpetual candidate generation, breeding and culling, quality-diversity preservation, null-result memory, Referee evaluation, and forward-observation machinery.
Failed research becomes structured information that steers future search instead of disappearing into notebooks or chat history.
A model is registered as shadow by default and cannot promote itself. Promotion is fail-closed by design: the console refuses the flip unless every bar passes, and the bar that has not been built yet refuses rather than waves through. The evidence-span bar is enforced today, at least sixty trading days of shadow readings with no long gaps. Research sets the thresholds and a separate platform action performs the flip, so the party that sets the bar is never the party that clears it.
The reusable cross-domain layer is architected. Productization, hardening, private deployment, and commercial validation remain the work ahead.
BOUNDARY / Predecessor systems show that the advancement philosophy on this page already exists in working software rather than only in a plan. They do not prove external product-market fit, a finished cross-domain product, or verified live trading alpha.
Run Kairos where sensitive data already lives, under customer security and governance rules.
Integrated local workstation or server profiles without turning Kairos into a hardware manufacturer.
The same Core logic where confidentiality, latency, and policy requirements permit hosted operation.