Mission Control
The clearest step-by-step introduction to calibrated decision support.
Open gameDecision Lab is not about teaching people to obey an AI system. It is about learning how to judge the evidence a model provides, combine it with consequences and available actions, and decide what should happen next.
Across four playable worlds, predictions meet costs, uncertainty, information, interventions, limited human attention and changing operating conditions. The ideas matter because they change decisions and consequences — not because a tutorial asks you to memorise terminology.
Prediction ≠ decision.
A prediction becomes decision support only when you understand what it means and combine it with uncertainty, consequences, possible actions, information cost, time and human capacity.
The games teach the main ideas through choices and consequences. Technical terminology appears only where it helps explain what you have already experienced in play.
Before acting, know what the model is actually telling you — and what it is not telling you.
Consequences, available actions and the value of more information determine what you should do.
Automation needs explicit policies, limited human attention and continued evidence that the system still behaves as expected.
The goal is behavioural: better decisions with uncertain model output, not vocabulary recall.
Interpret a calibrated probability as a probability forecast and combine it with uncertainty and consequences.
Turn predictions into decisions using consequences rather than a default 50% threshold.
Choose when to act, buy information, intervene or use scarce human review.
Use explanations and alternatives critically without confusing model reasoning with causality.
Design automation rules with explicit thresholds, exceptions and review routing.
Judge the system from repeated relevant evidence and recognise when changing conditions may invalidate old assumptions.
The clearest step-by-step introduction to calibrated decision support.
Open gameSequential decisions where information and interventions have different consequences over time.
Open gameAutomation, scarce human review, monitoring and calibration lifecycle under traffic pressure.
Open gameRegression, calibrated event probabilities, prediction intervals and long-run validity.
Open game“Primary lesson” means understanding the idea is central to doing well in that game. “Reinforces” means the mechanic matters, but it is not the game's main teaching focus.
Before acting, know what the model is actually telling you — and what it is not telling you.
A raw model score is not automatically a probability. Calibration gives a probability its operational meaning: among cases assigned probability p to an outcome, that outcome should occur with frequency p.
Level 2 puts an overconfident raw score beside the calibrated probability so you can see why the distinction changes decisions.
The farm advisor uses calibrated chances as decision inputs and later allows calibration to reflect soil context.
Clearance policies can use calibrated estimates and lower bounds, so probability quality becomes operational rather than merely descriptive.
Threshold and within-specification probabilities are calibrated explicitly and kept separate from the customer's batch-acceptance rule.
Two cases can have the same central estimate and still deserve different treatment when one is much more uncertain. A range can change whether you act, investigate or hand the case over.
Level 3 pairs similar probabilities with narrow and wide ranges. The lower end of the range can change the sensible launch decision.
Sensor quality, field conditions and purchased observations change how uncertain the advisor is before farm decisions.
The tutorial teaches recommendation → probability → range, and automation can require sufficiently strong evidence before CLEAR.
Prediction intervals show where hardness may land, while calibrated probability bounds accompany threshold and within-specification questions.
Consequences, available actions and the value of more information determine what you should do.
There is no universal 50% decision threshold. The right action depends on what success is worth, what failure costs, what alternatives exist and how much uncertainty you are willing to carry.
Level 1 makes this explicit: when a failed launch costs twice what a success earns, a simple GO/STOP recommendation is not enough.
Advice competes with crop value, cash, shared water, harvest capacity and delayed consequences.
CLEAR, HOLD and INSPECT trade incident risk against delay, inspection cost and harbour congestion.
Predicted quality is only one input; contract value, process cost and realised acceptance determine whether a batch was worth running.
Information changes what you know. An intervention changes the situation. Human review brings another decision-maker into the loop. Each costs time or resources, so the useful question is whether it can change an important decision enough to be worth its cost.
Scans and diagnostics improve information; charging, waiting or rerouting change the mission; specialist review consumes a limited resource.
Measurements change knowledge, while irrigation, mulch, ventilation, tarps and fertiliser change persistent field state.
Verifying a manifest improves an observation. HOLD and INSPECT change the operational path, while inspection consumes scarce team capacity.
Material information improves knowledge about the lot, while changing the furnace, soak, quench or tempering recipe changes the process you run.
WHY? can show what pushed the model's assessment up or down, but model reasoning is not automatically a causal explanation of the world. A promising what-if is not automatically an action either: it must be feasible, sufficiently supported by data and worth its cost and uncertainty.
Levels 4–7 move from WHY to actionable alternatives, then show why a higher estimate can still have greater uncertainty or weak data support.
WHY is paired with real field interventions, and the Reliability Lab checks whether proposed alternatives are well supported by comparable cases.
The tutorial explicitly says that WHY explains the model's assessment, not what causes problems at sea.
WHY changes with the decision question, while recipe alternatives let you compare model-supported process changes before committing a batch.
Automation needs explicit policies, limited human attention and continued evidence that the system still behaves as expected.
A predictive model does not decide which cases may be automated. The operating policy defines what evidence is sufficient for action and what happens when it is not: gather information, hold, defer or ask for human review. Human attention is limited, so sending every uncertain case to a person does not scale.
Levels 9–10 let policies use probabilities, uncertainty, reject signals, stakes, remaining specialist capacity, information and interventions.
The Automation Shed repeats rules every morning, while the agronomist remains scarce and the player can take over a policy decision.
Rising traffic makes manual handling impossible. Policy rules automate routine cases while inspection queues expose the cost of indiscriminate review.
The factory keeps recipe choice with the player; automated routing policy is intentionally not its main lesson.
One good or bad outcome proves very little. Trust comes from repeated evidence under relevant conditions. When the operating population or process changes, yesterday's calibration may stop representing today's operation, so monitoring, restriction and controlled lifecycle actions become part of decision quality.
Incident review reconnects costly or unusual outcomes to the evidence, explanations and routing signals available when the decision was made.
Season records connect observations, information, interventions, review and delayed consequences across a persistent world.
Decision Assurance monitors repeated outcomes, can pause automation, supports controlled recalibration with a fixed predictive model, and keeps checkpoints for restoration. It is an educational operational workflow, not a formal proof that a new calibration is valid.
The Model Record separates locked reference QA from the production recipes you chose, so long-run model evidence is not confused with selectively observed operational data.
Once the core decision ideas are clear, Precision Works makes several statistical distinctions explicit without changing the basic principle: the uncertainty quantity must match the decision question.
“Where will hardness land?”, “Will it clear a minimum?” and “Will it stay in specification?” need different forms of calibrated decision support.
A nominal prediction-interval level describes repeated coverage under the relevant assumptions. It is not the probability that this one unit or recipe will fall inside the interval.
The recipes you choose to run are not the same thing as an independent reference sample. The Model Record keeps those evidence streams separate.
Instead: For a well-calibrated forecaster, among cases assigned about 70% probability to an outcome, that outcome should occur about 70% of the time. That reliability is what makes 70% usable as a probability in a decision.
Instead: Decision thresholds come from consequences. If one type of error is much more costly than another, the sensible threshold can be far from 50%.
Instead: A wide range means the evidence supports a less precise assessment. That may justify more information or caution, but it does not by itself tell you the eventual outcome.
Instead: WHY tells you what influenced the model's assessment. Causal claims require additional assumptions or evidence.
Instead: Information has value when it can change a consequential decision enough to justify its cost, delay and opportunity cost.
Instead: Human review is limited. A useful policy reserves it for cases where review is likely to change an important decision.
Instead: Calibration is tied to the conditions under which it was established. If the operating population or process changes, old calibration evidence may no longer describe current behaviour.
Start with Mission Control for the core concepts. Move to Uncertain Acres for information versus intervention and sequential consequences. Use Port Authority for automation, scarce review and lifecycle assurance. Finish with Precision Works for the statistically most explicit treatment of calibrated regression, prediction intervals and long-run evidence.
Start with Mission Control