All gamesDecision Lab / learning map
Why these games exist

Learn to make trustworthy decisions with uncertain predictions

Decision Lab is not about teaching people to obey an AI system. It is about learning how to judge the evidence a model provides, combine it with consequences and available actions, and decide what should happen next.

Across four playable worlds, predictions meet costs, uncertainty, information, interventions, limited human attention and changing operating conditions. The ideas matter because they change decisions and consequences — not because a tutorial asks you to memorise terminology.

The central idea

Prediction ≠ decision.

A prediction becomes decision support only when you understand what it means and combine it with uncertainty, consequences, possible actions, information cost, time and human capacity.

No prior knowledge required

The games teach the main ideas through choices and consequences. Technical terminology appears only where it helps explain what you have already experienced in play.

The learning journey

Understand → Decide → Govern

1 / Understand

Understand the evidence

Before acting, know what the model is actually telling you — and what it is not telling you.

2 / Decide

Turn evidence into a decision

Consequences, available actions and the value of more information determine what you should do.

3 / Govern

Run the decision system responsibly

Automation needs explicit policies, limited human attention and continued evidence that the system still behaves as expected.

Learning objectives

What you should be able to do

The goal is behavioural: better decisions with uncertain model output, not vocabulary recall.

Interpret a calibrated probability as a probability forecast and combine it with uncertainty and consequences.

Turn predictions into decisions using consequences rather than a default 50% threshold.

Choose when to act, buy information, intervene or use scarce human review.

Use explanations and alternatives critically without confusing model reasoning with causality.

Design automation rules with explicit thresholds, exceptions and review routing.

Judge the system from repeated relevant evidence and recognise when changing conditions may invalidate old assumptions.

Four complementary worlds

Each game has a different teaching role

Teaching campaign

Mission Control

The clearest step-by-step introduction to calibrated decision support.

Open game
Persistent world

Uncertain Acres

Sequential decisions where information and interventions have different consequences over time.

Open game
Operations at scale

Port Authority

Automation, scarce human review, monitoring and calibration lifecycle under traffic pressure.

Open game
Statistical depth

Precision Works

Regression, calibrated event probabilities, prediction intervals and long-run validity.

Open game
Pedagogical walkthrough

Seven lessons across three stages

“Primary lesson” means understanding the idea is central to doing well in that game. “Reinforces” means the mechanic matters, but it is not the game's main teaching focus.

1 / Understand

Understand the evidence

Before acting, know what the model is actually telling you — and what it is not telling you.

Lesson 1

A probability should mean what it says

A raw model score is not automatically a probability. Calibration gives a probability its operational meaning: among cases assigned probability p to an outcome, that outcome should occur with frequency p.

You have learned it when: You can interpret a calibrated probability as a probability forecast and use it directly when weighing the consequences of different decisions.
Mission Control
Primary lesson

Level 2 puts an overconfident raw score beside the calibrated probability so you can see why the distinction changes decisions.

Uncertain Acres
Reinforces

The farm advisor uses calibrated chances as decision inputs and later allows calibration to reflect soil context.

Port Authority
Reinforces

Clearance policies can use calibrated estimates and lower bounds, so probability quality becomes operational rather than merely descriptive.

Precision Works
Primary lesson

Threshold and within-specification probabilities are calibrated explicitly and kept separate from the customer's batch-acceptance rule.

Lesson 2

The point estimate is not enough

Two cases can have the same central estimate and still deserve different treatment when one is much more uncertain. A range can change whether you act, investigate or hand the case over.

You have learned it when: You use uncertainty as decision information rather than treating it as decoration around a prediction.
Mission Control
Primary lesson

Level 3 pairs similar probabilities with narrow and wide ranges. The lower end of the range can change the sensible launch decision.

Uncertain Acres
Primary lesson

Sensor quality, field conditions and purchased observations change how uncertain the advisor is before farm decisions.

Port Authority
Primary lesson

The tutorial teaches recommendation → probability → range, and automation can require sufficiently strong evidence before CLEAR.

Precision Works
Primary lesson

Prediction intervals show where hardness may land, while calibrated probability bounds accompany threshold and within-specification questions.

2 / Decide

Turn evidence into a decision

Consequences, available actions and the value of more information determine what you should do.

Lesson 3

A prediction is not a decision

There is no universal 50% decision threshold. The right action depends on what success is worth, what failure costs, what alternatives exist and how much uncertainty you are willing to carry.

You have learned it when: You can separate model output from the decision rule and explain why the same probability can justify different actions under different consequences.
Mission Control
Primary lesson

Level 1 makes this explicit: when a failed launch costs twice what a success earns, a simple GO/STOP recommendation is not enough.

Uncertain Acres
Primary lesson

Advice competes with crop value, cash, shared water, harvest capacity and delayed consequences.

Port Authority
Primary lesson

CLEAR, HOLD and INSPECT trade incident risk against delay, inspection cost and harbour congestion.

Precision Works
Primary lesson

Predicted quality is only one input; contract value, process cost and realised acceptance determine whether a batch was worth running.

Lesson 4

Information, intervention and review solve different problems

Information changes what you know. An intervention changes the situation. Human review brings another decision-maker into the loop. Each costs time or resources, so the useful question is whether it can change an important decision enough to be worth its cost.

You have learned it when: You choose deliberately between acting now, buying information, changing the situation and using scarce human review.
Mission Control
Primary lesson

Scans and diagnostics improve information; charging, waiting or rerouting change the mission; specialist review consumes a limited resource.

Uncertain Acres
Primary lesson

Measurements change knowledge, while irrigation, mulch, ventilation, tarps and fertiliser change persistent field state.

Port Authority
Primary lesson

Verifying a manifest improves an observation. HOLD and INSPECT change the operational path, while inspection consumes scarce team capacity.

Precision Works
Reinforces

Material information improves knowledge about the lot, while changing the furnace, soak, quench or tempering recipe changes the process you run.

Lesson 5

Explanations and alternatives must be interpreted critically

WHY? can show what pushed the model's assessment up or down, but model reasoning is not automatically a causal explanation of the world. A promising what-if is not automatically an action either: it must be feasible, sufficiently supported by data and worth its cost and uncertainty.

You have learned it when: You can inspect a model's reasoning without turning association into causation, and distinguish an attractive hypothetical from an option you can responsibly act on.
Mission Control
Primary lesson

Levels 4–7 move from WHY to actionable alternatives, then show why a higher estimate can still have greater uncertainty or weak data support.

Uncertain Acres
Primary lesson

WHY is paired with real field interventions, and the Reliability Lab checks whether proposed alternatives are well supported by comparable cases.

Port Authority
Reinforces

The tutorial explicitly says that WHY explains the model's assessment, not what causes problems at sea.

Precision Works
Primary lesson

WHY changes with the decision question, while recipe alternatives let you compare model-supported process changes before committing a batch.

3 / Govern

Run the decision system responsibly

Automation needs explicit policies, limited human attention and continued evidence that the system still behaves as expected.

Lesson 6

Trustworthy automation needs an explicit decision policy

A predictive model does not decide which cases may be automated. The operating policy defines what evidence is sufficient for action and what happens when it is not: gather information, hold, defer or ask for human review. Human attention is limited, so sending every uncertain case to a person does not scale.

You have learned it when: You can describe an automation rule in operational terms and explain why its thresholds, exceptions and review routing exist.
Mission Control
Primary lesson

Levels 9–10 let policies use probabilities, uncertainty, reject signals, stakes, remaining specialist capacity, information and interventions.

Uncertain Acres
Reinforces

The Automation Shed repeats rules every morning, while the agronomist remains scarce and the player can take over a policy decision.

Port Authority
Primary lesson

Rising traffic makes manual handling impossible. Policy rules automate routine cases while inspection queues expose the cost of indiscriminate review.

Precision Works
Not a focus

The factory keeps recipe choice with the player; automated routing policy is intentionally not its main lesson.

Lesson 7

Trust must be maintained over time

One good or bad outcome proves very little. Trust comes from repeated evidence under relevant conditions. When the operating population or process changes, yesterday's calibration may stop representing today's operation, so monitoring, restriction and controlled lifecycle actions become part of decision quality.

You have learned it when: You judge patterns rather than anecdotes, recognise when old evidence may no longer apply, and keep decision records that make important actions reconstructable.
Mission Control
Reinforces

Incident review reconnects costly or unusual outcomes to the evidence, explanations and routing signals available when the decision was made.

Uncertain Acres
Reinforces

Season records connect observations, information, interventions, review and delayed consequences across a persistent world.

Port Authority
Primary lesson

Decision Assurance monitors repeated outcomes, can pause automation, supports controlled recalibration with a fixed predictive model, and keeps checkpoints for restoration. It is an educational operational workflow, not a formal proof that a new calibration is valid.

Precision Works
Primary lesson

The Model Record separates locked reference QA from the production recipes you chose, so long-run model evidence is not confused with selectively observed operational data.

Advanced track

Precision Works goes deeper into uncertainty

Once the core decision ideas are clear, Precision Works makes several statistical distinctions explicit without changing the basic principle: the uncertainty quantity must match the decision question.

Different questions, different outputs

“Where will hardness land?”, “Will it clear a minimum?” and “Will it stay in specification?” need different forms of calibrated decision support.

Long-run coverage is not a case guarantee

A nominal prediction-interval level describes repeated coverage under the relevant assumptions. It is not the probability that this one unit or recipe will fall inside the interval.

Your decisions change the data you observe

The recipes you choose to run are not the same thing as an independent reference sample. The Model Record keeps those evidence streams separate.

Common traps

Intuitions the games are designed to challenge

“70%” is just an arbitrary confidence score.

Instead: For a well-calibrated forecaster, among cases assigned about 70% probability to an outcome, that outcome should occur about 70% of the time. That reliability is what makes 70% usable as a probability in a decision.

The natural decision threshold is 50%.

Instead: Decision thresholds come from consequences. If one type of error is much more costly than another, the sensible threshold can be far from 50%.

A wide uncertainty range means the prediction is wrong.

Instead: A wide range means the evidence supports a less precise assessment. That may justify more information or caution, but it does not by itself tell you the eventual outcome.

WHY tells me what caused the outcome.

Instead: WHY tells you what influenced the model's assessment. Causal claims require additional assumptions or evidence.

More information is always safer.

Instead: Information has value when it can change a consequential decision enough to justify its cost, delay and opportunity cost.

The safe option is to send every uncertain case to a human.

Instead: Human review is limited. A useful policy reserves it for cases where review is likely to change an important decision.

Once a model is calibrated, its probabilities stay trustworthy.

Instead: Calibration is tied to the conditions under which it was established. If the operating population or process changes, old calibration evidence may no longer describe current behaviour.

Suggested route

Start simple, then add operational complexity

Start with Mission Control for the core concepts. Move to Uncertain Acres for information versus intervention and sequential consequences. Use Port Authority for automation, scarce review and lifecycle assurance. Finish with Precision Works for the statistically most explicit treatment of calibrated regression, prediction intervals and long-run evidence.

Start with Mission Control