Methodology

Make the assumptions visible. Then test them.

The Scenario Lab turns an uncertain, consequential decision into a governed computational experiment. Its purpose is not to produce a theatrical forecast. It is to compare interventions consistently, reveal failure paths and identify choices that remain acceptable when the world changes.

Agent-based modellingBounded LLM agentsSystems scienceParticipatory review

Lab charter

Fitness before sophistication.

A simulation is useful only when interaction, adaptation, feedback or path dependence materially affect a real decision. If a simpler analysis, forecast or facilitated workshop can answer the question, we use that instead.

Mechanism before spectacleA plausible animation is not a causal explanation.
Evidence before synthetic plausibilityGenerated behaviour is never promoted to observed fact.
Ranges before point predictionsUncertainty is part of the result.
Distribution before averagesEffects on different groups and ecosystems stay visible.
Participation is not simulationModelled stakeholders do not replace affected stakeholders.
Reproducibility by defaultInputs, versions, prompts, parameters and runs are traceable.

The suitability gate

First decide whether the decision should be modelled.

Before development, we screen for decision value, evidence fitness and possible harm. The result may be proceed, proceed as exploratory only, use a simpler method or do not model.

  1. Is there a defined decision and accountable owner?The options, deadline, consequences and authority to act must be explicit.
  2. Do interaction and feedback materially change the answer?Simulation must add something a spreadsheet or static forecast cannot.
  3. Is the evidence adequate for the intended claim?Exploratory questions can tolerate more uncertainty than high-impact quantitative claims.
  4. Could modelling create false legitimacy or material harm?We reject coercive, discriminatory or manipulative applications and restrict sensitive targeting.
  5. Can domain experts and affected stakeholders challenge the design?Real participation is required where legitimacy, lived experience or contested values matter.

Hybrid architecture

Rules where rules are reliable. AI where context matters.

Each layer has a distinct job. No component gets authority simply because it is sophisticated.

CAUSAL SPINE

Agent-based model

Controls actors, resources, networks, constraints, timing, state transitions and aggregate outcomes. The mechanisms are explicit, testable and reproducible.

BOUNDED CONTEXT

LLM layer

May interpret language, choose from an approved action set, negotiate inside limits or generate hypotheses. Outputs are structured, logged, versioned and benchmarked against simpler rules.

REALITY CHECK

Evidence and people

Observed data, research, experts and stakeholder participation anchor the design and challenge the interpretation. Simulated people never replace real people.

Evidence architecture

Every assumption carries a label.

A versioned evidence ledger records the source, date, owner, confidence, sensitivity and update requirement for important rules and parameters. LLM-generated material remains hypothetical unless independently supported.

Model claim ruleA model is decision-grade only for the question, context and permitted use for which it has been validated.
O

Observed

Directly measured, documented or reliably recorded for the system in scope.

E

Estimated

Statistically inferred or calibrated from declared data and methods.

L

Elicited

Provided by named experts or stakeholders with context and confidence captured.

H

Hypothesized

Plausible but not yet substantiated; tested transparently rather than presented as fact.

Experiment design

Keep worlds, choices and assumptions separate.

Scenarios are external worlds the decision-maker does not control. Strategies are available actions. Assumptions are uncertain mechanisms or parameters. Separating them makes comparisons interpretable.

SCENARIOS

What could surround us?

Current trajectory, credible shocks, compound events, tail risks, institutional change and alternative external conditions.

ASSUMPTIONS

What might work differently?

Behavioural rules, adoption rates, network effects, thresholds, elasticities and alternative model structures.

Repeated stochastic runs reveal outcome distributions, pathways and failure conditions—not one polished number pretending to be the future.

Verification and validation

A ladder of checks before a claim.

The validation standard rises with the impact of the decision. A failed gate lowers the claim, changes the route or stops the study.

1. Conceptual validityDoes the structure reflect the relevant theory and domain understanding?
2. Code verificationDoes the implementation behave as the documented design requires?
3. Data validationAre sources, transformations, coverage and limitations correct?
4. Micro-validationAre agent-level behaviours plausible and evidence-supported?
5. Macro-validationCan the model reproduce relevant aggregate patterns?
6. CalibrationDo selected parameters fit agreed empirical targets?
7. BacktestingCan the model reproduce withheld or historical periods where appropriate?
8. Sensitivity and uncertaintyWhich inputs, mechanisms or structures drive the conclusion?
9. Stress and adversarial testingWhere does the model break, mislead or become unsafe?
10. Independent challengeCan domain experts and stakeholders identify important omissions?

Delivery gates

Human accountability at every consequential handoff.

  1. G0

    Suitability and harm screen

    Set the permitted claim level and governance needs—or decide not to model.

  2. G1

    Decision contract

    The sponsor approves the decision, system boundary, options, consequences, permitted use and stop conditions.

  3. G2

    Evidence and model specification

    Domain and data owners approve assumptions, access, unresolved gaps and the minimum credible structure.

  4. G3

    Validation and claim review

    The evidence supports the proposed claim—or the claim is reduced before scenario experiments proceed.

  5. G4

    Decision rehearsal

    The accountable owner reviews counter-evidence, trade-offs, fragile findings and monitoring triggers before acting.

  6. G5

    Operation, change and retirement

    Production models require monitoring, version control, review dates, rollback and an explicit end of life.

Implementation routes

From a portable study to a governed living lab.

The delivery surface follows the decision, security boundary and need for reuse.

Comparison of Scenario Lab implementation routes
RouteBest suited toImplementationPrimary outcome
Scenario SprintEarly or urgent decisions with limited dataPortable prototype, bounded scenarios and facilitated reviewDecision map and initial robust or fragile options
Decision ModelMaterial one-off strategy, policy or transitionCalibration and validation proportionate to available evidence, with a defined experiment programmeAuditable decision evidence and trigger plan
Living LabRecurring decisions in a changing environmentApproved data refresh, monitoring, model registry and operating cadenceA governed, repeatable capability refreshed on an agreed cadence
Participatory LabPublic, contested or distributionally sensitive decisionsCo-design workshops, explicit contested assumptions and accessible explorerLegitimate trade-off analysis and shared understanding
Independent ReviewAn existing internal or third-party modelVerification, sensitivity, bias and governance assessmentAssurance evidence and remediation priorities

Responsible-use boundary

The model advises. People remain accountable.

Scientific and ethical caveatResults are conditional on the data, mechanisms and uncertainties represented. They must not be the sole basis for decisions affecting rights, safety, access to essential services or protected groups. Models of stakeholder behaviour do not replace consent, participation or lived experience.

We do not offer guaranteed forecasts, legal or regulatory conclusions, accredited assurance, automated high-impact individual decisions or evidence that does not exist. Sensitive projects require proportionate privacy, security, ethics and domain review.

A decision worth modelling?

Start with the question, not the technology.

Describe the decision, deadline, consequences and affected system. We will identify the smallest credible route—or tell you when another method is more appropriate.