Construct methodology

A plausible answer is not yet valid behaviour.

We evaluate AI as a situated system: a model acting inside a task, environment, population, interaction history and defined decision boundary.

The complete process

Ten controlled steps from request to decision.

The scope stays bounded because responsibilities, evidence and stop conditions are explicit before testing begins.

01 · FIT

Non-confidential review

We identify one working system, one decision and the smallest viable evidence boundary.

02 · CALL

20-minute scope call

We confirm intended use, deadline, owners, available cases and whether the track is general or behavioural.

03 · TERMS

NDA, DPA and order

When viable, the written B2B order fixes scope, fee, VAT, access, processing roles and schedule.

04 · CLAIM

Evidence protocol

We record version, environment, construct or capability, populations, criteria and non-claims.

05 · INPUTS

Cases and references

Representative, boundary and adversarial cases are linked to rubrics or observable outcomes.

06 · BASELINE

Repeatable execution

The runner, configuration and evidence capture are versioned; agreed conditions are run repeatedly.

07 · STRESS

Controlled change

We vary tools, ambiguity, language, group, resource, sequence, memory or feedback when relevant.

08 · FINDINGS

Failure explanation

Each finding records severity, condition, reproduction, likely cause, uncertainty and owner/action.

09 · RETEST

Bounded correction

When included, up to three agreed fixes are applied and the full protocol is rerun.

10 · HANDOFF

Decision and reuse

You receive the harness, evidence brief, limitations, recommendation and a live handoff.

The delivery clock starts only after the runnable environment, agreed cases and access are complete. New data collection, participant recruitment, expert panels or ethics review require a separate schedule.

Six evidence questions

Psychometrics becomes an engineering discipline.

These concepts are applied to the observed system within the agreed protocol. They do not turn an AI output into a psychological test or a synthetic persona into a real population.

Construct

Are we testing the claimed capability?

Tasks, rubrics and indicators must represent the intended behaviour—not an easy proxy.

Reliability

How stable is the pattern?

Repeated runs expose dispersion, stochastic failures and sensitivity to configuration.

Comparison

Can conditions be compared?

Group, language and context comparisons require equivalent tasks and explicit interpretation limits.

Sensitivity

Does context change behaviour appropriately?

A valid system should neither ignore relevant context nor become arbitrary when context shifts.

Trajectory

Do memory and adaptation remain coherent?

We test continuity, contamination, learning and path dependence across interactions.

External reference

What disciplines our interpretation?

Human data, theory, documentation, expert rubrics or a structured Delphi cycle can provide contrast.

About DIF and invariance

A differential signal is a reason to investigate, not automatic proof of bias, discrimination or invalidity. Classical psychometric claims require an appropriate instrument, sample and qualified analysis; Construct reports only what the agreed system evidence supports.

From isolated output to observable system

Context → agent → interaction → adaptation → evaluation.

When behaviour emerges across turns or actors, we inspect the relationship—not just the final sentence.

ContextPopulation · environment · taskAgentModel · rules · memoryInteractionRounds · feedback · conflictAdaptationUpdate · learning · trajectoryEvaluationTheory · experts · metrics
Theory

Interpretation criteria

Psychological, educational or domain theory defines what a defensible response pattern would mean.

Adaptation

Change across time

The system is tested before and after feedback, memory or environmental change.

Ground truthing

Independent confrontation

Experts or external observations discipline the analysis without eliminating uncertainty.

Situated example

Human–environment behaviour changes by group and event phase.

In a water-risk simulation, an “average agent” can hide resource constraints, unequal exposure and different capacities to adapt. A defensible protocol separates population conditions and trajectories.

Example conditions

Pre-event: risk perception and preparation · During: evacuation and coping · Post-event: recovery and adaptation.

A

Material protection

More room for anticipation, planning and delayed trade-offs.

B

Intermediate resources

Choices become conditional on competing costs and information.

C

Greater vulnerability

Decisions occur under urgency, restriction and recovery burden.

Who does what

No hidden handoffs.

The protocol makes ownership visible so evidence does not become an unreviewed vendor verdict.

StageYour teamConstructJoint decision
BoundaryOwn intended use, release and business risk.Challenge scope and translate the claim into observable evidence.Version, conditions, success criteria and non-claims.
InputsProvide authorised cases, environment and domain context.Minimise access, version the set and identify evidence gaps.Accept the sample and references as fit for this decision.
EvaluationRemain available for domain questions and access issues.Execute, reproduce, compare and document the protocol.Approve any change that would alter scope or interpretation.
ReleaseRetain product, research, educational and legal decisions.Deliver observations, limitations and a bounded recommendation.Review open findings and decide continue, conditional continue or hold.

The evidence can support

  • What was observed in a defined version and environment
  • Which failures were reproducible
  • Which differences appeared between agreed conditions
  • Whether the stated criteria were met
  • Which limitations affect the decision

It cannot establish by itself

  • Equivalence between synthetic people and real populations
  • Psychological or educational diagnosis
  • Exhaustive prediction of human behaviour
  • Certification, legal compliance or safety approval
  • Absence of all future bias, harm or failure

Ready to turn a claim into a protocol?

Start with one decision.

Describe the working system, the behaviour you need to trust and the reference evidence you already have.