Non-confidential review
We identify one working system, one decision and the smallest viable evidence boundary.
Construct methodology
We evaluate AI as a situated system: a model acting inside a task, environment, population, interaction history and defined decision boundary.
The complete process
The scope stays bounded because responsibilities, evidence and stop conditions are explicit before testing begins.
We identify one working system, one decision and the smallest viable evidence boundary.
We confirm intended use, deadline, owners, available cases and whether the track is general or behavioural.
When viable, the written B2B order fixes scope, fee, VAT, access, processing roles and schedule.
We record version, environment, construct or capability, populations, criteria and non-claims.
Representative, boundary and adversarial cases are linked to rubrics or observable outcomes.
The runner, configuration and evidence capture are versioned; agreed conditions are run repeatedly.
We vary tools, ambiguity, language, group, resource, sequence, memory or feedback when relevant.
Each finding records severity, condition, reproduction, likely cause, uncertainty and owner/action.
When included, up to three agreed fixes are applied and the full protocol is rerun.
You receive the harness, evidence brief, limitations, recommendation and a live handoff.
The delivery clock starts only after the runnable environment, agreed cases and access are complete. New data collection, participant recruitment, expert panels or ethics review require a separate schedule.
Six evidence questions
These concepts are applied to the observed system within the agreed protocol. They do not turn an AI output into a psychological test or a synthetic persona into a real population.
Tasks, rubrics and indicators must represent the intended behaviour—not an easy proxy.
Repeated runs expose dispersion, stochastic failures and sensitivity to configuration.
Group, language and context comparisons require equivalent tasks and explicit interpretation limits.
A valid system should neither ignore relevant context nor become arbitrary when context shifts.
We test continuity, contamination, learning and path dependence across interactions.
Human data, theory, documentation, expert rubrics or a structured Delphi cycle can provide contrast.
A differential signal is a reason to investigate, not automatic proof of bias, discrimination or invalidity. Classical psychometric claims require an appropriate instrument, sample and qualified analysis; Construct reports only what the agreed system evidence supports.
From isolated output to observable system
When behaviour emerges across turns or actors, we inspect the relationship—not just the final sentence.
Psychological, educational or domain theory defines what a defensible response pattern would mean.
The system is tested before and after feedback, memory or environmental change.
Experts or external observations discipline the analysis without eliminating uncertainty.
Situated example
In a water-risk simulation, an “average agent” can hide resource constraints, unequal exposure and different capacities to adapt. A defensible protocol separates population conditions and trajectories.
Pre-event: risk perception and preparation · During: evacuation and coping · Post-event: recovery and adaptation.
More room for anticipation, planning and delayed trade-offs.
Choices become conditional on competing costs and information.
Decisions occur under urgency, restriction and recovery burden.
Who does what
The protocol makes ownership visible so evidence does not become an unreviewed vendor verdict.
| Stage | Your team | Construct | Joint decision |
|---|---|---|---|
| Boundary | Own intended use, release and business risk. | Challenge scope and translate the claim into observable evidence. | Version, conditions, success criteria and non-claims. |
| Inputs | Provide authorised cases, environment and domain context. | Minimise access, version the set and identify evidence gaps. | Accept the sample and references as fit for this decision. |
| Evaluation | Remain available for domain questions and access issues. | Execute, reproduce, compare and document the protocol. | Approve any change that would alter scope or interpretation. |
| Release | Retain product, research, educational and legal decisions. | Deliver observations, limitations and a bounded recommendation. | Review open findings and decide continue, conditional continue or hold. |
Ready to turn a claim into a protocol?
Describe the working system, the behaviour you need to trust and the reference evidence you already have.