Does it repeat?
The same task should not produce a different decision for no reason.
Interactive demonstration · No account
Choose a use case, replay six synthetic cases and see what passed, what failed and what still needs evidence.
Local interaction · Synthetic cases · Your observations are not transmitted. Normal hosting logs still apply to the page load.
How to use the test
The guide follows the same controls you will use below. There is nothing to install and no technical setup.
Pick AI agent, Research & personas or Specialist workflow.
Select “Replay example” to load six synthetic observations.
Start with the plain-language finding, then check the supporting signals.
Open the failed cases and compare expected with observed behaviour.
01 · Choose a use case
How to read the result
Construct applies measurement science—including psychometric principles—to observable agent behaviour.
The same task should not produce a different decision for no reason.
Relevant information may change the answer; irrelevant labels should not.
The result shows which intended behaviours the cases actually covered.
A small sample has limits. We show them instead of hiding them in one score.
Synthetic profiles are controlled test conditions. They expose possible failure modes; they do not represent or predict real people, and they do not replace user research.
We first agree the claim, cases, allowed data and one controlled connection route. API access, expert analysis, timing and commercial terms are then defined privately in writing.