Conversational tutors
Test explanation, questioning, scaffolding, misconception handling and appropriate escalation.
- Multi-turn learning paths
- Help without answer leakage
- Memory boundaries
Education & learning systems
A factually correct output can still be pedagogically wrong. Construct evaluates whether a tutor, feedback system or content workflow supports the intended learning objective across learner contexts and interaction histories.
We evaluate the AI system—not a learner's intelligence, personality, diagnosis or future.
Systems in scope
The system must be runnable and one intended use must be isolated. High-stakes grading, admission, diagnosis or automated trajectory decisions need separate governance and may be declined.
Test explanation, questioning, scaffolding, misconception handling and appropriate escalation.
Check whether feedback is accurate, actionable, rubric-aligned and consistent for equivalent work.
Stress curriculum alignment, answer integrity, difficulty assumptions and accessibility across variants.
Test planning, summarisation and decision support while preserving human authority and uncertainty.
Evaluation dimensions
Responses and actions are mapped to the intended skill, level and pedagogical strategy.
We inspect correctness, clarity, examples, uncertainty and sources where required.
The system should support the next step without bypassing the learning task or over-helping.
Repeated runs and parallel tasks expose stochastic drift and rubric instability.
Matched cases can expose unjustified gaps while keeping interpretations proportional to available data.
We test continuity, outdated assumptions, contamination and appropriate response to progress or feedback.
Educator-only decisions, uncertainty and review points must remain visible and effective.
Without an appropriate human response sample and study design, Construct does not estimate human item difficulty, discrimination, DIF or learning gain. We evaluate the system and the agreed rubric within the tested environment.
Education protocol
Learning objective, audience, educator role and the decision the evidence must support.
Educators or subject experts define correctness, useful support, limits and critical failures.
Specify levels, languages, accessibility needs and common errors without personal records in the public request.
Representative, parallel and difficult cases include multi-turn trajectories with and without memory.
Measure correctness, feedback quality, consistency, context gaps, cost, latency and escalations.
Subject and education specialists review critical or disputed cases against the agreed rubric.
When included, agreed changes are applied and the same evidence protocol is rerun.
The evidence brief states suitable use, open risks, human oversight and continue/conditional/hold.
Deliverables
Required inputs
Do not submit student names, records, responses, health information or data about minors through the public form. Any necessary processing must be agreed separately.
Preparing a supervised education pilot?
We will identify the smallest defensible protocol, the educator input required and the decision boundary the evidence can support.