Education & learning systems

Test whether AI supports learning—not only answers.

A factually correct output can still be pedagogically wrong. Construct evaluates whether a tutor, feedback system or content workflow supports the intended learning objective across learner contexts and interaction histories.

We evaluate the AI system—not a learner's intelligence, personality, diagnosis or future.

Systems in scope

Evidence for controlled educational AI use.

The system must be runnable and one intended use must be isolated. High-stakes grading, admission, diagnosis or automated trajectory decisions need separate governance and may be declined.

01 / TUTOR

Conversational tutors

Test explanation, questioning, scaffolding, misconception handling and appropriate escalation.

  • Multi-turn learning paths
  • Help without answer leakage
  • Memory boundaries
Evaluate a tutor →
02 / FEEDBACK

Formative feedback

Check whether feedback is accurate, actionable, rubric-aligned and consistent for equivalent work.

  • Writing and problem solving
  • Rubric reliability
  • Educator review
Evaluate feedback →
03 / CONTENT

Activities and item generation

Stress curriculum alignment, answer integrity, difficulty assumptions and accessibility across variants.

  • Equivalent forms
  • Distractor and solution quality
  • Language/context checks
Evaluate content →
04 / COPILOT

Educator copilots

Test planning, summarisation and decision support while preserving human authority and uncertainty.

  • Source grounding
  • Approval boundaries
  • Workload and latency
Evaluate a copilot →

Evaluation dimensions

Seven questions beyond “is the answer right?”

Learning alignment

Does the interaction serve the objective?

Responses and actions are mapped to the intended skill, level and pedagogical strategy.

Explanation quality

Is the reasoning useful and accurate?

We inspect correctness, clarity, examples, uncertainty and sources where required.

Scaffolding

Does help preserve learner agency?

The system should support the next step without bypassing the learning task or over-helping.

Consistency

Are equivalent cases treated comparably?

Repeated runs and parallel tasks expose stochastic drift and rubric instability.

Context comparison

Do language and profile alter quality?

Matched cases can expose unjustified gaps while keeping interpretations proportional to available data.

Memory & adaptation

Does the learning path remain controllable?

We test continuity, outdated assumptions, contamination and appropriate response to progress or feedback.

Human boundary

Does the system know when to escalate?

Educator-only decisions, uncertainty and review points must remain visible and effective.

About formal psychometrics

Without an appropriate human response sample and study design, Construct does not estimate human item difficulty, discrimination, DIF or learning gain. We evaluate the system and the agreed rubric within the tested environment.

Education protocol

From learning objective to deployment boundary.

01 · INTENT

Define the use

Learning objective, audience, educator role and the decision the evidence must support.

02 · RUBRIC

Agree the criteria

Educators or subject experts define correctness, useful support, limits and critical failures.

03 · PROFILES

Model contexts safely

Specify levels, languages, accessibility needs and common errors without personal records in the public request.

04 · CASES

Build equivalent tasks

Representative, parallel and difficult cases include multi-turn trajectories with and without memory.

05 · RUN

Repeat and compare

Measure correctness, feedback quality, consistency, context gaps, cost, latency and escalations.

06 · REVIEW

Ground with experts

Subject and education specialists review critical or disputed cases against the agreed rubric.

07 · RETEST

Measure bounded fixes

When included, agreed changes are applied and the same evidence protocol is rerun.

08 · DECIDE

Set the boundary

The evidence brief states suitable use, open risks, human oversight and continue/conditional/hold.

Deliverables

Evidence an education team can review.

  • Learning-objective and criteria map
  • Versioned profile, context and interaction suite
  • Reusable expert rubric and runner
  • Correctness, consistency and context analysis
  • Ranked pedagogical and technical failure map
  • Expert review record when included
  • Before/after comparison and deployment boundary

Required inputs

Bring educational intent, not learner records.

  • One learning objective and intended audience
  • A runnable controlled environment
  • Sample activities or representative tasks
  • Existing rubrics or curriculum references
  • An available educator or subject expert
  • Known governance or accessibility constraints
Minors and learner data

Do not submit student names, records, responses, health information or data about minors through the public form. Any necessary processing must be agreed separately.

Appropriate evidence decisions

  • Pilot with a small supervised group
  • Choose between two configurations
  • Hold a feature until critical failures are fixed
  • Define educator escalation and usage limits
  • Prepare a later efficacy study

Not established by this Sprint

  • Accreditation or test certification
  • Learning efficacy or causal impact
  • Automated grading or admission approval
  • Diagnosis of a learner
  • Replacement of educator judgement

Preparing a supervised education pilot?

Bring one objective and one learning workflow.

We will identify the smallest defensible protocol, the educator input required and the decision boundary the evidence can support.