Validation for AI agents and LLM workflows

Validate your AI agent before users find the failures.

Construct tests tool use, RAG, memory, guardrails and end-to-end LLM workflows from the first runnable version to release. Get metrics through our API or a complete expert-reviewed evidence report.

Versioned cases · Repeat runs · Pass, stability, latency, token and cost metrics

See how Construct works — 37 secRun in your environment → send observations → receive metrics. Managed Validation adds the expert report.
Read the video transcript
  1. Run the AI agent or LLM workflow in your own environment; your model, tools, RAG and data stay with you.
  2. Open a prepaid Metrics API session for €50; the consecutive window lasts 60 minutes.
  3. Send status, latency, tokens and structured checks—not prompts, responses, commands, URLs or credentials.
  4. Receive pass, error, timeout, p50/p95 latency, stability and tool-check metrics in JSON.
  5. Choose Metrics API for metrics only or Managed Validation for test design, failure analysis and an expert-written report.

Two product modes

Choose metrics or interpretation.

Use the API when your team already owns the test protocol. Choose managed validation when you need us to design, challenge and interpret it.

Private betaYou call Construct

Metrics API

Run your own AI agent or LLM tests, send only structured observations and receive machine-readable metrics.

  • Pass, error and timeout rates
  • Latency p50/p95, reported tokens and cost
  • Repeat stability, tool checks and payload hash
€50per 60-minute access window

Prepaid. The consecutive window starts with POST /sessions. Your model-provider charges remain yours. VAT may apply.

Expert serviceConstruct tests with you

Managed Validation

We design the protocol, connect through one controlled route, investigate failures and deliver the complete evidence report.

  • Normal, boundary and adversarial cases
  • Failure analysis and bounded fixes where included
  • Human-reviewed release, fix or hold recommendation

Evidence Baseline

5 business days · one workflow · up to 30 agreed cases · reusable tests, failure map and decision note.

€2,500

Release Sprint

7–10 business days · up to 50 cases · bounded fixes, full rerun and before/after evidence.

€4,900

Metrics API returns metrics only from customer-reported observations. Its payload hash confirms identity, not factual correctness. It includes no written analysis, root-cause diagnosis, release recommendation, legal opinion or certification.

Managed validation · connection routes

Three ways we can test your system.

These routes describe how Construct reaches your test system during managed validation. They are different from the Metrics API, which your team calls.

A

API or test environment

You provide a temporary, limited-access token for an API or staging endpoint. We run the agreed cases, record responses, tool calls, errors, cost and latency, then you revoke the token.

Best when an endpoint already exists.
B

Run it in your environment

We prepare a small runner or container. Your team runs it locally or in your cloud and shares only the agreed result bundle. Code and data stay with you.

Best when data or code cannot leave.
C

Guided test session

No API? We use a temporary test workspace, supervised remote session, batch export or anonymised traces. The lower automation is stated in the report.

Best when integration is not available.
Access stays limitedNo production passwords or raw personal data. Credentials are exchanged only after written scope, by a method agreed in the SOW, and removed after testing.

What happens next

Same cases. More than once.

The test is small enough to understand and strict enough to rerun.

01

Define

Choose one workflow, version, intended use and success criteria.

02

Prepare

Create normal, boundary and repeat cases with expected outcomes.

03

Run

Execute the cases, vary relevant context and capture failures, latency and cost.

04

Decide

Deliver the runner, evidence and a release, fix or hold recommendation.

Open the complete technical method →

Applications

One method, four kinds of questions.

The core test stays the same. The cases and reference criteria change with the domain.

For behavioural work, we add construct, reliability, context, memory and independent-reference checks. Synthetic personas support hypothesis testing; they do not replace participants, representative samples or ethics review.

Test one agent. Learn exactly where it breaks.

Start with raw metrics through the API or ask us to design and interpret the complete validation.