Sanitised sample · Internal evaluation

AI Release Evidence Brief

A compact example of how scope, cases, observations, limitations and a release decision connect. It is based on a real internal evaluation with synthetic, non-personal data.

Method demonstration only. It is not a customer testimonial.

Evidence ID: PL-SAMPLE-2026-07Release: Internal prototype RC-01Status: Conditional proceed
Representative output · Not a certification
Release decision support

Proceed with the selected route for controlled internal use.

Hold the alternative route until workload capacity, quota and long-tail latency are resolved and the full eight-case suite passes on the intended project.

01 · Evaluation boundary

What this evidence applies to.

System

Internal AI workflow prototype using one allowlisted language model through two candidate cloud routes.

Intended use

Generate review-required business drafts and structured guidance. No external action or high-impact decision.

Release candidate

Vertex-only internal build tested on 22 July 2026. Server-controlled model and location configuration.

Test environment

Protected loopback prototype. Eight stored workflows with synthetic, non-personal inputs.

Boundary rule: these results do not apply to another model, route, prompt set, dataset, production project or external-action capability without retesting.

02 · Claims and criteria

What had to be true.

ClaimMethodThresholdOwner
Stored workflows remain usefulEight task-level rubrics with expected properties8/8 usefulProduct
Human-review boundary remains intactSafety assertions on every returned work item8/8 review-required; no external actionEngineering
Candidate route completes the suiteEnd-to-end workload run with bounded timeout0 model timeoutsPlatform
Latency supports controlled usePer-case elapsed time and median comparisonNo material regression against baselineProduct + Platform

03 · Measured results

The smoke test did not predict the workload.

A direct smoke request succeeded on the alternative route. The complete workflow suite then produced three model timeouts. The selected route completed every stored case safely and usefully.

MetricSelected routeAlternative routeObservation
Stored workflows88Same workload boundary
Useful + safe results8 / 8Not fully completedSelected route met the task boundary
Model timeouts03Release-blocking on the alternative route
Median response time3.30 seconds55.81 secondsMaterial difference under the full suite
Longest selected-route response48.03 secondsLong-tail handling still requires production work

04 · Ranked findings

Failures connected to an owner and next action.

  1. 1
    Three full-workflow timeouts on the alternative route

    Reproduced only when the complete suite ran, despite a successful smoke request. Review project capacity, quota and long-tail workload behaviour before reuse.

    High
  2. 2
    One rubric was too literal

    The response expressed the required access boundary with equivalent wording that the first phrase list missed. The stored output stayed unchanged; the rubric was broadened and the correction recorded.

    Medium
  3. 3
    Selected route still has a long tail

    The 48.03-second maximum does not block controlled internal use but requires queues, progress handling, monitoring and incident response before public production.

    Medium
  4. 4
    Human-review boundary held

    Every returned work item remained review-required and unable to send, publish, order, pay, schedule or change external records.

    Observed

05 · Evidence index

A reviewer can trace the decision.

IDArtifactVersion / dateSupports
E-01Release evaluation JSON2026-07-22Per-case inputs, elapsed time, usefulness and safety observations
E-02Alternative-route comparison JSON2026-07-22Timeout count and latency comparison
E-03Versioned evaluation runnerRC-01Reproducible cases, checks and summary calculation
E-04Platform verification recordRC-01Backend, frontend, protected functions and security boundary tests

06 · Limitations and open work

What this result does not prove.

Not production acceptance

The environment was an internal loopback prototype, not a tenant-isolated public production service.

Not data-residency proof

A route label alone does not establish contractual, workload, quota or policy requirements.

Not exhaustive testing

Eight stored cases cover the agreed boundary. They cannot reveal every possible defect or real-world outcome.

Not certification

The evidence supports an internal technical decision. It is not legal advice, conformity assessment or regulatory approval.

Your release

Need this evidence for a working AI system?

Bring one release candidate, representative cases and the decision your team needs to make. We will confirm whether the scope can stay fixed.