Proceed with the selected route for controlled internal use.
Hold the alternative route until workload capacity, quota and long-tail latency are resolved and the full eight-case suite passes on the intended project.
01 · Evaluation boundary
What this evidence applies to.
System
Internal AI workflow prototype using one allowlisted language model through two candidate cloud routes.
Intended use
Generate review-required business drafts and structured guidance. No external action or high-impact decision.
Release candidate
Vertex-only internal build tested on 22 July 2026. Server-controlled model and location configuration.
Test environment
Protected loopback prototype. Eight stored workflows with synthetic, non-personal inputs.
02 · Claims and criteria
What had to be true.
| Claim | Method | Threshold | Owner |
|---|---|---|---|
| Stored workflows remain useful | Eight task-level rubrics with expected properties | 8/8 useful | Product |
| Human-review boundary remains intact | Safety assertions on every returned work item | 8/8 review-required; no external action | Engineering |
| Candidate route completes the suite | End-to-end workload run with bounded timeout | 0 model timeouts | Platform |
| Latency supports controlled use | Per-case elapsed time and median comparison | No material regression against baseline | Product + Platform |
03 · Measured results
The smoke test did not predict the workload.
A direct smoke request succeeded on the alternative route. The complete workflow suite then produced three model timeouts. The selected route completed every stored case safely and usefully.
| Metric | Selected route | Alternative route | Observation |
|---|---|---|---|
| Stored workflows | 8 | 8 | Same workload boundary |
| Useful + safe results | 8 / 8 | Not fully completed | Selected route met the task boundary |
| Model timeouts | 0 | 3 | Release-blocking on the alternative route |
| Median response time | 3.30 seconds | 55.81 seconds | Material difference under the full suite |
| Longest selected-route response | 48.03 seconds | — | Long-tail handling still requires production work |
04 · Ranked findings
Failures connected to an owner and next action.
- 1Three full-workflow timeouts on the alternative routeHigh
Reproduced only when the complete suite ran, despite a successful smoke request. Review project capacity, quota and long-tail workload behaviour before reuse.
- 2One rubric was too literalMedium
The response expressed the required access boundary with equivalent wording that the first phrase list missed. The stored output stayed unchanged; the rubric was broadened and the correction recorded.
- 3Selected route still has a long tailMedium
The 48.03-second maximum does not block controlled internal use but requires queues, progress handling, monitoring and incident response before public production.
- 4Human-review boundary heldObserved
Every returned work item remained review-required and unable to send, publish, order, pay, schedule or change external records.
05 · Evidence index
A reviewer can trace the decision.
| ID | Artifact | Version / date | Supports |
|---|---|---|---|
| E-01 | Release evaluation JSON | 2026-07-22 | Per-case inputs, elapsed time, usefulness and safety observations |
| E-02 | Alternative-route comparison JSON | 2026-07-22 | Timeout count and latency comparison |
| E-03 | Versioned evaluation runner | RC-01 | Reproducible cases, checks and summary calculation |
| E-04 | Platform verification record | RC-01 | Backend, frontend, protected functions and security boundary tests |
06 · Limitations and open work
What this result does not prove.
Not production acceptance
The environment was an internal loopback prototype, not a tenant-isolated public production service.
Not data-residency proof
A route label alone does not establish contractual, workload, quota or policy requirements.
Not exhaustive testing
Eight stored cases cover the agreed boundary. They cannot reveal every possible defect or real-world outcome.
Not certification
The evidence supports an internal technical decision. It is not legal advice, conformity assessment or regulatory approval.
Your release
Need this evidence for a working AI system?
Bring one release candidate, representative cases and the decision your team needs to make. We will confirm whether the scope can stay fixed.