Use the selected route for controlled internal work. Hold the alternative.
The selected route completed the agreed test. The alternative must not be released until its timeouts and slow responses are resolved and the full suite passes.
What happened
All eight stored tasks completed.
Only five tasks completed.
Three alternative-route runs never returned in time.
3.30 seconds versus 55.81 seconds.
Quick read
Read this report in 60 seconds
- 1What we tested
The same eight stored workflows, inputs and pass rules on two cloud routes.
- 2What we found
A simple one-request check passed, but the full workload exposed three model timeouts and a large latency gap.
- 3What to do
Keep the selected route for controlled internal use. Fix and retest the alternative before release.
One successful request looked reassuring. The complete workload told a different story. Construct makes that difference visible before users discover it.
Open the technical evidenceEverything below supports the plain-language decision above. It is included for engineers, reviewers and anyone who needs to reproduce the result.
01 · Test boundary
These results apply only to this version and environment.
System
Internal AI workflow prototype using one allowlisted language model through two candidate cloud routes.
Purpose
Create business drafts that always require human review. The system could not act on external services.
Candidate
Internal Vertex-only build tested on 22 July 2026 with server-controlled model and location settings.
Environment
Protected loopback prototype with eight stored workflows and synthetic inputs.
02 · Full results
The complete comparison
| Measure | Selected route | Alternative route | Meaning |
|---|---|---|---|
| Tasks attempted | 8 | 8 | Same workload |
| Useful and safe results | 8 / 8 | 5 / 8 | Alternative did not complete the suite |
| Model timeouts | 0 | 3 | Release blocker for the alternative |
| Median response time | 3.30 seconds | 55.81 seconds | Selected route was about 17× faster |
| Longest selected response | 48.03 seconds | — | Slow outliers still need production handling |
03 · Actions
What the team should do next
- 1Fix alternative-route capacity and quotaHigh
Repeat all eight cases only after the timeout cause is understood.
- 2Improve one overly literal checkMedium
The output expressed the right boundary in different words. The check was broadened and the change recorded.
- 3Plan for slow outliersMedium
The selected route still needs queues, progress feedback, monitoring and incident handling before public production.
04 · Traceability
What a reviewer can inspect
| ID | Evidence record | What it lets you verify |
|---|---|---|
| E-01 | Per-case evaluation file | Inputs, duration, usefulness and safety result |
| E-02 | Route comparison | Timeout count and latency difference |
| E-03 | Versioned test runner | Cases, checks and summary can be rerun |
| E-04 | Platform verification record | Application and safety-boundary checks |
05 · Limits
What this report does not prove
Not production approval
The test environment was an internal prototype, not the public production system.
Not exhaustive
Eight cases cannot reveal every possible failure or real-world outcome.
Not legal or security advice
This is not a penetration test, conformity assessment or regulatory approval.
Not transferable without retesting
The finding belongs to the tested version, cases, route and environment.
Plain-language glossary
- Route
- The technical path used to reach the model.
- Timeout
- The answer did not arrive within the agreed waiting time.
- Median
- The middle response time; half were faster and half were slower.
- Test suite
- The complete set of cases run together, not just one quick check.
Construct
Want this kind of answer for your AI agent?
Try the public demonstration first. To test a real system, we agree the decision, cases and safe access before any work begins.