Preview: Workflow-level readiness instead of one aggregate score
A product preview of showing release evidence by workflow and condition, with ready, caution, blocked, and insufficient-data states.

This preview describes how Voxeval is approaching readiness at the workflow level instead of reducing a voice-agent release to one aggregate score. It is a product direction, not a universally available dashboard or a standard readiness threshold for every team.
This capability is being developed with early-access and design-partner teams; availability depends on the voice stack and evaluation scope. In a focused engagement, the participating team and Voxeval first agree on the workflows, conditions, evidence, and decision criteria the readiness view will represent.
The release question is specific
An agent is rarely ready or unready in the abstract. It may complete appointment scheduling in clean audio but miss a required confirmation after an interruption. A billing workflow may pass for one tool state and fail when escalation is required. Combining those results into a single number can hide the exact condition that should change a release plan.
The preview organizes evidence around named workflows and relevant conditions. Each intersection can carry a readiness state and a reason. That structure lets a technical team examine the release in the same units it uses to reason about operational risk.
Four states, with reasons
The current design language uses ready, caution, blocked, and not enough data. “Ready” means the agreed evidence supports the workflow under the tested condition. “Caution” keeps a material concern visible without treating it as an automatic stop. “Blocked” identifies a requirement that prevents the planned release decision. “Not enough data” means the evaluation cannot support a conclusion.
Those labels are not independent certifications. Their meaning comes from the criteria established for the scoped evaluation, and a participating team remains responsible for its release decision. The system should show the evidence and blocker behind a state so reviewers can challenge or refine the conclusion.
Avoid score compensation
Averages can allow many easy successes to outweigh a small number of critical failures. Workflow-level readiness is intended to prevent that compensation when a missed policy step or wrong business action is release-blocking. A result in one workflow should not erase a blocker in another simply because both contribute to the same total.
The same principle applies to evidence coverage. Repeated clean-audio calls do not establish behavior under an untested caller condition. When coverage is missing, the view should say so. The goal is a legible boundary around what was tested, not an impression of completeness.
Connect the matrix to action
A readiness view becomes useful when a team can move from a state to the underlying cases, evidence, and attributed blocker. The product direction is therefore tied to multi-layer grading and failure attribution rather than a standalone reporting surface.
For early-access work, the initial matrix is part of a bounded Workflow Readiness Sprint centered on one agent and five important workflows. Teams interested in shaping the representation can join the early-access list with a work email. Joining requires only that email; the form does not collect agent, workflow, stack, or release details. If we follow up, we will ask for the relevant context and constraints before discussing conditions, release decisions, support, timing, or scope.