Preview: Grade conversation, tools, policy, and business outcome together
An engineering preview of evaluating several evidence layers against one workflow without collapsing missing evidence into a pass.

This engineering preview describes a multi-layer grading direction for voice-agent workflows. It is not a released guarantee that every evidence type can be collected or graded for every agent, and it does not turn an incomplete trace into a complete readiness decision.
This capability is being developed with early-access and design-partner teams; availability depends on the voice stack and evaluation scope. The evidence available for a run is established during scoping and may differ across integrations, workflows, and team permissions.
One conversation can contain several decisions
A voice interaction can sound fluent while failing elsewhere. The transcript may show that the agent gathered the right details, while a tool trace shows the wrong argument. A tool may succeed, while a required disclosure was never given. The conversation may end politely, while the intended business action remains incomplete.
Grading those layers together means keeping separate questions visible: What happened in the audio and conversation? Which tool calls occurred? Were policy requirements satisfied? Did the intended business outcome occur? The answers can then be considered as evidence for one workflow without assuming that success in one layer cancels a blocker in another.
Evidence before judgment
The proposed grading path starts with observable evidence. Depending on the supported integration, that could include audio, transcript, trace, tool state, policy checkpoints, and a verifiable result. Each grader should identify the requirement it applies, the evidence it inspected, and the result it produced.
Some requirements are deterministic, such as whether a particular tool argument matched an agreed value. Others may require a model-assisted judgment or human review. The system should preserve that distinction. A model judgment is not the same thing as a directly observed tool result, and uncertain evidence should not be presented as certainty.
Missing data is a result
Multi-layer grading is useful only if absent evidence remains explicit. If a business outcome cannot be observed, the outcome layer should not pass by inference from a friendly transcript. If a policy definition is ambiguous, the evaluation should surface the ambiguity instead of silently selecting a convenient interpretation.
That approach supports four practical states at the workflow level: ready, caution, blocked, and not enough data. The final state still depends on criteria agreed for the evaluation. This preview does not establish universal thresholds or claim that a single grading policy fits every business process.
Building a reviewable record
The goal is to give technical teams a record they can inspect: the case that ran, the conditions applied, the evidence captured, the grading result for each requirement, and the reason a workflow reached its status. That record can support debugging and later regression work when the underlying evidence is available and approved for retention.
Teams interested in this direction can join the early-access list with a work email. Joining requires only that email; the form does not collect agent, workflow, stack, or release details. If we follow up, we will ask for the relevant context and constraints before discussing the integration path, evidence relationships, requirements, review needs, or supported scope.