← All News & Updates

Preview: Generate domain-specific evals from business context

A product preview of turning workflow requirements, policies, tools, and caller conditions into a scoped evaluation plan.

Plum and amber business-context paths resolve into a structured network of voice evaluation scenarios

This preview describes the direction we are taking for turning business context into domain-specific voice-agent evaluations. It is not a generally released feature or a promise that arbitrary documents can be converted automatically.

This capability is being developed with early-access and design-partner teams; availability depends on the voice stack and evaluation scope. Today, evaluation scope is established through the focused early-access and design-partner process, where the team and Voxeval define the workflows and evidence that matter for a release.

Begin with requirements, not generic prompts

A useful evaluation needs more than a collection of plausible caller questions. It needs to reflect the steps of an important workflow, the policy constraints around those steps, the tools the agent may call, the caller and audio conditions that can change behavior, and the business result that should follow.

The product direction is to organize those inputs into explicit evaluation cases. A scheduling workflow, for example, may require confirmation before a tool submission. A billing workflow may require identity verification before account details are discussed. These examples explain the model; they do not describe customer results or current universal coverage.

Keep the source context visible

The preview is designed around a traceable path from a requirement to a scenario and then to the evidence used for grading. That path matters because technical teams need to understand why a case exists, what behavior it expects, and whether a failed run points to the agent, the test definition, or missing context.

Inputs can include workflow steps and success criteria, policies and tool requirements, caller conditions, and approved incident evidence. Each input still needs review. Ambiguous requirements should remain visible as ambiguity rather than being silently converted into a confident test.

Scope before generation

In early access, the practical starting point is one agent and five important workflows. The team identifies what makes each workflow ready, blocked, or uncertain. Voxeval can then use that bounded context to shape a domain-specific evaluation plan and the calls needed to examine it.

The word “generate” in this preview does not mean the system can independently discover a company’s correct policy or business outcome. The participating team remains responsible for providing accurate requirements and deciding which evidence supports its release decision. Human review is especially important when policies conflict, a tool contract is incomplete, or an outcome cannot be observed directly.

What we are learning

The design question is how to reduce the distance between operational requirements and repeatable evaluations without hiding judgment. We are exploring which inputs can be structured reliably, which need clarification, and how to preserve a readable connection between each case and its source context.

Teams interested in shaping that workflow can join the early-access list with a work email. Joining requires only that email; the form does not collect agent, workflow, stack, or release details. If we follow up, we will ask for the relevant context and constraints before discussing participation, timing, or supported scope.