Voxeval blog

Evidence for voice-agent engineering.

Practical guidance for evaluating real workflows, learning from failures, and deciding when a voice agent is ready to release.

From the desk

Recent

  1. Why voice-agent evaluation has to happen at the conversation level11 min read
  2. Hybrid voice AI wins the edge cases: design automation lanes and human exits11 min read
  3. Instrument voice-agent latency as a trace, not a stopwatch11 min read
  4. Your barge-in test stops too early: evaluate what happens after the interruption11 min read
  5. The repeat-yourself rate: the voice AI handoff metric buyers should demand11 min read
  6. The customer changed the order mid-sentence. Did your voice agent change the cart?10 min read
  7. Restaurant voice AI needs an order-state engine, not a clever prompt11 min read
  8. Text normalization is on the critical path: keeping streaming TTS fast and correct10 min read
  9. Build a streaming TTS torture test for dates, money, phone numbers, and promo codes11 min read
  10. Stop selling voice AI on cost per call. Model cost per resolved outcome.11 min read
  11. Always-on VAD or push-to-talk? Build for noisy Indian streets6 min read
  12. The cost of conversational silence: latency budgets for Indian voice agents7 min read
  13. Seamless voice handoffs: WebRTC, SIP, and the multi-model constellation7 min read
  14. The PIN code nightmare: fixing digit capture on Indian phone calls8 min read
  15. RBI's collection-call curfew: put the 8 AM to 7 PM rule in code7 min read
  16. Beyond vibe-testing: regression gates for Indic vernacular voice agents6 min read
  17. Stochastic prompts, rigid banking APIs: close the IFSC and UPI payload gap7 min read
  18. Vapi vs. Retell AI: an India DPDP data localization review6 min read
  19. Voice biometrics plus OTP? Build payment AFA without trusting the voice channel6 min read
  20. Can an AI judge grade a phone call? A calibration guide12 min read
  21. Building compliant voice agents: HIPAA and PCI data flows13 min read
  22. From SOP to voice-agent eval suite: a practical conversion guide12 min read
  23. Hindi-English voice agents: how to test code-switching12 min read
  24. How to evaluate a voice agent: 10 metrics beyond word error rate12 min read
  25. How to run an STT bake-off on your own voice-agent calls11 min read
  26. Build a LiveKit voice agent, then add 25 tests before launch12 min read
  27. From production call failure to reproducible regression case10 min read
  28. Testing AI-to-human handoffs: escalation, context, and recovery11 min read
  29. Vapi vs Retell vs LiveKit vs Pipecat: which voice AI stack fits?13 min read
  30. Full-duplex voice changes the eval plan9 min read
  31. The 2026 voice-agent benchmark gap10 min read
  32. Production voice agents need release gates, not just low latency10 min read