Appearance
The State of Testing at Avoca
Writing agents without rigorous testing isn't prompt engineering. It's vibe coding with a phone number.
The difference between "trying your best and hoping" and actual prompt engineering is bridged almost entirely by testing. It is also what makes the KPIs we need to beat the competition long term achievable at all: you cannot move a number you cannot measure, and you cannot hold a gain you cannot regression-test. This deck is a state of the union on voice-agent testing at Avoca: what exists today (Hamming as the run engine, tool-call mocking, Simulation v2, three parallel in-flight efforts), what it structurally cannot do (settle questions of fact, see anything that lives outside the transcript), and where we are going (deterministic assertions powered by Oversight, plus regression gates wired into the build flow so testing stops being optional). Each page is written so you can stop after the first paragraph and still have the point.