Skip to content

AI Evaluation & Reliability

Building a Regression Testing Pipeline for Production AI Agents

How we moved from production-driven prompt changes to a release process that evaluates every change against real failure cases before deployment.

Daniel Dayto

Forward Deployed Engineer · Co-Founder & Technical Lead at Playmaker

Published
Reading time
12 min read
Diagram of a release gate: a prompt change flowing into a regression suite of scored calls, with a pass path to production and a blocked path back to review

Our voice agent at Playmaker answers the phone for a Texas HVAC and plumbing company and books jobs directly into ServiceTitan. Since March 2026 it has handled over 1,100 production calls. Every call is recorded, transcribed, and resolved to an explicit outcome: booked, escalated, callback, lead, or failed with context.

This article documents the evaluation and release pipeline we built around that system, and the production problem that made it necessary. I made the initial decision to automate prompt changes without a regression suite. That was a mistake. Most of what follows is the engineering that came out of correcting it.

The original feedback loop

Because every call resolves to a recorded outcome, production tells us exactly where the agent falls short. Call review surfaced failures like a mishandled reschedule, a caller asked to repeat their address, or a missed after-hours rule. Each confirmed finding became a system-prompt or rule change. This worked well, so I automated the flow: findings from production review fed prompt updates on a regular cadence, and the agent improved week over week.

Where it failed

I automated prompt updates before we had regression testing in place. That created a predictable problem: a change could fix the production failure it targeted while degrading behavior elsewhere.

We started seeing regressions in behaviors unrelated to the original change. Nothing crashed and no error rate moved. The agent simply handled certain call types slightly worse than it had before. Our review process was oriented toward finding new failures rather than re-verifying old behavior, so these regressions could sit unnoticed behind the improvements.

At that point, regression detection was still reactive. Some failures were identified from subsequent production calls rather than before deployment. For a system that books real appointments on a business's main phone line, that ordering is wrong.

Why prompt changes regress unrelated behavior

When you fix a bug in one service, you don't expect an unrelated service to break, because code has boundaries: modules, interfaces, and tests that localize a change. A system prompt has none of those. It is a single shared artifact in which every behavioral rule competes with every other rule for the model's attention.

Adding emphasis to one rule changes the relative weight of the rules around it. An edit that reads as local in the diff is in practice a global change to one function that handles every call. The same property applies to any system that modifies its own behavior through a shared artifact, including agent configurations and auto-tuned rules.

Requirements for the evaluation system

The correction was not to stop making prompt changes. Production review remained our main source of improvements. What we needed was verification that kept pace with the changes. Concretely, the system had to:

  • Score every production call consistently, on defined criteria, so quality is a number with evidence rather than an impression.
  • Detect regressions per dimension, not just in an overall average.
  • Run automatically against a candidate change before release, and block the release on a decline.
  • Produce evidence a human can audit: which turns, which tool calls, which deductions.
  • Keep the weekly improvement cadence intact.

Regression-suite architecture

The data layer already existed. Call transcripts and per-turn latency metrics come from ElevenLabs. Telephony events come from Twilio. Tool-call results, including customer lookups, availability checks, and job_created confirmations with ServiceTitan job IDs, are recorded by our Python/FastAPI services and stored in PostgreSQL alongside each call's transcript and final state.

The regression suite is built from declarative scenarios: seeded CRM state, a request, and expected outcomes. Each scenario drives the real request handlers against fake CRM adapters that record every external write to a ledger. Assertions run against both the response and the write ledger.

Testing at the write layer was a deliberate choice. What the agent says it did and what it actually wrote can differ, and the write ledger records the latter. Asserting on writes catches the failure class where the agent completes a workflow fluently but performs the wrong operation, which language-level checks tend to miss.

Scenarios are curated from production rather than written speculatively: one for each task type, each integration path, and each previously fixed failure. A scenario can reproduce a bug in its pre-fix and post-fix states, which is how we verify that a fix actually addressed the mechanism and not just the symptom.

Scoring production calls

The suite needs a measurement to assert on, and "did the call go well" is not one. Each production call is scored 0–100 across six weighted dimensions. Task completion carries the largest weight because a call that fails its task is a failed call regardless of how smooth the conversation was. The remaining dimensions exist to explain why a call scored the way it did.

DimensionWeightWhat it measures
Outcome / task execution40%Whether the agent correctly completed what the caller was trying to do
Conversation quality / customer effort20%Friction the caller experienced, measured from observable events
Operational reliability15%Whether tool calls, the ServiceTitan integration, and telephony functioned correctly throughout
Escalation quality10%Whether handoffs, transfers, and callbacks were correctly identified, routed, and completed
Latency / responsiveness10%Response speed to the caller, including tool-related delays, against defined thresholds
Policy compliance5%Whether the agent followed business rules and gave accurate information
The six dimensions. Weights redistribute proportionally when a dimension doesn't apply to a call.

Task execution is evaluated intent-first. The evaluator identifies the caller's primary task, then applies criteria specific to it. A booking is checked on customer identification, service selection, valid availability, appointment creation, and confirmation. A reschedule is checked on finding the correct existing appointment and completing the change. A callback is checked on capturing the right customer, reason, and destination. The failure modes differ by task: an agent can run a flawless booking flow against the wrong appointment, and only intent-specific criteria catch that.

When a dimension doesn't apply, its weight redistributes proportionally across the others. Escalation quality is only scored when a transfer or callback was actually appropriate. Scoring it on calls where it wasn't would hand out free points and compress the differences between good and bad calls.

Deterministic checks vs. LLM evaluation

The evaluator combines deterministic checks with an LLM judge, and the boundary between them follows one rule: outcomes available from system state are not inferred from language. Whether ServiceTitan created a job is deterministic; the job ID either exists or it doesn't. Whether the booking API needed retries is deterministic. Whether the agent unnecessarily repeated a question requires interpreting the conversation, so that goes to the judge.

The judge is constrained to observable events: the agent repeats a question, the caller repeats information or corrects the agent, unnecessary confirmations, conversation loops, failure to acknowledge something already said, explicit frustration. It must cite the specific turns behind every deduction. Anchoring on countable events instead of sentiment is what makes two evaluation runs agree with each other, and what lets a human audit any score in about a minute.

Reliability is scored from system records rather than the transcript. A call can complete its task and still lose reliability points if the booking write needed multiple retries. The caller didn't notice, but recoverable failures at current volume are the ones that stop being recoverable at higher volume, so they need to show up in the score.

Response timeRating
Under 1.5sExcellent
1.5 – 2.5sGood
2.5 – 4sAcceptable
4 – 6sPoor
Over 6sSevere friction
Latency thresholds (10% weight). Includes both conversational response time and tool-related delays such as lookups, availability checks, and booking confirmation. Values come from ElevenLabs' per-turn metrics.

Critical-failure handling

A weighted average has a specific weakness: a serious failure can hide under strong performance in the other dimensions. A call that hallucinates a price but is otherwise fluent, fast, and reliable could still average out to an acceptable number. To prevent that, certain failures cap the maximum score regardless of everything else:

FailureMaximum score
P0: critical compliance or security failure20
Materially incorrect booking or reschedule40
Failed primary task50
Ignored an explicit request for a human50
Hallucinated pricing or policy50
Score ceilings applied on top of the weighted model.

Score bands are 90–100 excellent, 80–89 good, 70–79 acceptable, below 60 failed. The caps exist so that a call that booked the wrong job, or talked past a caller asking for a human, can never read as acceptable on a dashboard. Every score ships with its evidence and deductions, so the number is a benchmark and the breakdown identifies the specific engineering or operational change the call is asking for.

The release gate

A candidate prompt or rule change runs against the regression suite before release. Results aggregate into an overall score and per-dimension scores, compared against the current baseline. If the candidate falls below the baseline overall, or exceeds the allowed regression for an individual dimension, CI blocks the release.

Per-dimension tolerances exist because the overall score alone is insufficient. A change can improve the aggregate while materially degrading one dimension, for example better latency alongside worse escalation handling. The aggregate would pass; the per-dimension check catches it.

python
def evaluate_release(candidate_prompt: str, baseline: Baseline) -> GateResult:
    """Score the candidate against the regression suite."""
    results = []
    for scenario in regression_suite():      # curated production scenarios
        outcome = run_scenario(candidate_prompt, scenario)
        results.append(score_call(outcome))  # 0-100 weighted model

    candidate = aggregate(results)           # overall + per-dimension means

    if candidate.overall < baseline.overall:
        return GateResult.blocked("overall score declined", candidate)
    for dim, score in candidate.dimensions.items():
        if score < baseline.dimensions[dim] - TOLERANCE[dim]:
            return GateResult.blocked(f"regression in {dim}", candidate)

    return GateResult.passed(candidate)      # promote as the new baseline
The gate's essential shape (illustrative, not production source).

Calibrating the evaluator against humans

An LLM judge is itself a model that can be wrong, and a release gate built on an unreliable evaluator is worse than no gate, because it converts noise into confident block/pass decisions. Before machine scores were trusted at volume, we compared them against human reviews of the same calls: agreement rates, precision and recall, and in particular the false-pass rate on the criteria that drive score caps. A judge that occasionally scores conversation quality a few points high is tolerable. A judge that passes a call that should have been capped for a hallucinated price is not.

Automated scoring sits behind an explicit acceptance gate on those calibration metrics. Calibration also isn't a one-time step: human-review queues and promoted gold-label calls continue to check the judge against people as prompts, rules, and traffic change.

How production failures enter the suite

Post-call data arrives by webhook and is scored. Low-scoring and flagged calls go to review. When a failure is confirmed and fixed, its scenario is added to the regression suite permanently. Scenarios come from production because production is where the surprising failures are; speculative test cases mostly re-verify behavior we already believed in. The permanent addition matters for a specific reason: without it, nothing stops a later prompt change from quietly reintroducing a failure we already paid to find.

What changed operationally

The release process for agent behavior now works like a release process for code. A change is proposed with a hypothesis, verified against the suite, and either promoted with a measured improvement attached or blocked with a specific, cited regression. The weekly improvement cadence continued, without the undetected regressions that used to come with it.

The scoring model also changed how we prioritize. A QA audit of 157 production calls measured booking conversion on eligible calls at 38.8%. The useful part was the decomposition: each missed booking maps to a scored dimension and a cited cause, which turns "improve conversion" into a ranked list of specific fixes, each verifiable against the suite after it ships.

The system has real limitations. Scenario coverage is only as good as our curation, and a failure class we haven't seen yet isn't in the suite. The LLM-judged dimensions carry judge error inside the bounds we measured during calibration, not zero error. The latency thresholds are operating heuristics, not derived constants. I'd rather state that than imply the gate catches everything.

Lessons

  • Build the evaluation harness before automating changes to agent behavior. We retrofitted it after regressions reached production, which cost more than building it first would have.
  • Score from system records where possible, use LLM judgment only where interpretation is genuinely required, and calibrate that judgment against human review before trusting it in a gate.
  • Track per-dimension baselines with tolerances, not just an overall score. The regressions that mattered most in our system were single-dimension declines that an aggregate would have hidden.

About the author

Daniel Dayto

Forward Deployed Engineer · Co-Founder & Technical Lead at Playmaker

Daniel Dayto builds and deploys production conversational AI systems for customer operations. His work spans voice agents, RAG assistants, CRM and dispatch integrations, multi-tenant infrastructure, and workflow automation.

Related articles

Deploying AI into a real operation?

I work with teams shipping voice agents, RAG systems, and workflow automation into production. Open to Forward Deployed Engineering, Applied AI, and founding technical roles.