QA for voice and chat agents

Hone the edge. Then watch the live blade.

Stropel generates synthetic customers, simulates conversations at scale, scores whether the agent hit its goal, and monitors production calls. A strop is the practice of putting an edge on, again and again.

The loop

Four strokes on the same leather.

One rubric across staging and live. Diff the prompt, prove the fix in simulation, then ship.

01

Simulate

Accents, noise, interruptions, policy traps. Load the floor before a customer does.

02

Evaluate

Goal completion, safety, tone. Criteria you define. Human review to set ground truth.

03

Observe

Production calls scored the same way. File issues. Alert to Slack or a webhook.

04

Hone

Diff the prompt, prove the fix in simulation, ship. The suite runs on every change.

Floors

Same platform. Rubrics you load.

Customer support, healthcare, financial services, logistics. Industry chips until named stories clear legal.

Customer support Healthcare Financial services Logistics
1,800 Simulated minutes on Edge each month
$0.12 Extra simulated minute after the pack
1 rubric Staging scores match live scores

From the bench

What the suite catches.

“We run the suite on every prompt change.” Voice PM, support agent team
“Simulation caught the identity skip before launch.” QA lead, healthcare intake line

Put the blade on the leather.

Start on Hone with 90 simulated minutes. Move to Edge when production scoring and CI gates join the loop.

Start honing