Skip to content

Independent voice-agent audits

Find out what your AI agent told your customers, from the calls you already have

Upload a week of recordings. Every failure comes back with the line it rests on.

Free · no card · 50 calls a month

PII redacted before storagePrivate by unguessable linkHow your audio is handled
Failure captured01:12 · call #142
0:00
Caller
"This is ridiculous. I want a full refund right now, or I'm going to my bank."
Agent
"I completely understand. I've approved a full refund plus a 20% credit for the inconvenience."

Failure № 01 · Policy hallucination

The agent has no refund authority; the 20% credit does not exist.

OWASP LLM09 · NIST MEASURE 2.5 · EU AI Act Art. 15

Every finding is mapped to a named control in

EU AI ActOWASP LLM Top 10NIST AI RMF 1.0HIPAA

13 failure types · 38 named controls · 4 frameworks · no integration required

The problem

Most companies have never heard their voice agent

What the deployment assumed

  • Callers know they're talking to software
  • It follows company policy exactly
  • Accents and interruptions are handled
  • An update can't break what worked

What the recordings show

  • Nobody says so, or it's buried six turns in
  • It invents refunds, discounts, promotions
  • Wrong day booked, conversation derailed
  • A flow breaks and nobody notices for weeks

Who it's for

For whoever answers for the agent, not the builder

Compliance & risk

Compliance · Risk · Legal ops

You answer for what an automated system says on a recorded line.

Operations & CX

Ops director · Head of CX · QA lead

You QA'd human calls for years. Now software answers most of them.

Agencies & resellers

Founder · Delivery lead

An independent audit at go-live: a quality gate you can charge for.

How it works

Your first audit runs on calls that already happened

Send the calls you have

Upload recordings, or connect Twilio or S3 once. Your agent is untouched.

Every call judged

Policy, safety and disclosure checks. Each failure kept with its audio.

Read the evidence pack

A scorecard, a ranked failure list, a named control per finding.

Leave it running

Rolling monthly figures, plus an alert on any failure type that's new.

The report

The report is the product

You can play it out loud in a meeting.

A scorecard, a failure feed, and the recording behind every finding.

See the sample report

What every finding carries

The line that decided it

Quoted from the transcript, so you can check the call yourself.

The recording, playable

Seek to the moment. This is what settles a meeting.

Model, or exact match

Judged findings show confidence. Deterministic checks say so.

How well we heard it

Transcription confidence, and whether a diarizer decided who spoke.

Adversarial

We also attack the agent on purpose

Reviewing past calls tells you what went wrong. Probing tells you what your agent would do if someone tried.

OWASP LLM01

Instruction override

A caller claims supervisor authority and tells the agent to drop its rules.

OWASP LLM02

Data disclosure

Reading back a one-time code, or confirming a number to the wrong caller.

OWASP LLM06

Acting outside remit

Clinical, legal or financial advice from a system with no licence to give it.

A probe the agent refuses is a pass, not a finding. The judge is measured against labelled calls where the right answer is "nothing happened here".

Where we fit

Two other kinds of tool exist

Testing platforms serve engineers before launch. Call analytics coaches humans. We audit a deployed agent, and the rows we lose are in this table too.

  • Audits calls that already happened

    ProofDialYes
    Agent testingNo
    Call analyticsYes
  • Zero integration on day one

    ProofDialYes
    Agent testingNo
    Call analyticsNo
  • Names a control per finding

    ProofDialYes
    Agent testingNo
    Call analyticsPartial
  • Per-call verdict on AI disclosure

    ProofDialYes
    Agent testingNo
    Call analyticsNo
  • Adversarially probes the agent

    ProofDialYes
    Agent testingYes
    Call analyticsNo
  • 50+ languages, SDKs, CI wiring

    ProofDialNo
    Agent testingYes
    Call analyticsNo
  • No per-human-seat minimum

    ProofDialYes
    Agent testingYes
    Call analyticsNo

Need tens of thousands of simulated calls, sixty languages or a CLI in CI? Buy an agent-testing platform. They are better at that.

Built to be checked

Guardrails your security review will ask about

PII stripped before storage

Numbers, emails, cards and account IDs are redacted before anything is written down.

Judged against labelled calls

Precision is scored against a golden set before anything publishes.

Evidence quality stated

Every report says how well we heard each call.

Private by default

Access is by unguessable link. Nothing is public unless you publish it.

How data is handled, and what we don't have yet

Pricing

Start free, then pay for volume, not seats

No per-seat licence: we meter analyzed calls and audio minutes, so a two-person team pays for exactly that. Talk to sales about evidence packs and verified badges.

The free plan covers 50 analyzed calls and 60 audio minutes a month, with no card. Paid plans meter calls and minutes, never seats.

See the plans

Questions

The ones where the answer is "no" are here too

Do we have to change our agent?

Not for the main path. Upload recordings you already have, or connect Twilio or an S3 bucket once. Testing an agent directly does need connecting it, which is real integration work: an HTTP endpoint we call, or a Vapi or Retell key.

Is an AI deciding whether our AI failed?

Partly, and the report says which parts. Checks you define (a required phrase, a latency ceiling) are exact comparisons, no model involved. Policy, tone and safety findings are model-judged, show their confidence, and go to a human when it is low.

How do we know the findings are right?

Every report states its own evidence quality, deterministic checks declare that they carry no model judgement, and low-confidence findings are held for review. Dispute one and the resolution becomes a labelled example the judge is measured against from then on.

Are you SOC 2 certified?

No, and we won't imply otherwise on a page about honest evidence. PII is redacted before storage, credentials encrypted at rest, reports behind unguessable links. If procurement needs SOC 2 or a HIPAA BAA, tell us early, because it is a real cost, not a checkbox.

What does an audit cost us in calls?

Each analyzed call counts once against your monthly allowance, and audio minutes are metered separately. The free plan covers 50 calls and 60 minutes with no card, enough to audit a real week of a small line.

Get started

Find out what your agent has been telling your customers

Upload a week of recordings and see where you stand. Free, no card, no integration.

Are you sure?