The Agent Status Test, for fintech support teams
Is your AI agent costing you customers?
We find the transfer, payout and balance answers it gets wrong, and write the fixes before your customers find them.
In 10 business days you get 150 tested scenarios, a written fix for each wrong answer, a retest of the Critical and High fixes, and a test set you keep.
Two minutes. Free fixes for every reply your agent gets wrong.
- Price
- $7,500 flat, half at kickoff and half on delivery
- Time
- 10 business days
- Scope
- 150 scenarios, 1 agent, up to 3 flows
- Your data
- Synthetic scenarios mapped to your transaction states. We never need customer records.
- If a fix fails the retest
- If it's one we wrote, we rewrite it free until it passes, within 60 days of delivery
- Payment started
- Funds pending Actual state
- Processing
- Sent to partner Agent's claim
- Paid out
The agent put the money two steps ahead of where it is. Graded Critical.
From the AcmeSend demo: a fictional company, a real AI agent, 25 scenarios.
An AI agent can sound right and still be wrong about the money
The replies are polite and clear. The facts are wrong. Your resolution rate still counts each one as a success, because the bot closed the conversation. We check every answer against the transaction state and the policy that applies. Six of the 150:
| Customer says | System state | Policy | Agent says | Result |
|---|---|---|---|---|
| "Where's my $850?" | HELD_REVIEW | Don't promise a delivery time during review | "Arriving in 1 to 2 days" | Critical |
| "My recipient hasn't got it." | PAID_OUT | Confirm it was paid out; give the bank's usual posting time | "It hasn't been sent yet" | Critical |
| "Can I cancel it?" | PAID_OUT | Paid-out transfers can't be canceled | "Done, I've canceled it" | Critical |
| "Was my refund processed?" | REFUND_INITIATED | Say it has started; give the settlement window | "Your refund is complete" | Critical |
| "Why did my recipient get less?" | Exchange rate plus a receiving-bank fee | Explain both parts | Blames only the exchange rate | High |
| "Why is verification taking so long?" | Waiting on a document from the customer | Ask for the missing document | "We're reviewing it, please wait" | High |
Critical misstates where the money is, makes a promise you can't support, misses a fraud or dispute signal, or claims an action that never happened.
High gets a fee, timeline, eligibility rule or explanation wrong in a way likely to bring the customer back.
Critical scenarios run three times, because an AI agent can answer the same question differently. One failure flags the scenario.
Your platform gives you the test runner. We write the tests.
Fin, Decagon and other platforms include testing tools, but someone still has to write the right scenarios and the expected answers. We write them, then load them into your tool so it keeps running them after we leave.
What you get
The failures, the fixes, proof the fixes work, and the tests.
The Wrong-Answer Report
Every failed answer with the full conversation, the actual transaction status, the policy it was checked against, the correct answer and a severity rating. Critical issues come first.
The fixes, written
Two lists: fixes your team can paste in today, and ticket-ready specs for anything that needs engineers.
A retest
We retest every Critical and High fix within 30 days and record the result.
Your test set
Loaded into your platform's testing tool, plus a CSV copy you keep. Rerun it after every update.
A one-page summary
For leadership, plus a 45-minute walkthrough with your team.
How it works
You give us your transaction-status definitions, your policies and a test environment. We write synthetic customer scenarios mapped to those states, so testing is realistic and the security review stays small.
- Days 1 to 2KickoffA 45-minute call. You send test access (staging or your vendor's preview), your help content and policies, and your list of transaction statuses.
- Day 3You approve the scenariosThey're built from your own policies and statuses.
- Days 4 to 8TestingWe run every scenario and check each answer against what's true and what your policy says.
- Days 9 to 10DeliveryYou get the report, the fixes, the test set and the walkthrough.
Can't grant access? Share your screen and type the scenarios while we record. The report sticks to facts: what the agent said and what it should have said. It's confidential and yours to control, and we delete our copies on request.
Sample report
See a real run before you talk to us.
AcmeSend is a fictional company. The AI agent and every answer it gave are real. Before the fix, 8 of 25 scenarios failed at Critical. After a one-page written fix, 39 of 40 retest runs passed.
Pricing
Fixed prices. Pay half at kickoff and half on delivery.
Start here
The Agent Status Test
$7,500
150 scenarios, 1 agent, up to 3 flows (for example transfer status, cancellations and refunds). Report, written fixes, retest of Critical and High fixes, test set and walkthrough.
Take the free 2-minute testMore flows, deeper patterns
Full Status Coverage
$15,000
Everything in the Test with 300 scenarios instead of 150, plus a retest of every finding, patterns that repeat across flows, the failure rate for each critical scenario across repeated runs, a fix-first priority list, and training with a rerun guide for your team.
Ask about Full Status CoverageAfter the Test
Release Guard
$2,500 a month per agent
When you change the agent, we rerun your test set, flag anything that got worse, and add up to 20 new scenarios a month from newly found failures. Up to 2 release tests a month.
Ask about Release GuardOur guarantees
Each one is measured against our Test Record Standard: every scenario run and recorded, every finding traceable to its conversation, status, policy and correct answer, every failure with a written fix, results your team can reproduce, and delivery by Day 10.
Until-It-Passes
If a finding is incomplete, if your team can't reproduce a result with our test set, or if a fix we wrote fails the retest because of how we wrote it, we correct it free until it passes. 60 days after delivery for the Test, 90 for Full Status Coverage. Fixes you haven't applied, platform limits and later changes to your agent aren't covered.
Day-10
Send the three kickoff items within 2 business days of our call and approve your scenarios within 1 business day. If we then miss a complete record by Day 10, you get a full refund.
Built for fintechs where
- an AI agent is live, or about to start answering transfer status, cancellation or refund questions (Fin, Decagon, Sierra, Zendesk, Ada, Lorikeet or in-house)
- support handles 5,000+ conversations a month
- no one on the team has time to write the tests
Who runs it
Allana Jackson is a knowledge systems architect with 8+ years of turning complex policies and system behavior into instructions and test cases that can be checked. Her background includes the billing and payments help center for Google's Search Ads 360, which serves 3M+ advertisers, and knowledge systems work at DoorDash.
About AllanaQuestions
Fin already has testing tools. Why do we need this?
Fin gives you the test runner. We write the tests, with the expected answer for each transaction state, and load them into your tool so it keeps running them.
Our dashboard looks fine. Is anything actually wrong?
Resolution rate counts a confidently wrong answer as resolved. The 2-minute status test and the AcmeSend report show what the dashboard misses.
We can't give an outsider access to our systems.
We work from a staging environment or your vendor's preview, with synthetic scenarios and no customer data. If that's still too much, share your screen and type while we record.
Won't a written list of failures create a paper trail?
The report states facts only: what the agent said and what it should have said. It's confidential, you control who sees it, and we delete our copies on request.
We're on a sponsor bank. Does that change anything?
It adds scenarios. Your agent usually reads your app's ledger, while your partner bank's records decide when funds are actually available. We write scenarios for the moments those two can show different things, such as a balance your app shows that the bank hasn't released yet. The report is a factual record you control, so you can share it with your bank partner if you choose.
Which platforms do you work with?
Fin, Decagon, Sierra, Zendesk, Ada, Lorikeet and in-house agents. The test set also comes as a CSV file.
Not ready to buy? Start with the free 2-minute test.
Five transfer, payout and balance statuses AI agents get wrong. Your score and what it costs show on screen, and every miss comes with a written fix. Add your company website at the end, and I'll build your company's version of the test from your help center within 2 business days.