AI Voice Agent Cost Simulator
Work out what an AI voice agent costs before you build it. Add up infrastructure, text to speech, the language model, telephony and add-ons per minute, then multiply by your real call volume, and see where the estimate and the invoice diverge.

Most voice agent budgets are built by multiplying a headline per-minute rate by expected call minutes. That method is wrong twice over, and in opposite directions. This simulator assembles the rate from its published components, then applies the billing rules that decide what the invoice says.
What this tool does and who it is for
Give the simulator a call volume, an average call length and a stack, and it assembles a monthly cost from the published component rates rather than from a headline figure. It is for the person deciding whether to build a voice agent at all, and for the agency quoting one to a client and needing the number to survive the first invoice.
Two honest notes before the maths. First, this tool asks for an email address and is capped at twenty five runs a day, which is unusual on this site because most of the utilities here are open. The reason is that the output is a costing you can put in front of a client, and it tends to generate follow-up questions I would rather answer than leave you guessing at. If you do not want to hand over an email, everything below is static: the component table, the model and the transfer sensitivity are all here in full, and you can run the same arithmetic in a spreadsheet in ten minutes.

Second, every figure on this page is in US dollars, because that is the currency the platforms bill in. There is no exchange rate anywhere in this page, deliberately, because any rate I typed here would be wrong within a week.
The finding, before the working: the two most common ways of estimating voice agent cost are both wrong, and they miss in opposite directions. Multiplying total minutes by the platform's headline rate understates the bill, because the headline excludes telephony and add-ons. Multiplying total minutes by a correctly assembled blended rate overstates it, because a transferred call stops being an AI call the moment a human picks up. On the model below, the first method lands 16.7% low and the second 27.4% high, on the same call volume.
How to read the output, with a worked example
The rates below come from the Retell AI pricing page as published on 22 August 2026. Retell is used here because it publishes a full component breakdown, which most platforms do not; the same method applies to any platform that itemises. If you want the cross-platform version of the layer model, that is a different exercise and it lives in AI voice agent cost per minute, which builds the number up across five layers on five vendors. This page is about what happens to that number once billing rules are applied.
Step one: assemble the per-minute rate
Retell publishes a range of $0.07 to $0.31 a minute for voice agents. That range is the agent, not the bill. Here is a realistic receptionist configuration, built line by line:
| Component | Choice | Rate |
|---|---|---|
| Voice infrastructure | Retell voice infra | $0.055 / min |
| Text to speech | Retell platform voice | $0.015 / min |
| Language model | GPT 5 mini, standard tier | $0.012 / min |
| Telephony | Twilio through Retell | $0.015 / min |
| Knowledge base | opening hours, prices, services | $0.005 / min |
| Safety guardrails | on | $0.005 / min |
| Assembled rate | $0.107 / min |
Two things about that total. It sits inside the published range, which is reassuring. And it is 53% above the bottom of the range, which is the number people quote when they are estimating quickly.
Worth noting what the range does not appear to cover. A speech-to-speech configuration using GPT Realtime is listed at $0.345 a minute for the model alone; add the $0.055 infrastructure line and you are at $0.400 before telephony, which is above the top of the published $0.07 to $0.31 band. So if you are pricing a realtime speech-to-speech agent, the headline range is not a ceiling and you should build the number from components.
Step two: apply the billing rules
Now the part that decides the invoice. Retell's own pricing FAQ states three rules that a naive model ignores.
Calls are metered to the nearest second, with no per-call rounding up. Good news, and rarer than you would think.
You are billed during silence and hold time, because the speech-to-text engine stays active and listening. So the number that matters is connected duration, not talk time. A greeting that runs three seconds long and two thinking pauses of five seconds each add thirteen seconds to every call, and you pay for all of it.
Once a call is transferred, the AI fee stops and only telephony continues. This is the rule that breaks blended-rate models. Post-transfer minutes bill at $0.015 rather than $0.107, which is a little over seven times cheaper.
And two more from the same page: a call that fails to connect is not billed at all, while a call that reaches voicemail is billed for however long the agent stays on the line. That pair matters far more for outbound campaigns than for inbound, because connection rate is the variable you cannot control.
Step three: the model
A single-location business, inbound receptionist agent, 600 connected calls a month.
- 450 calls handled end to end by the agent, average connected duration 3.4 minutes, which is 1,530 minutes at $0.107 = $163.71
- 150 calls transferred to a human 1.2 minutes in, average total call length 5.0 minutes. That is 180 agent minutes at $0.107 = $19.26, plus 570 telephony-only minutes at $0.015 = $8.55
- One Retell phone number at $2.00 a month = $2.00
Total: $193.52 a month, across 2,280 minutes on the line.
Now the two wrong answers. Multiplying all 2,280 minutes by the $0.07 headline gives $159.60, which is $31.92 low, or 16.7% under the real variable cost. Multiplying all 2,280 minutes by the assembled $0.107 gives $243.96, which is $52.44 high, or 27.4% over. The second error is larger than the first, which is not what anyone expects, because the second method is the careful one.
Benchmark: how much the transfer rate moves the bill
Holding everything else constant and varying only the proportion of calls that reach a human:
| Transfer rate | Agent minutes | Telephony-only minutes | Real variable cost | Blended-rate estimate | Error |
|---|---|---|---|---|---|
| 0% | 2,040 | 0 | $218.28 | $218.28 | none |
| 10% | 1,908 | 228 | $207.58 | $228.55 | +10.1% |
| 25% | 1,710 | 570 | $191.52 | $243.96 | +27.4% |
| 40% | 1,512 | 912 | $175.46 | $259.37 | +47.8% |
Read the third column before the fifth. The more often the agent escalates, the less the platform bill comes to. Going from a 0% transfer rate to 40% takes the variable cost from $218.28 to $175.46, a fall of $42.82 a month or $513.84 a year, because every transferred minute is priced as a phone call rather than as an AI call.
That is the opposite of how most people frame escalation. A transfer is usually treated as the agent failing, and the instinct is to reduce transfers by making the agent handle more. On the platform bill, reducing transfers costs you money. The cost of a transfer lands somewhere else entirely, on the staff member who now has to take the call, and that cost is real and larger. Which means the honest version of this decision is not "how do I reduce transfers" but "which transfers are worth a human minute", and that is a question about your margins rather than about your voice stack. The escalation mechanics themselves are covered in Retell AI human transfer.
The telephony decision, which is usually made by default
Retell charges a flat $0.015 a minute for telephony through its bundled Twilio integration, and charges nothing for telephony if you bring your own SIP trunk or Twilio account. So the obvious question is whether bringing your own is worth it, and the answer is a number rather than an opinion.
From Twilio's UK voice pricing, fetched the same day:
| Twilio UK, pay as you go | Rate |
|---|---|
| Receive on a local number | $0.0100 / min plus $3.50 / month |
| Receive on a mobile number | $0.0100 / min plus $2.50 / month |
| Receive on a toll-free number | $0.0798 / min plus $2.70 / month |
| Call a UK landline | $0.0158 / min |
| Call a UK mobile | $0.0305 / min |
| Call a UK personal number | $0.5577 / min |
| Call a UK premium service number | $1.0479 / min |
For an inbound agent on a local number, the comparison is $0.015 a minute plus a $2.00 Retell number against $0.010 a minute plus a $3.50 Twilio number. Those are equal at 300 minutes a month. Above that, your own Twilio number is cheaper. On the 2,280 minutes in the model above it saves $9.90 a month, which is $118.80 a year, and is almost certainly not worth adding a moving part for a single business.
At agency scale it stops being trivial. Forty clients on the same volume is 91,200 minutes a month, where the half-cent saving is $456 and the extra number fees cost $60, netting $396 a month, or $4,752 a year. That is the point at which the telephony decision deserves half an hour rather than a shrug.
Outbound inverts it. Retell's flat $0.015 covers a call to a UK mobile that Twilio bills directly at $0.0305, so for an outbound campaign dialling mobiles the bundled rate is less than half the carrier rate. Both of those statements are true of the rates published today, and rates move, which is the entire argument for building the model from components rather than memorising a number. The wiring for either route is in the Twilio voice setup guide.
Look at the bottom two rows of that table before you build any outbound dialler. A UK premium service number bills at $1.0479 a minute, roughly 66 times the landline rate. A list with a handful of mistyped numbers in it is not a data quality problem, it is a line on your invoice.
Common mistakes
Modelling talk time instead of connected time. Billing runs for the whole connected duration, silence included. If your model uses the length of the transcript rather than the length of the call, it is short by however long the pauses were, and pauses are where latency shows up. Latency is therefore a cost problem as well as an experience problem.
Forgetting the monthly lines. Phone numbers at $2.00, verified numbers at $10.00 each, SMS at $20.00, concurrency at $8.00 per concurrent call per month above the twenty included, and knowledge bases at $8.00 each above the first ten. None of these is large. All of them are invisible in a per-minute model, and for an agency running many clients on one account they compound.
Sizing concurrency on average volume. Concurrency is provisioned against your peak, not your mean. 2,280 minutes a month across business hours averages well under one simultaneous call, so a single business will never approach the twenty free channels. An agency will, and not gradually: because every client's peak is Monday morning, the peaks coincide rather than smoothing each other out. Model the busiest hour, not the month.
Pricing outbound on minutes alone. Connected minutes are only part of it. Batch dialling adds $0.005 per dial whether anyone answers, branded caller ID adds $0.10 per outbound call, and voicemails bill for however long the agent talks to an answering machine. On a campaign with a low answer rate, the per-dial charges can exceed the per-minute ones.
Comparing platform headline rates against each other. Different platforms bundle different layers into their headline, so the headline numbers are not comparable quantities at all. The only fair comparison is an assembled rate for the same call shape on each platform, which is what Best AI voice agent platforms in 2026 does across five of them.
Costing the agent without costing the alternative. A $193 monthly bill is either expensive or trivially cheap depending entirely on what the missed calls were costing before, and that is a separate calculation with separate inputs. The missed call revenue calculator does that side, and the two numbers only mean something next to each other.
Frequently Asked Questions
Why does this tool ask for my email when the other tools here do not?
Because the output is a costing rather than a lookup, and a costing that is wrong costs somebody real money. It is gated at twenty five runs a day for the same reason. If that is not a trade you want to make, the component table, the model and the transfer sensitivity table on this page contain every number the tool uses, and a spreadsheet will get you the same answer.
Is the assembled rate the same as the platform's headline rate?
No, and that gap is the point of the page. The configuration above assembles to $0.107 a minute against a published floor of $0.07. The headline is the agent alone at its cheapest settings; the assembled rate includes the voice you actually chose, the telephony, and the add-ons you switched on. Neither number is dishonest. They measure different things.
Does a cheaper language model make a meaningful difference?
It can, but usually less than the voice does. In the component list above, moving from GPT 5 mini at $0.012 to GPT 4.1 at $0.045 adds 3.3 cents a minute, while moving from a platform voice at $0.015 to an ElevenLabs voice at $0.040 adds 2.5 cents. Together those two dropdowns take the assembled rate from $0.107 to $0.165, a 54% increase, on 1,920 minutes a month that is a difference of $111.36. The lever nobody pulls is call length, and it is usually the biggest one available.
How accurate is this before I have built anything?
The rates are exact and the arithmetic is exact. The inputs are guesses, and one of them matters more than the rest: the transfer rate, which moves the total by nearly 20% across the range in the table above and which you cannot know until the agent has taken a few hundred real calls. Model a range rather than a point, and treat the first month of live data as the thing that replaces the estimate.
Do these numbers hold for Vapi, Bland or ElevenLabs?
The method holds. The numbers do not. Every platform bundles a different set of layers into its headline rate, so the assembled figure has to be rebuilt for each one, and the billing rules differ too. Whether the AI fee stops on transfer, whether silence is metered, and whether failed connections are billed are three questions worth asking any vendor before you commit, and the answers are not always on the pricing page. Vapi vs Retell AI works through one of those comparisons in full.
What about the build cost, as opposed to the running cost?
Different question, and the answer is not per-minute. Prompt design, call flow, CRM integration, testing against real recordings and the inevitable second pass after the first week of live calls are a project cost, not a usage cost. The project cost calculator is the right tool for that side of it.
Should I simulate the conversation as well as the cost?
Yes, and they are separate exercises. A cost model tells you whether the thing is worth building. A conversation walkthrough tells you whether it will work, and it is where you find out that your agent cannot handle a caller who interrupts, gives a date as "week after next", or asks a question the knowledge base does not cover. The five things worth testing before launch are intent recognition, entity extraction, interruption handling, escalation timing and whether the outcome actually reaches the CRM. Every one of those is cheaper to fix in a script than in production.
Where this tool stops
The simulator prices minutes. That is genuinely useful and it is also the smaller half of the decision.
What it cannot do is price the other side of a transfer. The table above shows that escalating more often reduces the platform bill, and that is arithmetically true and practically misleading, because every transferred call spends a staff minute that costs far more than eleven cents. The real optimisation is deciding which calls are worth a human, and that depends on your average order value, your margin and how much of the call a human actually needs to hear. No cost model can supply those.
It also cannot tell you your transfer rate, your average call length or how much dead air your agent will produce, and those are the three inputs the answer is most sensitive to. They come from a built agent taking real calls. Which means the correct use of this page is to establish whether the idea survives a range of plausible inputs, and then to get a real number from a small live pilot rather than a larger model.
That pilot is the work I do: a working inbound agent on your own number, with the call outcomes writing back to whatever CRM you already use, deliberately narrow so the first month produces honest numbers instead of a demo.
If you have run the numbers above and they look plausible, the useful next step is a short call about which calls you would actually let an agent take. Tell me what your phones do on a Monday morning and we can work out whether this is worth building before either of us spends money on it.

Want this built against your real numbers?
A 30-minute call to scope the workflow, agent, or automation you actually need.
More ai tools
All tools
AI Agent Builder Demo
Sketch a multi-step AI agent in your browser, then test it against the arithmetic that decides whether it survives production: per-step reliability, retries, and how many boxes you can afford to add.

AI Automation Consultant
Chat with an AI automation consultant and get a prioritised strategy for your business

AI Chatbot Builder Demo
Configure a custom AI chatbot, feed it your own FAQs and product information, and have a live conversation with it in the browser — no code and no sign-up.

AI Knowledge Base Generator
Turn any document, FAQ, or email log into a structured AI knowledge base instantly

AI Outreach Message Generator
Generate personalised cold outreach emails and multi-touch follow-up sequences, with a worked rewrite showing what makes a cold email answerable and a note on UK and EU compliance.

Prompt Library & Tester
Test and iterate on AI prompts until the output is consistent enough for production. Includes a worked before-and-after rewrite, structured-output tips and the failure modes that only appear at scale.
Have a workflow that's burning hours every week?
Bring me one real bottleneck. I'll tell you whether it's worth automating, and what it would take.