- 01After-hours is revenue recovery, not cost reduction. The baseline is nobody answering, not a cheaper human.
- 02On cost substitution the break-even is simple: containment must exceed the AI's cost per attempt divided by a human's cost per call.
- 03That is why after-hours is the easiest voice case to justify and daytime overflow is the hardest.
- 04Deflection and resolution are different numbers. A system optimised for deflection counts an abandoned call as a win.
- 05You cannot know what missed callers would have done. Run a holdout, or you are reporting an assumption.
Method. The models below are arithmetic — they hold regardless of vendor and need no citation. Where market context is mentioned it comes from secondary sources listed at the end and should be re-verified before it goes in a business case. Nothing here is a claim about what any particular product achieves; the point is to give you the formula so you can test a claim someone else makes.
The wrong baseline sinks most business cases
This single substitution changes which number you are optimising. Against a human baseline you are trying to reduce cost per call and the ceiling on your value is the wage bill. Against a voicemail baseline you are trying to recover revenue that currently leaks, and the ceiling is however much business walks away between 6pm and 8am.
“Compared to a person, an AI answering a call saves money. Compared to nobody, it makes money. Those are different orders of magnitude and most business cases pick the wrong one.”
It also explains a pattern that otherwise looks strange: the least technically demanding voice deployments are often the most valuable, and the most technically demanding ones are often the hardest to justify. After-hours answering is easy — the bar is voicemail. Daytime overflow is hard — the bar is a trained human who is already quite good.
Two different models, chosen by baseline
| Cost substitution | Revenue recovery | |
|---|---|---|
| Baseline | A person answers | Voicemail, or a ring-out |
| When | Business hours, overflow, queue | Nights, weekends, holidays, lunch |
| You are optimising | Cost per resolved call | Captured conversion value |
| The binding test | Break-even containment rate | Value per recovered caller × volume |
| Ceiling on value | The wage bill you displace | All the business currently lost |
| Difficulty | High — the incumbent is good | Low — the incumbent is a beep |
Almost nobody is purely in one case, so split the call volume by hour and model the two halves separately. Adding them together and computing one blended number is how a strong after-hours case gets diluted by a weak daytime one until the whole thing looks marginal.
The break-even containment rate
The derivation is short enough to do in full, and worth doing because the result is the single most useful number in a voice business case.
Let C_h = fully-loaded cost of a human handling one call
C_a = cost of one AI attempt (platform + telephony + tokens)
r = containment rate (share of calls the AI finishes alone)
Every call costs the AI attempt. The (1 − r) that escalate ALSO cost a human:
blended cost per call = C_a + (1 − r) · C_h
Automation is cheaper than human-only when:
C_a + (1 − r) · C_h < C_h
C_a < r · C_h
r > C_a / C_h ← the break-even containment rateThe term teams omit is the first one. They model escalated calls as costing what they always cost, forgetting that the failed AI attempt was also paid for — and on a voice call it was paid for in per-minute telephony while the caller was talking to something that could not help them.
| If AI costs this share of a human call | You need containment above |
|---|---|
| 5% | 5% |
| 10% | 10% |
| 25% | 25% |
| 50% | 50% |
| 75% | 75% — very hard; reconsider |
Two implications worth stating. First, the table is why cheap-per-attempt systems are so much easier to justify than good-per-attempt ones: halving cost per attempt halves the containment you need. Second, break-even is not the goal — it is the floor. A project that lands exactly at break-even has spent implementation effort to achieve nothing, so the number to target is comfortably above it.
Where the AI cost actually goes
C_a is usually underestimated because people count the subscription and forget the rest. On a voice call the components are telephony per minute, speech-to-text per minute, model tokens for every turn, text-to-speech per character, and whatever platform fee sits on top — and the token cost scales with conversation length, which scales with how badly the call is going. A failing call is more expensive than a succeeding one, which is a mildly perverse property worth knowing about.
The revenue-recovery model
Let N = after-hours calls per month
p_ans = P(caller converts | call was answered)
p_vm = P(caller converts | reached voicemail) ← the term nobody measures
V = average value of a conversion
C = total monthly cost of the answering system
monthly value = N · (p_ans − p_vm) · V − C
Sensitivity, in order of how much they move the answer:
1. p_vm — if callers reliably call back, p_vm is high and the case collapses
2. V — a £4,000 job and a £40 job are not the same business
3. N — easy to measure, and usually smaller than people assume
4. p_ans — bounded above by your close rate, not by the AINotice that p_vm — the probability a caller converts anyway after hitting voicemail — is doing most of the work, and it is unobservable without an experiment. If your callers are patient and you are the only plumber in town, p_vm is high and answering at 2am buys you very little. If they are shopping and will simply call the next number, p_vm is near zero and every missed call is a lost job.
- The one-question screen
- Before modelling anything, ask: if we do not answer, does this caller call back tomorrow, or call a competitor? The answer is the difference between a marginal project and an obvious one, and the business owner usually knows it without any data at all.
Why the last term is bounded by you, not the AI
p_ans cannot exceed the rate at which your team converts a live, answered enquiry. An AI that books appointments flawlessly still cannot convert a caller who was never going to buy. This is worth being explicit about because vendor cases sometimes model the AI as improving conversion above the human baseline, which it does not — it changeshow many callers reach a baseline at all.
Metrics that mislead
| Metric | What it counts | How it lies |
|---|---|---|
| Deflection | Calls that did not reach a human | An abandoned call deflects perfectly |
| Containment | Calls the AI completed without escalating | Includes calls it completed wrongly |
| Resolution | Calls where the need was actually met | Requires an outcome you can verify later |
| Abandonment | Callers who hung up mid-conversation | The honest counterweight to the other three |
The pairing that cannot be gamed is containment against resolution: if containment rises while resolution stays flat, the system is finishing more calls without helping more callers. That is the signature of a system that has learned to end conversations rather than solve problems, and reporting containment alone hides it completely.
The fourth deserves emphasis: divide by resolved calls, not by total calls. Cost per call always improves with automation because the denominator includes everything the system touched. Cost per resolution is the number that can get worse, which is exactly why it is the one to report.
Measuring it honestly: run a holdout
p_vm is unobservable, the only honest way to measure recovered revenue is to withhold coverage from a random slice of after-hours calls for a fixed period and compare conversion between covered and uncovered groups. Every figure produced without a holdout is an assumption with a decimal point on it.This is uncomfortable to propose — you are deliberately not answering some calls — which is why almost nobody does it, and why almost every published ROI figure in this category is unfalsifiable. It is also cheap: a few weeks at a 10–20% holdout is usually enough to separate a large effect from nothing.
Randomise at the CALL level, not the day level.
→ day-level assignment confounds with day-of-week and weather
Hold out 10–20% of after-hours calls to voicemail as before.
Measure, for both arms, over the same window:
• conversion within 7 days (the primary outcome)
• revenue per call (conversion × value; the number leadership wants)
• callbacks the next day (this is your direct estimate of p_vm)
Run long enough to see the conversion window close — for a considered
purchase that is weeks, not days. Stopping when the numbers look good
is how a holdout becomes a marketing exercise.The third measurement is the prize. Callbacks in the held-out arm give you a direct, empirical estimate of p_vm — the term the whole model hinges on — and once you have it for your business you can model every future coverage decision without another experiment.
A softer alternative if a holdout is genuinely unacceptable: compare after-hours conversion before and after launch, and accept that the result is confounded by season, marketing and everything else that changed. Weak evidence honestly labelled beats strong evidence that is fabricated.
When it doesn’t pay
| Disqualifier | The check | Why it kills the case |
|---|---|---|
| Low after-hours volume | Pull the call log by hour for a month | N is too small for any uplift to cover fixed cost |
| Low conversion value | Average order or job value | Even perfect capture cannot pay a monthly fee |
| High callback rate | How many voicemails call again next day? | p_vm is high, so the uplift term collapses |
| Regulated intake | Does a human have to take this? | Cost is irrelevant if the answer is not permitted |
Run the first check before anything else. It takes ten minutes with a phone log, it is the disqualifier that applies most often, and it is embarrassing to discover after a procurement process. A business receiving four after-hours calls a week is not a candidate at any price.
The fourth is a different kind of no. Where a human is required for regulatory or clinical reasons, the economics do not enter into it — and the right design is an agent that takes structured intake and escalates rather than one that attempts to resolve. That is an authority question rather than a cost one; see when an agent should ask.
The worksheet
AFTER-HOURS VOLUME
Calls outside staffed hours, per month ................. N = ______
(from your phone log, by hour — not an estimate)
VALUE
Average value of a conversion .......................... V = ______
Conversion rate on ANSWERED enquiries .................. p_ans = ____ %
Callback rate after voicemail (≈ p_vm) ................. p_vm = ____ %
(if you have never measured this, run the holdout first)
COST
Platform fee per month ................................. F = ______
Cost per AI attempt (telephony + STT + tokens + TTS) ... C_a = ______
Average call minutes ................................... ______
Fully-loaded human cost per escalated call ............. C_h = ______
Expected containment ................................... r = ____ %
RESULT
Recovered revenue = N · (p_ans − p_vm) · V
Total cost = F + N · C_a + N · (1 − r) · C_h
Monthly value = recovered revenue − total cost
SANITY CHECKS
□ Is r comfortably above C_a / C_h? (the break-even floor)
□ Is p_vm measured, or assumed? (if assumed, say so out loud)
□ Is cost per RESOLUTION improving? (not cost per call)
□ Would N alone kill this? (check first, it usually does)A note on how to read the result. If the monthly value is large and driven mostly by V and a low p_vm, the case is real and robust. If it is large only because p_answas set optimistically, it is not a business case — it is the vendor’s brochure with your logo on it.
For market context: the incumbent human answering-service market is substantial and being displaced, with established providers serving thousands of businesses and AI entrants publishing monthly pricing well below traditional per-call rates. That context explains why the category is crowded; it does not tell you whether it works for your call volume, which only the worksheet does.
And if the numbers do work, the implementation question is a different article — the thing that determines whether callers tolerate the system is not cost but latency and turn-taking. See the voice agent stack for why the delay a caller notices is mostly a configured wait rather than inference.
Frequently asked questions
How much does an after-hours answering service cost?
Published pricing for AI answering and receptionist services generally runs from tens to a few hundred dollars a month, while human answering services typically price per call or per minute. The more useful number is not the subscription but the cost per resolved call, which depends on how often the system finishes the job without a human.
Is after-hours answering a cost saving or a revenue gain?
For genuinely after-hours calls it is a revenue gain, because the baseline is not a human answering more cheaply — it is nobody answering at all. That distinction matters because it changes which number you compare against: captured revenue rather than saved labour cost, and the two justify very different spending.
What containment rate do you need for AI answering to pay for itself?
On a pure cost-substitution basis, containment must exceed the ratio of the AI's cost per attempt to a human's cost per call. If the AI costs a tenth of a human interaction you need better than roughly 10% containment to break even, which is easy; if it costs half as much you need better than 50%, which is not.
What is the difference between deflection and resolution?
Deflection means the call did not reach a human. Resolution means the caller's problem was actually solved. They diverge whenever a caller gives up, and a system optimised for deflection will happily count an abandoned call as a success — which is why deflection alone is a misleading measure of an answering system.
How do you measure the value of answering calls you previously missed?
With a holdout. You cannot know what fraction of missed callers would have converted, because you never spoke to them — so you have to withhold coverage from a random slice of after-hours calls for a period and compare conversion between the covered and uncovered groups. Any figure produced without a holdout is an assumption wearing a decimal point.
When does an AI answering service not pay for itself?
When after-hours volume is low, when the value of a conversion is small relative to the subscription, when callers reliably call back the next day so little is actually lost, or when the intake is regulated or clinically sensitive enough that a human is required. Low volume is the most common of these and the easiest to check first.
Sources
- 01Answering services market size — Kentley Insights, 2024
- 02Smith.ai receptionist review and scale — GetVoIP, 2025