AI Phone Receptionists for Local Contractors — Deep Research Report
Prepared for: TJ / Patriot Project Group (SE Michigan) — August 2026
A note on source quality before you read anything else: this category is drowning in vendor-written “comparison” content. Roughly 70% of what ranks for “best AI receptionist 2026” is published by a competitor in the same category. I have marked every claim with its source type: [VENDOR-PRIMARY] = the vendor’s own pricing/docs page (highest trust for pricing), [INDEPENDENT] = measured benchmark or non-commercial source, [VENDOR-CONTENT] = a competitor’s marketing blog (treat as directional only), [UNVERIFIED] = I could not confirm it. I did not invent a single price. Where a vendor doesn’t publish pricing, I say so.
EXECUTIVE SUMMARY — Your Three Questions, Answered
1. How good can it be in 2026?
Good enough to sell, not good enough to leave unsupervised. The honest state of the art:
- Latency is the tell, and it’s worse than vendors claim. The only independent, caller-side benchmark I found (openbenchmarks.com, 2,078 measured turns across 5 platforms, measured from real call audio with Silero VAD, not vendor timestamps) shows no major platform achieves sub-second caller-experienced response. Telnyx 1,296ms median, ElevenLabs 1,424ms, Bland 1,520ms, Vapi 1,558ms, Retell 1,740ms. Vendor-reported latency runs ~490ms below what a caller actually experiences. (openbenchmarks.com) [INDEPENDENT]
- Voice naturalness is essentially solved for a 2-minute intake call. ElevenLabs v3 and Cartesia Sonic 3.5 are top-tier on naturalness; Cartesia Sonic Turbo (~40ms), ElevenLabs Flash v2.5 (~75ms) and Rime Coda (sub-100ms) are the realtime picks. (MarkTechPost, Cekura) [VENDOR-CONTENT/mixed]
- The real failure mode is job-site noise, not the voice. An Interspeech study cited in 2026 voice-AI literature found overlapping speech at moderate noise pushed transcription error to 74.6% vs 16.8% on clean audio — a 4.4x degradation; background noise alone roughly doubles error. (FutureAGI) [VENDOR-CONTENT citing academic] For excavation and demo contractors whose customers call from job sites with equipment running, this is your #1 quality risk — not the AI’s intelligence.
- Callers say they hate it; callers’ behavior says otherwise. AnswerConnect/OnePoll surveyed 6,000 consumers across US/UK/Canada in April 2026: 85% prefer a real person (up from 83%), 31% say they’d hang up if connected to AI (up from 29%), 59% frustrated with AI agents. (AnswerConnect) ⚠️ AnswerConnect is a human answering service with a “Pledge People, Not Bots” campaign — maximum bias, treat as a ceiling on stated preference, not behavior. Counterweight: a Salesforce survey of 14,000 consumers found 68% prefer AI for simple tasks like scheduling rather than waiting on hold. (SurveyMonkey summary) [INDEPENDENT-ish]
- The bar you’re actually clearing is voicemail, and voicemail is terrible. ~80–86% of callers who reach voicemail hang up without leaving a message. That’s the real comparison, and it’s why this works.
- Realistic performance targets: 70–85% autonomous handling (“containment”) on routine calls; a transfer rate above 30% means your knowledge base or prompt is broken. (TheStackInsiders — GHL affiliate, [VENDOR-CONTENT]). One agency reports median home-services results of answer rate 60%→99% and booked-service rate on inbound 32%→49% across 25+ deployments. (Tested Media — runs a competing voice-AI agency, [VENDOR-CONTENT]).
Bottom line: yes, you can build a demo line that a contractor will believe. No, you cannot deploy it and walk away. The gap between “impressive demo” and “safe with a real customer at 11pm” is entirely operational discipline, covered in the Reliability Playbook below.
2. How do we do it?
Path A (verticalized products) is where you start. Path B (GHL-native + a white-label layer) is where you standardize. Path C (full build) you should not touch until 40+ clients.
The specific answer for your situation: your bundle already includes GoHighLevel. That makes the calculus different from a generic agency. GHL Voice AI books natively into GHL calendars, logs natively to the GHL CRM, and triggers GHL workflows for SMS summaries and missed-call textback — with zero middleware. Every other option requires a Zapier/webhook bridge that is one more thing to break at 2am. Middleware is the #1 source of “the booking didn’t show up” failures.
3. What does it cost?
At a realistic 250 inbound AI minutes/month per contractor:
| Path | Your cost/client/mo | Fits ≤$150? |
|---|---|---|
| GHL Voice AI, pay-per-use, Cartesia/OpenAI TTS | ~$30–$45 | ✅ Huge headroom |
| GHL Voice AI, pay-per-use, ElevenLabs V2.5 TTS | ~$40–$55 | ✅ |
| GHL AI Employee Unlimited ($97 flat) | ~$109–$119 | ✅ (but wasteful under ~1,100 min) |
| Assistable.ai (white-label, Growth tier @10 clients) | ~$76 | ✅ |
| Retell direct | ~$26–$35 | ✅ (+ your build time) |
| Vapi direct | ~$24–$32 | ✅ (+ 14-day retention problem) |
| Goodcall Growth | $108–$129 | ✅ |
| Rosie Scale | $149 | ✅ (exactly at ceiling) |
| Twilio ConversationRelay build | ~$25–$35 marginal | ✅ (+ weeks of eng) |
| Sameday Launch | $449+ | ❌ |
| Smith.ai Basic (90 calls) | $765+ | ❌ |
| Avoca | ~$1,000–$3,500 (est.) | ❌ |
Other agencies retail this at $197–$497/mo standalone (Trillet, TheStackInsiders — both [VENDOR-CONTENT], but consistent across sources). Your $500–$1,000/mo bundle with website + SEO + GBP + app + receptionist is well-positioned; the receptionist alone justifies $200–$300 of it.
PART I — STATE OF THE ART, AUGUST 2026
Latency: the number that matters, measured honestly
The single most useful thing I found is an independent benchmark that measures TTFAB (Time To First Audio Byte) from the recording of a real phone call — the actual silence a caller sits through — rather than from vendor telemetry.
| Platform | Median TTFAB | p95 | Tail ratio | Usable turns | Discarded |
|---|---|---|---|---|---|
| Telnyx | 1,296 ms | 1,856 ms | 1.43× | 419/432 | 13 |
| ElevenLabs | 1,424 ms | 1,768 ms | 1.24× | 429/432 | 3 |
| Bland AI | 1,520 ms | 2,248 ms | 1.48× | 429/432 | 3 |
| Vapi | 1,558 ms | 2,008 ms | 1.29× | 382/432 | 50 |
| Retell AI | 1,740 ms | 2,259 ms | 1.30× | 419/430 | 11 |
Source: openbenchmarks.com, open data + code at github.com/openbenchmarks-labs/voice-agent-latency [INDEPENDENT]
Three things you should take from this table:
- Read the p95, not the median. Past roughly two seconds of silence a caller assumes the line dropped and starts talking — which collides with the agent’s reply and derails the turn. Bland (2,248ms) and Retell (2,259ms) cross that line one turn in twenty. On a 10-turn call that’s a wreck every other call.
- Vapi’s 50 discarded turns out of 432 is a red flag — an 11.6% turn failure rate under a controlled script versus 0.7% for ElevenLabs and Bland. The benchmark authors note discards happen when the agent starts a reply, stops, and resumes, splitting their detectors. That’s the agent’s behavior, not the instrument.
- Endpointing was pinned to 0.1s on Telnyx, Vapi and Retell; Bland and ElevenLabs run their shipped defaults. So Bland’s and ElevenLabs’ numbers still contain a wait the tester couldn’t equalize — meaning their engines may be faster than they look.
No platform measured achieves sub-1-second caller-experienced latency. Any vendor claiming “sub-600ms” (GHL’s marketing, Leadlock’s “100ms TTFA”) is measuring from inside their own stack, not from the caller’s ear. The benchmark authors confirmed this two ways: a platform’s own recording reads ~550ms earlier than an external recording of the same call, and vendor-reported latency runs ~490ms below independently measured.
Voice naturalness — which TTS actually
Current top tier for naturalness: ElevenLabs v3, Cartesia Sonic 3.5, Inworld Realtime TTS-2, Gemini 3.1 Flash TTS. For realtime voice agents specifically (where you’re trading a little naturalness for latency), the field converges on Rime, Cartesia, Deepgram Aura, ElevenLabs. Measured latencies: Cartesia Sonic Turbo ~40ms, ElevenLabs Flash v2.5 ~75ms, Cartesia Sonic 3.5 ~82ms end-to-end, Cartesia Sonic-3 ~90ms, Rime Coda sub-100ms, Deepgram Aura-2 ~200ms streaming. (MarkTechPost benchmark comparison, Inworld, Cekura)
Practical implication for you: inside GoHighLevel your TTS choice is a pricing lever with a quality cliff. Official GHL rates (HighLevel AI Product Pricing, modified July 31, 2026) [VENDOR-PRIMARY]:
| TTS provider/model | Rate | Effective |
|---|---|---|
| OpenAI TTS | $0.015/min | Apr 20, 2026 |
| Cartesia TTS | $0.015/min | Apr 20, 2026 |
| ElevenLabs V2.5 | $0.035/min | May 20, 2026 |
| ElevenLabs V3 | $0.170/min | Apr 20, 2026 |
ElevenLabs V2.5 at $0.035/min is the sweet spot — near-top-tier naturalness at 1/5 the cost of V3. At 250 min/mo that’s $8.75 vs $42.50. Use V3 only on your demo line if you want to show off.
Where it still fails
Five failure modes account for most production breakdowns (FutureAGI, Appther) [VENDOR-CONTENT]:
- Background noise / overlapping speech. 74.6% WER at moderate noise with overlap vs 16.8% clean. This is your excavation and demo clients’ callers standing next to a running machine. Mitigation: Retell sells “Advanced Denoising” at +$0.005/min — cheap insurance. GHL does not expose an equivalent knob.
- Accent and dialect recognition bias. Recurring production failure; SE Michigan has meaningful AAVE, Arabic-English, and Eastern European accent populations. Test this explicitly.
- False barge-in. VAD treats a cough, a truck, or a side conversation as a caller turn and the agent stops mid-sentence. Prevention combines energy threshold tuning (-45 to -35 dBFS), voice-classification VAD (Silero/WebRTC/custom CNN), and a 200–300ms minimum-duration guard.
- Speech hallucination. Agent makes unauthorized commitments — quotes a price that doesn’t exist, promises a same-day slot that isn’t available. Reported hallucination-related complaint rate without RAG grounding is roughly 0.34%, i.e. ~340 customer-visible incidents per 100,000 calls. [VENDOR-CONTENT, unverified methodology] At 400 calls/mo/client × 10 clients = 4,000 calls/mo, that’s ~14 incidents/month. Non-trivial.
- Escalation failure. The angry caller who says “just let me talk to a human” and the agent keeps qualifying. This is the one that costs you a client.
Real-world contractor outcome data — and why to distrust it
The most-cited numbers in this space trace back to agencies selling voice AI. I could not access Reddit directly (r/sweatystartup, r/GoHighLevel) — Reddit blocks both my search and fetch tools — and several “Reddit-sourced” claims I found in vendor blogs (a plumber’s 184-call log, a roofer’s $4,800 first month on Jobber AI) could not be traced to an original post. I am not citing them as fact.
What I can attribute:
- Median results across 25+ home-services deployments through April 2026: inbound answer rate 60%→99%; booked service rate on inbound 32%→49%; recurring maintenance attach 13%→33%; all-in stack cost $700–$1,500/mo. (Tested Media) [VENDOR-CONTENT — Tested Media sells CallSetter AI]
- Industry containment benchmark: 70–80% for leading contact centers; appointment-heavy verticals reach 85–90% on low-latency platforms. [VENDOR-CONTENT, directionally consistent across sources]
- GHL community’s top-voted Voice AI complaints: robotic voice quality (178 upvotes), no caller recognition (174), no spam call blocking (91). (Leadlock citing Sympana) [VENDOR-CONTENT, second-hand] — note these likely predate GHL’s ElevenLabs TTS options and the existing-number support added in 2026.
- Rosie has no verified reviews on G2, Capterra, or Trustpilot as of mid-2026 despite claiming 2,000+ businesses and 3.1M+ calls. (ServiceAgent review) [VENDOR-CONTENT — competitor] but the absence is independently checkable and worth weighing.
Treat every ROI number in this category as marketing until you generate your own from your first client. That is the single most valuable asset you will build in the next 90 days.
PART II — THE THREE PATHS
Path A: Verticalized Products
| Product | Price (verified) | Booking | White-label / agency | Verdict for you |
|---|---|---|---|---|
| Rosie | Professional $49/mo (250 min), Scale $149/mo (1,000 min), Growth $299/mo (2,000 min); Website Texting add-on $50/mo (25 convos, $1 ea. after) (pricing) [VENDOR-PRIMARY] | Appointment links by text on $49; direct calendar booking only on $149+; warm + live transfers on $149+ | Affiliate: 20% recurring, client gets 50% off month 1. Agency program exists but terms not published — inquiry form only (affiliate page) | Best “just works” baseline. Not white-label. CRM via Zapier only. English + Spanish included on all plans, spam detection, iOS/Android apps, 10+ voices. Hits your ceiling exactly at $149. |
| Goodcall | Starter $79/mo/agent ($66 annual), Growth $129/mo ($108 annual), Scale $249/mo ($208 annual) (pricing) [VENDOR-PRIMARY] | Flows + forms; Zapier for CRM | “Performance viewer” role explicitly designed so agencies can give clients read-only access without exposing the build. No branded portal. | Best billing model in the category. Bills by unique customer (unique phone number that actually speaks), not minutes or tokens. Robocalls, blocked numbers, and silent hangups don’t count. Growth = 250 unique customers/mo, $0.50 each after. This structurally eliminates your seasonal-spike cost risk. |
| Sameday | Launch from $449/mo (500 min), Scale from $789/mo (1,000 min, voice cloning), Enterprise custom (pricing) [VENDOR-PRIMARY] | Native ServiceTitan / pro scheduling | Reseller/partner program (contact only) | 3x over your ceiling. Purpose-built for home services and genuinely good at it, but priced for $3M+ shops. Park it. |
| Avoca | Not published. Third-party estimates $1,000–$3,500/mo, per-minute billing. Raised $125M+ at $1B valuation April 2026, 800+ customers, targets $3M+ rev HVAC on ServiceTitan/HCP (PRNewswire, ServiceAgent) | ServiceTitan / Housecall Pro | No agency program found | Wrong customer. Out of range by 10x. |
| Smith.ai | Starter $285/mo (30 calls, $10.50 over), Basic $765/mo (90 calls, $9.50 over), Pro $1,950/mo (300 calls, $7.50 over). Add-ons per call: booking $1.50, SMS/Slack $0.50, recording+transcript $0.25, complex routing $1.50, extended intake $1.50, Spanish line $1.00. Extra transfer destination $15/mo (pricing) [VENDOR-PRIMARY] | Calendly/Acuity etc. at $1.50/call | No | Human-staffed, not AI-first. 3,000+ businesses, no charge for spam, 24/7. Useful as a premium/overflow tier you upsell, not the base offer. |
| Slang.ai | Core from $399/location, Premium from $599 (pricing) [VENDOR-PRIMARY] | OpenTable/SevenRooms/Yelp | Technology partner program | Restaurants only now. Their entire site is reservations. Remove from shortlist. |
| Numa | Not published | DMS-integrated | — | Automotive dealerships only. (Numa) Remove from shortlist. |
| Signpost | Could not verify current AI-receptionist product, pricing, or status for 2026. | — | — | [UNVERIFIED] — I found no current evidence of a competitive contractor AI receptionist offering. Do not plan around it. |
| Jobber AI Receptionist | $29/mo add-on (30 conversations, $0.79 each after) on select plans; $99/mo Grow-plan add-on (100 conversations); included unlimited on Plus plan. (Jobber help center, Jobber pricing) [VENDOR-PRIMARY, two prices published in different places — verify at signup] | Native Jobber | No agency program | Only relevant if the client already runs Jobber. Pulls services, pricing, availability straight from their account; ~30 min setup. Great, but it’s their tool, not yours — no agency leverage. |
| Housecall Pro CSR AI | Not published. Confirmed to exist as a paid add-on. HCP base plans $59–$329/mo (HCP pricing, CSR AI overview) | Native HCP | No | Same as Jobber — client’s tool, not yours. |
| ServiceTitan Contact Center Pro Voice Agents | Not published; enterprise (ServiceTitan) | Native ST | No | Wrong segment entirely. Note: ServiceTitan is a listed Vapi customer (Vapi pricing page logos) — interesting signal about what enterprise home-services builds on. |
| Allo (withallo.com) | Solo (annual only), Business $32/user/mo annual or $45 monthly, Ultra $100/user/mo annual or $120 monthly; 10DLC $24 one-time. AI Receptionist included on all paid plans. Partner/reseller program exists (pricing) [VENDOR-PRIMARY] | AI Appointment Booker; Google/Outlook | Partner program | ⚠️ This is a mismatch on your shortlist. Allo is an AI phone system — an Aircall alternative for sales teams (power dialer, CRM sync, call analytics), priced per user. It has an AI receptionist feature but is not a home-services answering product. Per-user pricing also breaks your model. Drop it. |
Path A verdict: Rosie Scale and Goodcall Growth are the only two that clear your ≤$150 ceiling and do the job. Neither is white-label. Both would work as a “we manage it for you” offering where your brand is on the service and theirs is on the portal you don’t show the client.
Path B: Agent Platforms (control + white-label)
B1. GoHighLevel native Voice AI — the structural favorite
All figures from HighLevel’s official AI Product Pricing doc (modified July 31, 2026) and Phone System Pricing & Billing Guide (modified July 20, 2026). [VENDOR-PRIMARY]
Three billing modes:
- Pay-per-use: Voice Engine $0.045/min (effective May 20, 2026) + TTS ($0.015–$0.170/min per table above) + LLM tokens + phone charges
- AI Employee Growth $50/mo per enabled location: 100 Voice AI minutes/mo (inbound + outbound + widget combined), 1,000 Conversation AI responses, unlimited Reviews AI + Content AI. Overages at pay-per-use.
- AI Employee Unlimited $97/mo per enabled location: unlimited Voice AI (inbound, outbound, and widget), unlimited Conversation AI, subject to fair use.
⚠️ Discrepancy flagged: an affiliate site claims Unlimited covers inbound only and that outbound/widget/prompt-optimizer are “AI Employee Plus” billed separately (TheStackInsiders). HighLevel’s own doc explicitly says Unlimited includes “inbound, outbound, and widget.” Trust the official doc, but confirm in-app before you quote a client — this exact line changed within the last quarter.
LLM rates (per 1M tokens): GPT-5 Nano $0.05 · Gemini 2.5 Flash Lite $0.10 · Gemini 2.0 Flash $0.10 · GPT-4.1 Nano $0.10 · GPT-4o Mini $0.15 · GPT-5 Mini $0.25 · Gemini 2.5 Flash $0.30 · GPT-4.1 Mini $0.40 · Claude 3.5 Haiku $0.80 · GPT-5 $1.25 · GPT-4.1 $2.00 · GPT-4o $2.50 · Claude 4.5 Sonnet $3.00
Telephony (LC Phone, matches Twilio rates): local number $1.15/mo · inbound to web/mobile/deskphone $0.01165/min · inbound via forwarding number $0.02/min · SMS $0.00747/segment in and out + carrier fees ($0.0025–$0.01/msg by carrier) · call recording $0.0025/min · recording storage $0.0005/min/mo · spam detection $0.005/test · name lookup $0.01/lookup. Calls billed in full minutes, rounded up. A fixed 5% markup applies to pass-through categories (A2P fees, carrier fees, verified caller ID) at the sub-account level regardless of rebilling.
Booking — the reason this path wins for you. GHL Voice AI has a native Appointment Booking action (official guide) that:
- Books on a single calendar or routes across multiple eligible calendars using AI intent-based selection (e.g. “I need an estimate” → Estimate Calendar; “my basement’s flooding” → Emergency Calendar)
- Honors calendar availability, buffers, minimum notice, and conflict settings
- Prompts for missing details (e.g. email) before confirming
- Supports configurable fallback behavior when no calendar matches
- Lets you control offering days, slots per day, and hours between slots
- Saves recordings/transcripts/summaries to Voice AI call logs, qualification answers to contact custom fields, appointment context to appointment notes via workflow
Known limitation: GHL “Services” is not supported in Voice AI booking yet — only Single Calendar and Multiple Calendars.
Agency/white-label economics: GHL Unlimited $297/mo (unlimited sub-accounts, rebill phone/email at no markup) or Agency Pro $497/mo (SaaS Mode, rebill with markup). You must be on the $497 plan to rebill AI Employee usage — but since you don’t need margin, on the $297 plan you simply absorb the $97/location (or pay-per-use) at agency level. Desktop web app white-labeling is included; white-label mobile app is $497/mo; Branded Client Portal App $49/mo per sub-account. (GHL pricing)
Honest weaknesses: slowest measured latency of the GHL-adjacent options per third-party claims (~1,000ms, [VENDOR-CONTENT] — not in the independent benchmark, which didn’t test GHL); the top community complaints (robotic voice, no caller recognition, no spam blocking) are real, though ElevenLabs TTS options and Number Intelligence spam detection at $0.005/test partially address two of the three.
B2. Retell AI
Official pricing page [VENDOR-PRIMARY] — the most transparent rate card in the category.
- Retell Voice Infra $0.055/min (platform fee, no monthly minimum, $10 free credits, 20 free concurrent calls)
- TTS: Retell Platform / Minimax / Fish / Cartesia / OpenAI $0.015/min; ElevenLabs $0.040/min
- LLM per minute: GPT-5 nano $0.003 · GPT-5 mini $0.012 · Gemini 2.5 Flash Lite $0.006 · GPT-4.1 mini $0.016 · Claude 4.5 haiku $0.025 · Gemini 3.0 Flash $0.027 · GPT-5 $0.040 · GPT-4.1 $0.045 · Claude 4.5/4.6 Sonnet $0.080 · GPT-5.5 $0.160 (Fast Tier ~2×)
- Speech-to-speech: GPT Realtime mini $0.07/min · GPT Realtime $0.345/min
- Telephony $0.015/min (Twilio/Telnyx US); $0 for BYO SIP
- Reliability add-ons that matter to you: Knowledge Base +$0.005/min · Advanced Denoising +$0.005/min · Safety Guardrails +$0.005/min · PII Removal +$0.01/min · AI Quality Assurance $0.10/min (first 100 min free)
- Subscriptions: phone number $2/mo · concurrency $8/concurrent/mo beyond 20 · knowledge base $8/mo beyond first 10 · verified phone number $10/mo
- Billing behavior: billed to the nearest second, no per-call rounding. AI fee stops on transfer, telephony continues. You are charged during silence/hold. Not charged for calls that fail to connect.
- Has a GoHighLevel integration page and a Solution Partner Program with a partner directory. No published white-label/multi-tenant reseller tier — enterprise starts around $8,000 per third-party reports [VENDOR-CONTENT].
B3. Vapi
Official pricing [VENDOR-PRIMARY]
- Build tier: $0.05/min Vapi hosting + model provider costs at cost (or $0 with your own API keys); SMS/chat $0.005/msg; 10 concurrent included, $10/line/mo beyond
- ⚠️ 14-day call history retention on Build. For a contractor who disputes what a customer said six weeks ago, this is disqualifying without your own archival pipeline.
- ⚠️ No SOC 2, SSO, or RBAC on Build. HIPAA +$2,000/mo, Zero Data Retention +$1,000/mo.
- Scale tier = annual contract, fixed platform fee + committed volume. Official GHL calendar tools exist (Get Contact, Create Contact, Check Availability, Create Event).
- Highest turn-discard rate in the independent latency benchmark (50/432).
B4. Bland — disqualified for your use case
Official pricing [VENDOR-PRIMARY]: Start $0.14/min ($0 platform fee, 10 concurrent, includes an inbound number); Build $0.12/min + $299/mo; Scale $0.11/min + $499/mo. Flat rate includes LLM + STT + TTS with no token charges — genuinely the cleanest pricing model. SOC 2 Type I/II, HIPAA-eligible with BAA, GDPR, PCI DSS.
But read the feature comparison table carefully. On Start, Build, and Scale, these are all marked “—” (Enterprise only):
- Warm Transfers
- Live Transfers
- Guardrails (Protected Calls)
- Alarm & Monitoring
- Appointment Scheduling Node
- SMS Node (in-call texts)
Transfer-to-owner, in-call booking, guardrails, and monitoring are four of your six hard requirements. Bland is out unless you go Enterprise.
B5. Synthflow — disqualified
Official pricing [VENDOR-PRIMARY]: “Enterprise contracts start at $30,000 annually.” They have removed self-serve tiers entirely and moved upmarket. That’s $2,500/mo before per-client usage. Out.
B6. ElevenLabs Agents
Official pricing [VENDOR-PRIMARY]: Free $0 (15 min, 4 concurrent) · Starter $6 (75 min, 6) · Creator $22 (275 min, 10) · Pro $99 (1,238 min, 20) · Scale $299 (3,738 min, 30, 3 seats) · Business $990 (12,375 min, 40, 10 seats). Additional call minutes $0.080/min; burst pricing $0.160/min; text messages $0.003/msg. LLM and telephony are external/at cost.
Best independent p95 in the latency benchmark (1,768ms, 1.24× tail ratio) and best-in-class voice. But there is no agency/multi-tenant/white-label program — you’d be running one workspace per client on consumer-grade plans with no client portal. Good component, bad platform for an agency.
B7. Phonely
Official pricing [VENDOR-PRIMARY]: Free $0 (100 min) · Starter $50/mo ($33 annual) · Professional $150/mo ($100 annual) · Enterprise “as low as 5¢/min”. SOC2 & GDPR on all tiers; HIPAA BAA + PCI on Enterprise. Post-transfer minutes $0.02/min. Unlimited concurrency on all tiers.
⚠️ Their own pricing page contradicts itself: the plan cards say Starter = 200 min and Pro = 650 min; the comparison table below says 250 and 750. Overage is listed as $0.25/min monthly but $0.35/min annual for Starter. Get a written quote before committing. That kind of sloppiness on a pricing page is itself a reliability signal.
B8. Sindarin — [UNVERIFIED]
I could not confirm current pricing, product status, or an agency program for Sindarin in 2026 searches. Do not plan around it without direct contact.
B9. Assistable.ai — the white-label option you’re missing
Official pricing [VENDOR-PRIMARY] — this wasn’t on your shortlist and it should be. It is purpose-built for GHL agencies.
Agency plans (14-day free trial, annual = 2 months free):
| Tier | Price | Sub-accounts |
|---|---|---|
| Startup | $225/mo | up to 3 |
| Growth | $450/mo | up to 10 |
| Agency | $975/mo | unlimited |
All tiers include: unlimited hybrid AI agents, white-label client portal, re-billing & reselling on all costs, 24/7 support, MCP/API/CLI/Python SDK, reporting & analytics, $10 free calling credit.
Usage on top, fully itemized:
- Voice orchestration $0.070/min
- LLM: Qwen 235B / Kimi K2.6 $0.020 · GPT-5 / 5.1 / 5.2 / 4.1 $0.040 · Claude Haiku 4.5 $0.040 · GPT-4o $0.050
- Telephony: Assistable $0.015/min, or $0 with your own SIP trunk
- Add-ons: PII removal $0.005/min · Knowledge base $0.010/query · Observation/QA $0.020 per observer per call
- Chat: $0.020/message
- Their calculator: $0.125/min all-in (orchestration + GPT-5.2 + Assistable telephony)
This is the only option I found that gives you a genuinely branded client portal, rebilling, and GHL-ecosystem integration at a price that works.
Also in this niche (less verified): MyAIFrontDesk white-label — wholesale starting at $500/month, unlimited client sub-accounts, custom domain, Stripe rebilling, per-client feature gating, branded login. Per-minute usage rates are not published. [VENDOR-PRIMARY for the $500 floor, UNVERIFIED for usage]
Path C: Full Build
Twilio ConversationRelay (Twilio handles STT + TTS + orchestration; you bring the LLM over a WebSocket):
- ConversationRelay: $0.07/minute of active AI agent session time — confirmed in Twilio’s own docs billing example: “Conversation Relay minutes: $.07 per minute × 5 minutes = $0.35”. Real-time transcription is included in that rate. [VENDOR-PRIMARY]
- Twilio US inbound local voice: $0.0085/min (Twilio Pricing API docs); toll-free inbound $0.022/min. US local number $1.15/mo.
- Conversation Intelligence language operators (e.g. call summarization): $0.0035/min — storing ConversationRelay transcripts in Conversation Intelligence is free; you only pay incrementally for operators.
- Your LLM cost on top (text tokens, not audio): roughly $0.01–$0.05/min depending on model.
Alternative: Twilio Media Streams + OpenAI Realtime (true speech-to-speech):
- gpt-realtime-2.1: $32/1M audio input tokens, $64/1M audio output tokens, cached audio input $0.40/1M. Mini model $10/$20. (multiple 2026 pricing analyses, TokenCost) [VENDOR-CONTENT — OpenAI’s own page not directly verified]
- Practical per-minute: $0.18–$0.46/min uncached; $0.05–$0.10/min with prompt caching enabled. Prompt caching is the whole ballgame here — a 4x cost difference.
What you gain: total control of the prompt, the fallback logic, the barge-in tuning, the archival, the data. No platform’s roadmap becomes your roadmap. Marginal cost of ~$25–$35/client/month.
What you lose: the entire operational layer. No admin UI for a non-technical VA to tune a client’s FAQ. No call review dashboard. No client portal. No status page. No vendor to call at 2am when the WebSocket drops. You become a 24/7 telephony operator for other people’s businesses. You also lose the GHL calendar integration you get for free on Path B — you’d rebuild it against GHL’s API.
Realistic effort: 3–6 weeks to a working inbound agent for a competent developer; 3–6 months to something you’d trust with a plumber’s emergency line at 11pm on a Saturday. The gap is entirely in edge cases: dropped WebSockets mid-call, DTMF handling, transfer failure fallbacks, calendar API rate limits, timezone bugs, concurrency spikes.
PART III — RELIABILITY ENGINEERING PLAYBOOK
This section applies regardless of which path you choose. It is the actual difference between a demo and a business.
1. Prompt scoping — write rules, not essays
Structure the system prompt in four blocks, not one wall of text:
- Persona — “You are Sarah, the receptionist for [Business]. Friendly, efficient, local.”
- Core rules — always / never. Never quote a price not in the knowledge base. Never promise a specific technician. Never commit to work outside the license scope. Never discuss competitors. Never guess at an address.
- Booking rules — which calendar, what must be collected before booking, what to do if the caller won’t give an email.
- Escalation triggers — enumerated, explicit.
A tight prompt under ~500 words with specific instructions outperforms a 2,000-word essay (TheStackInsiders) [VENDOR-CONTENT but consistent with general practice]. Everything factual goes in the knowledge base, not the prompt — that’s what makes it retrievable and auditable instead of hallucinable.
2. Never let the model free-form a booking
All state changes go through function calls with typed parameters. The model’s job is to fill a struct, not to write a calendar entry. On GHL that’s the native Appointment Booking action. On Retell/Vapi that’s a tool call to a webhook that validates and writes. The model should never produce a confirmation number, a price, or a date it computed itself — it reads back what the function returned.
Concretely: the agent says “let me check” → tool call returns three real slots → agent reads those three slots → caller picks → tool call books → tool returns confirmation → agent reads confirmation. If the tool errors, the agent says “I’m having trouble with the schedule — let me have someone call you back within the hour” and fires the fallback. It never invents a slot.
3. Hallucination prevention
- Ground everything in a knowledge base / RAG. Reported hallucination complaint rate drops sharply with grounding. On Retell that’s the Knowledge Base add-on (+$0.005/min); on GHL it’s the native knowledge base; on Assistable it’s $0.010/query.
- Explicit “I don’t know” behavior. “That’s a good question and I want to get it right — let me have [Owner] call you back.” Callers accept this. They do not accept a wrong price.
- Price policy per trade. For excavation/demo/roofing where every job is custom: the agent books a site visit, it does not quote. For plumbing/HVAC where there’s a fixed diagnostic fee: the agent quotes only the diagnostic fee, verbatim from the KB.
- License scope guard. SE Michigan contractors carry specific licenses. The agent must not commit to electrical work for a plumber. Put the license scope in the “never” block.
- Retell sells Safety Guardrails as an add-on at +$0.005/min; Bland gates “Guardrails (Protected Calls)” to Enterprise. If your platform doesn’t have them, you implement them in the prompt + tool validation.
4. Graceful fallback + transfer rules
Define these as hard triggers, not vibes:
| Trigger | Action |
|---|---|
| Caller says any variant of “person / human / real / manager” | Transfer immediately. No qualifying first. This is non-negotiable. |
| Emergency keyword (gas, flooding, no heat, sewage, electrical, collapse) | Skip qualification → warm transfer to on-call, or same-day slot + SMS alert to owner |
| 2 consecutive ASR failures (agent asks caller to repeat twice) | “I’m having trouble hearing you — let me have someone call you right back” → capture callback number → end |
| Caller frustration / raised voice / profanity | Transfer, or take message + immediate owner SMS |
| Topic outside knowledge base | Take message, don’t improvise |
| Call duration > 5 min without a booking | Offer transfer or callback |
| Tool call fails (calendar API error) | Take message + owner SMS with full detail. Never say “you’re booked” if the booking didn’t return success. |
Use warm transfers, not cold. “I’m connecting you with Mike now — I’ll let him know you’re calling about a driveway excavation quote in Sterling Heights so you don’t have to repeat yourself.” Rosie Scale includes warm transfers; Rosie Growth adds waterfall transfers (tries multiple numbers until someone answers) which is genuinely valuable for a 2-truck operation where the owner is in a hole.
Always have a terminal fallback that captures the lead. Every path — hallucination, tool failure, transfer failure, caller hangs up mid-booking — must end with name + number + what they wanted in the CRM. If you get that 100% of the time, the worst case is “the AI was a bit awkward but they called me back in 5 minutes.” That’s survivable. Losing the lead entirely is not.
5. Testing & simulation tooling — current names verified
| Tool | Status / pricing | Notes |
|---|---|---|
| Coval | Enterprise sales motion, no public pricing (coval.ai) | Autonomous-vehicle testing methodology applied to voice; CI/CD-oriented. Best for engineering-led teams. Also publishes independent-ish STT and TTS benchmarks. |
| Hamming AI | No public pricing, sales conversation required (hamming.ai) | Simulation, regression, and production QA/monitoring. |
| Cekura | Listed at $30/month with 7-day trial; voice testing consumes 5 credits/minute (cekura.ai) [VENDOR-PRIMARY-ish, credit cost not prominently disclosed] | Full-lifecycle voice QA: gibberish detection, interruption/latency tracking, red teaming. Cheapest entry point. |
| Bluejay, Maxim | Custom | Also active in this space. |
| Name appears retired. | I found no active voice-agent-QA product marketed as “Vocera AI” in 2026. (Vocera is a Stryker clinical-communications brand.) Cekura is the name to use. |
Built into platforms: Retell includes Simulation Testing on the free/PAYG tier and sells AI Quality Assurance at $0.10/min (first 100 minutes free). Assistable sells Observation/QA at $0.020 per observer per call. GHL has Voice AI performance reporting (call volume, duration, booking rate, transfer rate, caller sentiment) in Labs.
For your first 10 clients, you don’t need Coval or Hamming. Cekura at $30/mo or Retell’s built-in QA is plenty. What you need is discipline, not tooling.
6. Monitoring & call review workflow
A 4-layer evaluation model is the industry consensus (Hamming, Cekura):
- Infrastructure — audio quality, latency (target p95 under 800ms internally, knowing caller-experienced will be ~1.3–1.8s), component uptime
- Execution — did the tool calls fire, did the booking land, did the transfer connect
- User reaction — sentiment, interruptions, hangups, “let me talk to a person” rate
- Business outcome — booked rate, containment rate, lead capture rate
Your actual weekly workflow for the first 90 days:
- Days 1–14: listen to every single call. All of them. There will be 20–40. This is where you find the failure modes no test suite would have generated.
- Weeks 3–8: listen to every transferred call, every call over 4 minutes, every call with a hangup in the first 20 seconds, and a random 10% sample.
- Week 9+: weekly review of the same buckets + a monthly full-sample audit.
- Alert immediately on: booking tool failures, transfer failures, calls where the caller said “human” more than once, and any call where sentiment flags negative.
Track these five numbers per client, weekly: answer rate, containment rate (% resolved without transfer), booking rate (% of qualified calls that book), transfer rate, and lead-capture rate (% of all calls where name + number + intent landed in CRM). Lead-capture rate is the one that must be ~100%. The others can be mediocre and you still win.
7. Versioned prompt changes
- Never edit a live client’s prompt directly. Duplicate the agent, edit the copy, run your test suite against the copy, then swap.
- Keep the prompt, KB, and tool schema in a versioned file (a git repo or even a dated Google Doc). Every change gets a date, a one-line reason, and the metric you expected to move.
- Change one thing at a time. If you change the voice, the prompt, and the escalation rules in the same week, you will never know which one broke it.
- Regression-test after every model change. When GHL or Retell swaps the underlying LLM version — and they will, silently — your prompt’s behavior changes. Re-run your 12-scenario suite monthly regardless of whether you changed anything.
8. Agency QA gate before go-live
Run 20–30 test calls minimum covering the scenario list below before a single real customer touches it. Then a staged rollout:
- Week 1: after-hours and overflow only. Real callers, but a human safety net covers business hours. This is the single highest-leverage risk-reduction move available to you, and it’s free.
- Weeks 2–4: expand to overflow during peak. Confirm CRM writes and calendar bookings are landing every time.
- Month 2: full 24/7, if and only if the metrics hold.
Tell the client this is the plan up front. It reframes “we’re being careful” as professionalism rather than doubt, and it gives you two weeks of cover if something goes sideways.
PART IV — COST MODEL
Assumptions: 250 inbound AI minutes/month (~100 calls at 2.5 min avg), 1 local number, call recording on, ~200 SMS segments/mo for confirmations and summaries, one A2P Low Volume campaign.
Detailed per-client monthly economics
| Path | Platform / seat fee | Per-minute usage @250 min | Telephony + number | SMS + A2P | Total |
|---|---|---|---|---|---|
| GHL pay-per-use, Cartesia/OpenAI TTS, GPT-5 Mini | $0 | Engine $11.25 + TTS $3.75 + LLM ~$2.50–$7.50 | Number $1.15 + inbound ~$2.91 + rec. $0.63 | ~$1.49 + carrier ~$0.80 + A2P $1.50–$10.50 | $26–$40 |
| GHL pay-per-use, ElevenLabs V2.5 TTS | $0 | Engine $11.25 + TTS $8.75 + LLM ~$2.50–$7.50 | ~$4.69 | ~$4–$13 | $31–$45 |
| GHL pay-per-use, ElevenLabs V3 TTS | $0 | Engine $11.25 + TTS $42.50 + LLM ~$5 | ~$4.69 | ~$4–$13 | $67–$76 |
| GHL AI Employee Unlimited | $97 | included | ~$4.69 | ~$4–$13 | $106–$115 |
| GHL AI Employee Growth ($50, 100 min) | $50 | 150 min overage @~$0.06–$0.09 = $9–$14 | ~$4.69 | ~$4–$13 | $68–$82 |
| Assistable Growth @10 clients | $45 (=$450÷10) | $0.125 × 250 = $31.25 | included in $0.125 (number rental not published) | via GHL or Assistable | ~$76–$90 |
| Assistable Startup @3 clients | $75 (=$225÷3) | $31.25 | " | " | ~$106–$120 |
| Retell direct (Cartesia TTS + GPT-5 Mini + KB + denoising + guardrails) | $0 | Infra $13.75 + TTS $3.75 + LLM $3.00 + KB $1.25 + denoise $1.25 + guardrails $1.25 + telephony $3.75 = $28.00 | Number $2.00 | separate | ~$30–$38 |
| Vapi Build (BYO keys) | $0 | Vapi $12.50 + STT ~$1.10 + LLM ~$3 + Cartesia TTS $3.75 + Twilio $2.13 | Number $1.15 | separate | ~$24–$32 |
| Bland Start | $0 | $0.14 × 250 = $35.00 | included number | separate | ~$35–$45 ❌ missing transfers/booking |
| ElevenLabs Agents Creator (275 min incl.) | $22 | included; $0.080/min over | LLM + telephony at cost ~$5–$10 | separate | ~$30–$40 ❌ no agency structure |
| Phonely Starter | $50 ($33 annual) | 250 min may exceed included; overage $0.25/min | included | separate | ~$50–$75 |
| Twilio ConversationRelay build | $0 (+ hosting ~$10–25 shared) | CR $17.50 + inbound $2.13 + LLM ~$2.50–$7.50 + summary operator $0.88 | Number $1.15 | separate | ~$25–$35 marginal ❌ + months of eng |
| Goodcall Growth | $129 ($108 annual) | unlimited minutes; 250 unique customers incl. | included | separate | $108–$129 |
| Rosie Scale | $149 | 1,000 min incl. | included | +$50 if website texting | $149–$199 |
| Sameday Launch | $449 | 500 min incl. | included | included | $449 ❌ |
| Smith.ai Basic (90 calls) | $765 | $9.50/call over 90 | included | + $1.50/booking, $0.25/recording | $900+ ❌ |
Break-even: GHL pay-per-use vs. $97 Unlimited
| TTS | Effective $/min (engine + TTS + LLM) | Break-even vs $97 |
|---|---|---|
| Cartesia / OpenAI | ~$0.075–$0.085 | ~1,150–1,300 min/mo |
| ElevenLabs V2.5 | ~$0.095–$0.105 | ~925–1,020 min/mo |
| ElevenLabs V3 | ~$0.23–$0.24 | ~405–420 min/mo |
A typical contractor at 100–400 min/mo will not hit any of these. On the cheap TTS stack, pay-per-use is 2–3x cheaper than the flat $97. Switch a client to Unlimited only when they cross ~1,000 min/mo — which is a good problem, and by then they should be paying you more.
What agencies retail this for (sanity check on your bundle)
- $297–$497/mo is the reported standard for HVAC voice AI (Trillet) [VENDOR-CONTENT]
- Common tiers: Bronze $197 (200 min) / Silver $297 (500 min) / Gold $497 (unlimited) (Trillet) [VENDOR-CONTENT]
- GHL-focused agencies package AI receptionist at $200–$400/mo (TheStackInsiders) [VENDOR-CONTENT]
- General market range for AI receptionists: $29–$500+/mo, with most small businesses paying $99–$249/mo (multiple 2026 pricing guides) [VENDOR-CONTENT]
Your $1,000 setup + $500–$1,000/mo bundle is priced correctly. The receptionist is worth $250–$300 of that on its own at market rates, and it costs you $30–$120. The margin you’re “giving up” on the receptionist is more than covered by the fact that it makes the rest of the bundle sticky — a client will cancel SEO before they’ll cancel the thing that answers their phone.
PART V — COMPLIANCE
Call recording consent in Michigan
Michigan is effectively one-party consent for a participant to the call.
MCL 750.539c is written as an all-party-consent eavesdropping statute, but Michigan courts have recognized a participant exception since Sullivan v. Gray (1982): a party to a conversation is a participant, not an eavesdropper. The Sixth Circuit reaffirmed this in Fisher v. Perron, 30 F.4th 289 (6th Cir. 2022) — a case arising directly from recorded phone calls — dismissing all MCL 750.539c claims on participant-exception grounds. (Varnum LLP, Butzel Long, Recording Law)
Exposure if you get it wrong: MCL 750.539c violation is a felony punishable by up to 2 years and $2,000; civil liability includes actual damages, punitive damages up to $5,000, and attorney fees.
Practical guidance: your client’s business is a party to every inbound call, so recording is lawful in Michigan without announcing it. Announce it anyway. Reasons: (a) callers from Illinois, Pennsylvania, Florida, California and eight other all-party states dial Michigan contractors constantly and the safe rule is the strictest state on the call; (b) it costs you two seconds of a greeting; © it doubles as your AI disclosure hedge. Standard line: “Thanks for calling [Business] — this call may be recorded for quality. How can I help?”
AI / bot disclosure laws
Michigan has no enacted AI disclosure requirement. Senate Bill 760 (chatbot disclosure) passed the Michigan Senate 20-17 in May 2026 and moved to the House, where a Republican majority and federal-preemption pressure make its path uncertain. You face no Michigan compliance deadline today, but this could change within a legislative cycle. (STACK Cybersecurity — Livonia MI firm tracking this)
Other states matter because these laws apply based on where the caller is located, not where the business is. Enacted landscape as of mid-2026:
| State | Law | Effective | What it requires | Applies to a contractor’s receptionist? |
|---|---|---|---|---|
| California | SB 243 | Jan 1, 2026 | Companion-chatbot disclosure, minor protections, crisis protocols. $5,000/violation/day | Likely no — transactional customer service chatbots are expressly exempted. ⚠️ But AB 1609, which would extend disclosure to customer service chatbots, advanced out of committee spring 2026. Watch it. |
| Maine | LD 1727 | Sept 24, 2025 | Disclose when a reasonable consumer couldn’t tell it’s not human. Enforced under Unfair Trade Practices Act — AG enforcement AND private suits | Probably yes if a caller couldn’t tell. Broadest trigger currently in force. |
| Utah | AI Policy Act SB 149 / SB 226 | May 1, 2024 (amended May 7, 2025) | Disclose when a consumer clearly and unambiguously asks. For regulated occupations (state license required), disclose orally at the start of the interaction. Fine up to $2,500/violation. Sunsets July 1, 2027 | Possibly yes for licensed trades — plumbing, electrical, HVAC and mechanical contractors are licensed occupations. Utah callers are rare for you, but the “disclose if asked” rule is trivial to implement and should be your default everywhere. |
| Iowa | SF 2417 | July 1, 2026 | Disclosure, minor protections, frequent in-session AI reminders | Narrow/transactional likely excluded |
| Washington | HB 2225 | Jan 1, 2027 | Capability-based; companion chatbots | Likely excluded |
| Oregon | SB 1546 | Jan 1, 2027 | Private right of action, $1,000/violation statutory damages | Behavior-based definition excludes narrow customer service bots — but this is the model to watch |
| Idaho | S 1297 | July 1, 2027 | Conversational AI disclosure | Likely excluded |
| Connecticut | SB 5 (AIRT Act) | Jan 1, 2027 | Disclosure at session start + reminders every 3 hours | Likely excluded |
| Nebraska | LB 525 | July 1, 2027 | Disclosure; excludes apps “primarily designed for commercial use by business entities” and “narrow and discrete” topics | Likely excluded |
| Georgia | SB 540 | July 1, 2027 | Disclosure, minors, crisis protocols; no platform carve-out | Likely excluded |
Source for the table: STACK Cybersecurity state chatbot law guide (June 2026), cross-checked against Henson Legal and Davis Polk on Utah SB 226.
FCC / TCPA — and why inbound is your friend
- FCC Declaratory Ruling, February 8, 2024: AI-generated voices — including real-time conversational AI — are “artificial or prerecorded voice” under the TCPA. Consent requirements attach regardless of how human it sounds and regardless of whether an ATDS is used. (Wilson Sonsini)
- ⭐ The TCPA’s restrictions do not extend to technologies used to answer inbound calls. (Henson Legal) Your entire use case is inbound. This removes the single largest legal risk in voice AI. Guard it — the moment you add outbound AI calling for a client, you inherit the whole consent regime, $500–$1,500 per call, no cap.
- FCC September 2024 NPRM (not finalized as of mid-2026) proposes: AI disclosure at the start of every AI-generated call, disclosure at consent capture, and an automated opt-out mechanism within two seconds. Aimed at outbound, but it’s where the wind is blowing.
- Fifth Circuit, Bradford v. Sovereign Pest Control of Texas, Feb 25, 2026: TCPA text requires only “prior express consent,” not “prior express written consent,” for artificial-voice calls. Binds TX/LA/MS only — the other 47 states, including Michigan, still apply the FCC’s written-consent rule.
A2P 10DLC for your SMS follow-ups
Missed-call textback and SMS summaries mean you’re sending A2P messages and must register. HighLevel rates (official A2P/LC Phone guide) [VENDOR-PRIMARY]:
| Brand type | Daily segment limit | Monthly campaign fee | One-time registration |
|---|---|---|---|
| Sole Proprietor | 3,000 | up to $2.10/campaign | up to $23.475 (incl. $3 Fast Track) |
| Low Volume | 600,000 | $1.50–$10.50/campaign | up to $23.475 |
| High Volume | 600,000 | $10.50/campaign | $68.625 |
Notes: submitting the campaign starts both fees regardless of review outcome; the $3 Fast Track fee is non-refundable; A2P resubmissions became free February 1, 2026 (the old $15 resubmission fee is gone). Carrier surcharges apply on top of message rates ($0.0025–$0.01/msg depending on carrier). A fixed 5% markup applies to A2P and carrier fees at the sub-account level whether or not you have rebilling on. Twilio direct is comparable: $4 one-time US sole-proprietor brand registration, $15 campaign vetting fee.
Most contractors register as Low Volume Standard. Budget ~$25 one-time + $2–$11/mo per client and build it into your setup fee.
Practical “disclose or not” — what I’d actually do
Disclose softly, always. Don’t say the word “AI” in the greeting.
Recommended opening: “Thanks for calling Patriot Excavating, this is Sarah — this call may be recorded. How can I help you today?”
Then in the rules block: if the caller asks in any form whether they’re speaking to a person, answer truthfully and immediately, then offer a human. “I’m an automated assistant — I can get your details and book you in, or I can connect you with Mike right now. Which would you rather?”
Why this is the right posture:
- It satisfies Utah’s “disclose when asked” rule and Maine’s “reasonable consumer” trigger without a robotic disclaimer that tanks your containment rate.
- It handles the recording consent question in one clause.
- It’s honest, which matters more than compliance — the reputational downside of a caller feeling deceived is worse than any current fine.
- It leaves you a clean upgrade path if Michigan SB 760 or California AB 1609 passes: you change one line in a prompt, not your whole product.
What I would not do: lead with “You are speaking with an AI assistant.” The independent survey data says 31% would hang up if connected to AI. Leading with it converts a soft signal into a hard one and costs your client leads for no legal benefit in Michigan today.
PART VI — RECOMMENDATION
What to do this week
Run a 3-way bake-off. Budget: under $500. Timeline: 10 days.
| # | Contender | Cost to trial | Why it’s in |
|---|---|---|---|
| 1 | GoHighLevel native Voice AI (ElevenLabs V2.5 voice, GPT-5 Mini, native Appointment Booking action, knowledge base) | Already paying for GHL; pay-per-use ~$10–20 in test minutes | Zero integration surface. Books natively into GHL calendars, writes natively to GHL CRM, triggers GHL workflows for SMS. One vendor, one bill, one support path. Cheapest at your volume. If this passes, everything else is a distraction. |
| 2 | Assistable.ai Startup — $225/mo, 3 sub-accounts, 14-day free trial | $0 during trial | The white-label insurance policy. Branded client portal, rebilling, GHL-ecosystem native, fully itemized $0.125/min. If GHL native isn’t good enough, this is the answer at ~$76/client at 10 clients. |
| 3 | Rosie Scale — $149/mo, 7-day free trial | $0 during trial | The known-good benchmark. 3.1M+ calls handled, direct calendar booking, warm + live transfers, spam detection, English/Spanish, mobile apps. You are not shipping this necessarily — you are using it to calibrate “what good sounds like.” If GHL native can’t match Rosie on the test script, you have your answer. |
Optional 4th if you want a minutes-risk-free fallback: Goodcall Growth at $129/mo — the only vendor billing by unique customer rather than minutes, which structurally immunizes you against a client’s seasonal spike. 14-day free trial.
Explicitly excluded and why: Bland (transfers/booking/guardrails are Enterprise-only), Synthflow ($30k/yr minimum), Sameday ($449 floor), Avoca (~$1k–$3.5k), Smith.ai (human-staffed, $765+ at your volume), Slang.ai (restaurants only), Numa (auto dealers only), Allo (per-user sales phone system, not a receptionist product), Signpost (unverifiable), ElevenLabs Agents (no agency structure), Vapi (14-day retention, no SOC 2 on Build).
The test script — 14 scenarios, run on every contender
Build one agent profile per contender for a fictional excavation/demo company (matches your hardest vertical: custom pricing, no fixed rates, safety-critical). Then call each one with these, from a cell phone, at least half from a noisy environment:
| # | Scenario | PASS criteria |
|---|---|---|
| 1 | Clean happy path. “I need a quote on removing a concrete driveway in Sterling Heights.” | Captures name, phone, address/city, job type, timeline. Books a site visit. Booking appears in GHL calendar within 60s. Contact created with all fields. |
| 2 | Job-site noise. Same as #1, with a leaf blower / equipment audio playing 3 feet away. | Completes the booking, or cleanly says “I’m having trouble hearing you” and captures a callback number. Fails if it transcribes garbage and books the wrong thing. |
| 3 | Heavy accent. Have someone with a strong non-native or regional accent run scenario #1. | Same as #1. This is the one most agencies skip and most often fails in production. |
| 4 | Barge-in. Interrupt the greeting mid-sentence with “yeah I need someone out today.” | Stops cleanly, picks up the intent, doesn’t restart the greeting. |
| 5 | Price pressure. “Just give me a ballpark, what’s it gonna cost?” — ask three times, escalating. | Never quotes a number. Redirects to a site visit each time without sounding like a broken record. |
| 6 | Emergency. “There’s water pouring into my basement, the line’s broken.” | Skips qualification, warm-transfers or offers same-day, fires an SMS alert to the owner. Time-to-transfer under 20 seconds. |
| 7 | “Get me a human.” Say it at second 8. | Transfers immediately. Does not ask one more qualifying question. Any hesitation = FAIL. |
| 8 | Angry caller. “This is the third time I’ve called, nobody’s called me back.” | Doesn’t argue, doesn’t over-apologize in a loop, captures the issue, escalates. |
| 9 | Spam / robocall. Call from a known spam-flagged number, or play a robocall recording into it. | Detects and terminates without creating a CRM contact or burning 3 minutes. |
| 10 | Complex scheduling. “I can do Tuesday or Thursday but not before 10, and not the week of the 15th.” | Handles the constraint correctly or gracefully hands to a human. Fails if it books Tuesday at 8am. |
| 11 | Out-of-scope work. “Can you also rewire my panel while you’re there?” | Says it’s outside their license/scope, doesn’t commit, offers to note it for the owner. |
| 12 | Caller won’t give email. “I don’t do email.” | Books anyway with phone-only, or clearly explains and offers alternative. Fails if it loops. |
| 13 | Rambler. Talk for 90 seconds about your neighbor’s dog before mentioning you need a stump removed. | Stays patient, extracts the intent, doesn’t cut off mid-story. |
| 14 | Tool failure simulation. Point the booking action at a calendar with zero availability. | Says something true (“I don’t have anything open this week — let me have Mike call you”), captures the lead, alerts the owner. Catastrophic fail if it says “you’re all set.” |
Scoring — the pass/fail gate
Score each contender 0/1 on each scenario. Then apply these gates:
Hard gates (any single failure = eliminated):
- Scenario 7 (immediate human transfer) — must pass
- Scenario 14 (never claims a booking that didn’t happen) — must pass
- Scenario 5 (never invents a price) — must pass
- Lead capture on all 14 scenarios — every call must end with name + number + intent in the CRM, including the ones where the agent “failed.” This is the floor.
Quality gates:
- ≥12/14 overall
- Scenarios 2 and 3 (noise, accent) — at least one clean pass, and no garbage-booking on either
- Booking-to-calendar latency under 60 seconds, with the appointment landing on the GHL calendar (or via a bridge you’ve verified twice)
Tiebreakers, in order: (1) how it sounds on scenario 2, (2) total cost at 400 min/mo, (3) how fast a non-technical person could tune the FAQ, (4) whether the client-facing portal carries your brand.
Also do the thing you actually care about: record all three demo lines and play them for one real contractor you trust. Ask which one he’d let answer his phone. His gut is worth more than your scorecard.
What to standardize on for the first 10 clients
Primary: GoHighLevel native Voice AI, pay-per-use, ElevenLabs V2.5 voice.
Rationale:
- ~$31–$45/client/month, all-in. Roughly 4% of a $1,000/mo bundle, 7% of a $500 bundle. You have enormous headroom, which means you can afford to over-engineer the reliability side.
- Zero integration surface. Native calendar booking with intent-based multi-calendar routing, native CRM writes to contact custom fields, native workflow triggers for SMS summary + missed-call textback. Every competing option needs a Zapier or webhook bridge, and bridges are where bookings go to die.
- One vendor, one invoice, one support ticket. When something breaks at 9pm you make one call, not three.
- White-label is already handled — GHL’s desktop web app white-labeling is included in your plan, so the client logs into your brand to see call logs and recordings.
- Prompt/KB/calendar config is portable enough that if GHL disappoints in month 6, you migrate to Assistable in a week, not a quarter.
Backup, provisioned before you need it: Rosie Scale ($149/mo) as a per-client escape hatch. For any client where GHL native fails the bake-off criteria in production, or who specifically wants a mobile app, or who’s bilingual-heavy — move that one client to Rosie. You still make money; you just don’t make as much. Having a proven fallback you’ve already tested is worth more than squeezing $100/mo.
If GHL native fails the bake-off outright: Assistable.ai Growth ($450/mo for 10 sub-accounts + $0.125/min ≈ $76/client). Branded portal, rebilling, GHL-ecosystem native. Still comfortably under your ceiling.
Standardize these five things across all 10 clients — this is what makes it scalable:
- One master agent template with the four-block prompt structure, cloned per client. Client-specific facts live in the knowledge base only.
- Two calendars per client: “Estimate / Site Visit” and “Emergency Same-Day.” Let the intent router pick.
- A fixed escalation table (the one in the Reliability Playbook), identical for every client, with only the transfer number swapped.
- A standard post-call workflow: SMS confirmation to the caller → SMS summary to the owner → CRM contact + opportunity created → tagged by intent → missed-call textback if the caller hung up before booking.
- The same 14-scenario test suite, run at go-live and monthly thereafter.
Roll out after-hours-only for the first two weeks on every single client. No exceptions. It’s free risk reduction and it’s the best possible answer to “what if it screws up with my customers.”
When building on Retell / Vapi / ConversationRelay makes sense
Not before 40 clients. Realistically, never — for this business.
Do the arithmetic honestly. Building on Retell saves roughly $10–$20/client/month versus GHL pay-per-use, and roughly $40–$50 versus Assistable.
- At 10 clients: saves $100–$500/mo. Costs you weeks of build plus permanent on-call. Obviously wrong.
- At 25 clients: saves $250–$1,250/mo. Less than a part-time engineer’s monthly cost. Still wrong.
- At 50 clients: saves $500–$2,500/mo. Now it’s a real number — roughly a half-time technical hire. Arguable, if and only if you already have that person.
- At 100+ clients: saves $1,000–$5,000/mo and the operational leverage of controlling your own stack starts to matter. Now it’s a real decision.
The volume threshold is not the real trigger, though. Build only when one of these is true:
- A capability you can’t buy becomes a sales differentiator. E.g. you discover that job-site noise handling is what wins deals in excavation, and Retell’s Advanced Denoising plus custom VAD tuning measurably beats everything off-the-shelf. That’s a moat. $30/mo of savings is not.
- A vendor materially breaks your business — GHL changes Voice AI pricing 3x, or an acquisition kills the product. Then you build because you must, and you’ll be glad you kept the prompt and KB in version control.
- You have a dedicated technical person who is not you. The hidden cost of Path C is not the build — it’s being the 24/7 on-call for ten contractors’ phone lines. That job has to belong to someone.
If and when you do build: use Twilio ConversationRelay, not raw Media Streams + Realtime. ConversationRelay at $0.07/min gives you Twilio-grade telephony, STT, TTS, and orchestration with real-time transcription included, and you bring only the LLM over a WebSocket. That’s 80% of the control at 20% of the engineering. And you keep Twilio’s status page, SLA, and support — which is the thing you’re actually paying for.
Middle path worth remembering: if you outgrow GHL native but don’t want to build, Assistable at $975/mo unlimited sub-accounts is $19.50/client at 50 clients plus $0.125/min usage. That’s within a rounding error of a self-built stack, with none of the on-call. For an agency your size, that’s almost certainly where this story ends.
APPENDIX — Confidence and gaps
High confidence (vendor’s own pricing/docs page, verified August 2026): Rosie, Goodcall, Sameday, Smith.ai, Slang.ai, Bland, Vapi, Retell, Synthflow, ElevenLabs Agents, Phonely, Assistable, Allo, GoHighLevel (platform, AI, and LC Phone pricing), Twilio ConversationRelay and inbound voice rates, Jobber (with the noted $29/$99 discrepancy).
Medium confidence (independent measurement, single source): the openbenchmarks latency table. It publishes open data and code, discloses its method and limits, and is the only caller-side measurement I found. It does not include GoHighLevel.
Low confidence — treat as directional only: all ROI/lift statistics (answer rate 60→99%, booking 32→49%, containment 70–90%, hallucination rate 0.34%), all “agencies retail at $X” figures, the GHL community complaint upvote counts, and the “GHL native latency ~1,000ms” claim. Every one of these traces to a company selling something in the category.
Explicit gaps I could not close:
- Reddit is inaccessible to my tools (blocked for both search and fetch). Several vendor blogs cite specific r/sweatystartup posts that I could not verify and have therefore not repeated as fact. You should go read r/sweatystartup, r/GoHighLevel, and r/HVAC yourself — it’s 30 minutes and it’s the highest-value verification left on this list.
- Housecall Pro CSR AI pricing is not published anywhere I could find, including their own site.
- ServiceTitan Contact Center Pro Voice Agent pricing — not published.
- Avoca pricing — not published; the $1,000–$3,500 range is third-party estimate only.
- Rosie’s agency program terms — not published; inquiry form only. Worth a direct email given it’s your #1 fallback.
- Sindarin — could not verify current status, product, or pricing.
- Signpost — could not verify a current contractor AI receptionist offering.
- MyAIFrontDesk per-minute usage rates — only the $500/mo wholesale floor is published.
- GHL AI Employee Unlimited inbound-vs-outbound coverage — official doc says all three (inbound/outbound/widget); an affiliate source says inbound only. Confirm in-app before quoting.