Skip to main content

What This Example Shows

  • How to treat voice as I/O — Vapi runs the realtime audio loop, Swarms runs the brain
  • A FastAPI webhook that receives caller transcripts and fires a MultiAgentRouter swarm
  • Four specialist voice agents (Booking, Billing, Emergency Triage, FAQ) tuned for spoken replies
  • An end-of-call summary sub-agent that writes a structured record to Airtable or HubSpot
  • A real per-turn latency budget and per-call cost breakdown you can quote to a customer
Vapi, Retell, and 11Labs handle the realtime audio loop — STT, TTS, barge-in detection, and the WebRTC plumbing. Swarms handles the brains: routing, specialist reasoning, structured output, and CRM logging. This pattern works with any voice infrastructure that supports a webhook on user turn.

Why This Matters

A single voice receptionist costs around $35K/yr fully loaded and works 40 hours a week. A voice agent stack runs 24/7 for roughly $0.30 per call. Most voice startups shipping today wire a single LLM call to each turn and call it done — that hits a ceiling fast because one prompt can’t be specialist, structured, and fast all at once. A swarm gives you proper intent routing, parallel reasoning when you need it, and CRM logging on the same turn budget. You ship a real product, not a chatbot with a phone number.

The Architecture

Step 1: Setup

Vapi-side configuration (creating the assistant, attaching a phone number, picking a voice) is out of scope for this tutorial — follow the Vapi docs at vapi.ai to get an assistant running with a webhook URL pointing at your FastAPI server.

Step 2: Configure the Vapi Webhook

In your Vapi assistant config, point the serverUrl at your FastAPI endpoint. Vapi will POST a payload on every user turn and again at end-of-call. The exact schema is documented by Vapi — what matters for this pattern is that the payload carries the caller’s transcript, a stable call_id, and the caller’s phone number.
The Vapi assistant’s own model is intentionally minimal — it just keeps the audio loop alive. The actual reasoning happens in your FastAPI handler when Vapi calls back with the user transcript. This is the pattern voice-AI builders use when they want full control over routing and tools.

Step 3: Define the Specialist Agents

These agents are tuned for speech, not text. Short sentences. No bullet lists. No markdown. The caller is on a phone — they can’t see your formatting.

Step 4: The MultiAgentRouter Endpoint

This is the per-turn handler. Vapi posts the running transcript on every user turn; we send it through the router and return the reply in the format Vapi expects.
The timeout=20 matters. If Swarms takes longer than the turn budget, Vapi will stall the caller — pick fast models (gpt-4.1-mini, claude-haiku-4.5) for routing and only escalate to a larger model on flagged calls.

Step 5: End-of-Call Summary → CRM

When the caller hangs up, Vapi sends an end-of-call-report event with the full transcript, duration, and any tool calls made during the call. Route that to a single summary agent and POST the result to your CRM.
For HubSpot, swap write_to_airtable for a POST to https://api.hubapi.com/crm/v3/objects/calls with an OAuth bearer — the structured-summary shape is the same.

Latency Budget

A real-feeling voice call needs round-trip turn latency under ~1.5 seconds. Here is where the time goes: Keep the router on gpt-4.1-mini or claude-haiku-4.5. Only escalate to gpt-4.1 or claude-sonnet-4.5 on flagged calls (emergencies, high-value accounts). Cap max_tokens aggressively — voice replies are short anyway, and a 200-token cap shaves real milliseconds off the response.

Real Cost

A representative five-turn call (caller books an appointment): Now compare: You are not eliminating the human — you are taking the after-hours, overflow, and routine-routing load off the front desk so they can do the work that actually needs a human voice.

Next Steps