The starting point: voice as a tool, not a project
audioMCP.ai exposes voice-agent creation as MCP tools, so the AI agent itself handles provisioning, configuration, and deployment. You own the conversation, the prompts, and the business logic. audioMCP.ai runs the telephony, speech-to-text, text-to-speech, turn-taking, barge-in handling, latency management, and scaling. That boundary is fixed and clear.
Provision the voice agent
Your AI agent makes a single MCP tool call to request a voice agent. audioMCP.ai allocates telephony and speech infrastructure immediately, without a separate account setup, vendor portal, or manual provisioning step. The agent gets infrastructure on demand.
Configure voice, prompt, tools, and routing
With infrastructure in place, your AI agent sets the voice agent's behavior through another MCP tool call. Voice selection, system prompt, tool access, and call routing are all set in plain language or structured parameters, directly by the agent. There is no separate dashboard configuration step.
Deploy to a number, channel, or workflow trigger
One more MCP tool call takes the voice agent from configured to live. The agent can be deployed on a phone number for inbound support, triggered outbound for qualification calls or reminders, or embedded inside your own product as a scoped voice assistant. Workflow-triggered calls require no human in the loop: an event fires, and the call is placed or answered automatically.
Observe: receive transcripts, outcomes, and call metadata
After every call, transcripts, outcomes, and call metadata stream back to your agent as structured data. Your agent can read those results and act on them, closing the feedback loop without any manual review or export step. The calling agent gets information it can use, not just a record of a phone call.
What does the AI agent handle mid-conversation?
During a live call, the voice agent handles sub-second turn-taking and barge-in, so callers can interrupt mid-sentence without the conversation breaking. The voice agent can also invoke external tools, such as a calendar or CRM lookup, live during the call. Streaming STT and TTS run in real time rather than in batch, keeping latency low enough for natural conversation.
What does my team need to do to integrate?
If your stack already uses an MCP-compatible client or agent framework, audioMCP.ai drops in as a first-class tool with low integration lift. You do not need to learn a new protocol. If your stack does not currently use MCP-compatible tooling, audioMCP.ai is designed specifically for that ecosystem and may not be the right fit today.
How does pricing work?
audioMCP.ai uses usage-based, per-minute billing. There are three tiers. The Developer tier provides API key access, MCP tool setup, and a sandbox demo flow, structured for building and testing. The Scale tier covers production calls and includes telephony, STT, TTS, and infrastructure with no hidden fees, described as an all-inclusive pilot rate. The Enterprise tier offers custom volume pricing, SLA review, security review, and dedicated support. There are no seat fees at any tier.
Start building
The Developer tier is the starting point. Get your API key, connect audioMCP.ai to your agent as an MCP tool, and run the four-call sequence in a sandbox environment. For production scale or enterprise requirements, including SLA and security review, you can talk to an engineer.