Real-time by design
Streaming in both directions, over HTTP or WebSocket. Transcripts keep pace with the speaker and synthesized speech starts playing almost immediately, so conversations never stall.
Low-latency speech-to-text and text-to-speech APIs with transcripts you can rely on, proven in clinical, contact-center and voice-agent workflows.
Hi, I'm Sofia. I can confirm your appointment, answer billing questions, or connect you with a nurse.
Unedited output from the API. Pick a voice and press play.
Why xVoice
When a transcript goes into a patient record or a voice agent is on the line with a customer, speed and accuracy are not features. They are the job.
Streaming in both directions, over HTTP or WebSocket. Transcripts keep pace with the speaker and synthesized speech starts playing almost immediately, so conversations never stall.
Drug names, dosages, lab values and clinical findings transcribed correctly. Healthcare-grade accuracy that holds up on real calls, not just clean demo audio.
Proven in real enterprise workflows, with the controls operations teams ask for: traceable requests, predictable limits, spend caps and a full audit trail.
Workflows
xVoice runs where errors are expensive and delays are noticed: in clinics, contact centers and customer-facing voice agents.
Ambient scribing and dictation that turn visits into accurate notes, with medications, dosages and findings captured correctly.
Live agent assist while the customer is still talking, and quality review on every call instead of a sample.
Scheduling, intake and support agents that hear the caller and answer without the pauses that make people hang up.
Searchable transcripts of recorded calls and consultations for audits, disputes and regulatory review.
Products
Transcribe live or recorded audio, and generate speech that streams as it is produced. One key, one bill, one set of logs.
"Your order shipped this morning."
POST /v1/audio/speech
Clear, natural speech returned as a whole file or streamed while it is generated, so playback starts right away.
Speech docsThanks for calling. I can see the refund went through on Monday,
so it should reach your account within three days.
POST /v1/audio/transcriptions
Upload a recording, get an accurate transcript back. Streamed as it is produced if you want it, billed by the second of audio.
Transcription docsSo the plan is to ship the beta on Friday.
WS /v1/realtime
Transcribe calls, visits and browser microphones while the audio is still being captured, over a single WebSocket session.
Realtime docsSpeech to speech
One model that listens and answers in voice, without a transcription step in between. For the most natural, lowest-latency conversations.
Get early accessSecurity and compliance
Healthcare and enterprise teams can put xVoice into production without a compliance detour.
Process protected health information through xVoice. We sign a Business Associate Agreement with covered entities and their partners.
BAA available
Our security controls are independently audited. The report is available to customers and prospects under NDA.
Developer experience
The xVoice API follows the OpenAI audio API, so there is no new SDK to learn. Already on OpenAI? Change the base URL and the key, and your code runs on xVoice.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.openai.com/v1",
api_key=os.environ["OPENAI_API_KEY"],
base_url="https://api.xvoice.dev/v1",
api_key=os.environ["XVOICE_API_KEY"],
)
with open("call.wav", "rb") as audio:
transcript = client.audio.transcriptions.create(
model="xvoice-stt1",
file=audio,
)
print(transcript.text)import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.openai.com/v1",
apiKey: process.env.OPENAI_API_KEY,
baseURL: "https://api.xvoice.dev/v1",
apiKey: process.env.XVOICE_API_KEY,
});
const speech = await client.audio.speech.create({
model: "xvoice-tts2",
voice: "sofia",
input: "Hello from xVoice.",
});
await fs.promises.writeFile("hello.mp3", Buffer.from(await speech.arrayBuffer()));curl https://api.xvoice.dev/v1/audio/speech \
-H "Authorization: Bearer $XVOICE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "xvoice-tts1",
"voice": "neutral-female",
"input": "Hello from xVoice.",
"response_format": "mp3"
}' \
-o hello.mp3Platform
The dashboard your team needs on day one: keys, logs, usage and spend, all scoped by organization and project.
Separate keys per app and environment, test or live. Revoke one without touching the rest.
Every call has a request id that ties an invoice line to a row in your logs, errors included.
Seconds of audio and characters of text, grouped by project, model or key, over any range.
Pay as you go from a balance. Auto-recharge keeps it topped up; nothing bills you by surprise.
Cap what a month can cost. Rate limit headers on every metered call tell you where you stand.
Invite people by email with roles. Every change to keys, members or spending is recorded.
Requests
clinic-intake · live
Requests
12,480
Audio
41.2 h
Characters
2.9M
| POST | /v1/audio/speech | xvoice-tts2 | 200 | 312 ms |
| WS | /v1/realtime | xvoice-stt1 | 101 | 14m 02s |
| POST | /v1/audio/transcriptions | xvoice-stt1 | 200 | 1.84 s |
| POST | /v1/audio/speech | xvoice-tts1 | 200 | 688 ms |
| POST | /v1/audio/speech | xvoice-tts1 | 400 | 9 ms |
Pricing
Prepaid credits, metered from our own measurements: seconds of audio in, characters of text out.
xvoice-stt1$0.36per hour of audio
Files and realtime sessions, metered per second.
xvoice-tts1$15per 1M characters
Five audio formats, adjustable speed.
xvoice-tts2$10per 1M characters
Lowest time to first audio.
Quickstart
No SDK of ours to learn, no new concepts to read about before the first response.
Sign up with email or Google. An organization and a first project come with the account.
app.xvoice.dev/signupVerify your email for free starting credit, then create a test key. It is shown once, so store it as a secret.
export XVOICE_API_KEY=sk_test_…Install the openai client you would use anyway, set the base URL, and send audio.
pip install openaiFAQ
Yes. You can process protected health information through xVoice, and we sign a Business Associate Agreement with covered entities and their business associates.
Yes, SOC 2 Type II. The report is available to customers and prospects under NDA.
Built for live conversation. Realtime sessions return partial transcripts while the speaker is still talking, and text-to-speech streams audio as it is generated, so playback starts before the whole response exists.
Transcription is built to get drug names, dosages, lab values and numbers right, on real calls rather than clean studio audio. The free starting credit is enough to run your own recordings through it before you commit.
No. The API follows the OpenAI audio API, so the official Python and Node clients work unchanged. Point them at the xVoice base URL with an xVoice key.
Prepaid credit, deducted per second of audio transcribed and per character of speech generated. Verifying your email adds $5 of free credit. Top-ups run from $5 to $1,000, auto-recharge is optional, and a monthly spending cap protects you from runaway usage.
Limits are per organization, covering requests per minute, concurrent requests, characters and audio seconds. Every metered response carries headers showing where you stand, and limits can be raised for production workloads.
Soon. They listen and answer in voice in a single model, with no transcription step in between. Sign up to get early access when they launch.
Still have questions? Read the docs →
Start with free credit and make your first call in minutes, or talk to us about a BAA and volume pricing.