Speech-to-speech models coming soon

Give your product a voice.

Low-latency speech-to-text and text-to-speech APIs with transcripts you can rely on, proven in clinical, contact-center and voice-agent workflows.

  • HIPAA compliant
  • SOC 2 Type II
  • BAA available

Hi, I'm Sofia. I can confirm your appointment, answer billing questions, or connect you with a nurse.

Conversational · xvoice-tts2

Unedited output from the API. Pick a voice and press play.

Why xVoice

Built for conversations that matter

When a transcript goes into a patient record or a voice agent is on the line with a customer, speed and accuracy are not features. They are the job.

Real-time by design

Streaming in both directions, over HTTP or WebSocket. Transcripts keep pace with the speaker and synthesized speech starts playing almost immediately, so conversations never stall.

Accurate where it counts

Drug names, dosages, lab values and clinical findings transcribed correctly. Healthcare-grade accuracy that holds up on real calls, not just clean demo audio.

Production-grade from day one

Proven in real enterprise workflows, with the controls operations teams ask for: traceable requests, predictable limits, spend caps and a full audit trail.

Workflows

Proven in real-world enterprise workflows

xVoice runs where errors are expensive and delays are noticed: in clinics, contact centers and customer-facing voice agents.

Clinical documentation

Ambient scribing and dictation that turn visits into accurate notes, with medications, dosages and findings captured correctly.

  • Realtime transcription
  • Transcription

Contact centers

Live agent assist while the customer is still talking, and quality review on every call instead of a sample.

  • Realtime transcription
  • Transcription

Voice agents

Scheduling, intake and support agents that hear the caller and answer without the pauses that make people hang up.

  • Realtime transcription
  • Text to speech

Compliance and review

Searchable transcripts of recorded calls and consultations for audits, disputes and regulatory review.

  • Transcription

Products

Every direction voice travels

Transcribe live or recorded audio, and generate speech that streams as it is produced. One key, one bill, one set of logs.

"Your order shipped this morning."

POST /v1/audio/speech

Text to speech

Clear, natural speech returned as a whole file or streamed while it is generated, so playback starts right away.

Speech docs
support-call.wav04:12

Thanks for calling. I can see the refund went through on Monday,

so it should reach your account within three days.

POST /v1/audio/transcriptions

Transcription

Upload a recording, get an accurate transcript back. Streamed as it is produced if you want it, billed by the second of audio.

Transcription docs
live session

So the plan is to ship the beta on Friday.

WS /v1/realtime

Realtime transcription

Transcribe calls, visits and browser microphones while the audio is still being captured, over a single WebSocket session.

Realtime docs

Speech to speech

Speech-to-speech models

One model that listens and answers in voice, without a transcription step in between. For the most natural, lowest-latency conversations.

Get early access

Security and compliance

Ready for regulated data

Healthcare and enterprise teams can put xVoice into production without a compliance detour.

HIPAA compliant

Process protected health information through xVoice. We sign a Business Associate Agreement with covered entities and their partners.

BAA available

SOC 2 Type II

Our security controls are independently audited. The report is available to customers and prospects under NDA.

Controls built into the platform

  • Role-based access for owners, admins and members
  • Audit log of every key, member and spending change
  • Separate, revocable keys per project and environment
  • Every request traceable by id, from API call to invoice
  • Monthly spending caps and per-organization rate limits

Developer experience

Works with the client you already use.

The xVoice API follows the OpenAI audio API, so there is no new SDK to learn. Already on OpenAI? Change the base URL and the key, and your code runs on xVoice.

  • The official Python and Node clients, unchanged
  • Same endpoints, request bodies and streaming
  • Errors come back in a shape your code already handles
Browse runnable examples →
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://api.openai.com/v1",  
    api_key=os.environ["OPENAI_API_KEY"],  
    base_url="https://api.xvoice.dev/v1",  
    api_key=os.environ["XVOICE_API_KEY"],  
)

with open("call.wav", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model="xvoice-stt1",
        file=audio,
    )

print(transcript.text)

Platform

Everything around the API, already built

The dashboard your team needs on day one: keys, logs, usage and spend, all scoped by organization and project.

Projects and scoped keys

Separate keys per app and environment, test or live. Revoke one without touching the rest.

Request logs

Every call has a request id that ties an invoice line to a row in your logs, errors included.

Usage you can slice

Seconds of audio and characters of text, grouped by project, model or key, over any range.

Prepaid credits

Pay as you go from a balance. Auto-recharge keeps it topped up; nothing bills you by surprise.

Spending and rate limits

Cap what a month can cost. Rate limit headers on every metered call tell you where you stand.

Teams and audit log

Invite people by email with roles. Every change to keys, members or spending is recorded.

Pricing

Pay for what you use

Prepaid credits, metered from our own measurements: seconds of audio in, characters of text out.

Speech to textxvoice-stt1

$0.36per hour of audio

Files and realtime sessions, metered per second.

Text to speechxvoice-tts1

$15per 1M characters

Five audio formats, adjustable speed.

Text to speechxvoice-tts2

$10per 1M characters

Lowest time to first audio.

  • Free starting credit when you verify your email
  • No subscription or seat fees
  • Top up from $5, auto-recharge optional
  • Monthly spending cap you control

Quickstart

Your first call in under five minutes

No SDK of ours to learn, no new concepts to read about before the first response.

  1. 1

    Create an account

    Sign up with email or Google. An organization and a first project come with the account.

    app.xvoice.dev/signup
  2. 2

    Issue an API key

    Verify your email for free starting credit, then create a test key. It is shown once, so store it as a secret.

    export XVOICE_API_KEY=sk_test_…
  3. 3

    Make your first call

    Install the openai client you would use anyway, set the base URL, and send audio.

    pip install openai

Follow the full quickstart →

FAQ

Questions, answered

Is xVoice HIPAA compliant?

Yes. You can process protected health information through xVoice, and we sign a Business Associate Agreement with covered entities and their business associates.

Do you have a SOC 2 report?

Yes, SOC 2 Type II. The report is available to customers and prospects under NDA.

How fast is it?

Built for live conversation. Realtime sessions return partial transcripts while the speaker is still talking, and text-to-speech streams audio as it is generated, so playback starts before the whole response exists.

How accurate is it on medical and domain-specific audio?

Transcription is built to get drug names, dosages, lab values and numbers right, on real calls rather than clean studio audio. The free starting credit is enough to run your own recordings through it before you commit.

Do I need a new SDK?

No. The API follows the OpenAI audio API, so the official Python and Node clients work unchanged. Point them at the xVoice base URL with an xVoice key.

How does billing work?

Prepaid credit, deducted per second of audio transcribed and per character of speech generated. Verifying your email adds $5 of free credit. Top-ups run from $5 to $1,000, auto-recharge is optional, and a monthly spending cap protects you from runaway usage.

What are the rate limits?

Limits are per organization, covering requests per minute, concurrent requests, characters and audio seconds. Every metered response carries headers showing where you stand, and limits can be raised for production workloads.

When are speech-to-speech models coming?

Soon. They listen and answer in voice in a single model, with no transcription step in between. Sign up to get early access when they launch.

Still have questions? Read the docs →

Ready for production on day one.

Start with free credit and make your first call in minutes, or talk to us about a BAA and volume pricing.