ElevenLabs 语音模型上线 OpenRouter:9 个文本转语音与 2 个语音转文本模型
ElevenLabs is now on OpenRouter
ElevenLabs 的 9 个 Text to Speech 模型和 2 个 Speech to Text 模型现已接入 OpenRouter,用户用 OpenRouter API key 调用 POST /api/v1/audio/speech 和 POST /api/v1/audio/transcriptions 即可,无需单独的 ElevenLabs 计划。
ElevenLabs is now on OpenRouter with nine Text to Speech models and two Speech to Text models.
If your app already talks to a language model through OpenRouter, it can now speak and listen too. Narrate long-form content with direction over tone and delivery with Eleven v4, give your agent faster spoken replies with v4 Turbo, or transcribe conversations with speaker labels with Scribe v2. You don’t need a separate ElevenLabs plan. Just call POST /api/v1/audio/speech and POST /api/v1/audio/transcriptions with your OpenRouter API key.
Every ElevenLabs model is 50% off OpenRouter’s list price for all OpenRouter customers through October 19, 8am PT
Which model to use
Start with three models and reach for the rest when a job needs them:
- Narration, characters, product videos: Eleven v4. Audio tags let the script carry the performance.
- Voice agents and live replies: Eleven v4 Turbo, ElevenLabs built it for real-time use, at half the v4 rate.
- Transcripts, meeting notes, captions: Scribe v2, with speaker labels and word timestamps.
Use Multilingual v2 when long-form audio needs speed control, and Flash v2.5 for high-volume speech of up to 40,000 characters per request. Scribe v2 Medical is fine-tuned for medical terms.
Text to Speech
| Model | OpenRouter ID | Max characters per request | Pick it for |
|---|---|---|---|
| Eleven v4 | elevenlabs/eleven-v4 | 10,000 | ElevenLab’s highest-quality, most emotive speech. Expressive narration and character work, with audio tags like [whispering] and [laughing] |
| Eleven v4 Turbo | elevenlabs/eleven-v4-turbo | 10,000 | low-latency replies for voice agents |
| Eleven v3 | elevenlabs/eleven-v3 | 5,000 | Expressive/dramatic narration and character work with audio tags |
| Eleven v3 Conversational | elevenlabs/eleven-v3-conversational | 5,000 | Back-and-forth dialogue at half the v3 rate. Has dramatic range, tuned for live dialogue. |
| Eleven Multilingual v2 | elevenlabs/eleven-multilingual-v2 | 10,000 | Stable long-form audio across many languages, with speed control. |
| Eleven Flash v2.5 | elevenlabs/eleven-flash-v2.5 | 40,000 | Fast, low-cost multilingual speech at high volume. 32 languages available. |
| Eleven Flash v2 | elevenlabs/eleven-flash-v2 | 30,000 | Fast English-only speech |
| Eleven Turbo v2.5 | elevenlabs/eleven-turbo-v2.5 | 40,000 | ElevenLab’s First-generation low-latency model, now superseded by Flash |
| Eleven Turbo v2 | elevenlabs/eleven-turbo-v2 | 30,000 | ElevenLab’s first-generation low-latency model, now superseded by Flash |
ElevenLabs prices its models in two tiers. Eleven v4, v3, and Multilingual v2 bill the full per-character rate, and v4 Turbo, v3 Conversational, Flash, and Turbo bill half of it. Current per-character prices are on each model page in the Text to Speech collection.
Speech to Text
| Model | OpenRouter ID | Pick it for |
|---|---|---|
| Scribe v2 | elevenlabs/scribe-v2 | General transcription with word timestamps, speaker labels, and audio-event tags in 90+ languages |
| Scribe v2 Medical | elevenlabs/scribe-v2-medical | Medical terms with 35% fewer errors on clinical audio compared to Scribe v2 |
Both Scribe models use the same request shape and the same per-second price. Realtime Scribe v2 over WebSocket isn’t part of this launch.
Walkthrough: from a first MP3 to a voice agent
We’ll build a small product tour in three steps. First Eleven v4 narrates it. Then a voice agent answers questions about it, and finally Scribe v2 turns the team’s review meeting into notes that show who said what. If a request doesn’t behave the way you expect, check the tips at the end.
Set your key once:
export OPENROUTER_API_KEY="your-api-key"Step 1: Narrate a product tour with Eleven v4
Start with one line of narration:
curl https://openrouter.ai/api/v1/audio/speech \
--fail-with-body \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs/eleven-v4",
"input": "[cheerfully] Welcome to the tour. [whispering] Let me show you the shortcut first.",
"voice": "george",
"response_format": "mp3"
}' \
--output intro.mp3voice takes any of ElevenLabs’ 21 premade voices by name (george here); names aren’t case-sensitive. Note the explicit response_format — see Tips for your first requests for both.
The bracketed words are audio tags. Eleven v4 and v3 read tags such as [whispering], [laughing], [sighs], and [curious] as delivery directions instead of speaking them. Tags count toward your billed characters.
A full tour script is longer than one request allows (10,000 characters on v4, 5,000 on v3), so split it on paragraph boundaries. ElevenLabs-only settings go under provider.options.elevenlabs. Pass the neighbouring text as previous_text and next_text so intonation stays consistent across the joins, and fix a seed so reruns are more repeatable:
import os
import requests
API = "https://openrouter.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"}
NARRATOR = "george"
def split_paragraphs(text: str, limit: int = 9000) -> list[str]:
chunks, current = [], ""
for para in text.split("\n\n"):
if current and len(current) + len(para) + 2 > limit:
chunks.append(current)
current = para
else:
current = f"{current}\n\n{para}" if current else para
return chunks + [current] if current else chunks
chunks = split_paragraphs(open("tour-script.txt").read())
with open("tour.mp3", "wb") as out:
for i, chunk in enumerate(chunks):
options = {"seed": 7}
if i > 0:
options["previous_text"] = chunks[i - 1][-500:]
if i + 1 < len(chunks):
options["next_text"] = chunks[i + 1][:500]
r = requests.post(
f"{API}/audio/speech",
headers=HEADERS,
json={
"model": "elevenlabs/eleven-v4",
"input": chunk,
"voice": NARRATOR,
"response_format": "mp3",
"provider": {"options": {"elevenlabs": options}},
},
)
r.raise_for_status()
out.write(r.content)Step 2: Answer questions with a voice agent
Now let users ask the tour questions out loud. A voice agent chains three models: Scribe v2 for the question, a chat model for the answer, and Eleven v4 Turbo for the reply. ElevenLabs builds v4 Turbo for real-time use, and it bills at half the v4 rate.
import base64
def transcribe(path: str) -> str:
with open(path, "rb") as f:
audio_b64 = base64.b64encode(f.read()).decode()
r = requests.post(
f"{API}/audio/transcriptions",
headers=HEADERS,
json={
"model": "elevenlabs/scribe-v2",
"input_audio": {"data": audio_b64, "format": "wav"},
},
)
r.raise_for_status()
return r.json()["text"]
def answer(question: str) -> str:
r = requests.post(
f"{API}/chat/completions",
headers=HEADERS,
json={
"model": "openai/gpt-5-mini",
"messages": [
{"role": "system", "content": "You are the product tour guide. Answer in one or two sentences."},
{"role": "user", "content": question},
],
},
)
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"]
def speak(text: str, out_path: str) -> None:
r = requests.post(
f"{API}/audio/speech",
headers=HEADERS,
json={
"model": "elevenlabs/eleven-v4-turbo",
"input": text,
"voice": NARRATOR,
"response_format": "mp3",
},
)
r.raise_for_status()
with open(out_path, "wb") as f:
f.write(r.content)
speak(answer(transcribe("question.wav")), "reply.mp3")The transcription response is JSON with the transcript in text and a usage object that reports the audio seconds and the dollar cost of the request. Swap openai/gpt-5-mini for any chat model in the catalog without touching the audio calls, and keep the same voice so the agent sounds like the narrator from step 1.
Step 3: Turn the review meeting into speaker-labelled notes
When your team meets to review the tour, Scribe v2 can tell you who said what. Upload the recording as multipart, ask for verbose_json with word timestamps, and set diarize under provider.options.elevenlabs. Every word in the response then carries a speaker. If you know how many people were in the meeting, pass num_speakers to improve the split.
curl https://openrouter.ai/api/v1/audio/transcriptions \
--fail-with-body \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-F file="@tour-review.mp3" \
-F model="elevenlabs/scribe-v2" \
-F response_format="verbose_json" \
-F "timestamp_granularities[]=word" \
-F provider='{"options": {"elevenlabs": {"diarize": true, "num_speakers": 3, "tag_audio_events": true}}}'Each entry in words has word, start, end, confidence, speaker (an integer), and speaker_label (speaker_0, speaker_1, and so on). tag_audio_events adds entries such as (laughter) alongside the words. Group consecutive words by speaker to get turns, then send the turns to a chat model for a summary and action items:
def to_turns(words: list[dict]) -> list[str]:
turns: list[tuple[int, list[str]]] = []
for w in words:
if turns and turns[-1][0] == w.get("speaker"):
turns[-1][1].append(w["word"])
else:
turns.append((w.get("speaker"), [w["word"]]))
return [f"Speaker {s}: {' '.join(ws)}" for s, ws in turns]Uploads are capped at 25 MB, which covers about 27 minutes of 128 kbps MP3.
For longer recordings, pass a public URL and ElevenLabs downloads the file directly, so the 25 MB cap doesn’t apply:
"input_audio":{ "url": "https://example.com/tour-review.mp3" }Very long files can still hit the 180-second upstream request timeout, so split recordings that run long.
Tips for your first requests
- Pick a voice by name or ID.
For
voice, pass any of ElevenLabs’ 21 premade voices by name:george,sarah,adam,alice,bella,brian,charlie,daniel,jessica,roger,willand more (full list on each model page). Names aren’t case-sensitive. To use any other voice, pass its ElevenLabs voice ID asvoice— copy it from the ElevenLabs voice library. Voices you clone or design live in your own ElevenLabs account, so use those with BYOK. - Ask for
mp3when you want a file you can play If you leave outresponse_format, our speech endpoint returns raw PCM (16-bit little-endian mono at 24 kHz). That suits an audio pipeline, but most players can’t open it. Set"response_format": "mp3"to get a playable 44.1 kHz, 128 kbps file. - Use Multilingual v2 or Flash for speed control
Multilingual v2 and Flash take
speedfrom 0.7 to 1.2. Eleven v4 and v4 Turbo have no speed control and reject a non-defaultspeed, so leave it out on those models and change the pacing with audio tags instead. - Split long scripts
Each model has a per-request character limit (see the model table). Audio tags count toward the limit and toward your billed characters. Split on paragraph boundaries as in step 1, and pass
previous_textandnext_textso the joins sound natural. - ElevenLabs-specific settings go under
provider.options.elevenlabsText to Speech acceptsseed,previous_text,next_text,language_code,apply_text_normalization,apply_language_text_normalization,pronunciation_dictionary_locators,voice_settings(stability,similarity_boost,style,use_speaker_boost,speed), andoutput_format(an ElevenLabs format name such asmp3_22050_32orpcm_16000; it must match yourresponse_formatcodec). Speech to Text acceptsdiarize,num_speakers,diarization_threshold,tag_audio_events,no_verbatim, andseed.num_speakersanddiarization_thresholdneeddiarize: true, and you can set one or the other, not both. Use the field names exactly as listed. Keys that aren’t on these lists are ignored; a listed key with a bad value returns a 400. - Scribe returns three-letter language codes
You can send a two-letter
languagehint (en). The detected language comes back as an ISO 639-3 code (eng).
Getting started
- Create an OpenRouter API key.
- Make your first request with the curl example in Step 1 above.
- For every parameter, see the Text to Speech and Speech to Text guides.
Text to Speech is billed per character of input, counted as Unicode code points, so an emoji or a CJK character counts once. Audio tags count as characters. Scribe is billed per second of audio, and the usage.cost field on every transcription response tells you what that request cost. We pass through ElevenLabs pricing without markup; see each model page for the current rate.
Already have an ElevenLabs account?
Add your key with BYOK to use your own cloned voices, pronunciation dictionaries, and higher-quality formats such as mp3_44100_192.
Create an OpenRouter API key, or try a model first in the Playground. The Text to Speech and Speech to Text guides have the full request reference.
来源:OpenRouter:Announcements · openrouter.ai