Realtime Speech-to-Text (ASR) API
Overview
Real-time speech-to-text streaming over a single WebSocket connection.
This endpoint is a dedicated transcription stream: you push raw audio frames in and receive incremental transcriptions and, optionally, translations back as JSON.
Palabra API client
Consider using the Palabra API Python client and checking the code example with it.
Key Setup Steps
Step 1: Authentication
Create an API Key to authenticate requests.
Step 2: WebSocket Connection
Connect to the endpoint with the required parameters:
wss://stream.palabra.ai/asr/v1/speech-to-text/stream?token=<API_KEY>&language=en&format=pcm_s16le&sample_rate=16000
Step 3: Audio Transmission
Send audio as raw binary WebSocket frames, with 320ms chunks recommended.
Step 4: Message Reception
The server returns JSON text frames with transcription data, identified by message_type.
Supported Parameters
Apply values as WebSocket query parameters
| Parameter | Required | Purpose |
|---|---|---|
api_key | Yes | API Key |
format | Yes | Audio format specification |
sample_rate | Conditional | Hz rate for PCM formats |
language | No | Source language code (defaults to auto) |
translate_languages | No | Comma-separated target languages |
enable_filler_filter | No | Filter control (true by default) |
Supported Languages
13 languages are supported. Automatic language detection operates in experimental mode.
| Code | Language |
|---|---|
ar | Arabic |
de | German |
en | English |
es | Spanish |
fr | French |
hi | Hindi |
it | Italian |
ja | Japanese |
ko | Korean |
nl | Dutch |
pt | Portuguese |
ru | Russian |
zh | Chinese |
Audio Formats
| format | sample_rate | Notes |
|---|---|---|
pcm_s16le | only if ≠ 16000 | 16-bit signed little-endian PCM. Recommended |
pcm_f32le / pcm_f32be | required | 32-bit float PCM |
pcm_s32le / pcm_s32be | required | 32-bit signed PCM |
mulaw / alaw | required | G.711 |
webm / mp3 / aac / ogg / flac / wav | not used | Container formats; rate is read from the stream |
Message Types
Transcription Messages
Transcriptions deliver incremental text with segment timing and delta hints showing newly-added text.
{
"message_type": "transcription",
"transcription_id": "a1b2c3d4",
"language": "en",
"is_eos": false,
"segment": {
"text": "Hello world how are",
"start_time": 0.32,
"end_time": 1.84
},
"delta": {
"text": "how are",
"start_time": 1.20,
"end_time": 1.84
}
}
| Field | Description |
|---|---|
transcription_id | Stable id for the segment. All messages of one segment share the same id. A new id means a new segment has started |
language | Detected (or configured) source language of this segment |
is_eos | false — partial; the segment is still being updated. true — the segment is committed and final |
segment.text | The full text of the segment so far |
segment.start_time / end_time | Segment timing, in seconds relative to session start |
delta | Incremental hint: the text added since the previous partial of the same segment (see below) |
Working with delta
When the filler filter is disabled, delta.text is append-only: each transcription message carries exactly the text appended since the previous partial, so you can concatenate deltas directly.
With the enable_filler_filter enabled, the recognizer's tail might be rewritten mid-segment, which breaks the append relationship. In that mode treat segment.text as authoritative and overwrite the current segment on each message; use delta only as a hint.
Translated Transcription Messages
Translated Transcription appear only when translate_languages is set, once per target language, after each final (is_eos: true) transcription.
{
"message_type": "translated_transcription",
"transcription_id": "a1b2c3d4",
"language": "es",
"is_eos": true,
"segment": {
"text": "Hola mundo, ¿cómo estás?",
"start_time": 0.32,
"end_time": 1.84
}
}
transcription_id matches the id of the source transcription (the is_eos: true one) this translation was produced from — use it to correlate a translation back to its original segment. language here is the target language, and is_eos is always true (translations are produced only for finalized segments).
Errors
Authentication and routing failures are reported as HTTP status codes during the WebSocket upgrade, before the connection is established:
| HTTP status | Meaning |
|---|---|
401 | Missing or invalid API Key / token |
409 | A session is already active for this identity |
After a successful upgrade, the server does not send application-level error messages over the wire — it closes the connection with a standard WebSocket close frame.
Complete example
Streams microphone audio and prints transcriptions (and translations, if
PALABRA_LANGUAGE targets are configured).
pip install pyaudio websockets
export PALABRA_API_KEY=... # from Step 1
export PALABRA_LANGUAGE=en # source language
import json
import os
import asyncio
import threading
import queue
import pyaudio
import websockets
WS_URL = "wss://stream.palabra.ai/asr/v1/speech-to-text/stream"
LANGUAGE = os.environ.get("PALABRA_LANGUAGE", "en")
SAMPLE_RATE = 16000
CHANNELS = 1
CHUNK = 5120 # samples ≈ 320 ms at 16 kHz (recommended chunk size)
def mic_reader(audio_queue: queue.Queue, stop_event: threading.Event):
pa = pyaudio.PyAudio()
stream = pa.open(
format=pyaudio.paInt16,
channels=CHANNELS,
rate=SAMPLE_RATE,
input=True,
frames_per_buffer=CHUNK,
)
print("Microphone open, speak now...")
try:
while not stop_event.is_set():
audio_queue.put(stream.read(CHUNK, exception_on_overflow=False))
finally:
stream.stop_stream()
stream.close()
pa.terminate
async def stream(key: str):
url = (
f"{WS_URL}?token={key}&language={LANGUAGE}"
f"&format=pcm_s16le&sample_rate={SAMPLE_RATE}"
)
audio_queue: queue.Queue = queue.Queue()
stop_event = threading.Event()
threading.Thread(
target=mic_reader, args=(audio_queue, stop_event), daemon=True
).start()
async with websockets.connect(url) as ws:
print("Connected")
async def send_audio():
loop = asyncio.get_event_loop()
while True:
data = await loop.run_in_executor(None, audio_queue.get)
await ws.send(data) # raw binary frame
async def receive():
async for message in ws:
msg = json.loads(message)
msg_type = msg.get("message_type")
if msg_type == "transcription":
text = msg["segment"]["text"]
tid = msg.get("transcription_id", "")
if msg.get("is_eos"):
print(f"\n[EOS] {text} [{tid}]")
else:
# segment.text is the source of truth — render it whole
print(f"\r {text}", end="", flush=True)
elif msg_type == "translated_transcription":
lang = msg.get("language", "?")
tid = msg.get("transcription_id", "")
print(f"\n[{lang}] {msg['segment']['text']} [{tid}]")
try:
await asyncio.gather(send_audio(), receive())
finally:
stop_event.set()
if __name__ == "__main__":
try:
asyncio.run(stream(os.environ["PALABRA_API_KEY"]))
except KeyboardInterrupt:
print("\nStopped.")