GPT Realtime 2.1
OpenAI logo

GPT Realtime 2.1

gpt-realtime-2.1llms.txt
OpenAI
New

Pricing

PricingInput TextInput AudioInput Text Cached Input Audio Cached Output TextOutput Audio
$0$0
------

Input Modalities

  • Text
  • Vision
  • Audio

Output Modalities

  • Text
  • Audio

Context length

  • 128K tokens

Max output

  • 32K tokens

Capabilities

  • Thinking
  • Streaming
  • Tool calling
  • Web search
  • URL context
  • Code interpreter
  • Computer use
  • File search
  • Memory tool
  • Audio output
  • Structured outputs
  • Structured decision
  • Citations
  • Prompt caching
  • Background mode
  • Server-side sessions

Try this model

Python
# pip install websockets
import asyncio
import base64
import json
import os
import websockets

# Realtime conversation is a WebSocket session: model must be in the handshake URL.
URL = "wss://aihubmix.com/v1/realtime?model=gpt-realtime-2.1"

async def main():
    # websockets >= 13 uses additional_headers; older versions use extra_headers
    async with websockets.connect(
        URL, additional_headers={"Authorization": "Bearer " + os.environ["AIHUBMIX_API_KEY"]}
    ) as ws:
        # 1) Configure the speech-to-speech session (voice + server VAD; no input transcription)
        await ws.send(json.dumps({
          "type": "session.update",
          "session": {
            "type": "realtime",
            "audio": {
              "input": {
                "format": {
                  "type": "audio/pcm",
                  "rate": 24000
                },
                "turn_detection": {
                  "type": "server_vad"
                }
              },
              "output": {
                "format": {
                  "type": "audio/pcm",
                  "rate": 24000
                },
                "voice": "marin"
              }
            }
          }
        }))

        # 2) Stream your mic as raw PCM16 / 24kHz / mono in ~100ms chunks. Server VAD
        #    detects when you stop talking and starts the reply automatically —
        #    no commit / response.create needed.
        async def send_audio():
            with open("audio_pcm16_24k.raw", "rb") as f:
                pcm = f.read()
            chunk = 24000 * 2 * 100 // 1000  # 100ms of 16-bit mono samples
            for i in range(0, len(pcm), chunk):
                await ws.send(json.dumps({
                    "type": "input_audio_buffer.append",
                    "audio": base64.b64encode(pcm[i:i + chunk]).decode(),
                }))
                await asyncio.sleep(0.1)  # simulate realtime pacing

        asyncio.create_task(send_audio())

        # 3) Receive the reply: audio streams as base64 PCM16 (24kHz) — write it to a file you
        #    can play; the transcript of what the model says arrives as text deltas.
        reply = open("assistant_reply_pcm16_24k.raw", "wb")
        async for msg in ws:
            evt = json.loads(msg)
            etype = evt.get("type", "")
            if etype == "input_audio_buffer.speech_started":
                # Barge-in: you started talking — stop/flush local playback here
                print("\n[listening…]")
            elif etype.endswith("audio_transcript.delta"):
                print(evt.get("delta", ""), end="", flush=True)
            elif etype.endswith("audio.delta"):
                reply.write(base64.b64decode(evt.get("delta", "")))
            elif etype.endswith("response.done"):
                print("\n[reply complete]")
            elif etype == "error":
                print("\n[error]", evt.get("error"))
                break
        reply.close()

asyncio.run(main())

Frequently asked questions

What is the context length of GPT Realtime 2.1?

GPT Realtime 2.1 has a 128,000 token context window. It supports up to 32,000 output tokens.