Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 10 additions & 36 deletions fern/openai-realtime.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -17,13 +17,13 @@ OpenAI’s Realtime API enables developers to use a native speech-to-speech mode

## Available models

OpenAI offers three realtime models, each with different capabilities and cost/performance trade-offs:
The model picker offers three OpenAI realtime options:

| Model | Status | Best For | Key Features |
|-------|---------|----------|--------------|
| `gpt-realtime-2025-08-28` | **Production** | Production workloads | Production Ready |
| `gpt-4o-realtime-preview-2024-12-17` | Preview | Development & testing | Balanced performance/cost |
| `gpt-4o-mini-realtime-preview-2024-12-17` | Preview | Cost-sensitive apps | Lower latency, reduced cost |
| Dashboard option | Model ID | Guidance |
|------------------|----------|----------|
| GPT Realtime Cluster | `gpt-realtime-2025-08-28` | Existing assistants |
| GPT Realtime Mini | `gpt-realtime-mini-2025-12-15` | Cost-sensitive assistants |
| GPT Realtime 2 | `gpt-realtime-2` | Recommended for new assistants |

## Voice options

Expand Down Expand Up @@ -66,7 +66,7 @@ The tool's `body` schema defines the `location` argument the model supplies. For
{
"model": {
"provider": "openai",
"model": "gpt-realtime-2025-08-28",
"model": "gpt-realtime-2",
"messages": [
{
"role": "system",
Expand Down Expand Up @@ -112,7 +112,7 @@ const vapi = new VapiClient({ token: apiKey });
const assistant = await vapi.assistants.create({
model: {
provider: "openai",
model: "gpt-realtime-2025-08-28",
model: "gpt-realtime-2",
messages: [{
role: "system",
content: "You are a concise, friendly weather assistant. If the caller has not provided a location, ask for one. If the city is ambiguous, ask for the missing region or country before using getWeather. Call getWeather for each new current-weather request, including a request for another city. Pass the complete location, preserving any region/state and country the caller supplied. Use only the latest successful result for the requested location and report the returned location with the weather. If the returned city, region, or country conflicts with the request, clarify before reporting weather. Differences in spelling or formatting alone are not a location mismatch. If the lookup fails or current-weather data is missing, explain that current weather is unavailable. Do not invent weather or reuse an earlier result after a failed lookup."
Expand Down Expand Up @@ -154,7 +154,7 @@ vapi = Vapi(token=os.getenv("VAPI_API_KEY"))
assistant = vapi.assistants.create(
model={
"provider": "openai",
"model": "gpt-realtime-2025-08-28",
"model": "gpt-realtime-2",
"messages": [{
"role": "system",
"content": "You are a concise, friendly weather assistant. If the caller has not provided a location, ask for one. If the city is ambiguous, ask for the missing region or country before using getWeather. Call getWeather for each new current-weather request, including a request for another city. Pass the complete location, preserving any region/state and country the caller supplied. Use only the latest successful result for the requested location and report the returned location with the weather. If the returned city, region, or country conflicts with the request, clarify before reporting weather. Differences in spelling or formatting alone are not a location mismatch. If the lookup fails or current-weather data is missing, explain that current weather is unavailable. Do not invent weather or reuse an earlier result after a failed lookup."
Expand Down Expand Up @@ -311,7 +311,7 @@ Transitioning from standard STT/TTS to realtime models:
{
"model": {
"provider": "openai",
"model": "gpt-realtime-2025-08-28" // Changed from gpt-4
"model": "gpt-realtime-2"
}
}
```
Expand All @@ -336,32 +336,6 @@ Transitioning from standard STT/TTS to realtime models:

## Best practices

### Model selection strategy

<AccordionGroup>
<Accordion title="When to use gpt-realtime-2025-08-28">
**Best for production workloads requiring:**
- Structured outputs for form filling or data collection
- Complex function orchestration
- Highest quality voice interactions
- Responses API integration
</Accordion>

<Accordion title="When to use gpt-4o-realtime-preview">
**Best for development and testing:**
- Prototyping voice applications
- Balanced cost/performance during development
- Testing conversation flows before production
</Accordion>

<Accordion title="When to use gpt-4o-mini-realtime-preview">
**Best for cost-sensitive applications:**
- High-volume voice interactions
- Simple Q&A or routing scenarios
- Applications where latency is critical
</Accordion>
</AccordionGroup>

### Performance optimization

- **Temperature settings**: Use 0.5-0.7 for consistent yet natural responses
Expand Down
2 changes: 2 additions & 0 deletions fern/providers/model/openai.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,8 @@ For additional API configuration options, review the [`OpenAIModel` fields](/api
| GPT 4.1 | `gpt-4.1` |
| GPT 4.1 Mini | `gpt-4.1-mini` |
| GPT 4o Mini | `gpt-4o-mini` |
| GPT Realtime Cluster | `gpt-realtime-2025-08-28` |
| GPT Realtime Mini | `gpt-realtime-mini-2025-12-15` |
| GPT Realtime 2 | `gpt-realtime-2` |
| o3 | `o3` |

Expand Down
Loading