Frontier AI & Vision Rate Card
OpenAI API Pricing Sheet & Model Token Calculator 2026
OpenAI models are billed per 1 million tokens, partitioned into Input Tokens and Output Tokens. Cached input tokens receive a 50% discount, and asynchronous Batch API jobs receive an automatic 50% price reduction across all models.
OpenAI Flagship & Reasoning Models Token Pricing
| Model Identifier | Input Rate / 1M | Cached Input / 1M | Output Rate / 1M | Batch API (50% Off) | Context Window |
|---|---|---|---|---|---|
| o1 (Flagship Reasoning) | $15.00 | $7.50 | $60.00 | $7.50 / $30.00 | 200,000 tokens |
| o3-mini (Fast Reasoning) | $1.10 | $0.55 | $4.40 | $0.55 / $2.20 | 200,000 tokens |
| o1-mini (Legacy Reasoning) | $1.10 | $0.55 | $4.40 | $0.55 / $2.20 | 128,000 tokens |
| GPT-4o (Omni Flagship) | $2.50 | $1.25 | $10.00 | $1.25 / $5.00 | 128,000 tokens |
| GPT-4o-mini (Lightweight) | $0.15 | $0.075 | $0.60 | $0.075 / $0.30 | 128,000 tokens |
| GPT-4 Turbo (Legacy) | $10.00 | $5.00 | $30.00 | $5.00 / $15.00 | 128,000 tokens |
| GPT-3.5 Turbo (Legacy) | $0.50 | $0.50 | $1.50 | $0.25 / $0.75 | 16,384 tokens |
OpenAI Audio, Vision & Embedding Specialized Models
| Model / Capability | Standard Pricing | Unit Metric | Use Case & Capabilities |
|---|---|---|---|
| Whisper (STT) | $0.0060 | Per minute of audio | Speech-to-text audio transcription and translation |
| TTS (Text-to-Speech Standard) | $15.00 | Per 1M characters (~$0.015/1K) | Natural AI speech generation (6 voices) |
| TTS HD (High Definition Voice) | $30.00 | Per 1M characters (~$0.030/1K) | Studio quality voice synthesis |
| text-embedding-3-small | $0.020 | Per 1M tokens | 1536-dim dense vector embedding for RAG |
| text-embedding-3-large | $0.130 | Per 1M tokens | 3072-dim high accuracy vector search embedding |
| DALL-E 3 (Standard 1024x1024) | $0.040 | Per generated image | Square standard resolution image generation |
| DALL-E 3 (HD 1024x1792) | $0.080 | Per generated image | High-definition widescreen image synthesis |
Official Developer SDK / API Integration Example
Python SDK
# OpenAI Python SDK (GPT-4o & o1 Reasoning)
from openai import OpenAI
client = OpenAI(api_key="sk-...")
# o1 reasoning model with hidden internal chain-of-thought tokens ($60/M out)
response = client.chat.completions.create(
model="o1",
messages=[{"role": "user", "content": "Analyze and solve this algorithm problem..."}]
)
print("Output:", response.choices[0].message.content)
Frequently Asked Billing & Pricing Questions
How does OpenAI Batch API reduce costs by 50%?
The Batch API allows submitting requests in bulk via JSONL files. Responses are returned asynchronously within 24 hours at an automatic 50% discount on both input and output tokens, perfect for data labeling and backfilling.
Why are reasoning tokens billed on o1?
Reasoning models generate internal chain-of-thought tokens to deliberate on complex math, logic, and code. While invisible in standard API completions, these tokens are billed at full output rate.
What is OpenAI Prompt Caching?
For prompts over 1,024 tokens, OpenAI automatically caches the prefix in memory. Cache hits receive an automatic 50% discount ($1.25/M on GPT-4o) with zero configuration required.