← Master Hub
Frontier AI & Vision Rate Card

OpenAI API Pricing Sheet & Model Token Calculator 2026

OpenAI models are billed per 1 million tokens, partitioned into Input Tokens and Output Tokens. Cached input tokens receive a 50% discount, and asynchronous Batch API jobs receive an automatic 50% price reduction across all models.

OpenAI Flagship & Reasoning Models Token Pricing

Model IdentifierInput Rate / 1MCached Input / 1MOutput Rate / 1MBatch API (50% Off)Context Window
o1 (Flagship Reasoning)$15.00$7.50$60.00$7.50 / $30.00200,000 tokens
o3-mini (Fast Reasoning)$1.10$0.55$4.40$0.55 / $2.20200,000 tokens
o1-mini (Legacy Reasoning)$1.10$0.55$4.40$0.55 / $2.20128,000 tokens
GPT-4o (Omni Flagship)$2.50$1.25$10.00$1.25 / $5.00128,000 tokens
GPT-4o-mini (Lightweight)$0.15$0.075$0.60$0.075 / $0.30128,000 tokens
GPT-4 Turbo (Legacy)$10.00$5.00$30.00$5.00 / $15.00128,000 tokens
GPT-3.5 Turbo (Legacy)$0.50$0.50$1.50$0.25 / $0.7516,384 tokens

OpenAI Audio, Vision & Embedding Specialized Models

Model / CapabilityStandard PricingUnit MetricUse Case & Capabilities
Whisper (STT)$0.0060Per minute of audioSpeech-to-text audio transcription and translation
TTS (Text-to-Speech Standard)$15.00Per 1M characters (~$0.015/1K)Natural AI speech generation (6 voices)
TTS HD (High Definition Voice)$30.00Per 1M characters (~$0.030/1K)Studio quality voice synthesis
text-embedding-3-small$0.020Per 1M tokens1536-dim dense vector embedding for RAG
text-embedding-3-large$0.130Per 1M tokens3072-dim high accuracy vector search embedding
DALL-E 3 (Standard 1024x1024)$0.040Per generated imageSquare standard resolution image generation
DALL-E 3 (HD 1024x1792)$0.080Per generated imageHigh-definition widescreen image synthesis
Official Developer SDK / API Integration Example Python SDK
# OpenAI Python SDK (GPT-4o & o1 Reasoning)
from openai import OpenAI

client = OpenAI(api_key="sk-...")

# o1 reasoning model with hidden internal chain-of-thought tokens ($60/M out)
response = client.chat.completions.create(
    model="o1",
    messages=[{"role": "user", "content": "Analyze and solve this algorithm problem..."}]
)
print("Output:", response.choices[0].message.content)

Frequently Asked Billing & Pricing Questions

How does OpenAI Batch API reduce costs by 50%?
The Batch API allows submitting requests in bulk via JSONL files. Responses are returned asynchronously within 24 hours at an automatic 50% discount on both input and output tokens, perfect for data labeling and backfilling.
Why are reasoning tokens billed on o1?
Reasoning models generate internal chain-of-thought tokens to deliberate on complex math, logic, and code. While invisible in standard API completions, these tokens are billed at full output rate.
What is OpenAI Prompt Caching?
For prompts over 1,024 tokens, OpenAI automatically caches the prefix in memory. Cache hits receive an automatic 50% discount ($1.25/M on GPT-4o) with zero configuration required.