RAG & Vector Retrieval Pricing
Embedding Models & Vector Dimensions Pricing Matrix 2026
Vector embeddings represent text passages as high-dimensional numerical vectors for semantic search and RAG. Pricing is billed per million tokens processed, but downstream vector database costs scale directly with embedding dimensionality (e.g. 512 vs 1536 vs 3072 dimensions).
Frontier Text Embedding Models Benchmark & Pricing
| Model Identifier | Provider | Price / 1M Tokens | Dimensions | MTEB Benchmark Score | Max Input Tokens |
|---|---|---|---|---|---|
| text-embedding-3-small | OpenAI | $0.020 / 1M | 1,536 dims (shortenable to 512) | 62.3 (MTEB Average) | 8,191 tokens |
| text-embedding-3-large | OpenAI | $0.130 / 1M | 3,072 dims (shortenable to 1024) | 64.6 (MTEB Average) | 8,191 tokens |
| voyage-3 | Voyage AI | $0.120 / 1M | 1,024 dims | 67.8 (Top Ranked) | 32,000 tokens |
| voyage-code-3 | Voyage AI | $0.120 / 1M | 1,024 dims | State-of-the-Art Code Retrieval | 32,000 tokens |
| embed-english-v3.0 | Cohere | $0.100 / 1M | 1,024 dims | 64.5 (High Recall) | 512 tokens |
| text-embedding-004 | Google Gemini | $0.025 / 1M | 768 dims | 66.3 (Excellent Value) | 2,048 tokens |
| bge-large-en-v1.5 | Self-Hosted / Open | $0.000 (Open Weights) | 1,024 dims | 64.1 (Free on GPU) | 512 tokens |
Frequently Asked Technical Questions
How does embedding dimensionality impact vector database storage bills?
A 3,072-dimension vector requires twice the RAM and storage of a 1,536-dimension vector. Shortening vectors using Matryoshka representation learning (e.g. text-embedding-3 down to 512 dims) cuts Pinecone and Qdrant storage costs by up to 66% with under 2% retrieval accuracy loss.
Why is Voyage AI ranked higher than OpenAI on MTEB?
Voyage AI specializes exclusively in retrieval embeddings with tailored tokenizers and specialized domain models (code, finance, law) that outperform general-purpose models on long-context technical documentation.
When should I self-host embedding models like BGE?
If you process over 500 million tokens per month (e.g. daily web crawls or multi-terabyte data lakes), running BGE-large or Nomic-embed on a single $150/mo GPU instance is 10x cheaper than paying API rates.