turbopuffer v3
turbopuffer v3 is a major overhaul of turbopuffer's storage architecture that will allow us to serve many more query plans at even greater scale. As of 09-05-2026, 100% of CI passes on turbopuffer v3. As expected, performance regressed significantly from production, and we've been grinding query plans since. As we write up what changed, we'll update the charts so you can watch the lines move from day zero until they catch up to today. Read the dev log for updates.
Benchmark
v3/v2. below parity is faster than current production.
Benchmarks forthcoming
Date
| Benchmark | v3 p90 (ms) | v2 p90 (ms) | v3/v2 (lower = better) | |
|---|---|---|---|---|
Full-text search hot | 2,264 | 18 | 126 x | |
Order by attribute hot | 1,642 | 17 | 97 x | |
Hybrid search hot | 2,289 | 40 | 57 x | |
Full-text search cold | 4,086 | 381 | 11 x | |
Order by attribute, filtered hot | 1,070 | 231 | 4.63 x | |
Vector search hot | 17 | 17 | 1 x | |
Vector search cold* | 1,499 | 1,806 | 0.83 x |
*cold query latency is inherently noisy; we do not expect v3 to significantly reduce latency for cold vector queries
Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector]
}
// cache disabledBenchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["text", "BM25", query]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["text", "BM25", query]
}
// cache disabledBenchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["popularity", "desc"]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["popularity", "desc"],
"filters": ["category", "Eq", category]
}Benchmarks forthcoming
Date
Query syntax
{
"queries": [
{
"top_k": 10,
"rank_by": ["text", "BM25", query]
},
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector]
}
],
"rerank_by": ["RRF"],
"consistency": {"level": "eventual"}
}Dev Log
Methodology
Each point in the benchmark charts is turbopuffer v3 p90 latency divided by turbopuffer v2 p90 latency for the same workload, measured nightly. Below 1, v3 is faster.
Test namespaces live in gcp-us-central1 and hold 10 million documents. A c4a-highmem-32 in us-central1-c sends 8 queries per second for 10 minutes. Hot runs wait until the cache reaches a 100% hit ratio. Cold runs send queries with the cache disabled.
Vector workloads use 1024-dimensional Cohere embed-multilingual-v3 embeddings from the Cohere Wikipedia dataset. Full-text and hybrid workloads use MS MARCO, with 1024-dimensional Cohere embed-english-v3 embeddings for the vector query. Attribute ordering uses synthetic documents with a numeric popularity attribute and 100 category values.
Generic query syntax for each workload can be found under its corresponding chart.