turbopuffer v3
turbopuffer v3 is a major overhaul of turbopuffer's storage architecture that will allow us to serve many more query plans at even greater scale. As of 09-30-2026, tpuf v3 passes 100% of CI, but as expected, performance has regressed from production turbopuffer. We're grinding query plans in public until v3 is ready for primetime. Read the log for updates.
Benchmark
v3/v2. below parity is faster than current production.
Benchmarks forthcoming
Date
| Benchmark | v3 p90 (ms) | v2 p90 (ms) | v3/v2 (lower = better) | |
|---|---|---|---|---|
Full-text search hot | 2,264 | 18 | 126 x | |
Order by attribute hot | 1,642 | 17 | 97 x | |
Hybrid search hot | 2,289 | 40 | 57 x | |
Full-text search cold | 4,086 | 381 | 11 x | |
Order by attribute, filtered hot | 1,070 | 231 | 4.63 x | |
Vector search hot | 17 | 17 | 1 x | |
Vector search cold | 1,499 | 1,806 | 0.83 x |
Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector],
"disable_cache": true
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["text", "BM25", query]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["text", "BM25", query],
"disable_cache": true
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["popularity", "desc"]
}Benchmarks forthcoming
Date
Query syntax
{
"top_k": 10,
"rank_by": ["popularity", "desc"],
"filters": ["category", "Eq", category]
}Benchmarks forthcoming
Date
Query syntax
{
"queries": [
{
"top_k": 10,
"rank_by": ["text", "BM25", query]
},
{
"top_k": 10,
"rank_by": ["vector", "ANN", vector]
}
],
"rerank_by": ["RRF"],
"consistency": {"level": "eventual"}
}Log
Methodology
Each point in the benchmark charts is turbopuffer v3 p90 latency divided by turbopuffer v2 p90 latency for the same workload, measured nightly. Below 1, v3 is faster.
Test namespaces live in gcp-us-central1 and hold 10 million documents. A c4a-highmem-32 in us-central1-c sends 8 queries per second for 10 minutes. Hot runs wait until the cache reaches a 100% hit ratio. Cold runs send queries with the cache disabled.
Vector workloads use 1024-dimensional Cohere embed-multilingual-v3 embeddings from the Cohere Wikipedia dataset. Full-text and hybrid workloads use MS MARCO, with 1024-dimensional Cohere embed-english-v3 embeddings for the vector query. Attribute ordering uses synthetic documents with a numeric popularity attribute and 100 category values.
Generic query syntax for each workload can be found under its corresponding chart.