turbopuffer v3

turbopuffer v3 is a major overhaul of turbopuffer's storage architecture that will allow us to serve many more query plans at even greater scale. As of 09-05-2026, 100% of CI passes on turbopuffer v3. As expected, performance regressed significantly from production, and we've been grinding query plans since. As we write up what changed, we'll update the charts so you can watch the lines move from day zero until they catch up to today. Read the dev log for updates.

Benchmark

p90 query latency amplification

v3/v2. below parity is faster than current production.

×

Benchmarks forthcoming

Date

Benchmarkv3/v2 (lower = better)
Full-text search
hot
126 x
Order by attribute
hot
97 x
Hybrid search
hot
57 x
Full-text search
cold
11 x
Order by attribute, filtered
hot
4.63 x
Vector search
hot
1 x
Vector search
cold*
0.83 x

*cold query latency is inherently noisy; we do not expect v3 to significantly reduce latency for cold vector queries

Vector search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector]
}
Vector search
cold*
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector]
}
// cache disabled
Full-text search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query]
}
Full-text search
cold
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query]
}
// cache disabled
Order by attribute
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"]
}
Order by attribute, filtered
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"],
  "filters": ["category", "Eq", category]
}
Hybrid search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "queries": [
    {
      "top_k": 10,
      "rank_by": ["text", "BM25", query]
    },
    {
      "top_k": 10,
      "rank_by": ["vector", "ANN", vector]
    }
  ],
  "rerank_by": ["RRF"],
  "consistency": {"level": "eventual"}
}

Dev Log

  1. RIP, vector databaseturbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.Dan Harrison (Engineer)

Methodology

Each point in the benchmark charts is turbopuffer v3 p90 latency divided by turbopuffer v2 p90 latency for the same workload, measured nightly. Below 1, v3 is faster.

Test namespaces live in gcp-us-central1 and hold 10 million documents. A c4a-highmem-32 in us-central1-c sends 8 queries per second for 10 minutes. Hot runs wait until the cache reaches a 100% hit ratio. Cold runs send queries with the cache disabled.

Vector workloads use 1024-dimensional Cohere embed-multilingual-v3 embeddings from the Cohere Wikipedia dataset. Full-text and hybrid workloads use MS MARCO, with 1024-dimensional Cohere embed-english-v3 embeddings for the vector query. Attribute ordering uses synthetic documents with a numeric popularity attribute and 100 category values.

Generic query syntax for each workload can be found under its corresponding chart.