turbopuffer v3

turbopuffer v3 is a major overhaul of turbopuffer's storage architecture that will allow us to serve many more query plans at even greater scale. As of 09-30-2026, tpuf v3 passes 100% of CI, but as expected, performance has regressed from production turbopuffer. We're grinding query plans in public until v3 is ready for primetime. Read the log for updates.

Benchmark

p90 query latency amplification

v3/v2. below parity is faster than current production.

×

Benchmarks forthcoming

Date

Benchmarkv3/v2 (lower = better)
Full-text search
hot
126 x
Order by attribute
hot
97 x
Hybrid search
hot
57 x
Full-text search
cold
11 x
Order by attribute, filtered
hot
4.63 x
Vector search
hot
1 x
Vector search
cold
0.83 x
Vector search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector]
}
Vector search
cold
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector],
  "disable_cache": true
}
Full-text search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query]
}
Full-text search
cold
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query],
  "disable_cache": true
}
Order by attribute
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"]
}
Order by attribute, filtered
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"],
  "filters": ["category", "Eq", category]
}
Hybrid search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "queries": [
    {
      "top_k": 10,
      "rank_by": ["text", "BM25", query]
    },
    {
      "top_k": 10,
      "rank_by": ["vector", "ANN", vector]
    }
  ],
  "rerank_by": ["RRF"],
  "consistency": {"level": "eventual"}
}

Log

  1. RIP, vector databaseturbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.Dan Harrison (Engineer)

Methodology

Each point in the benchmark charts is turbopuffer v3 p90 latency divided by turbopuffer v2 p90 latency for the same workload, measured nightly. Below 1, v3 is faster.

Test namespaces live in gcp-us-central1 and hold 10 million documents. A c4a-highmem-32 in us-central1-c sends 8 queries per second for 10 minutes. Hot runs wait until the cache reaches a 100% hit ratio. Cold runs send queries with the cache disabled.

Vector workloads use 1024-dimensional Cohere embed-multilingual-v3 embeddings from the Cohere Wikipedia dataset. Full-text and hybrid workloads use MS MARCO, with 1024-dimensional Cohere embed-english-v3 embeddings for the vector query. Attribute ordering uses synthetic documents with a numeric popularity attribute and 100 category values.

Generic query syntax for each workload can be found under its corresponding chart.