turbopuffer v3

turbopuffer v3 is a major overhaul of turbopuffer's storage architecture that will allow us to serve many more query plans at even greater scale. As of 09-05-2026, 100% of CI passes on turbopuffer v3. As expected, performance regressed significantly from production, and we've been grinding query plans since. As we write up what changed, we'll update the charts so you can watch the lines move from day zero until they catch up to today. Read the dev log for updates.

Benchmark

p90 query latency amplification

v3/v2. below parity is faster than current production.

×

Benchmarks forthcoming

Date

Benchmarkv3/v2 (lower = better)
Order by attribute
hot
88 x
Full-text search
hot
19 x
Full-text search
cold
5.51 x
Hybrid search
hot
5.12 x
Order by attribute, filtered
hot
4.45 x
Vector search
hot
1 x
Vector search
cold*
0.71 x

*cold query latency is inherently noisy; we do not expect v3 to significantly reduce latency for cold vector queries

Vector search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector]
}
Vector search
cold*
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["vector", "ANN", vector]
}
// cache disabled
Full-text search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query]
}
Full-text search
cold
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["text", "BM25", query]
}
// cache disabled
Order by attribute
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"]
}
Order by attribute, filtered
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "top_k": 10,
  "rank_by": ["popularity", "desc"],
  "filters": ["category", "Eq", category]
}
Hybrid search
hot
×

Benchmarks forthcoming

Date

Query syntax
{
  "queries": [
    {
      "top_k": 10,
      "rank_by": ["text", "BM25", query]
    },
    {
      "top_k": 10,
      "rank_by": ["vector", "ANN", vector]
    }
  ],
  "rerank_by": ["RRF"],
  "consistency": {"level": "eventual"}
}

Dev Log

  1. On 09/08, we implemented batched reads in the v3 full-text engine, leading to a ~9x performance improvement. Sometimes, optimization is not about being clever, but simply avoiding mistakes. Postings are stored as columns, and we were reading them a row at a time. The fix here was obvious, we batched reads from storage. Notably, we're not even batching execution during scoring yet, that will come soon.

    Modern CPUs are happiest when their pipelines are full, and batching helps keep them full. Full-text search on tpuf v2 took advantage of this with a bespoke storage system and query engine to enable batching all the way through the pipeline, from reads to scoring.

    Unlike v2, v3 unifies every index that uses posting lists on the same internal storage format, so this fix benefits all of them, not just full-text.

    We're working on a deeper technical post about this, and we'll share it in the next few weeks.

    Xavier Denis (Engineer)
  2. RIP, vector databaseturbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.Dan Harrison (Engineer)

Methodology

Each point in the benchmark charts is turbopuffer v3 p90 latency divided by turbopuffer v2 p90 latency for the same workload, measured nightly. Below 1, v3 is faster.

Test namespaces live in gcp-us-central1 and hold 10 million documents. A c4a-highmem-32 in us-central1-c sends 8 queries per second for 10 minutes. Hot runs wait until the cache reaches a 100% hit ratio. Cold runs send queries with the cache disabled.

Vector workloads use 1024-dimensional Cohere embed-multilingual-v3 embeddings from the Cohere Wikipedia dataset. Full-text and hybrid workloads use MS MARCO, with 1024-dimensional Cohere embed-english-v3 embeddings for the vector query. Attribute ordering uses synthetic documents with a numeric popularity attribute and 100 category values.

Generic query syntax for each workload can be found under its corresponding chart.