Softwr
turbopuffer logo

turbopuffer

Closed-source vector and full-text search service built directly on object storage, with cold queries measured in seconds rather than milliseconds.

As of 30 August 2026, turbopuffer starts at $16/month. turbopuffer keeps indexes on object storage and caches them on memory and SSD, so idle data costs storage prices rather than RAM prices. Softwr lists it under Databases.

Overview

What turbopuffer does

turbopuffer is a hosted search database for vector and full-text retrieval. Its architecture inverts the usual design: instead of holding indexes in memory across a fleet of always-on nodes, it keeps the authoritative state in object storage and treats memory and local SSD purely as caches. Data is organised into namespaces, which are the unit of isolation, tenancy and throughput, and it supports dense and sparse vectors, BM25 full-text search, attribute filtering and hybrid queries. Published limits are unusually explicit: up to 128 billion documents and 256 TB per namespace, 500 million documents per shard, 64 MiB per document, 10,752 dimensions for dense vectors, roughly 10,000 writes per second and 32 MB/s per namespace, and 5,000 or more queries per second per namespace. It is offered multi-tenant by default, and also as single-tenant or BYOC deployments in a customer's own cloud. The commercial argument follows directly from the architecture. In a memory-resident vector database, cost scales with the total size of the corpus whether it is being queried or not, which is punishing for the common shape of a multi-tenant product where most customers' data is idle most of the time. Putting the index on object storage makes cold data nearly free to keep and pushes the cost onto queries and caching, which is why turbopuffer has been adopted by products with very large numbers of small, mostly dormant per-customer indexes. The published numbers are what makes this a design decision rather than a marketing one: cold queries have a p90 around 1,214 ms on a million documents, warm queries perform like an in-memory engine, and upserts have a p90 around 248 ms for a 512 KB batch because writes go to object storage directly. The buyers are engineering teams at products with per-customer search indexes and a long tail of inactive tenants. The trade-offs are latency and consistency, both documented honestly. Queries are eventually consistent by default, with the vendor stating that more than 99.8% return consistent data and that after more than 128 MiB of outstanding writes further writes remain invisible until indexed, which can take tens of seconds on a small namespace and tens of minutes on a large one. Strong consistency is available at a performance cost. And it is closed source with no community edition, so a self-hosted deployment is a commercial negotiation rather than a container image, and your contingency if the company changes direction is re-indexing your corpus somewhere else.

What people use it for

  • A product with one search index per customer and thousands of customers, most of whose data is idle on any given day
  • Very large corpora where holding every vector in memory is the dominant cost and occasional cold-query latency is acceptable
  • Hybrid retrieval combining BM25 and vector search where running and synchronising two separate systems is the problem being solved
  • Retrieval for agent and assistant products where indexes are created and destroyed frequently and per-index overhead must be near zero

The honest half

Where it falls short

Concrete and checkable, so you can decide whether any of them matter to you. This is the half of a review a vendor will not write about turbopuffer.

  • A cold namespace pays object storage latency on the first query, with a documented p90 around 1,214 ms on a million documents, so any interactive search box needs the data kept warm or the user waits about a second.
  • Queries are eventually consistent by default, and after roughly 128 MiB of outstanding writes new data is invisible until indexed, which the vendor puts at tens of seconds for small namespaces and tens of minutes for large ones, so a bulk re-index is not immediately queryable.
  • It is closed source with no community edition, so single-tenant or bring-your-own-cloud deployment is a commercial negotiation rather than a deployment choice, and there is no path to running it yourself if the relationship ends.
  • Per-namespace ceilings, roughly 10,000 writes per second, 32 MB/s and 500 million documents per shard, mean a single enormous index has to be sharded across namespaces by your application rather than by the service.
  • It is a search engine, not a database: there are no joins, no cross-document transactions and no SQL, so it sits beside a primary datastore and keeping the two in step is work that belongs to you.

Cross-shopped

What people choose instead of turbopuffer

Each pairing was judged by two reviewers asking whether a buyer would genuinely weigh the two against each other. The ones that failed were deleted rather than published.

Pricing

What turbopuffer costs

Taken from the vendor's own pricing page. Prices move, so check before you buy.

Launch

$16 /mo

  • All database features
  • Multi-tenancy deployment
  • SOC2 & GDPR-ready DPA

Scale

$256 /mo

  • Everything in Launch
  • HIPAA-ready BAA
  • Single Sign-On (SSO)
  • Audit Logs (+$128/mo for Audit Log Streams)
  • IP allowlisting
  • Private Slack support (8-5 hours)

Enterprise

$4,096 /mo

  • Everything in Scale
  • Single-tenancy & BYOC deployment options
  • Private networking
  • CMEK (Per Namespace)
  • 24/7 support with SLA
  • 99.95% uptime SLA

Capabilities

Features

  • Object storage architecture

    Authoritative index state lives in S3-class storage, with memory and SSD used as caches

  • Namespaces

    The unit of tenancy, isolation and throughput, designed for very large numbers of per-customer indexes

  • Vector search

    Dense and sparse vectors with up to 10,752 dimensions and reported recall at 10 between 90 and 100 per cent

  • Full-text search

    BM25 keyword search in the same system as vector retrieval, so hybrid queries need no second service

  • Attribute filtering

    Filters on document attributes applied within the query rather than after retrieval

  • Documented limits

    Published ceilings for namespace size, document size, write throughput and query rate rather than vague scalability claims

  • Configurable consistency

    Eventual consistency by default with a strong consistency option per query when correctness matters more than latency

  • Durable writes

    Writes are committed to object storage before the API returns

  • Multi-query requests

    Up to 16 queries batched into a single request

  • Deployment options

    Multi-tenant by default, with single-tenant and bring-your-own-cloud deployments available

Answered, with sources

Questions people ask

Each answer names the page it came from, so you can check it rather than take our word for it.

Can I self-host turbopuffer?

There is no open source or community edition. Single-tenant and bring-your-own-cloud deployments exist as commercial arrangements, but there is no way to run it independently of the vendor.

How fast is it really?

Warm queries perform comparably to in-memory search engines. Cold queries, where data is not cached, have a documented p90 around 1,214 ms on a million documents. Write p90 is around 248 ms for a 512 KB upsert because writes go straight to object storage.

Is it consistent?

Eventually consistent by default, with the vendor reporting that over 99.8% of queries return consistent data. Strong consistency can be requested per query at a latency cost. Large write bursts have a longer visibility delay while indexing catches up.

What is it best at?

Large numbers of namespaces where most are idle. The architecture makes cold data cheap to keep, which is exactly the shape of a multi-tenant product with a long tail of inactive customers.

What are the hard limits?

Up to 128 billion documents and 256 TB per namespace, 500 million documents per shard, 64 MiB per document, 10,752 dense vector dimensions, roughly 10,000 writes per second per namespace and a maximum result set of 10,000.

Share

Keep looking

Where to go from turbopuffer

Best Databases software for

Compare turbopuffer with

Other Databases software

  • The world's most advanced open source relational database

    Free plan11 researched notes
  • Create apps that perfectly fit your team's needs

    Free, then $20/month per editor11 researched notes
  • The cloud-native distributed SQL database

    Free plan14 researched notes
  • MySQL and PostgreSQL-compatible relational database built for the cloud

    Free plan12 researched notes
  • Apache 2.0 vector and full-text search engine that runs as an embedded library, a single server or a distributed cloud service.

    Free plan14 researched notes
  • Google Cloud's serverless analytical warehouse, billed either by bytes scanned per query or by reserved compute slots.

    Free plan13 researched notes
  • SQL query engine and lakehouse layer over Iceberg tables in object storage

    Free, then $0.2/hour12 researched notes
  • MIT-licensed analytical SQL database that runs inside your process, with no server, no dependencies and one writer at a time.

    Free plan14 researched notes
  • Open-source typo-tolerant search engine as an Algolia alternative

    Open source9 researched notes
  • High-performance Redis-compatible in-memory datastore with 25x better throughput

    Free plan10 researched notes
  • Embedded retrieval library over the Apache 2.0 Lance columnar format, with proprietary Cloud and Enterprise tiers for serving at scale.

    Pricing on request14 researched notes
  • Database caching and optimization that reduces infrastructure costs 30-70%

    Free plan10 researched notes
  • Open-source in-memory data store forked from Redis

    Open source7 researched notes
  • MPP analytical database with a MySQL wire protocol and sub-second aggregation on wide tables

    Open source12 researched notes
  • Multi-model database for graph, document, and search

    Free plan9 researched notes
  • Industrial process historian with published per-tag pricing and no client licence fees

    From $3,000/yr12 researched notes
  • Open-source typo-tolerant search engine as an Algolia alternative

    Open source9 researched notes
  • Enterprise search platform built on Apache Lucene

    Open source10 researched notes

Softwr does not host reviews and shows no star rating for turbopuffer, because a rating we did not collect is not ours to publish. What is here is the pricing and platform detail from the vendor’s own pages, limitations we could state concretely, and alternatives a reviewer confirmed people weigh against it. Tell us if any of it is wrong.

More on turbopuffer

Best Databases software alternatives

Industrial process historian with published per-tag pricing and no client licence fees

Per tag band per year

NetApp-owned managed service for Cassandra, Kafka, OpenSearch, PostgreSQL and Cadence with a bring-your-own-cloud model

quote

Erlang MQTT broker for large IoT fleets, relicensed to BSL with production free use limited to one node

Per month by connection and session volume

Attribute-based access control and masking applied inside Snowflake, Databricks and BigQuery

quote

MPP analytical database with a MySQL wire protocol and sub-second aggregation on wide tables

Open source, no licence fee

Streaming database that maintains incremental materialised views in SQL instead of Flink jobs

Per RisingWave Unit hour

Centralised data access governance from the creators of Apache Ranger, now rebranding as Trust3 AI

quote

The Meta-lineage distributed SQL query engine, distinct from the Trino fork

Open source, no licence fee

Compare turbopuffer with alternatives