Turbopuffer Shifts to Secondary Vector Indexing in Version 3
Turbopuffer announced a major architectural shift in version 3, moving from primary vector indexing to secondary vector indexing. This transition resolves severe write amplification caused by keying data directly on Approximate Nearest Neighbor (ANN) addresses. As AI workloads expand, primary vector indexing causes excessive reindexing overhead and storage rewrites during updates. Decoupling vector indexes from raw data positioning aligns vector search engines with traditional database designs to improve write throughput and scalability. In version 3, raw rows reside in stable data fragments while the ANN index points to them secondarily, preventing row relocation whenever the vector graph rebalances. The technical trade-off shifts Turbopuffer from a PostgreSQL-like lookup pattern to a MySQL-like secondary index pattern, trading slightly higher query lookup latency for vastly reduced write amplification.
## BACKGROUND
Vector databases use Approximate Nearest Neighbor (ANN) indexing to enable rapid similarity search across high-dimensional embeddings generated by AI models. Write amplification occurs when a single physical database update triggers multiple underlying writes to storage, often due to index compaction and graph rebalancing.