AI & MLDeep
Intermediate

Vector Databases Explained

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Static Playbook
Static guide & reference tables

Why regular databases can't do this

A traditional database is built to answer exact or range queries: find the row where id = 42, or all rows where price < 100. But once text, images, or audio get converted into embeddings — dense vectors of hundreds or thousands of floating-point numbers, as covered in the embeddings lesson — the question you actually want to ask changes: *which of these million vectors is most similar to this query vector?*

That's a nearest neighbor search problem, and it's fundamentally different from indexed lookups. A vector database is purpose-built infrastructure for storing embeddings at scale and answering similarity queries fast — this is the backbone that makes retrieval-augmented generation (RAG) practical at production scale.

The brute-force baseline — and why it doesn't scale

The simplest similarity search is exhaustive: compute the cosine similarity (or Euclidean/dot-product distance) between the query vector and every single vector in the database, then sort and return the top matches. This is exact — it always finds the true nearest neighbors — but it's $O(n)$ per query, where $n$ is the number of stored vectors.

For a few thousand vectors, brute force is fine and genuinely used in practice. But for a production RAG system indexing millions of document chunks, comparing every query against every stored vector in real time becomes too slow. This is the exact problem vector databases exist to solve.

Approximate Nearest Neighbor (ANN) search

Vector databases trade a small amount of accuracy for massive speed gains using Approximate Nearest Neighbor (ANN) algorithms. Instead of guaranteeing the exact top-k matches, ANN indexes are built so that queries return results that are *very likely* the true nearest neighbors — typically 95–99%+ recall — in a fraction of the time brute force would take, often sub-linear or even near-constant relative to database size.

The most widely deployed ANN algorithm today is HNSW (Hierarchical Navigable Small World), used internally by Pinecone, Weaviate, Qdrant, and pgvector's HNSW index type.

How HNSW works

HNSW builds a multi-layer graph structure over the stored vectors. The top layer contains a small number of nodes with long-range connections, acting like a highway system; each layer below has progressively more nodes and shorter, denser connections, down to the bottom layer which contains every vector.

A search starts at an entry point in the top layer and greedily moves to whichever neighbor is closest to the query, layer by layer, narrowing in — similar in spirit to how you'd navigate a city by first choosing the right highway, then the right arterial road, then the right street. This gives HNSW logarithmic-ish search complexity instead of linear, which is why it comfortably handles tens of millions of vectors with millisecond query latency.

text

Note

Every ANN index exposes parameters that trade accuracy for latency — HNSW's key ones are ef_construction (build-time search breadth) and ef_search (query-time search breadth). Higher values mean the algorithm explores more of the graph before returning results, improving recall at the cost of speed. Most vector database defaults land around 95–99% recall versus brute force, which is imperceptible for RAG but worth tuning if your use case is highly precision-sensitive.

Pinecone vs. Weaviate vs. pgvector

Pinecone is a fully managed, cloud-native vector database — no infrastructure to run, built specifically for vector search with features like metadata filtering, namespaces, and serverless scaling. It's a common default for teams that want vector search without managing servers.

Weaviate is open-source and can be self-hosted or used as a managed cloud service. It supports hybrid search (combining vector similarity with traditional keyword/BM25 search), built-in modules for generating embeddings, and a GraphQL query interface.

pgvector is a PostgreSQL extension that adds a vector data type and ANN indexing (IVFFlat and HNSW) directly inside Postgres. Its appeal is simplicity: if your application already uses Postgres, you get vector search without adding a new database to your stack, and you can join vector queries against your regular relational data in the same query.

Vector database options compared

OptionHostingBest forNotable trade-off
PineconeFully managed onlyTeams wanting zero infra overheadVendor lock-in, usage-based cost
WeaviateSelf-hosted or managedHybrid search, built-in embedding modulesMore operational surface to manage
pgvectorSelf-hosted (Postgres extension)Teams already on Postgres, relational + vector joinsScales less smoothly past tens of millions of vectors
Brute-force (e.g. NumPy)In-processSmall datasets, prototyping, exact recallLinear scan — doesn't scale

How vector databases power RAG

In a RAG pipeline, documents are split into chunks, each chunk is embedded into a vector, and all vectors are stored in a vector database alongside metadata (source document, page number, timestamps). At query time, the user's question is embedded with the same model, the vector database returns the top-k most similar chunks via ANN search, and those chunks are inserted into the LLM's prompt as context before it generates an answer.

The vector database is the retrieval engine that makes this fast enough to run on every user query — without it, RAG would require re-computing similarity against the entire document corpus on every single request, which is untenable at any meaningful scale or document count.

What's next

Vector databases are the infrastructure; RAG is the pattern built on top of them. If you haven't already, see the RAG lesson for the end-to-end pipeline, and Cosine Similarity for the math underlying every similarity comparison discussed here.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.