Talk to me

What is a vector database and how does it work?

A short learning video. A business owner and a software engineer talk it through, with pictures and quiet music.

Video. Nora, who runs a small renovation company, cannot find the warranty answer in two thousand files: her search for “leak” finds nothing. Sam, a software engineer, shows how a vector database searches by meaning, in five steps, and where it is not enough. 2 min 55 s, two voices, quiet music.

What this video shows

The question comes first. Nora runs a small home-renovation company. A customer calls: her shower is leaking, is it still under warranty? The answer is somewhere in two thousand files. Nora searches for “leak” and finds nothing, because the warranty paper never says leak. It says “water damage, from a failed seal”. Normal search looks for the same letters, so it misses different words with the same meaning.

Sam, a software engineer, explains the fix: a vector database searches by meaning, not by exact words. It works in five steps. One: every file is cut into small pieces, about a page each. Two: an AI model reads each piece and turns its meaning into a long list of numbers, often more than a thousand. That list is called a vector; think of it as an address on a map of meaning, where similar meanings sit close together. Three: the database keeps all these addresses in a smart index, like a library catalogue, so it stays fast. Four: the same AI model turns the question into an address on the same map. Five: the database takes the pieces closest to the question. The warranty paragraph is the closest, even though it never says leak.

Then the AI app gives just those few pieces to the AI model, which answers and shows the page it used. Sam is honest about the limit: for exact codes and names, such as a part number, keyword search is better, so good systems use both (hybrid search). And often no big new system is needed: normal databases such as Postgres can store vectors too. This is how I make the AI apps I build for daily work answer from a company’s own files.

What you will learn

  • Why normal search misses an answer that uses different words with the same meaning.
  • What a vector is: a long list of numbers that works like an address on a map of meaning.
  • The five steps: cut the files into pieces, turn each piece into an address, store them with an index, give the question an address, take the nearest pieces.
  • Why AI apps use a vector database: the AI model does not know your files and cannot read them all at once, so it gets the right few pages.
  • The honest limits: exact codes and names need keyword search too (hybrid search), and your own database can often do the job.

The numbers behind the picture

The video gives only a few numbers, in words. The exact numbers and their sources are here.

  • 2025: Big storage and database products added vector search as a normal feature. Amazon S3 Vectors became generally available on 2 December 2025: up to two billion vectors in one index, about 100 milliseconds for frequent searches, and up to 90% lower cost to upload, store and search. SQL Server 2025 shipped a vector data type the same month. (AWS, 2 December 2025; Microsoft, 18 November 2025)
  • 2026: The common embedding models turn each piece of text into 1,024 to 3,072 numbers by default: OpenAI text-embedding-3-small 1,536 and text-embedding-3-large 3,072, Google Gemini 3,072, Cohere Embed v5 2,048, Voyage 4 1,024. PostgreSQL, which can store vectors with the free pgvector extension, is used by 57.9% of developers in the 2026 Stack Overflow survey. Analysts put the vector database market at about 3 to 3.7 billion US dollars this year; their numbers differ because each firm counts the market differently. (Vendor documentation, read 10 October 2026; Stack Overflow Developer Survey 2026; Grand View Research and Meticulous Research, 2026)
  • 2027, my guess: Vector search becomes a normal part of the database a company already has, and mixing search by meaning with exact-word search becomes the default. Gartner predicts that by 2028, 80% of generative AI business apps will be built on companies’ existing data platforms. (my guess; Gartner, June 2025)

Questions people ask

What is the difference between a vector database and a normal database?

A normal database finds rows by exact values and words. A vector database stores the meaning of each piece of text as a list of numbers and finds the pieces closest in meaning to a question. In MongoDB’s words: “Unlike traditional full-text search which finds text matches, vector search finds vectors that are close to your search query in multi-dimensional space.” Today many normal databases can do both, for example PostgreSQL with pgvector, SQL Server 2025 and MongoDB Atlas.

What is an embedding?

It is the long list of numbers that an AI model makes from a piece of text, the “address on the map of meaning” in the video. OpenAI’s definition: “An embedding is a vector (list) of floating point numbers.” Texts with a similar meaning get similar numbers. One point to know: if you change the embedding model later, you must make the numbers again for all your files.

Is a vector database the same as RAG?

No. RAG is the whole method: find the right pages in your own files first, then let the AI answer from them. A vector database is the most common tool for the finding step. The video “What is RAG?” shows the whole method in one minute.

Why does it sometimes miss an exact part number or code?

Search by meaning is good at ideas, not at exact strings. Microsoft’s search documentation says product codes, special jargon, dates and people’s names “perform better with keyword search because it can identify exact matches.” That is why good systems run both searches and mix the results: hybrid search. In Anthropic’s own test, adding keyword search to search by meaning cut failed searches from 3.7% to 2.9%.

Do I need a separate vector database?

Often not. If you already use PostgreSQL, the free pgvector extension adds vector search to it, and many other databases now have it built in. For a very small set of files you may not need search at all: Anthropic notes that below about 500 pages you can give the whole text to the model.

Does the AI model behind it need to be the biggest one?

Not usually. Once the right few pages are found, answering from them is often a simple job. Start with a small model, test it on your real questions, and step up only if you need to. The video “Which AI model should I use?” shows this rule.

Further reading

  • MongoDB, “Vector Search overview” (search by meaning, in plain words). mongodb.com
  • Microsoft, “Integrated vector database” (what a vector database is, and using the one you already have). learn.microsoft.com
  • OpenAI, “Embeddings” guide. developers.openai.com
  • Google, Gemini “Embeddings”. ai.google.dev
  • pgvector, vector search for PostgreSQL. github.com
  • Microsoft, “Hybrid search overview” (meaning plus exact words). learn.microsoft.com
  • Anthropic, “Introducing Contextual Retrieval”, September 2024. anthropic.com
  • AWS, “Amazon S3 Vectors is now generally available”, December 2025. aws.amazon.com

The AI apps I build answer from a company’s own files: search by meaning plus exact-word search, often inside the database the team already has (how I can help). If your team wants that set up, tell me about the work.

Transcript

Every spoken line and every word shown on screen, in order. Two voices: Nora, who owns a small home-renovation company, and Sam, a software engineer. Quiet music plays under the voices.

Opening

  • Logo: vai SoftLab

The question

Picture: a dark screen with the question in very large words.

Nora: What is a vector database, and how does it work?

On screen: What is a vector database, and how does it work?

Part one: the problem

Picture: Nora’s small office, full of folders, binders and tile samples. She is on the phone, looking stressed.

Nora: A customer calls. Her shower is leaking. Is it still under warranty?

Nora: I know the answer is here. Somewhere, in two thousand files.

On screen: Two thousand files. One answer.

Picture: Nora types in a search box on her laptop; the results are empty. Sam walks in with two coffees.

Nora: I search for: leak. Nothing.

Sam: Because your warranty paper never says leak. It says: water damage, from a failed seal.

On screen: Search for “leak”: nothing found

Part two: why normal search misses it

Picture: Sam sits next to Nora and holds up two sheets: one with a water drop, one with a seal ring.

Sam: Normal search looks for the same letters. Different words with the same meaning? It misses them.

On screen: Normal search: same words only

Part three: search by meaning

Picture: Sam draws groups of dots on a glass whiteboard. Nora listens, coffee in hand.

Sam: A vector database fixes this. It searches by meaning, not by exact words.

Nora: How can a computer know what something means?

Sam: In five steps.

On screen: Vector database: search by meaning

Step one: cut the files into pieces

Picture: Sam cuts a stack of documents into small cards; Nora sorts the cards into rows.

Sam: Step one. Every file is cut into small pieces, about a page each. So later, we can find the exact part that answers a question.

On screen: Step 1: cut files into pieces

Step two: meaning becomes an address

Picture: the cards go into a navy machine; small points of light rise out of it. Nora watches, amazed.

Sam: Step two. An AI model reads each piece, and turns its meaning into a long list of numbers. Often more than a thousand.

Sam: That list is called a vector. Think of it as an address on a map of meaning.

On screen: Step 2: meaning becomes an address

Picture: a map of meaning on a dark screen. Labelled dots appear in three groups: wet floor tiles, warranty: failed seal, water damage and leaking pipe in one corner; invoice, price list and payment due in another; staff roster, new hire and holiday leave in a third.

Sam: Similar meanings get addresses close together.

Sam: Water damage, a leaking pipe, a failed seal: one corner. Invoices and prices: far away.

On screen: Near in meaning = near on the map

Step three: store them, with an index

Picture: a large room like a modern library, its shelves full of tiny points of light, with bright paths between them. Sam shows it to Nora.

Sam: Step three. The database keeps all these addresses in a smart index, like a library catalogue.

Sam: It is fast. One cloud service searches up to two billion pieces in about a tenth of a second.

On screen: Step 3: store them, with an index

Step four: the question gets an address

Picture: Nora types a question; it rises from her laptop as one gold point of light. Sam points at it.

Nora: So now I ask: is a leaking shower covered?

Sam: Step four. The same AI model turns your question into an address on the map.

On screen: Step 4: the question gets an address

Picture: the map again. The question “Is a leaking shower covered?” appears as a gold star and lands in the water corner.

Sam: Look where it lands. Right in the water corner.

On screen: It lands near the right pieces

Step five: take the nearest pieces

Picture: a circle around the star takes the nearest dots. “Warranty: failed seal” is the closest and lights up.

Sam: Step five. It takes the pieces closest to your question. The warranty paragraph is the closest.

Nora: Even though it never says leak.

Sam: Yes. Close in meaning. Close on the map.

On screen: Step 5: take the nearest pieces

The answer

Picture: Nora and Sam at the laptop. The screen shows a chat answer next to a page with one paragraph marked. Nora smiles.

Sam: Now the AI app gives just those pieces to the AI model. It answers, and shows the page it used.

Nora: Covered for five years. Here is the page.

On screen: The AI answers from your own files

Why AI apps use it

Picture: a huge, messy pile of folders; only three small glowing cards fly from it to the laptop.

Sam: This is why AI apps use them. The AI model does not know your files, and cannot read them all at once.

Sam: So the vector database hands it the right few pages.

On screen: The AI gets the right few pages

The limit: exact words

Picture: Nora holds up a small tap part with tiny marks. Sam holds a tablet where a gold beam and a blue beam join into one.

Nora: Does it also find an exact part number?

Sam: Not always. For exact codes and names, keyword search is better. Good systems use both. That is hybrid search.

On screen: Best: meaning + exact words

Do I need a big new system?

Picture: an ordinary server cupboard in the office with one new gold panel. Sam points at it; Nora looks relaxed.

Nora: Do I need a big new system for this?

Sam: Often not. Normal databases, like Postgres, can store vectors too. Start small.

On screen: Often your own database can do it

The happy ending

Picture: Nora on the phone at her desk, smiling, a page open on her laptop. Sam waves goodbye at the door.

Nora: Good news. Your shower is covered. I am sending you the page now.

On screen: Found by meaning, in seconds

To remember

Picture: a dark screen with large words.

Sam: A vector database finds what you mean, not just what you type.

On screen: It finds what you mean, not just what you type.

Closing

  • Logo: vai SoftLab
  • vaisoftlab.com

Have an app idea, a stuck app, or a daily job you want an AI agent to do?

Write three lines. I reply within one working day with a plain answer.