Talk to me

Why does AI sometimes make things up?

A short learning video. No voice, only titles and screens, with music.

Video. AI guesses the next word that sounds right and does not check it. One made-up answer about a bakery that does not exist, then three ways to stop it. 1 min 6 s, music only, no voice.

What this video shows

The answer comes first, in two lines. AI guesses the next word that sounds right. It does not check if it is true.

Then one example. Someone asks a chat: “What is the Blue Door bakery’s best seller?” The answer comes back calm and sure: “Its best seller is the sourdough loaf.” Beside the chat, the outline of a shop called Blue Door appears, and then fades away to nothing. There is no such bakery. The AI never saw it. It made the answer up, and the answer looks exactly as sure as a true one would.

The last part is what to do about it, and it is how I build every small AI app for daily work. Give it the source: a document called Menu goes into the chat. Ask where it found the answer: a gold line runs from the answer back to the line in the menu it came from. And let it say “I don’t know”: asked about something the menu does not cover, it answers “The menu does not say.”

What you will learn

  • Why AI makes things up, in one sentence: it guesses the next word that sounds right, and it does not check if it is true.
  • Why a made-up answer looks just like a true one: the guess is written in the same calm, sure way.
  • The word you will see for this everywhere: hallucination. The film does not use it. It means an answer that sounds sure and is made up.
  • Three ways to stop it: give it the source, ask where it found the answer, and let it say “I don’t know”.

The numbers behind the picture

Nothing with a digit is on screen. The numbers belong here.

  • 2025: on hard fact questions with nothing to look at, OpenAI’s o3 model answered wrongly about half the time in April, and GPT-5 Thinking two times in five in August. When the model was given the document to work from, the best models made something up in under one summary in a hundred (October). A database of court decisions involving AI-made-up citations held about 719 cases by December. (OpenAI system cards, April and August 2025; Vectara leaderboard, October 2025; Charlotin database via Social Science Space, January 2026)
  • 2026: with nothing to look at, the newest Anthropic models still give a confident wrong answer about one time in five on hard questions (June). In real ChatGPT use with web search on, OpenAI reports fewer than one answer in a hundred with a major error (from December 2025). The court-case database reached 2,149 cases on 5 October 2026, 112 of them in Australia. In October 2025 Deloitte repaid the final instalment of an Australian government report after references in it turned out to be made up. (Anthropic system card, June 2026; OpenAI system cards, December 2025 to September 2026; Charlotin database, October 2026; City AM, October 2025)
  • 2027, my guess: the rates keep falling on every official test, and no lab says they reach zero. The number of real incidents rises with use. So the three steps in the film stay the same, and for anything that matters a person still checks.

The pattern in both years is the one the film shows: with nothing to look at, the guess is often wrong; with the source in front of it, made-up answers become rare.

Questions people ask

Why does AI make things up?

Because it writes by guessing the next word that sounds right, and nothing in that step checks whether the sentence is true. When it knows the subject, the likely words are usually the right ones. When it does not, it still writes likely words, and they come out sounding just as sure. Anthropic’s own line is short: models “are always supposed to give a guess for the next word”.

Is the AI lying?

No. Lying needs an intent. The model has no idea whether the bakery exists; it has a sentence to finish, and it finishes it. That is why the wrong answer sounds exactly like a right one. The industry word is hallucination.

Do newer models still do it?

Less, but yes. On hard questions with nothing to look at, the best models of 2026 still give a confident wrong answer about one time in five. With the source in front of them, it becomes rare. That difference is the whole point of the last scene.

What can I do today?

The three steps in the film, in this order. Give it the document, page or data the answer should come from. Ask it to show where in the source it found the answer. Tell it that “I don’t know” is a good answer. Anthropic’s guidance for builders says the “I don’t know” permission alone “can drastically reduce false information”. For a bigger set of documents, the AI can look them up for you: that is the idea behind RAG, the topic of a coming video, “What is RAG?”. For anything that matters, a person still checks.

Further reading

Every small AI app I build does what the last scene shows: it answers from your own documents and data, it shows where the answer came from, and it is allowed to say it does not know. See how I can help. If your team wants answers it can check, tell me about the work.

Transcript

Every line of text shown on screen, in order. Nothing is spoken. Each label stays at the bottom of the screen for at least four seconds.

Opening

  • Logo: vai SoftLab (small, no title line)

Scene one: the question

Alone on a dark screen, the biggest text in the film:

  • Why does AI sometimes make things up?

Scene two: the answer

A single line fills with faded grey words, one after another, each one dropping into place. Nothing checks them.

Label: AI guesses the next word that sounds right.

Label: It does not check if it is true.

Scene three: one example

A chat window with a small grey tag:

  • Chat

Typed, letter by letter, then sent:

  • What is the Blue Door bakery’s best seller?

The answer writes itself word by word, calm and sure:

  • Its best seller is the sourdough loaf.

Label: It answers. It sounds sure.

Beside the chat, the outline of a small shop appears with its sign, then fades away to nothing:

  • Blue Door

Label: There is no such bakery.

The answer card turns a little paler.

Label: It made it up.

Scene four: how to stop it

The chat is empty again. A hand drags a small gold document into it:

  • Menu

Label: Give it the source.

A typed line:

  • From this menu: what is the best seller?

A gold answer card appears, and a thin gold line runs from it to one lit row in the Menu document.

Label: Ask where it found it.

A second typed line:

  • If the menu does not say, say so.

A small grey card answers:

  • The menu does not say.

Label: Let it say “I don’t know”.

Scene five: to remember

Dark screen. Three lines, one by one:

  • Give it the source.
  • Ask where it found it.
  • Let it say “I don’t know”.

Closing

  • Logo: vai SoftLab
  • vaisoftlab.com