Which AI model should I use?
A short learning video. Two engineers talk it through, with pictures and quiet music.
What this video shows
The answer comes first. Sara, a software engineer, says every AI app gives her a list of models. Daniel, a senior engineer, gives her the simple rule: small for simple jobs, big for hard jobs. The small model is fast and cheap; it can cost a tenth, or even a hundredth, of the big one. The big model is slower, but it thinks deeper.
Then one example. Sara needs to sort a thousand customer emails. That is a simple job, so the small model does it fast and gets it right. The big model makes exactly the same three piles, only later, and the bill is much longer. A hard job is different: for a big code change across many files, call the big one, because on long, hard work it does much better. Many models also have a thinking setting. On hard problems it gives better answers, but you wait longer and you pay for the thinking.
The last part is what to do. Start with the small one and test it on your real work. If it is good enough, keep it: you save time and money. If it is not, step up to the bigger model. The big AI labs give the same advice. Picking the right size for each step is a large part of how I keep the AI apps I build for daily work fast and affordable to run.
What you will learn
- What the model menu means: the same company offers a small, fast model and a big, slower one.
- That the small one is often just as good for simple, everyday jobs, and it costs much less.
- When the big one is worth it: long, hard work, such as a big code change across many files.
- What the thinking setting does: the model spends more time on a hard problem, so you wait longer and pay more.
- The simple rule the AI labs themselves give: start small, test it, and step up only when you need to.
The numbers behind the picture
The video says the numbers in words only. The exact numbers and their sources are here.
- 2025: Routing studies showed how much a business saves by sending easy requests to small models: one cut cost by 85% while keeping 95% of the quality, another by up to 60% with under 1% loss. OpenAI’s August 2025 launch added an automatic router that picked a model for each request. (LMSYS RouteLLM, July 2024; BEST-Route, ICML 2025; OpenAI, August 2025)
- 2026: Every big vendor now sells the same three choices: a small model, a big one, and a thinking dial. Per word, the small model costs about a tenth to a hundredth of the big one, depending on the vendor. On an independent speed test, small models write two to three times faster. On a test of hard science questions the gap is small (Anthropic’s own results: 85% for its small model, 91% for its middle and big ones); on hard, long coding tasks it is large (39% for the small model against 71% for the middle one). (Vendor price pages and model pages, read 8 October 2026; Artificial Analysis, October 2026; Anthropic model results)
- 2027, my guess: Small models keep getting better at simple work, so the rule in the video gets stronger, not weaker. More apps will choose the size for you; it will still pay to know that small is the right first choice. (my guess)
Questions people ask
Which ChatGPT, Claude or Gemini model should I use?
In October 2026 each one offers a small, fast model (for example Claude Haiku 5.5, GPT-6 Luna, Gemini 3.8 Flash), a bigger one (Claude Opus 5.5 or Fable 5.1, GPT-6 Astra, Gemini 3.1 Pro), and a thinking setting. For emails, summaries, sorting and short drafts, start with the small one. Anthropic’s own advice is to begin with its small model and “Upgrade only if necessary”; OpenAI’s is to “keep the lightest setting that meets your quality bar”.
Is the big model always better?
On hard, long jobs, yes, clearly. On simple jobs the difference is small and the big one is slower and costs more per word. On a test of hard science questions, Anthropic’s small model scored 85% and its middle and big ones 91%; on a hard coding test the small one scored 39% against 71% for the middle one.
What does thinking mode do?
It lets the model work through the problem before it answers. All three vendors count that extra thinking as extra words, so it takes longer and costs more. Use it for the hardest questions, not for everyday ones. At the highest setting the wait can be minutes.
Can I use a small model on my own computer?
Yes, some small models run on a laptop with no internet. The video “Can I run AI on my own computer?” shows how, and what you give up.
How does an AI app find the right page in my own files?
Usually with a vector database: it searches your files by meaning, not by exact words, and gives the model only the few pages it needs. A small model is often enough to answer from them. The video “What is a vector database and how does it work?” shows the five steps.
Further reading
- Anthropic, “Choosing a model”. platform.claude.com
- Anthropic, models overview (current lineup). platform.claude.com
- OpenAI, “Model selection” guide. developers.openai.com
- OpenAI, “Reasoning” guide (the thinking dial). developers.openai.com
- Google, Gemini models. ai.google.dev
- Google, Gemini “Thinking”. ai.google.dev
- Artificial Analysis, model leaderboard (speed and quality index). artificialanalysis.ai
- LMSYS, “RouteLLM”, July 2024 (sending easy requests to small models). lmsys.org
The AI apps I build use the small model for the everyday steps and the big one only where the job needs it, so they stay fast and affordable to run (how I can help). If your team wants that set up, tell me about the work.
Transcript
Every spoken line and every word shown on screen, in order. Two voices: Sara, a software engineer, and Daniel, a senior engineer. The two AI models in the pictures, a small white floating device and a tall dark android, do not speak. Quiet music plays under the voices.
Opening
- Logo: vai SoftLab
The question
Picture: a dark screen with the question in very large words.
Sara: Which AI model should I use?
On screen: Which AI model should I use?
Part one: the rule
Picture: Sara at her desk, unsure, with a long list of models on her laptop. Daniel stops by with a coffee.
Sara: Every AI app gives me a list of models.
Daniel: Here is the simple rule. Small for simple jobs. Big for hard jobs.
On screen: Small: simple jobs. Big: hard jobs.
Part two: the two models
Picture: the two AI models next to the desk, the small floating device and the tall android. Daniel points at them.
Daniel: The small model is fast and cheap. It can cost a tenth, or even a hundredth, of the big one.
Daniel: The big model is slower. But it thinks deeper.
On screen: Small: fast and cheap. Big: slower, deeper.
Part three: a simple job
Picture: the small model sorts a fast stream of emails into three neat piles. Sara watches.
Sara: I need to sort a thousand customer emails.
Daniel: That is a simple job. The small one does it fast, and gets it right.
On screen: Simple job: small model
Part four: the same job, big model
Picture: both models have made the same three piles; the big one finishes late. Sara holds up a very long bill.
Sara: And the big one?
Daniel: Same result. Just slower, and it costs much more.
On screen: Same result. Higher cost.
Part five: a hard job
Picture: the big model at a large wall screen with a complex system plan. Sara and Daniel study it.
Sara: What about a hard job? A big code change, across many files?
Daniel: Call the big one. On long, hard work, it does much better.
On screen: Hard job: big model
Part six: thinking
Picture: Sara turns a large dial on a console. Light lines glow around the big model’s head.
Daniel: Many models also have a thinking setting.
Sara: More thinking, better answers?
Daniel: On hard problems, yes. But you wait longer, and you pay for the thinking.
On screen: More thinking: better, but slower.
Part seven: start small
Picture: Sara checks the small model’s work on her screen against a checklist. Daniel nods.
Daniel: So, start with the small one. Test it on your real work.
Sara: And if it is good enough?
Daniel: Keep it. You save time and money.
On screen: Start small. Test it.
Part eight: step up
Picture: two wide steps, the small model on the low step and the big model on the top step. Sara walks up.
Sara: And if it is not good enough?
Daniel: Step up to the bigger model. The big AI labs give the same advice.
On screen: Not good enough? Step up.
To remember
Picture: a dark screen with large words.
Sara: Start small. Step up only when you need to.
On screen: Start small. Step up only when you need to.
Closing
- Logo: vai SoftLab
- vaisoftlab.com
Have an app idea, a stuck app, or a daily job you want an AI agent to do?
Write three lines. I reply within one working day with a plain answer.


