Smol Beep Boops
Hello folks. On August 22nd, near the bottom of issue 18, I wrote this:
And the tiny models? Still coming, and still headed for your light bulb. Next week, I promise, unless Anthropic goes and cures something else on a Tuesday.
That was two weeks ago. In my defence, Anthropic did not cure anything on a Tuesday, so I have no excuse whatsoever and we are just going to move past it. Today, at last: the tiny models. The little beep boops 🤖 that are about to end up inside everything you own.
Quick throat-clear on the week, because it was a big one. OpenAI shipped GPT-6 Astra on Wednesday, the frontier model they had been trailing since August, and it landed with an asterisk I have not seen before: OpenAI's own evaluation put it at the Critical cybersecurity threshold, the top rung of their Preparedness Framework, and the first model they have ever placed there. I want to spend proper time with it. Not this week though, because this week we are going the other direction entirely.
|
◆ Previously · Issue 17 In The 3 Comma Club we went one shelf down from the frontier, to the models you could realistically own and run on a good desktop or a Mac Studio. We also found that a model named “Small” was 119 billion parameters. Today we go down another shelf, to the ones that fit on a phone, and the naming gets no better. |
So how tiny is tiny
There is no committee, no standard, nobody signing off. The best answer I found is that small has stopped being a number and started being a place. A model is small if it runs on a thing you already own, fast enough to be worth talking to. Which means the line moves every time somebody ships a better phone.
So here is my working line for today, and you should feel free to argue with it. Under about a billion parameters, and small enough that the whole thing is a file you would not think twice about downloading. Google's Gemma 3 270M is 291 megabytes. That is a couple of podcast episodes. That is smaller than the photos on your phone from one decent holiday.
|
◆ Concept · A tiny model A language model small enough to live on the device instead of in a datacentre. No account, no internet, no per-question cost, and nothing you say gets sent anywhere. The trade is that it knows less and thinks worse. The bet is that for most of what you actually ask a machine to do, that is fine. |
The efficiency is genuinely silly. Google measured their 270M model on a Pixel 9 Pro and reported that it burned 0.75% of the battery across 25 conversations. And the way they explain who it is for is the least corporate sentence I have read from a big lab this year:
You wouldn't use a sledgehammer to hang a picture frame. The same principle applies to building with AI.
It does not need to know the capital of France
Here is the idea that makes these things interesting. A tiny model is not trying to be a smaller encyclopedia. Its job is to sit between your mouth and a switch. You say it's hot in here, and it turns that into set the thermostat to 72. Nobody is going to route a trivia question to the model living in your thermostat, so it does not need trivia. It needs to take a sloppy human sentence and produce one correct action.
The labs say this out loud, in their own documentation, which surprised me. Google shipped a 270M model called FunctionGemma whose card says flatly that it “is not intended for use as a direct dialogue model” and that models like it “are not knowledge bases”. Apple says the same thing about the roughly 3 billion parameter model sitting on your iPhone right now: it is “not designed to be a chatbot for general world knowledge”. Two of the largest companies alive, both saying please stop asking it things.
So I downloaded two of them and asked. Not on a workstation. On the cheapest computer in this house, a 2018 Pentium Silver mini-PC, four slow cores, no graphics card, the sort of machine you find in a drawer. The results are below and I have not tidied them up.
Look at what happened there, because it is the opposite of what I expected. The 291MB model knew the capital of France. It knew who wrote Moby-Dick. Then it told me 17 times 4 is 17, and that 12 times 3 is 3. It has the trivia and it cannot do the times tables (I did check, twice). And when I handed it three tools and two worked examples, it copied the first example back to me three times in a row, including for the laundry.
Then I went one rung up. Qwen 3 at 0.6 billion parameters, 522 megabytes, still free, still on the drawer computer. I typed it's hot in here and it replied with set_thermostat, 72 degrees. I did not feed it that number. I had literally written the 72-degree example into my notes for this issue a week earlier as a made-up illustration, and then the free half-gigabyte model did it. Kitchen too dark, lights on, correct. Then I said the towels were filthy and it announced that none of its tools were relevant, with a working start_wash sitting right there in the list.
Two out of three, on a machine worth less than a nice dinner, running a file you could email to someone. That is the whole story of this issue in one line. Not good enough to trust yet. Startlingly close for the money.
We have done this before
Cars used to have a handle you turned. Then somebody put a tiny motor and a tiny chip in the door, and now the window goes down because you touched a button. Nobody calls that a computer. Nobody thinks about it at all. There are somewhere between thirty and a hundred of those little chips in an ordinary car and you have never met one of them.
The last interface that ever needed instructions. Photo by Santeri Viinamäki, CC BY-SA 4.0.
That is the shape I think this takes. The tiny model becomes the window motor of the next twenty years: boring, everywhere, and invisible. You will not open an app to talk to your dishwasher. You will just say something at it while walking past, and it will either do the thing or it will not, and if it does you will stop noticing within a week.
What I find genuinely exciting is the tuning. Right now a machine gives you dials, and the dials assume you know what you want in the machine's units. Ninety degrees, 800 spin, programme 7. Ask me why my coffee tastes different today than yesterday and I have no idea, but the machine does know: it brewed five degrees cooler because the room was cold. A model in the middle could just tell you that. Or skip telling you and fix the next cup.
Twelve programmes, seven temperatures, seven spin speeds, four mystery icons. Quick, wash a jumper. Photo by Belchers Albert, CC BY-SA 3.0.
Yes, this is just Siri. That is the point.
You have been thinking it for four paragraphs, so let us say it. Talking at your house to make things happen is Siri, or Alexa, or the one in your car that mishears you. And the comparison is fair, and it is the most encouraging thing in this whole issue.
Because look at what Siri cost. It came out of CALO, a five-year DARPA programme run out of SRI International with $150 million behind it and more than 300 researchers across 22 institutions, which SRI still describes as the largest AI project in US history. Siri was spun out as a company in 2007, raised another $24 million, got bought by Apple in April 2010, and shipped on the iPhone 4S in October 2011. A defence research budget, three hundred researchers, then the richest company on earth. And the product is still the punchline of the category.
That was the price of entry. An expert team, hand-written intents, years of labelling what a sentence was supposed to mean, and you still ended up with something that sets a timer and gives up. This week I did a worse version of it in an afternoon, for free, on a computer from a drawer. The floor fell out of the whole category, and everybody who makes a physical object now gets to try.
Which is happening this week, in Berlin
While I was writing this, the appliance industry was at IFA in Berlin doing exactly this, which was rude of them but convenient for me. LG's headline appliance is a fridge with a language model behind the microphone, so you can ask it how to store something instead of learning which drawer means what. Samsung got there in March, putting an LLM-based Bixby into its 2026 washers, air conditioners, robot vacuums and water purifiers. The demo they keep showing people is a person saying “I'm going to wash jeans, so set the right course” and the machine picking the denim programme. Be prepared to talk to your washing machine, indeed.
The part you will never see is the routing, and it is the part that makes the whole thing hold together. Most of what you say is easy and gets handled where you stand. Some of it goes to a bigger model somewhere in the house. A little of it goes out to a frontier model that somebody bills you for. You will not know which, and mostly you will not care, right up until the day you care very much.
And now the $499 toothbrush
I want to leave you with the product that I think is the first real one of these, or at least the first one brave enough to put the price on the box. On Tuesday Dyson launched the CameraJet, a $499 toothbrush with a camera in it. Yes. Four hundred and ninety-nine dollars. It reads 28 images a second off your actual teeth, spots the gaps between them, and fires a jet of mouthrinse at one within about a tenth of a second of seeing it. Dyson says the thing was trained on 470,000 dental images, which is a sentence I have now typed and cannot untype.
Is it worth five hundred dollars? Almost certainly not (my toothbrush cost eleven). But it is the shape of the thing, and the shape is what I am watching. A cheap sensor, a model small enough to live inside the handle, and a device that adjusts to you instead of handing you a dial and wishing you luck. The first one is always ridiculous and overpriced. Then it is in everything and costs forty dollars and nobody remembers being impressed.
Below the fold I will get you running one of these yourself, which takes about ten minutes and costs nothing. It is genuinely a good time, mostly because of how badly the smallest one fails. Watching a model confidently tell you that 17 times 4 is 17 is worth the download on its own.
Until next week,
| ◆ Below the fold ◆ |
Ten minutes, no account, no card, nothing leaves your computer.
Run a tiny model tonight
Everything I did up there, you can do. You need Ollama, which is a free program that downloads models and talks to them. Mac, Windows and Linux. On Mac and Windows you grab the installer from the site like any other app. On Linux it is one line in a terminal:
|
curl -fsSL https://ollama.com/install.sh | sh |
Then open a terminal and type one more line. This downloads the 291MB model and starts talking to it. First run pulls the file, after that it is instant:
|
ollama run gemma3:270m |
That is it. You are now running an AI model on your own machine with the wifi off if you like. Say hello. Then go and be mean to it, because the fun of a model this small is finding the edges, and the edges are close. Things I would try:
- Some easy trivia. It will probably get it. This surprises everybody.
- Then two-digit multiplication. Watch it fall over. Ask it twice, it falls over the same way.
- Ask it to reply only in JSON, then see whether it obeys the format or just repeats your example back at you.
- Ask it something with a follow-up. Small models forget what you were talking about fast.
When that gets frustrating, and it will, go up one rung. ollama run qwen3:0.6b is 522MB and noticeably more competent, which is the one that got the thermostat right for me. If your machine is newer than mine (a low bar), try ollama run gemma3:1b at 815MB, or ollama run gemma3:4b at 3.3GB if you have 8GB of memory to spare. That last one is a real assistant and the difference will shock you.
|
◆ One honest warning These are not chatbots and judging them as chatbots will just make you sad. The 270M one exists to be trained on one narrow job, which is why Google's own version of it for tool calling says it “is not intended for use as a direct dialogue model”. Out of the box you are meeting the raw material, not the product. Enjoy it as an engine you have been handed rather than a car you have been sold. |
When you want more of them, two places. The Ollama library lists everything you can run with that one command, with the file size next to each, so sort by small and work down. And Hugging Face is where they all actually live: bigger, messier, and where the new ones land first. If you want the purpose-built end of the shelf, look at the SmolLM family, which is Hugging Face's own line of deliberately tiny models and where the smol spelling comes from in the first place. Yes, really. That is the actual name.
Tell me what you break. I want to hear the worst answer any of you gets out of the 270M one, because mine set the thermostat to 74 degrees when I told it I was too hot and I refuse to believe that is the record.
“But this was such a wonderfully small sigh, that she wouldn’t have heard it at all, if it hadn’t come quite close to her ear.”— narration, Through the Looking-Glass