A demo called MicroLLM Lab reached the Hacker News front page today with 225 points. It lets you load seven small language models into a browser tab and chat with them. There is no API key and no account, and your prompt never leaves your machine. The models range from 26 million to about 362 million parameters, and the largest download is 216 MB. If you have wondered whether on-device AI belongs in your product, this is a quick way to find out how good or bad it is.
What MicroLLM Lab actually is
The lab runs on WebGPU and falls back to WebAssembly and then plain JavaScript. It does not use one of the popular libraries. The author wrote a custom WebGPU decoder and a 4-bit (Q4) weight packer, building on an open research model called PetitGPT by GitHub user yangqi0. Weights are saved in your browser's IndexedDB, so a model you have loaded once doesn't need downloading again.
The seven models in the live catalogue are PetitGPT research-v1 (124.6M, 74 MB), SmolLM2 135M Instruct and SmolLM2 360M Instruct from Hugging Face (80 MB and 216 MB), L20-Edu 135M (80 MB), MiniMind2 104M and MiniMind2 Small 26M (62 MB and 15 MB), and OpenAI's GPT-2 124M from 2019 (77 MB). All are Apache-2.0 except GPT-2, which is MIT.
The lab also has a benchmark tab of simple objective checks, and you can write your own tests in JavaScript. The author's README reports results from an Apple M4 in Safari: PetitGPT at 115 tokens per second and SmolLM2 135M Instruct at 66 tokens per second. SmolLM2 scored 14 out of 20 on the author's expanded test set. The README is frank about the limits: "A 135M network will fail arithmetic and invent facts."
What Hacker News found
Commenters on Hacker News set about breaking it straight away, and mostly did. People posted answers such as "2 + 2 = 4 + 2", a claim that cats "are not mammals", and a population for California given as two different wrong figures. Several commenters were impressed by the speed, though, calling it "surprisingly snappy" and "perfect for quick demos without a backend".
One commenter summed up the fair reading: models this small have very little knowledge and weak reasoning. They are useful for sentiment analysis, text classification and entity extraction, not for being your assistant. Others reported that it would not start in Firefox on Linux, where WebGPU is not available, and several criticised the dense interface.
A related story on the same front page shows how far down the hardware can go: a GitHub project runs a sliced 0.5B-parameter, 1.58-bit (BitNet) model across a cluster of seven ESP32-S3 microcontrollers.
Why a SaaS founder should care
Running a model on the user's device changes three things.
Privacy. The text never reaches your server. That makes compliance conversations easier for features that touch drafts, notes or customer data.
Cost. The user's hardware does the inference, so each call costs you nothing. Your remaining costs are hosting the weights and the bandwidth to serve them. The MicroLLM author said the models are served from the author's own server, which slowed under front-page traffic, so put yours on a CDN.
Offline use and latency. Once cached, the model works without a connection and without a network round trip.
What tiny models are good enough for today
Treat a sub-billion-parameter model as a fast, narrow helper rather than a chatbot. Sensible jobs:
Classification and routing. Is this support ticket about billing or a bug? Is this message spam? Does this query need the expensive cloud model at all?
Extraction. Pulling names, dates or product mentions out of short text, ideally with a model fine-tuned for that task.
Autocomplete and short rewrites. Finishing a phrase or tidying a sentence, where a weak suggestion is easy to ignore.
What they are not good for: factual questions, arithmetic, multi-step reasoning, long summaries, or anything where a confident wrong answer hurts the user. The Hacker News thread is full of examples.
The tools to build with
You don't need to write your own WebGPU kernels as this author did. Three established options:
Transformers.js from Hugging Face runs models in the browser on ONNX Runtime. By default it runs on the CPU through WebAssembly, and it can use WebGPU as an option, though its docs still warn that the WebGPU API is experimental in many browsers. It covers text classification, named entity recognition, summarisation, embeddings and zero-shot classification. For most SaaS classification features, this is where to start.
WebLLM from the MLC project is an in-browser LLM engine on WebGPU with an OpenAI-compatible API, including streaming and JSON mode. It is the better fit if you want chat-style generation with model families such as Llama, Phi, Gemma, Mistral or Qwen.
ONNX Runtime Web is the lower-level layer. It gives you WebAssembly, WebGPU and WebNN execution for your own exported models. One catch from its docs: only a subset of operators is supported on the GPU back-ends.
Check browser support before you commit. Can I Use lists WebGPU as supported in Chrome and Edge, partial in desktop Safari, and disabled by default in Firefox. Always ship a WebAssembly or server fallback.
What to do this week
Open MicroLLM Lab and try the kind of prompts your product would actually send, not trivia. Pick one narrow feature where a wrong answer is cheap, such as tagging, routing or suggestions. Prototype it with Transformers.js and a small task-specific model, and measure accuracy on a hundred real examples before you remove the cloud call. Keep the cloud model as the fallback for anything the small model is unsure about.
The bottom line
MicroLLM Lab shows that tiny models now load quickly and run fast in an ordinary browser tab. It also shows, in public, that they are unreliable on facts and maths. For SaaS products, the useful role today is a private, free first pass: classify, route, extract and suggest on the device, and send the hard cases to a bigger model. If you design for that split, you can ship this now.