Small Models, Big Energy: When You Don't Need a Giant Hammer

There's a temptation in AI right now to use the largest model for everything. Need to check whether an email is spam? Giant model. Extract a date from an invoice? Giant model. Decide if a support ticket is about billing or shipping? Giant model, three paragraphs of reasoning and a cost that slowly makes your finance team pale.
It's like hiring a world-famous architect to hang a picture frame. They'll do a great job. You'll also get a great invoice.
Match the model to the task
Large models shine at open-ended reasoning, long complex documents, nuanced writing and tasks you can't easily specify. But lots of production work is narrow and repetitive:
- classify this text into one of 8 categories
- extract these 5 fields
- detect the language
- decide whether this needs a human
For tasks like these, a small model, or even a classic classifier, is often just as accurate.

Why small is attractive
- Cost. At millions of requests, a cheaper model is the difference between viable and not.
- Speed. Small models answer in milliseconds, which matters for autocomplete, moderation and anything interactive.
- Privacy. You can run them on your own servers, or even on a laptop or phone, so sensitive data never leaves.
- Control. You pin a version and it doesn't change under you.
- Energy. Less compute per request is a real environmental difference at scale.

A sensible strategy
- Prototype with a big model. It's the fastest way to learn what's possible and to create good examples.
- Build a test set with those examples (you did read the post about evaluation, right?).
- Try smaller models on the same test set. You'll often find one that's within a point or two of the big model's accuracy.
- Route. Send easy cases to the small model and only hard or uncertain cases to the big one.
- Consider fine-tuning a small model on the outputs you've collected. A small specialist frequently beats a large generalist on its narrow task.
def answer(ticket):
label, confidence = small_model.classify(ticket)
if confidence >= 0.85:
return label
return large_model.classify(ticket) # only the tricky ones
When to go big anyway
Complex reasoning, long messy documents, high-stakes outputs where an extra few percent of quality matters more than cost, and early exploration. Use the big hammer for nails that deserve it.
The engineering lesson
This is an old lesson in new clothes: don't use a database cluster for a to-do list, don't use Kubernetes for a static site, and don't use a frontier model to detect whether a message says "unsubscribe." Pick the smallest tool that does the job well. Your users will notice the speed, and your accountant will notice the rest.