October 24, 2024
Large Models, Small Models: A Practical Distinction for Firms
Not every task needs the biggest available model, and using the biggest one everywhere is often the more expensive mistake, not the safer one.
What size actually buys you
Larger models tend to handle ambiguous, open-ended tasks better, synthesizing a strategic recommendation from mixed signals, or writing in a specific voice consistently across a long document. Smaller models are often just as accurate on narrow, well-defined tasks, classifying a document, extracting a date, checking whether a name appears in a filing, and they run faster and cheaper.
Where firms overspend
Running every routine extraction task through the largest available model is common and usually wasteful. The task does not need the extra capability, and the extra cost and latency add up across thousands of small operations a week.
A simple rule of thumb
Use the smallest model that reliably gets the narrow task right, and reserve larger models for the steps that require judgment, synthesis, or a specific tone. Test both on a sample of real tasks before deciding, the right answer is empirical, not theoretical, and it changes as models improve.
A simple test to run internally
Take a batch of twenty recent routine extraction tasks and run half through the largest available model and half through a smaller, cheaper one. Compare accuracy directly. In most firms that run this test, the smaller model matches the larger one closely enough on narrow tasks that the cost difference is not justified, and the result is enough to change the default setting for that task category going forward.
Repeat the same test periodically, since the right choice shifts as models improve at different rates, a smaller model that lagged noticeably a year ago may have closed the gap since, and a default set once should not be assumed permanent.
Documenting the decision, not just making it
Once a test like this settles which model size to use for a given task, write the decision down along with the date and the sample results, rather than relying on someone remembering it a year later when a new hire asks why the firm uses a smaller model for extraction. A short, dated internal note prevents the same test from being re-litigated informally every time staff turns over.
A note on switching costs
Changing which model size handles a given task is usually a configuration change, not a rebuild, when the workflow is designed with task-based routing in mind from the start. Firms that build this flexibility in early avoid the larger cost of realizing, a year later, that every workflow needs to be redesigned to take advantage of a cheaper or better option that has since become available.
Where this leaves a firm
None of this is complicated in principle, which is exactly why it gets skipped under deadline pressure. The question worth returning to before treating matching the right model to the right task as settled is what a careful reader would actually notice if the firm got it right. On the point raised above under “what size actually buys you,” the answer is usually specific rather than clever: bigger is not automatically better for narrow, well-defined tasks. Firms that build this expectation into how they train new associates find it easier to sustain once experienced staff move on, because the standard lives in a documented habit rather than in one person's memory. The gap between a firm that talks about matching the right model to the right task and a firm that actually practices it shows up over several quarters, not in any single engagement, and it tends to show up most clearly in the small, unglamorous checks that a client never sees directly but benefits from anyway.
It also helps to name, plainly, who is responsible for keeping this working once the novelty of a new tool wears off. Someone should own the point raised under “where firms overspend,” check it periodically rather than assume it stays true on its own, and be the person a colleague asks when a new situation does not fit the pattern described here. Put simply: test on real tasks rather than assuming the largest model is the safe default. That kind of ownership, named and specific, is a small addition to a firm's process, and it is usually the difference between a good idea that is followed for a month and a standard that actually holds up over a year of real client work.
None of this needs to be elaborate to be effective. A short, dated note in a shared file, reviewed at the next quarterly check-in, is usually enough to keep the responsibility from quietly disappearing when the person who first cared about it moves on to something else.
Key takeaways
- Bigger is not automatically better for narrow, well-defined tasks.
- Smaller models are often more cost-effective for extraction and classification.
- Reserve larger models for synthesis, judgment, and tone-sensitive writing.
- Test on real tasks rather than assuming the largest model is the safe default.