What Changed in Model Capability Over the Last Two Years, and What Did Not

It is tempting to describe every new model release as transformative. Looking back honestly, some changes mattered a great deal for business development work, and others mattered less than the announcements suggested.

What genuinely improved

The ability to work reliably with long documents, to follow multi-step instructions without losing track partway through, and to produce writing that reads as considered rather than generic all improved meaningfully. Tasks that were unreliable enough to require heavy babysitting became reliable enough to run with a lighter review pass.

What improved less than advertised

The core problem of a model stating something false with full confidence did not go away. It became somewhat less frequent on well-documented, factual questions, and it remains a live risk on anything where the model has thin or ambiguous source material. No release changed the basic rule that unverified output needs a human check.

The practical implication for a firm

The right response to two years of steady improvement is not to relax review discipline, it is to keep the review discipline exactly as strict while letting the volume of work that discipline can cover grow. More gets automated at the front end; the back-end check does not get to shrink.

What to watch for going forward

The most useful thing a firm can do is not try to predict every future capability change, but keep a lightweight habit of periodically re-testing its own workflows against current tools, on the same terms as the internal test described earlier for model sizing. A workflow built two years ago on the assumption that heavy human review is needed at every step may be able to safely automate more of that review today, and a firm that never re-checks will keep paying for caution it no longer strictly needs.

The discipline that does not change, regardless of how capable the tools become, is that every client-facing claim still needs a traceable source.

Keeping this review lightweight

The re-check does not need to be a formal project. A short, recurring calendar reminder, once a quarter, an hour of a technical lead's time, to revisit whether current workflows still make sense given what tools can now do is enough to catch most of the value of staying current, without turning the exercise into a burden nobody has time for.

Where this leaves a firm

None of this is complicated in principle, which is exactly why it gets skipped under deadline pressure. The question worth returning to before treating matching the right model to the right task as settled is what a careful reader would actually notice if the firm got it right. On the point raised above under “what genuinely improved,” the answer is usually specific rather than clever: long-document handling and multi-step instruction-following improved substantially. Firms that build this expectation into how they train new associates find it easier to sustain once experienced staff move on, because the standard lives in a documented habit rather than in one person's memory. The gap between a firm that talks about matching the right model to the right task and a firm that actually practices it shows up over several quarters, not in any single engagement, and it tends to show up most clearly in the small, unglamorous checks that a client never sees directly but benefits from anyway.

It also helps to name, plainly, who is responsible for keeping this working once the novelty of a new tool wears off. Someone should own the point raised under “what improved less than advertised,” check it periodically rather than assume it stays true on its own, and be the person a colleague asks when a new situation does not fit the pattern described here. Put simply: improvement should expand what gets automated, not shrink the review discipline. That kind of ownership, named and specific, is a small addition to a firm's process, and it is usually the difference between a good idea that is followed for a month and a standard that actually holds up over a year of real client work.

None of this needs to be elaborate to be effective. A short, dated note in a shared file, reviewed at the next quarterly check-in, is usually enough to keep the responsibility from quietly disappearing when the person who first cared about it moves on to something else.

Key takeaways

  • Long-document handling and multi-step instruction-following improved substantially.
  • Confident, unverified claims remain a persistent risk regardless of model generation.
  • No amount of model improvement replaces a human verification step.
  • Improvement should expand what gets automated, not shrink the review discipline.