Skip to main content
Back to News Hub
🟢TechCrunch AI
June 9, 2026
Business

Can tech companies learn to love cheaper AI models?

Overview

Tech companies are starting to route most everyday AI work to cheaper models instead of defaulting to the most expensive frontier systems, a change that would reshape the economics of the field. The clearest example comes from legal AI firm Harvey, which cut inference costs by about 3x with no drop in quality by pairing an open-source model with Anthropic's Claude Opus and sending only the hardest tasks to the pricier system. Coinbase co-founder Brian Armstrong predicts 80 percent of AI workloads will run on far cheaper models within 12 to 18 months.

Key Takeaways

  • If those same AI workloads can be handled by cheaper models without affecting quality, it would mean a massive shift in the economics of AI.

    The AI boom has been built on a basic assumption: Bigger models are more powerful, and the most powerful models win.

  • "20% of workloads will still run on latest gen models where IQ maxing is important."

    It's hard to overstate what a significant shift it will be for the AI industry if Armstrong's prediction comes true.

  • In a recent test by the legal AI tool Harvey, the company was able to reduce inference costs by 3x without reducing quality.

    The result was a significantly lower load in terms of server time and overall cost.

  • There's an active price war going on between in-house inference from the big labs and independently served open-weight models.

    For the bigger question of small versus large, it doesn't really matter which kind of small model wins out.

  • They could just as easily economize by making fewer calls, using less context, or simply giving up on the least promising deployments.

Stats & Key Facts

  • #The clearest example comes from legal AI firm Harvey, which cut inference costs by about 3x with no drop in quality by pairing an open-source model with Anthropic's Claude Opus and sending only the hardest tasks to the pricier system.
  • #Coinbase co-founder Brian Armstrong predicts 80 percent of AI workloads will run on far cheaper models within 12 to 18 months.
  • #"[D]emand for intelligence is near infinite, but 80% of workloads will be running on 99% cheaper models within 12-18 months," Armstrong wrote on X .
  • #"20% of workloads will still run on latest gen models where IQ maxing is important."

If those same AI workloads can be handled by cheaper models without affecting quality, it would mean a massive shift in the economics of AI. The AI boom has been built on a basic assumption: Bigger models are more powerful, and the most powerful models win. Now, the industry is about to learn what happens if that assumption starts to break.

Mounting costs have already pressured users to give smaller and cheaper models a second look. This cost-conscious model-shopping is new and it's unclear how it will affect the industry, but the impact is likely to be significant. One prediction, laid out best by Coinbase co-founder Brian Armstrong, is that it will result in the vast majority of tasks shifting to cheaper models.

"[D]emand for intelligence is near infinite, but 80% of workloads will be running on 99% cheaper models within 12-18 months," Armstrong wrote on X . "20% of workloads will still run on latest gen models where IQ maxing is important." It's hard to overstate what a significant shift it will be for the AI industry if Armstrong's prediction comes true.

For more details please read the original article at TechCrunch AI.

Why It Matters for Business

Real business deployments are the most reliable signal of where AI is generating measurable ROI. Watching which sectors operationalize AI, what they pay for it, and how it changes their P&L tells you more than any vendor demo. These case studies are what serious buyers and investors triangulate on.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by TechCrunch AI
Read the original