BendersonMEDIA
Markets
NVDA$4,127.83+2.14%
AAPL$241.52-0.38%
BTC$97,412+3.21%
MSFT$478.90+0.67%
ETH$4,128+1.89%
GOOGL$182.34-0.52%
TSLA$312.67+4.23%
META$621.45+1.05%
S&P 500$6,142.80+0.31%
NASDAQ$20,847.50+0.78%
NVDA$4,127.83+2.14%
AAPL$241.52-0.38%
BTC$97,412+3.21%
MSFT$478.90+0.67%
ETH$4,128+1.89%
GOOGL$182.34-0.52%
TSLA$312.67+4.23%
META$621.45+1.05%
S&P 500$6,142.80+0.31%
NASDAQ$20,847.50+0.78%

Small Models Win While Big Tech Burns Billions

By Brandon Henderson·July 14, 2026·5 min read
Small Models Win While Big Tech Burns Billions
Image: TechCrunch | Source

“`html

Small Models Win While Big Tech Burns Billions

The companies spending $10 billion to train the next frontier model aren’t the ones printing cash right now. Meta’s free Llama models have been downloaded over 650 million times, according to Meta AI. Microsoft’s tiny Phi-3 Mini matches GPT-3.5 on most benchmarks. The AI race has already moved. Most people are still watching the wrong scoreboard.

The Frontier Is Winning Headlines, Not Profits

Since 2022, the AI story has been about who can build the biggest, smartest model. OpenAI, Google, Anthropic, and xAI have spent billions chasing benchmark records. According to Epoch AI, the compute used to train frontier models doubled roughly every six months between 2022 and 2025, pushing individual training runs into the hundreds of millions of dollars per cycle.

That arms race produced real breakthroughs. Nobody denies that. But the math stopped making sense for most businesses a long time ago. You don’t need GPT-4o to summarize an email. You don’t need Claude Opus to classify a support ticket. The frontier model is a Ferrari. Most business problems need a pickup truck.

In 2025, the pickup trucks got very good very fast. According to the Stanford 2025 AI Index Report, the performance gap between frontier models and open source models shrank by more than 50% over 18 months. Small models got faster, cheaper, and accurate enough for most real-world tasks.

Where the Real Race Is Happening

The winners in the next 24 months won’t be the labs training the biggest models. They’ll be the builders who figure out which model is cheapest for each specific task, then run that model at scale.

I call this the efficiency race. It’s already underway.

Google’s Gemma 2 at 9 billion parameters outperforms models ten times its size on coding and reasoning benchmarks, according to Google DeepMind. Microsoft’s Phi-4 Mini runs on a smartphone. Meta’s Llama 3.3 at 70 billion parameters beats GPT-4 class models on multiple standardized tests, and it costs nothing to download.

Meanwhile, the cost to run AI inference has dropped about 99% since 2022, according to a16z research. A query that cost $0.06 in early 2023 costs less than $0.001 today on models with comparable capability. That number matters more than any benchmark score ever will.

The rich mindset sees this clearly. The poor mindset keeps chasing the biggest model because bigger feels safer. It’s the same mistake people made in the mainframe era, buying IBM because nobody got fired for it, while the builders running smaller, cheaper infrastructure printed money.

Smart operators are doing three things right now. First, they’re benchmarking their actual tasks against cheap models before buying expensive API access. Second, they’re building with open source models they can host themselves or fine tune, cutting inference costs by 80 to 90 percent. Third, they’re using the savings to ship more product faster, not to pay for frontier model subscriptions they don’t actually need.

Content creators specifically have figured this out. Tools like InVideo AI run on efficient model layers rather than frontier APIs, which is how they can offer fast video generation at a price that makes business sense for individual creators and small teams. The model behind the tool doesn’t need to be a frontier model. It just needs to be good enough, fast enough, and cheap enough to justify the output.

What I Would Do Right Now

If I’m building a product or running a business that uses AI, I’m not asking “which is the best model?” I’m asking “what is the cheapest model that produces output I can actually ship?”

That question changes everything about how you spend money on AI.

Start by running your core AI tasks, content creation, customer support, data classification, code generation, against at least three models at different price points. You will almost always find that the midrange or open source model handles 80% of your workload just fine. The 20% that actually needs a frontier model is where you spend the money, and nowhere else.

Next, build your stack around the cheapest viable option. Upgrade specific tasks to frontier models only when the output quality gap is measurable and the economics justify it. Most businesses never reach that threshold.

For tools and software, look for options that have already baked in this efficiency thinking. AppSumo regularly surfaces AI tools with lifetime deals, meaning you pay once instead of a monthly subscription that grows with usage. When AI inference costs keep dropping, the tools built on top of efficient models get cheaper to run. That margin usually passes to early buyers who locked in before the price went up.

The businesses that win won’t be the ones with the biggest AI budgets. They’ll be the ones who got the same output for a fraction of the cost and reinvested the savings into distribution, product, and customer acquisition.

The Bottom Line

The frontier model race is real, but it’s being fought by companies with $100 billion balance sheets. That’s not your race. Your race is efficiency. The model that costs 90% less and gets the job done 85% as well beats the expensive one every time when you’re running at scale. The AI advantage isn’t going to the labs with the biggest models anymore. It’s going to the operators who run the cheapest model that works. Pick the right race.

Frequently Asked Questions

What does “frontier AI model” mean?

A frontier model is the most advanced AI available at a given time, typically from companies like OpenAI, Google, or Anthropic. These models score highest on benchmarks but cost the most to train and run. They represent the leading edge of capability, but that doesn’t mean they’re the right choice for most business tasks.

Are smaller AI models actually as good as frontier models?

For most business tasks, yes. According to the Stanford 2025 AI Index, the performance gap between frontier and open source models narrowed by more than 50% over 18 months. For specialized work like complex legal reasoning or advanced scientific research, frontier models still have an edge. For content, customer service, and code generation, smaller models perform comparably at a fraction of the cost.

Why is the AI race shifting away from the frontier?

Because the economics changed. Training frontier models costs hundreds of millions per run. But inference costs dropped 99% since 2022, according to a16z. Running smaller models at scale is now cheap enough to build a real business on. The money follows the margin, and the margin is in efficient deployment, not in training the biggest model.

How do I choose the right AI model for my business?

Benchmark your specific tasks against three to five models at different price points before committing. Most businesses find that a midrange or open source model handles the majority of their workload. Start cheap, measure quality, and only upgrade specific tasks to frontier models when the quality gap is worth the price difference.

Will frontier AI models still matter going forward?

Yes, but for fewer businesses than most people expect. Frontier models will matter for research, specialized reasoning, and tasks where accuracy justifies any price. For everyone else, the efficiency race determines who wins. The operators who master cheap, reliable AI deployment will outcompete the ones chasing the most impressive model every single time.

“`

Get stories like this in your inbox. Daily.

Free. No spam. The AI, tech, and finance stories that move money.

The Daily Brief

Sharper than your feed.

AI, finance, and tech stories that actually matter. One email, every weekday.

Free · No spam · Unsubscribe anytime