ANALYSIS

Local AI Just Got Good Enough to Leave the Cloud

A smartphone and laptop glowing with an AI interface while disconnected from any network, representing AI running entirely on-device
Capable models now run entirely on hardware people already own. Source: Santage
TLDR

A 2.6-billion-parameter model now does agent work on a phone

For three years the assumption underneath the AI business was that serious models live in a data center and you rent them by the token. That assumption is quietly breaking. A class of small models has crossed the line from toy to tool, and the machine in your hand is now a credible place to run them.

The clearest marker arrived this month. Liquid AI released LFM2.5-2.6B, an open-weight model with a 128,000-token context window that runs entirely on-device, calls tools, and handles multi-step agent workflows. Liquid says it competes with, and often beats, models up to four times its size on instruction following and tool use, while fitting in under 2.5 GB of memory. It is not a demo. It runs on laptops and phones today through standard runtimes.

The buzz this week around local-first AI, captured in a widely shared post from Dragonfly's Haseeb Qureshi, is not hype getting ahead of the technology. It is the market noticing that the technology already shipped.

Source: Haseeb Qureshi (@hosseeb), on X.

Good enough, offline, erodes the cloud's pricing power

The reason this matters is not that local models are better. They are not. The frontier still lives in the cloud, and the largest reasoning models will run in data centers for years. The reason it matters is that a large share of real AI work does not need the frontier. It needs a competent model, a fast answer, and privacy, and it needs those things cheaply and repeatedly.

The shift in numbers
Parameters in LFM2.5-2.6B, running agent tasks on-device2.6 billion
Context window, on a model under 2.5 GB of memory128K tokens
Projected edge AI market by 2033, from $30B in 2026$118.7 billion
Larger models it matches on tool use and instruction followingUp to 4x
Source: Liquid AI and Grand View Research, 2026.

Once a model is good enough and runs on hardware the user already owns, the per-token meter that funds the cloud AI business stops running. There is no inference bill, no round trip to a server, no data leaving the device, and no rate limit. For a bounded but growing set of tasks, the marginal cost of intelligence falls to the price of the electricity already flowing through a laptop.

Bar chart showing the global edge AI market growing from $24.9 billion in 2025 to $30.0 billion in 2026 and a projected $118.7 billion by 2033, a 21.7 percent compound annual growth rate
The edge AI market is compounding at 21.7% a year as capable models shrink below data-center size. Source: Grand View Research, 2026.

The performance table that makes the case

The specifics are what separate this from previous local-AI enthusiasm. A model that runs at reading speed on a phone and near-instant speed on a laptop, with a real context window and working tool calls, covers a wide band of everyday tasks.

Where it runsThroughputWhat that enables
Apple M5 Max laptop220 tokens/secInteractive agents, coding help, and document work with no latency
AMD Ryzen AI Max+ 395113 tokens/secFull desktop assistant workloads, offline
Modern phone30 tokens/secReading-speed responses in the palm of a hand
Memory footprintUnder 2.5 GBFits alongside other apps on consumer devices

LFM2.5-2.6B on-device performance across consumer hardware. Source: Liquid AI, August 2026.

What a local default does to the frontier labs' model

The strategic consequence is not that the cloud loses. It is that the cloud loses the easy, high-volume, low-complexity tier that padded the token counts. When a phone can draft, summarize, classify, translate, and run a simple agent for free, those requests never reach a paid API. What remains in the cloud is the genuinely hard work: the long-horizon reasoning, the largest context, the tasks where a small model still fails. That work is more valuable, but there is less of it, and it is exactly the tier where the labs compete most fiercely on price.

The cloud will keep the hardest problems. It is losing the profitable ordinary ones, and those were subsidizing the bill.

This is also why the download numbers matter more than the benchmarks. Recent small open models have been pulled hundreds of thousands to more than a million times each within weeks of release, a demand signal that the metered cloud cannot observe because it happens on other people's machines. Every one of those downloads is a set of tasks that left the rented data center for good.

The frontier of AI is still measured in the cloud, in bigger models and longer reasoning. But the center of gravity for volume is moving toward the device, where the model is free, the data stays home, and the meter never starts. The labs that treated inference revenue as the foundation of the business are about to learn which of their workloads were only ever renting because there was nowhere else to run.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.