Hardware
Inference silicon's insurgents take real share from the training giants
Purpose-built inference chips now serve a meaningful slice of global AI traffic, competing on tokens-per-dollar where training champions are overbuilt.
By Daniel Reyes, Chips & Infrastructure Editor — SANTA CLARA
SANTA CLARA — The AI chip market has split in two, and the smaller half is where the insurgents are winning. Purpose-built inference processors — stripped of training-oriented flexibility, ruthlessly optimised for serving tokens — now carry a meaningful share of global AI traffic, according to new market data.
The economics diverge cleanly: training rewards the general-purpose giants' ecosystems, while inference at scale rewards whoever delivers the most tokens per dollar per watt on a frozen model — a narrower, winnable game.
The incumbents are responding with inference-tuned variants and aggressive pricing, but cloud providers — eager for any second source — are underwriting the challengers with multi-year capacity deals.
Enable JavaScript to read the full story on Neural Daily News.