Fireworks Raised $1.5B at $17.5B While Serving 40 Trillion Tokens a Day

Fireworks closed a $1.5 billion Series D at a $17.5 billion valuation with Nvidia participating. It says it passed $1 billion in annualized revenue, and 95% of its traffic runs specialized rather than general models.

Fireworks Raised $1.5B at $17.5B While Serving 40 Trillion Tokens a Day

Fireworks announced a $1.505 billion Series D at a $17.5 billion valuation, led by Atreides Management, Index Ventures, and TCV, with Nvidia, Lightspeed, Bessemer, and Menlo Ventures participating, per the company. Fireworks runs inference infrastructure, the layer that actually serves AI models in production. It says it has passed $1 billion in annualized revenue run rate and serves more than 40 trillion tokens a day, per CNBC.

The figure that carries the mark is that $1 billion of annualized revenue, which puts the round at roughly 17 times run-rate. That is a real multiple on real revenue, which stands out in a cohort where valuations routinely price a thesis instead. The more interesting number is that more than 95% of those 40 trillion daily tokens come from models specialized on customers' own data. Fireworks is monetizing the assumption that most production AI work runs better on a smaller model tuned to a specific job and served cheaply than on the largest general model available.

For the cap table, Nvidia's participation is the strategic signal. Nvidia has been investing across the inference layer to keep demand pointed at its silicon, and a stake here keeps a large token-serving platform inside that orbit. The competitive risk is the obvious one: hyperscalers and frontier labs both want the inference layer, and Together, Baseten, and the GPU clouds are chasing the same workloads. At 17 times revenue, the next mark needs the specialization thesis to keep compounding rather than getting absorbed into a hyperscaler's default stack.

A $17.5 billion valuation on $1 billion of run-rate is priced on actual revenue, and the 95% specialized-token figure is the whole thesis. Watch whether specialization holds as hyperscalers bundle inference into their default stacks.