A San Francisco infrastructure startup called Sapiom slashed one AI agent company's monthly token costs from $1.2 million to roughly $100,000—a tenfold reduction—by routing requests to cheaper open-weight models instead of expensive frontier offerings. The company announced it closed a $35 million Series A led by Dragonfly's Haseeb Qureshi, following a $15 million seed round led by Accel earlier in 2025. The customer, an AI agent startup named Polsia, had seen its projected annual revenue surge from $100,000 to $10 million while its Anthropic token expenses ballooned to what founder and CEO Ilan Zerbib described as unsustainable levels.

Sapiom operates open-weight models on server racks it owns in a San Jose data center and provides a single API key that directs each incoming request to whichever model can complete the task at the lowest cost. The company charges customers directly for compute rather than reselling access to other providers' APIs with a markup layered on top. Polsia, which has no employees and whose cost structure consists almost entirely of token spending, serves as a particularly clear example of the economics facing real AI agents running at scale rather than in demonstration environments.

According to founder Ilan Zerbib, quoted in Semafor's reporting, "in 95% of cases, it doesn't make sense to go to a very expensive frontier model." The report notes that Sapiom distinguishes itself from roughly 80 other routers currently operating in this market by owning its own compute infrastructure instead of acting as a middleman atop other providers' APIs. The original reporting cautions that the Polsia figures come from the CEO who delivered the savings and do not include details on which specific open-weight models Sapiom routes to, how output quality changed after the tenfold cost cut, which other investors joined the round beyond Dragonfly, or whether a board seat was granted.

The report frames the Sapiom case as the clearest recent indication that the economics of running real AI agents at scale are forcing a second infrastructure layer to emerge between applications and frontier labs. When a company like Polsia, whose cost line is almost entirely tokens, declares its bill "unsustainable," the ceiling on current pricing models becomes visible. The competitive landscape is crowded—OpenRouter alone processes 25 trillion tokens weekly, before accounting for Amazon Bedrock or Microsoft Azure—but the key question is whether owning inference capacity, rather than reselling it, creates a lasting advantage or merely a temporary arbitrage opportunity. If Zerbib's prediction that trillions of agents are coming proves accurate, the routers who own their own racks will capture the margin that frontier labs currently collect. For enterprises evaluating AI infrastructure, the distinction between reseller markups and direct compute ownership may increasingly determine which cost structures remain viable as agent workloads scale. The choice between convenience and margin compression will define the next phase of AI deployment economics.