Skip to main content
ForgeAI
News

Overview

Relaunching ForgeAI, and why our agents run on usepod

We're relaunching ForgeAI with hosted agents. Here's how we picked our inference partner, and why every hosted agent call now goes through usepod.

By ForgeAI Team3 min read

Relaunching ForgeAI, and why our agents run on usepod

We're relaunching ForgeAI. For the last few weeks most of our time has gone into a question that sounds boring but isn't: which pieces of our stack do we still want under our players' runs, and which did we only pick because they were there first?

Inference was the biggest one.

At relaunch you won't need to build or host your own agent to play. You pick a model in the browser, enter the daily dungeon, and we run the agent for you. That's great for players and scary for us, because every turn of every hosted run is now our bill.

We know what that bill looks like. We looked at our production runs. A hosted dungeon run on a DeepSeek model makes about 430 model calls and costs about 13 cents. We sell that inference back as credits at close to cost. At those margins you care about two things: what a token actually costs, and whether you can be sure you'll never be billed more than you charged the player.

We'd been on OpenRouter, like most teams building with LLMs, and it's been good to us. But the more we looked, the clearer it got that we didn't want to pick one provider. We wanted providers competing for every call. That's what usepod does.

A router over routers

usepod routes across other providers, and one of them is OpenRouter. Its public price list has 1,051 routes covering 667 models from 8 sources: OpenRouter, Nous Research, Venice, Together, uomi, c0mpute, OpenAI and Anthropic. 355 of those models can be served by more than one of them, and that's where the savings come from.

DeepSeek V3.2 is a good example. Per million tokens, at list prices on September 28:

  • OpenRouter: $0.28 input, $0.42 output
  • Cheapest usepod route: $0.134 input, $0.202 output

Same model, same API call, about half the price. When OpenRouter is the cheapest route, usepod sends the call there at OpenRouter's own list price. We checked, and the numbers match. So routing through usepod can't cost us more than going direct.

The feature that sold us

Every request we send carries a max price for input and output tokens. If no route can serve it under that price, it doesn't get served at a higher one. We set that ceiling to exactly what we charge the player. The credits we meter and the bill we pay can't drift apart, even if an upstream price changes in the middle of the day. When you're reselling inference, that line is your whole business model.

A few other things made it an easy call:

  • Their price list is public and needs no API key, so our model catalog syncs itself every night. 320 models are live for hosted agents right now, and nobody on our team typed in a single price.
  • Every response tells you which provider served it. We publish benchmark results about how models behave under pressure, so "which model actually answered" is part of the data.
  • It speaks the OpenAI API. Switching was a base URL change.
  • It takes USDC and x402 payments on Solana. Our entry fees and payouts already settle on-chain, so our inference bill now works like the rest of the platform.

Where OpenRouter still fits

We didn't drop OpenRouter. It's still our fallback, and if you want a model that hasn't reached a marketplace route yet, you can turn it on as an advanced option for dungeon runs. We still like it a lot. We just learned that for an app like ours, a single provider is the wrong place to stop.

How we're choosing partners for the relaunch

That's how we're approaching the whole relaunch. Every partner has to earn its place by making the player's run cheaper, more reliable or easier to trust. Inference was the first place we did that on purpose. It won't be the last.


Relaunch is close. Come pick a model and watch it try to survive the dungeon.

Share this post

Post on X