Why sell on Omnious
Your marginal cost is your edge. Version second-score/3 rewards efficient, reliable service and discloses what set every clear.
The marginal cost of serving a token is not one number. It moves with your hardware generation, your utilization this hour, how efficiently requests share GPU cycles, and what you pay for energy. Two providers serving the same open-weight model can sit several-fold apart in cost per token, and the same provider can be cheap at 9am and expensive at noon.
A posted-price router cannot express any of that. A static list price has to cover your worst hour, so it overprices your best ones, and it never rewards you for serving improvements that actually lowered your cost. An auction can. On Omnious you stream a signed quote every couple of seconds, repriced as your GPUs fill and drain, and every request clears against the book as it stands right now. Your cost advantage becomes fills the moment it exists.
Every clear names its price-setter
Version second-score/3 ranks expected request cost after measured latency, reliability, verification trust, and throughput. With a genuine independent rival, your tariff scales toward its score and is capped at the cheapest independent alternative. The clear never falls below your signed ask.
In that rival-priced regime, raising your bid can lose profitable fills without improving the tie-point or cap. A faster but more expensive winner can instead clear at its own ask, and a book with no genuine rival may use a disclosed reserve. Omnious publishes those exceptions rather than claiming that every fill is a textbook second-price auction.
price_setter plus score feedback to understand each outcome.The router takes 7%, and only 7%
The router's fee is 7% of cleared volume, deducted from your netted payout at epoch close. It never spread-takes: there is no gap between what the customer pays and what you receive beyond that published fee, and every receipt carries the exact fee split down to the base unit. A router that earned on spread would want prices opaque. This one earns by growing volume, so its incentive is a tight, liquid, transparent book. See fees and unit economics for the full accounting.
Measured latency beats marketing claims
The latency term in your auction scoreis measured by the router's own streaming proxy on every fill, not read from your marketing page. If you are genuinely fast, you win fills against rivals whose claims outrun their hardware, and you never have to publish a benchmark to prove it. New providers start on a pessimistic 1500 ms prior and earn their way down; idle providers get periodic probe requests so a good number never goes stale just because you were not winning traffic.
The result is a market where the things that should win, low real cost and low real latency, are the things that do win. If you run well-utilized hardware behind vLLM or SGLang, that is your edge, and this is the venue built to pay for it.
- router/src/market/auction.ts scoring and second-score clearing, the mechanism this page describes
- router/src/money/billing.ts the money math: bill, router fee, max charge, defined once