Buyer: GPU cloud operators — first IREN, then other NeoClouds. The
pitch: stop wholesaling GPU-hours to whoever owns the layer above; sell
inference under your own brand, to your own customers, on your own terms.
This is the full architecture from the
system overview: global control plane, N
regions, gateway routing, Stripe monetization, and the operated-service posture
(24/7, SLAs, escalation) that amazee.ai already practices in production.
A NeoCloud edition v0 exists today in all but name — amazee.ai runs it:
tokenfactory-portal, formerly MOAD) with usage dashboards ✅What's missing is the factory floor: today the "regions" route to upstream
LLM providers; the Token Factory work adds our own
runtime and
registry underneath, so an
operator's GPUs — not third-party APIs — produce the tokens.
| Phase | Adds | Why it sells |
|---|---|---|
| 1. Local serving | vLLM fleets behind existing regional gateways; GPU + runtime telemetry | Tokens produced on the operator's GPUs — the core value prop |
| 2. Catalog & registry | Curated open-weight catalog, artifact pipeline, weight distribution | A credible model catalog, fast cold-starts, multi-site scale |
| 3. Gateway routing | Latency/residency/capacity-aware routing, failover, dedicated regions hardening | 5-nines SLAs, sovereignty guarantees — the enterprise-workload unlock |
| 4. Scale-out runtime | llm-d or Dynamo disaggregated serving, follow-the-sun warm pools | Tokens/sec/$ leadership; the routing controls the market |
| 5. Federation | Cross-NeoCloud "surge to partner" | Hyperscaler-rivaling capacity without losing the customer relationship |
Deliberately parked until the token business stands on its own — we only care
about LLMs and tokens right now:
Neither appears in the capability model today. When one is unparked, it enters
src/data/capabilities.ts as a capability like everything else.
Everything customer-visible must be operator-brandable: Portal theming,
login/email themes (Keycloakify ✅), API domains, model catalog naming and
pricing (products/pricing tables ✅). IREN is the first integration test of
"amazee.ai as a productized stack another operator runs".