Skip to content
Markdown

Open-weight inference economics and GPU useful life

Scope: how open-weight models, self-hosted token cost, and A100-through-B200 rental data bear on whether older NVIDIA generations keep earning; covers the Ornn Data paper of 7 September 2026 (closed-access rationing, cost per task, cost per million output tokens from rent and throughput, forward-curve retention), with a runnable reproduction of its arithmetic and an audit of where its prose and tables disagree.

Overview

The paper (Ornn Data, 7 September 2026) argues that a new GPU generation does not by itself make its predecessor obsolete. Its chain: closed-model access is rationed through allowances the provider can reset, so the metered token rate is the real marginal price; open-weight models complete benchmark tasks for less; open weights are portable across providers and hardware; on a sparse model an A100 produces tokens more cheaply than an H100; and Ornn's rental marks show the A100 holding value at long tenors.

The publisher sells rental data and GPU rentals, which the paper discloses. Every number below is the paper's unless marked otherwise, and the raw extracts and scripts are not released, so most of it cannot be independently reproduced (see the footnotes and the validation section).

flowchart LR
  CL["Closed access: allowances the provider resets"] --> MP["Metered rate is the marginal price"]
  OW["Open weights: portable across providers and GPUs"] --> CT["Lower cost per task at some thresholds"]
  OW --> SH["Self-host: rent / throughput"]
  SH --> SP["Sparse model: A100 cheaper than H100"]
  SP --> RM["A100 5Y term = 80% of 1M term"]
  MP --> CT
  RM -. "consistent with, not proof of" .-> UL["Older GPU keeps earning"]

Core knowledge

Five access modes

Mode Who sets quantity Marginal price beyond allowance
Subscription Provider, per rolling five-hour window and per week, in an unpublished unit Metered rate via usage credits
Usage credits Subscriber Standard API rate (Anthropic) or a credit rate card (OpenAI)
Metered API Subscriber, up to rate-limit maxima that are not guaranteed minimums Same rate; 0.5x batch, 2x fast
Hosted open-weight API Operator across competing providers Same, or another provider's rate
Self-hosted Operator Rent of another GPU

The paper declines to convert a subscription fee into tokens because the allowance unit is unpublished. Its example: at the same $20 fee, the published Codex range for Plus falls from 10-100 local messages per five hours on GPT-5.6 Sol to 5-45 on GPT-6 Astra, because each Astra message draws 2.5 times the credits.

Cost per task (Artificial Analysis Intelligence Index v4.1.1)

Cost per task is Artificial Analysis' weighted benchmark cost, including cache and reasoning tokens, so list prices alone do not reproduce it. The cheapest qualifying configuration on each side (Table 4 of the paper):

Minimum Index Open, $/task Closed, $/task Closed / open
52 GLM-5.3-Flash 0.09 GPT-5.6 Luna max 0.05 0.56
53-57 GLM-5.3-Flash 0.09 GPT-5.6 Sol high 0.43 4.78
58-60 GLM-5.3 max 0.68 GPT-5.6 Sol max 0.95 1.40
61 none GPT-5.6 Sol max 0.95 n/a

The closed frontier is six Index points above the open one (66 against 60). On Index v4.2, which replaced v4.1.1 the same week and scores every model lower, no open model in the paper's selected panel clears 51, and the closed Sol configuration is cheaper than open alternatives at 49-50. The paper states it cannot separate the effect of GPT-6 Astra from the change of Index version.1

Self-hosted token cost

Equation (1): cost per million output tokens = GPU-hour price / (millions of tokens per GPU-hour x (1 - h) x u), with h the reserved headroom and u the productive utilisation. Base case u = 0.5, h = 0.15; full-use case u = 1, h = 0.

Hardware, gpt-oss-120b (5.1B active) Spot $/GPU-hr Full use Base case
A100 SXM4 (GPUStack, one device) 1.00 0.12 0.29
H100 SXM (GPUStack, one device) 2.83 0.27 0.64
H200 (MLPerf v6.0, Red Hat) 4.51 0.35 0.98
B200 (MLPerf v6.0, Red Hat) 6.40 0.15 0.47

For dense Llama-2-70B the ranking flips: A100 $0.75 against H100 $0.68 at base case, B200 $0.39. The A100 dense throughput is not measured; it is 0.323 times the H100 figure, taken from an NVIDIA comparison of A100 FP16 against H100 FP8. With same-precision ratios of 0.42-0.68 the A100 costs $0.36-$0.58, below the H100, so precision alone can reverse the dense ranking.

Against the hosted gpt-oss-120b output price, break-even utilisation with 15 percent headroom is 24 percent for the A100 and 53 percent for the H100 at $0.60 per million (provider median), and about 84 percent and 188 percent at $0.17 (lowest OpenRouter listing). An H100 cannot beat the $0.17 listing by self-hosting. The volume-weighted OpenRouter gpt-oss-120b output price reported was $0.39.

Rental evidence

  • A100 spot rose about 20 percent from 1 March to 1 September 2026; the other families rose 70 to 111 percent.
  • A100 occupancy (rented divided by listed capacity, tracked on-demand providers) went from 74 to 90 percent while listed capacity grew 13 percent, implying about 37 percent more rented capacity.
  • Term-price retention against the one-month price:
Family 6M 1Y 3Y 5Y Implied forward, months 37-60 Family age at end of 5Y contract
A100 SXM4 95.0% 93.1% 84.2% 80.2% 74% 11.3 years
H100 SXM 95.0% 86.2% 68.2% 59.8% 47% 9.4
H200 92.4% 80.2% 51.0% 43.7% 33% 7.8
B200 99.1% 97.9% 69.1% 53.8% 31% 7.5
B300 93.8% 81.3% 60.9% 53.8% 43% 6.5

Implied forward = (60 x F60 - 36 x F36) / 24 for flat term prices F36 and F60. On gpt-oss-120b, Table 10 indexes token cost to A100 = 100: H100 165 (3Y) and 152 (5Y); B200 119 and 97, falling to 60 at the implied forward. So the A100 advantage holds against H100 and H200 at every tenor, while the B200 roughly matches it at five years and beats it at the forward.2

Where the workload fits

Interactive serving values latency and newer bandwidth. Agent rollouts, batch evaluation, and parts of reinforcement learning tolerate latency and can move to whichever hardware is cheapest per completed task. The only demand-scale evidence is indirect: Modal reported over $300 million annualised revenue on 21 May 2026 with sandboxes above a third of it, OpenRouter reported agentic token volume passing human volume around 1 February 2026, and a Menlo Ventures survey put enterprise open-source spend at 11 percent in 2025, down from 19 percent in 2024.

Don't-miss checklist

  • Compare on cost per completed task at a fixed quality threshold, not on list price per token.
  • Price self-hosting at your real utilisation; below the break-even utilisation a hosted endpoint wins before any host, network, storage, or labour cost.
  • Add 25 to 60 percent for non-compute cost; the paper's scenario lifts A100 gpt-oss-120b to $0.36-$0.46 and still clears the $0.60 median but not the $0.17 listing.
  • Check the serving precision and engine behind any throughput ratio before ranking generations. See inference parallelism strategies.
  • Treat term marks as analyst indicators, not executable quotes. Do not read retention as residual value of a device: ages are family ages, not device ages.
  • Read open-weight licences: five of the eleven models carry revenue or user-count gates (Kimi K3, GLM-5.3, Qwen3.8, MiniMax-M3, Llama 4).

Failure modes

  • Ranking hardware on a dense benchmark and applying it to a sparse model, or the reverse. The paper's own A100 result flips sign between Llama-2-70B and gpt-oss-120b.
  • Cross-source throughput comparison: A100 and H100 gpt-oss-120b come from GPUStack single-device runs on different vLLM versions with unbounded concurrency (p99 time to first token about 54-60 seconds), while H200 and B200 come from Red Hat MLPerf v6.0 submissions. The paper calls Table 10 indicative only.
  • Mixing Index versions: v4.2 scores and costs are not comparable with v4.1.1 panels.
  • Reading a high A100 five-year retention as proof of earning life. Retention is a percentage of a starting rent, and low-priced families can retain more simply because they started lower.
  • Causal over-reach. The paper's own limitations list device retirements, dense inference, non-LLM work, financing, and scarcity as alternative explanations.

Open questions & validation

The script below reproduces equation (1), the break-even utilisations, the implied forward, and the owner-side arithmetic from the paper's printed inputs, with guard cases. Validated by execution; the paper's A100 and H100 base-case figures (0.29, 0.64) reproduce when Offline throughput is used at u = 0.5 and h = 0.15.

from dataclasses import dataclass

@dataclass(frozen=True)
class Gpu:
    name: str
    rent: float       # USD per GPU-hour, Ornn spot 1 Sep 2026
    offline: float    # output tok/s per GPU

def cost(rent: float, tok_s: float, u: float, h: float) -> float:
    assert 0 < u <= 1 and 0 <= h < 1 and tok_s > 0
    return rent / (tok_s * 3600 / 1e6 * u * (1 - h))

def breakeven_u(rent: float, tok_s: float, h: float, price: float) -> float:
    return rent / (tok_s * 3600 / 1e6 * (1 - h) * price)

gpus = [Gpu("A100", 1.00, 2276), Gpu("H100", 2.83, 2901),
        Gpu("H200", 4.51, 3585), Gpu("B200", 6.40, 11634)]
full = {g.name: round(cost(g.rent, g.offline, 1.0, 0.0), 2) for g in gpus}
assert full == {"A100": 0.12, "H100": 0.27, "H200": 0.35, "B200": 0.15}, full
print("full use", full)
assert round(cost(4.51, 3013, .5, .15), 2) == 0.98   # H200 Server throughput
assert round(cost(6.40, 8949, .5, .15), 2) == 0.47   # B200 Server throughput
for g, paper in ((gpus[0], 0.29), (gpus[1], 0.64)):
    c = cost(g.rent, g.offline, 0.5, 0.15)
    print(g.name, "base", round(c, 2), "paper", paper)
    assert abs(c - paper) < 0.01
for g in gpus[:2]:
    for p in (0.60, 0.17):
        print(g.name, f"vs ${p}: u* = {breakeven_u(g.rent, g.offline, 0.15, p):.2f}")
assert round(breakeven_u(1.00, 2276, .15, .60), 2) == 0.24
assert round(breakeven_u(2.83, 2901, .15, .60), 2) == 0.53
assert breakeven_u(2.83, 2901, .15, .17) > 1      # H100 cannot beat $0.17
fwd = lambda f36, f60: (60 * f60 - 36 * f36) / 24
assert round(fwd(.842, .802), 2) == 0.74 and round(fwd(.682, .598), 2) == 0.47
monthly = (1.00 * 0.80 - 0.042) * 730
print("monthly", round(monthly, 2), "5y", round(monthly * 60, 2))
assert round(monthly, 2) == 553.34
for bad in ((1, 100, 0, 0), (1, 100, 1.5, 0), (1, 0, 1, 0), (1, 100, 1, 1)):
    try:
        cost(*bad)
    except AssertionError:
        continue
    raise SystemExit("guard failed")
print("ok")

Executed output:

full use {'A100': 0.12, 'H100': 0.27, 'H200': 0.35, 'B200': 0.15}
A100 base 0.29 paper 0.29
H100 base 0.64 paper 0.64
A100 vs $0.6: u* = 0.24
A100 vs $0.17: u* = 0.84
H100 vs $0.6: u* = 0.53
H100 vs $0.17: u* = 1.88
monthly 553.34 5y 33200.4
ok

Not verified: Table 10 (the cross-family one-month price ratios it needs are not printed), the Artificial Analysis scores and costs (sources were not refetched), the catch-up statistics (eleven of fourteen closed-frontier scores matched after a median 3.9 months), occupancy and spot series, and all third-party throughput figures.

References

  • Ornn Data, "The Economics of Open-Weight Inference", 7 September 2026: https://data.ornn.com/the-economics-of-open-weight-inference.pdf (fetched and read in full; URL returned HTTP 200).
  • Artificial Analysis, model and leaderboard pages: https://artificialanalysis.ai/ (cited by the paper for Index v4.1.1 and v4.2; not refetched).
  • GPUStack, inference benchmark source for A100/H100 gpt-oss-120b: https://github.com/gpustack/gpustack (repository root only; the specific benchmark page the paper cites was not located).
  • MLCommons Inference: https://mlcommons.org/benchmarks/inference-datacenter/
  • OpenRouter: https://openrouter.ai/

Related: build-vs-rent GPU cost model · cloud, neoclouds and cost · agent loop economics · inference parallelism strategies · GLM infra agent and dense feedback


  1. The abstract says the cheapest qualifying open-weight model completes a task at "roughly one fifth" of the cost of a comparable closed model. Table 4 gives a closed/open ratio of 4.78 only for thresholds 53-57, 1.40 for 58-60, 0.56 at 52 (closed cheaper), and no open model at 61. The "one fifth" figure is the best range, not a general result. A related prose line says that below Index 49 on v4.2 GLM-5.3-Flash is "an order of magnitude cheaper" than any closed configuration in Panel C; the printed costs ($0.18 against $1.16 for Astra medium and $1.25 for Sol max) give a factor of about 6 to 7. ↩

  2. The abstract and Section 6 state the A100 is cheaper than the H100 at spot and at the three- and five-year prices. That holds against H100 and H200. Table 10 also shows B200 at 97 (five-year) against the A100's 100, and 60 at the implied forward, which Section 5.4 acknowledges ("B200 is cheapest at the implied forward"). The headline "reverses the hardware ranking" is therefore specific to Hopper against Ampere on one sparse model. ↩