Manifold

More useful output from existing GPU fleets.

Reduce repeated computation through state reuse and token drafting.

Reuse context. Increase output.

Designed to:

  • Park agent state outside GPU memory.
  • Reuse shared prefixes.
  • Place compatible state across the fleet.
  • Fast first tokens on cache hits.
  • Draft with spare host CPU and rack capacity.
  • Verify draft batches in the serving engine.

More output from your H100 fleet.

3.02× delivered output · 57.4% less energy per token · 4.97× warm DRAM capacity.

Modeled targets: Manifold rack versus H100 native.

64 GPUs in two 32-GPU racks per fleet. The Manifold rack adds one CPU rack.

Draft targets: 11 accepted of 16. Ready rounds: software 50% · rack 75%.

Rack target: 8% residual misses. Striped range: 8–12% sensitivity.

Modeled targets

Delivered output

H100 · native stack3,596
H100 · Manifold software7,124
H100 · Manifold rack10,855
B200 FP8 · native stack8,553

Output tokens/s

Modeled targets

Energy lower is better

H100 · native stack5.19
H100 · Manifold software2.70
H100 · Manifold rack2.21
B200 FP8 · native stack3.12

kWh per million output tokens

Modeled targets

DRAM parking capacity

Background CPU targets
64 host-only · 96 rack slots

H100 · native stack6,327
H100 · Manifold rack31,466
B200 FP8 · native stack11,266

Memory slots after drafting reservations

GLM-5.3-Flash · 128K retained / 1K new / 1K output.

Model details
  • Status: Modeled targets, not measurements.
  • Serving: FP8 weights. H100: TP8, batch 16, BF16 KV. B200: TP4, batch 64, FP8 KV.
  • Native stack: Dynamo-class host, NVMe and object reuse with MTP and local suffix lookup; 25% of output drafted at .80 acceptance.
  • Prefill: 23,700 input tokens/s per H100 TP8 group from published 8K time-to-first-token; 23,108 per B200 TP4 group at 1.95× per GPU, the published Blackwell-to-Hopper ratio for GLM-5.3.
  • Misses: Resumed context recomputed: native 30%, software 14%, rack 8%.
  • Drafting: 11 accepted of 16 per proposal plus one verified token, ready on 50% of rounds for software and 75% for the rack.
  • Power: IT AC, PUE excluded. H100 fleet 67.2 kW; with Manifold software 69.2 kW; with the Manifold rack 86.4 kW at 24 active sockets. B200 fleet 96.0 kW.
  • Transfers: Host 50, NVMe 12, rack attach 35, broadcast 450 GB/s.
  • Capacity: DRAM slots for complete 128K session state; HBM, NVMe and object tiers excluded.
  • Artifact: Inputs, sources, equations and sensitivities: public model JSON.

Formats

Manifold42U

Manifold Mini22U

Hardware design

  • Scale memory independently of GPUs.
  • Air-cooled design.
  • Buyer-supplied DDR5 memory.

SKU options by deployment.

Discuss your deployment.

victor@avorant.com