Manifold
More useful output from existing GPU fleets.
Reduce repeated computation through state reuse and token drafting.
Reuse context. Increase output.
Designed to:
- Park agent state outside GPU memory.
- Reuse shared prefixes.
- Place compatible state across the fleet.
- Fast first tokens on cache hits.
- Draft with spare host CPU and rack capacity.
- Verify draft batches in the serving engine.
More output from your H100 fleet.
3.02× delivered output · 57.4% less energy per token · 4.97× warm DRAM capacity.
Modeled targets: Manifold rack versus H100 native.
64 GPUs in two 32-GPU racks per fleet. The Manifold rack adds one CPU rack.
Draft targets: 11 accepted of 16. Ready rounds: software 50% · rack 75%.
Rack target: 8% residual misses. Striped range: 8–12% sensitivity.
Modeled targets
Delivered output
Output tokens/s
Modeled targets
Energy lower is better
kWh per million output tokens
Modeled targets
DRAM parking capacity
Background CPU targets
64 host-only · 96 rack slots
Memory slots after drafting reservations
GLM-5.3-Flash · 128K retained / 1K new / 1K output.
Model details
- Status: Modeled targets, not measurements.
- Serving: FP8 weights. H100: TP8, batch 16, BF16 KV. B200: TP4, batch 64, FP8 KV.
- Native stack: Dynamo-class host, NVMe and object reuse with MTP and local suffix lookup; 25% of output drafted at .80 acceptance.
- Prefill: 23,700 input tokens/s per H100 TP8 group from published 8K time-to-first-token; 23,108 per B200 TP4 group at 1.95× per GPU, the published Blackwell-to-Hopper ratio for GLM-5.3.
- Misses: Resumed context recomputed: native 30%, software 14%, rack 8%.
- Drafting: 11 accepted of 16 per proposal plus one verified token, ready on 50% of rounds for software and 75% for the rack.
- Power: IT AC, PUE excluded. H100 fleet 67.2 kW; with Manifold software 69.2 kW; with the Manifold rack 86.4 kW at 24 active sockets. B200 fleet 96.0 kW.
- Transfers: Host 50, NVMe 12, rack attach 35, broadcast 450 GB/s.
- Capacity: DRAM slots for complete 128K session state; HBM, NVMe and object tiers excluded.
- Artifact: Inputs, sources, equations and sensitivities: public model JSON.
Formats
Manifold42U
Manifold Mini22U
Hardware design
- Scale memory independently of GPUs.
- Air-cooled design.
- Buyer-supplied DDR5 memory.
SKU options by deployment.