AI
WEKA and Oracle Serve 10x Users on the Same GPUs
WEKA and Oracle’s H100 tests show long-context inference stalling on DRAM, not FLOPS, as NVMe cache lifted users 10x on 72 GPUs.
WEKA and Oracle Cloud Infrastructure ran a 72-GPU H100 test that served 10x more long-context users without adding cards. The extra sessions came from a larger key-value cache on NVMe, on the same silicon.
Amar Gowda, director of product management at OCI, and Dennis Kennetz, a senior machine learning engineer there, published the full benchmark methodology on May 13, 2026. WEKA repeated the figures in a June 9, 2026 announcement. The work is the second chapter of a 2025 OCI partnership, and it is aimed at buyers who keep treating long-context AI inference as a GPU-count problem.
A 72-GPU Cluster and a 600-User Cliff
The cluster was nine OCI bare-metal H100 nodes with eight GPUs each, 72 GPUs in all. DRAM-only serving, using standard vLLM, topped out at about 600 concurrent users. NeuralMesh with Augmented Memory Grid scaled past 5,000 in unbounded testing, which OCI and WEKA call a 10x gain.
Throughput moved the same way. The NVMe path reached about 2 million tokens per second, against under 200,000 for DRAM-only. In a one-hour run with 2,400 users, the expanded cache served 5 billion tokens against 700 million, and finished more than 47,000 requests against about 6,700.
THE OCI H100 SCORECARD
| Metric | DRAM-only baseline | Augmented Memory Grid | Stated gain |
|---|---|---|---|
| Max concurrent users | about 600 | more than 5,000 | 10x |
| Token throughput | under 200,000 per second | about 2 million per second | 10x |
| Tokens served, 2,400-user hour | 700 million | 5 billion | 7x |
| Requests finished, same hour | about 6,700 | more than 47,000 | 7x |
OCI’s write-up is blunt about where the curves split. Below the point where DRAM fills, the two stacks can look alike. Above it, cache misses rise, GPUs spend cycles rebuilding prefixes they already had, latency jumps around, and service-level targets slip. The NVMe tier is there to keep that cliff from showing up in the first place.
MiniMax-M2.5 Ran With 100,000-Token Inputs
The model was MiniMax-M2.5-NVFP4, served as several tensor-parallel-4 replicas. Each simulated user sent a 100,000-token input and got a 100-token reply. The tests were set to push cache hit rates as high as they could, so the comparison isolates offload against recompute rather than a mixed live mix of new and repeated prompts.
WHAT SAT IN THE RACK
- GPU layer: Nine bare-metal H100 nodes, eight GPUs per node, 72 GPUs total.
- Local flash: Sixteen Gen4 NVMe drives per node, 3.84 TiB each, pooled as a converged Augmented Memory Grid.
- Working set: About 8.64 TiB of DRAM in the baseline, against 287 TiB of usable NVMe with the grid on.
- Network: Two 200 Gb RDMA NICs per node, with NVIDIA GPUDirect Storage on the data path.
- Baselines: HBM plus DRAM only, HBM plus NVMe only, and the full HBM plus DRAM plus NVMe stack.
That 287 TiB figure is the whole argument in one number. You cannot buy another 33 times the DRAM in the same box at a price that looks like flash, and you cannot keep 5,000 long sessions in 8.64 TiB without eviction. The GPUs did not get faster. The cache they could see got much larger.
Why DRAM Saturation Looks Like a GPU Shortage
A transformer keeps a key-value cache so it does not rebuild every past token at every step. That cache grows with context length and with the number of live sessions. When HBM and DRAM fill, the runtime evicts blocks and pays a prefill tax to put them back. From the outside, utilization falls and queues grow, which looks like you need more GPUs.
WEKA’s own Llama-3.3-70B math at FP16 puts the cache at 326 KB per token. One 100,000-token session is about 32.6 GB before allocator waste. A few hundred of those sessions will blow past a DRAM pool that is only 8.64 TiB, which is exactly where the OCI baseline stopped adding users. Agentic runs make it worse, because OCI says a single session can burn 500,000 to several million tokens of retained state.
Enterprise AI workloads are pushing context windows and GPU utilization to new limits. These benchmarks show how WEKA’s NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks so customers can support larger, more demanding inference workloads without simply adding more GPUs.
Pablo Selem, Senior Director of Software Development, Oracle Cloud Infrastructure
Teams that stand up owned NVIDIA clusters for multi-model inference still hit this wall once the working set falls out of DRAM. Extra cards raise the FLOPS budget. They do not, by themselves, stop a 100,000-token prefix from being thrown away and rebuilt.
The Cache Grew From 8.64 TiB to 287 TiB
Augmented Memory Grid is WEKA’s name for a persistent token warehouse on NeuralMesh. KV blocks stream between GPU HBM and NVMe over RDMA and GPUDirect Storage, which is meant to skip a CPU copy on the hot path. On this cluster the warehouse used the drives already in the GPU nodes, the converged layout, rather than a separate storage fleet.
WEKA first showed the feature at NVIDIA GTC in 2025 and put it on Oracle Cloud Marketplace at SC25, with OCI as the exclusive cloud launch partner. The May 2026 run was the production-density follow-up: concurrency, sustained tokens, cache persistence, and SLO shape under load, not a synthetic time-to-first-token demo.
HOW THE OCI STACK GOT HERE
- April 19, 2022: WEKA says its platform is available on OCI bare metal for GPU and HPC jobs.
- March 2025: Augmented Memory Grid debuts at GTC as a KV-cache extension off GPU memory.
- September 18, 2025: An earlier OCI proof on eight BM.GPU.H100.8 nodes (64 GPUs) reports nearly 20x faster time to first token at 128,000 tokens versus baseline vLLM.
- November 18, 2025: WEKA makes the grid generally available on NeuralMesh and on the Oracle Marketplace, with OCI as exclusive cloud launch partner.
- May 13, 2026: Gowda and Kennetz post the nine-node, 72-GPU, 100,000-token density results.
- June 9, 2026: WEKA issues the 10x users, 10x throughput, and 7x tokens figures to a wider audience.
- July 21, 2026: NeuralMesh 6 ships and cites the same OCI H100 numbers as production proof.
The 1000x capacity line that still sits on WEKA’s product pages is a separate, older claim about stretching KV cache from gigabytes toward petabytes. This OCI job measured 287 TiB against 8.64 TiB, about 33 times more usable cache in that rack, not 1000 times.
NVIDIA, LMCache, and the Same Memory Wall
The interesting part of the 10x is that it is not a private trick. Inference stacks have been trying to offload KV cache through NVIDIA Dynamo, NIXL, and open tools such as LMCache. NVIDIA’s own September 18, 2025 engineering note says providers otherwise evict and recompute, cap the prompt, or buy more GPUs. Offload is the fourth option, and it only pays when reuse beats the cost of moving blocks.
In that same note, Vast moved cache to a single H100 at 35 GB/s with a GPU Direct Storage plugin. WEKA’s Dynamo lab on an eight-GPU DGX H100 hit read throughput up to 270 GB/s across the box, using an open NIXL plugin. Dell later published its own LMCache plus NIXL runs on four H100s and said token throughput rose as much as 2.8 times against HBM-only vLLM. Google Cloud showed tiered HBM, CPU RAM, and local SSD cache on eight H100s, with the biggest lift at 50,000- and 100,000-token prompts.
WHERE THE READINGS SPLIT
- OCI and WEKA: On this 72-GPU job, NVMe-backed cache is a production density lever, 10x users and 10x tokens per second versus DRAM-only vLLM.
- NVIDIA’s Dynamo guidance: Offload helps long sessions, high concurrency, and shared prefixes; it is overhead when prompts are unique or the transfer cannot hide behind compute.
- The hardware caveat: Flash does not replace HBM for the inner decode loop. The win is retrieving a prefix faster than rebuilding it, then staging it back into GPU memory before the next turn.
NVIDIA has since talked about a pod-level flash tier for inference context on BlueField-4, aimed at later 2026 systems, and storage vendors including WEKA have lined up around that design. Val Bercovici, WEKA’s chief AI officer, said in January 2026 that the company had already been building and sizing augmented GPU memory for a year while that market heated up. The OCI numbers are what that year looks like on H100s you can rent now.
The Test Maximized Cache Hits on Purpose
Gowda and Kennetz state the limit in one line: the workload was configured to maximize potential cache hit rate. That is the honest frame for the 10x. A coding agent that rereads the same 100,000-token project, or a fleet of users who share a system prompt, will see this shape of gain. A stream of brand-new 100,000-token documents with no overlap will still pay prefill, and the NVMe warehouse will look more like a buffer than a multiplier.
WHAT WE KNOW
- The footprint: 72 H100s, nine nodes, local Gen4 NVMe, RDMA, GPUDirect Storage, MiniMax-M2.5-NVFP4 at 100,000-token inputs.
- The cliff: DRAM-only serving failed to add users past about 600; the NVMe path kept taking load past 5,000.
- The SKU: NeuralMesh with Augmented Memory Grid is generally available to WEKA customers and through the NeuralMesh listing on Oracle Marketplace, with OCI as exclusive cloud launch partner.
WHAT IS UNCONFIRMED
- Live traffic mix: How often production prompts would hit at the rates this test forced.
- Cost per token: No dollar TCO, GPU-hour price, or WEKA subscription figure was attached to the 7x token count.
- Other clouds: The November 2025 launch promised more clouds later; this density proof is OCI H100 only.
Buyers who copy the architecture still have to land Kubernetes, GDS, RDMA, vLLM, and the warehouse on the same path, which OCI flags as a full-stack problem rather than a drive upgrade. The payoff they are chasing is not a faster chip. It is a 100,000-token prefix that is still there when the next turn arrives.
On the one-hour, 2,400-user chart, DRAM response time climbs at the saturation point and stays jumpy, while the NVMe line holds and the 7x gap in finished requests opens from there. That is the second-order bill for treating memory as an afterthought.
Frequently Asked Questions
How Did the 2026 OCI Test Differ From the 2025 Run?
The 2025 OCI proof used eight BM.GPU.H100.8 nodes, 64 GPUs, and focused on time to first token. At a 128,000-token context, baseline vLLM sat at 39.40 seconds and the WEKA path at 2.00 seconds, about 20x. The May 2026 job added a ninth node, 72 GPUs, MiniMax-M2.5-NVFP4, and measured users, tokens per second, and completed requests at 100,000-token inputs rather than first-token latency alone.
How Large Is the KV Cache for a 70B Model at 100,000 Tokens?
WEKA’s Llama-3.3-70B figure at FP16 is 326 KB of cache per token, so one 100,000-token sequence is about 32.6 GB before runtime overhead. That is why a few hundred long sessions exhaust a DRAM pool measured in single-digit tebibytes, even when the model weights already sit on the GPUs.
Where Can You Deploy NeuralMesh With Augmented Memory Grid?
WEKA says the feature is generally available to NeuralMesh customers and on Oracle Cloud Marketplace, and that OCI is the exclusive cloud launch partner. The nine-node recipe in the May 2026 blog is the public reference architecture for that listing, not a custom lab one-off.
Does KV Cache Offload Help If Every Prompt Is Unique?
Only at the margins. NVIDIA’s Dynamo guidance says moving KV to CPU, SSD, or remote storage pays when reuse outweighs transfer, especially in multi-turn and shared-prefix traffic. The OCI 10x run was built to maximize hits, so a shop whose prompts never repeat should expect something closer to the DRAM baseline plus the cost of the extra software path.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
