Connect with us

AI

ScitiX Puts RadixArk’s Heaviest NVIDIA Jobs in Production

ScitiX put a multi-model inference layer on owned NVIDIA B200 clusters, with RadixArk’s SGLang team as the named production customer and a Chinese open-weight menu.

Published

on

ScitiX said on Aug. 16, 2026 that its inference platform processes over 1 trillion tokens a day. The Singapore GPU operator runs that traffic on company-owned NVIDIA B200, H200, and H100 clusters, with about one second to first token and 99.9% uptime.

The named production customer is RadixArk, the commercial team behind SGLang. That testimonial is doing more work than the uptime number.

RadixArk Parks Its Heaviest NVIDIA Jobs Here

ScitiX said RadixArk “runs its heaviest scenarios” on the new layer. SGLang is the open serving engine that RadixArk maintains, and RadixArk itself said on May 5, 2026 that the project already powers trillions of tokens daily for Google, Microsoft, NVIDIA, Oracle, AMD, LinkedIn, xAI, Thinking Machines Lab, and humans&.

As inference gets more complex, the underlying infrastructure becomes the differentiator. ScitiX delivers the responsiveness and reliability we depend on.

RadixArk, commercial team behind SGLang, in ScitiX’s Aug. 16 statement

RadixArk is not a quiet design partner. It launched with $100 million in Seed funding at a $400 million post-money valuation, a round led by Accel and co-led by Spark Capital, with NVentures, AMD, MediaTek, and Databricks among the names on the post.

It also does not live only on ScitiX metal. A DigitalOcean customer note says RadixArk serves DeepSeek V4 at 3.5K+ tokens per second per GPU on AMD Instinct MI350X machines, and RadixArk has said it is bringing SGLang to Google TPUs. ScitiX is the NVIDIA lane for a team that already splits work across vendors.

That is a useful customer and a narrow one. The engine authors need reliable B200 and H100 time. ScitiX needs a name that already ships at the volume its launch cites.

A 2025 Singapore Cloud With No Public Funding

The Aug. 16 statement is datelined San Francisco. The company listing is not. ScitiX was founded in 2025, sits at 250 North Bridge Road, Raffles City Tower, in Singapore, and lists 17 people, inside an 11-50 headcount band.

Tracxn lists the firm as unfunded. That is a long way from the inference hosts that have been raising nine-figure rounds to fill halls with GPUs. Groq, which is not a customer here, said on Jun. 22, 2026 that it had raised $650 million, runs 13 data centers, and serves more than 5 million developers.

ScitiX still claims “10+ yrs” of HPC work on its site, and it sells like a cluster shop. Action Community for Entrepreneurship in Singapore lists a 12-month GPU cloud contract with the 12th month waived, a standard 8-GPU node, and a free 7-day test before a team commits.

SCITIX ON PAPER

  • Founded: 2025, with headquarters listed at Raffles City Tower in Singapore.
  • Headcount: 17 people named on its listing, inside the 11-50 band.
  • Funding: Tracxn lists the company as unfunded.
  • Metal: Company-owned NVIDIA B200, H200, and H100 clusters.

The launch copy calls the product a “neutral” routing layer. The GPUs are not neutral. They are ScitiX-owned NVIDIA parts, and the firm’s own statement says enterprises should not have to become GPU operators. The control sits one layer up, on someone else’s cluster.

What a Token Costs on the ScitiX Menu

ScitiX publishes listed per-token prices for 17 models on its B200 and H100 clusters, in U.S. dollars, on an OpenAI-compatible API. GLM-5.3 Flash input is $0.15 per million tokens and DeepSeek V4 Flash 0731 input is $0.14. Cache hits cost less, and batch jobs are billed at 50% of the listed price.

The paid text menu is almost entirely Chinese open weights: DeepSeek, Z.ai’s GLM, Moonshot’s Kimi, Qwen, and Tencent’s hy3, plus Boson speech. Marketing posts also name Llama-4-Maverick. That name is not on the price table that was live when the page was read.

SAMPLE PER-MILLION TOKEN RATES

Model Input Output Cache hit Context
GLM-5.3 Flash $0.15 $0.50 $0.030 1M
DeepSeek V4 Flash 0731 $0.14 $0.28 $0.028 1M
DeepSeek V4 Pro 0813 $1.32 $3.96 $0.44 1M
Kimi k2.6 $0.95 $4.00 $0.16 262K
Qwen3.5-397B-A17B $0.48 $2.88 $0.24 262K

A newer DeepSeek V4.1 Flash SKU sits at $0.30 input and $1.20 output per million tokens, with cache hits at $0.006. Boson ASR is $0.006 per minute of audio. Boson TTS is $0.050 per minute.

The product the statement actually sells is routing, fallback, session cache, retries, dedicated tenancy, zero retention, and telemetry. That is the same instinct that routes AI tasks to cut spend inside a data cloud, applied here to a multi-model API on rented-looking endpoints that sit on owned GPUs.

Serverless inference on the marketing site starts at $0.01 per 1 million tokens. That floor is a come-on rate, not the GLM-5.3 or DeepSeek V4 Pro bill a production app will see.

Cache Hits Now Drive the Bill

On Sep. 4, 2026, ScitiX posted the figures it wants buyers to remember: 1 second average time to first token, a 93.9% KV-cache hit rate, 99.9% uptime, and 72% average cost savings, plus session-aware orchestration and 3-tier failover. The 72% savings figure is the company’s own claim and is not broken out by model or workload.

The Aug. 16 statement had used a looser “exceeding 90%” cache line. The Sep. 4 post is the later, tighter KV number. A separate 24-hour test on Sep. 10 is higher still, and it is a test, not the production average.

ScitiX’s own Aug. 16 copy already said token prices are falling while total spend is not, because one user turn can fan out into many internal calls. The work it published after the launch is almost all cache and storage, not a new model.

WHAT SHIPPED AFTER AUG. 16

  1. Aug. 31, 2026: DeepSeek-V4-Pro-0813 goes live on ScitiX clusters with PD-separated serving.
  2. Sep. 1, 2026: The company posts KV footprint notes, putting compact MLA at the floor, deep full attention at about 10x, and classic GQA at 100x+ at iso-scale.
  3. Sep. 3, 2026: A metrics panel adds TTFT, latency, prompt buckets, and top-model breakdowns, with account tiering rolled out.
  4. Sep. 4, 2026: ScitiX posts the 93.9% KV-cache hit rate, 72% average cost savings claim, and 3-tier failover line.
  5. Sep. 7, 2026: It describes agent-aware KV eviction during remote tool calls that can take 10+ seconds.
  6. Sep. 10, 2026: An 80-GPU GLM-5.2 PD-disaggregated job runs 24 hours at a 99.4% warm KV cache hit rate, with L3 success at 100% and zero L3 timeouts.

On that Sep. 10 post, ScitiX said it raised the SGLang L3 read batch from 128 to 512, added Mooncake load balancing for SSD saturation, and got 1M-context full SSD reads 30%+ faster at peak. A follow-up said prefetch_threshold moved from 256 to 64. There is almost no public argument under those posts. The company is talking to itself in public, at kernel depth.

An internal tool called SiEval is the other number in the Aug. 16 statement. ScitiX said it saw up to 10.5x acceleration on evaluation-heavy pipelines and 7.22x end-to-end on large leaderboard workflows, with the biggest gains on LLM judges, sandboxed code, and long context. Those results are the company’s own tests.

The Open-Source Control Plane Called Arks

Under the API sits Arks, ScitiX’s cloud-native inference framework on Kubernetes. The repo is Apache 2.0, has a Chinese README, 181 commits, 50 stars, and 6 forks. That is a working cluster tool, not a crowded public project.

Arks speaks vLLM, SGLang, and Dynamo. The gateway is Envoy. Tenants get API tokens, quotas, and rate limits in tokens per minute and requests per minute. Models are cached and shared so cold starts hurt less. The quickstart spins a Qwen app on SGLang and answers through an OpenAI-style chat completions call.

WHAT ARKS HANDLES IN THE CLUSTER

  • Engines: vLLM, SGLang, and Dynamo run as Kubernetes workloads, with multi-node jobs and mixed hardware.
  • Gateway: Envoy plus ArksToken, ArksEndpoint, and quota objects sit in front of every call.
  • Autoscaling: Horizontal pod autoscaling is tied to stated SLOs, with live weight changes on endpoints.
  • Models: ArksModel objects cache weights locally and share them across apps so nodes do not all pull the same file.

Needed cluster add-ons include Envoy Gateway v1.2.8, LeaderWorkerSet v0.7.0, and RBGS v0.5.0-alpha.4. The README asks for Kubernetes v1.20 or newer. This is how a 17-person shop can claim a full inference layer: the control plane is code they already run, and SGLang is the engine their marquee customer wrote.

B200 Hours Price Near the Market Median

The same site that sells tokens also sells hours. ScitiX lists B200 NVL at $6.33 per hour, H200 NVL at $4.03, H100 SXM5 at $2.88, and A100 80G PCIe at $1.71, with reserved discounts and no egress fees on its object store.

A 35-provider survey on Sep. 13, 2026 put the median on-demand B200 at $6.62 per GPU-hour, with the cheapest listed at $3.75. ScitiX’s $6.33 is a shade under that median, not a dump-the-market rate. The inference pitch is utilization, cache, and routing, not a freak GPU sticker.

SCITIX GPU HOUR RATES

SKU ScitiX list Note
B200 NVL $6.33 /hr Under the $6.62 median on-demand B200 in a 35-provider survey
H200 NVL $4.03 /hr Listed beside B200 on the same pricing block
H100 SXM5 $2.88 /hr Hopper generation still on the menu
A100 80G PCIe $1.71 /hr Older card, still sold by the hour

NVIDIA’s DGX B200 page describes a box with eight NVIDIA Blackwell GPUs, 1,440 GB of GPU memory, and 15x the inference performance of the prior generation, at about 14.3 kW. ScitiX sells B200 NVL time. It does not say those hours are DGX-branded machines, and it should not be read that way.

Dedicated tenancy and zero retention are the enterprise clauses. Prompts and outputs, ScitiX said, do not persist past the transaction. Private environments are on offer for residency rules. Those are contract features. They are also how a small APAC cluster shop talks to banks and labs that will not send prompts to a US model API.

Tool Calls Leave GPUs Sitting Idle

The sharper note after the launch is not a new SKU. It is dead time. On Sep. 7, ScitiX said local tools return in milliseconds while remote search and APIs can take 10+ seconds, and that during the wait the GPU and its KV cache sit in HBM doing nothing.

The experiment is agent-aware KV eviction: kick the cache out while the tool runs, pull it back before the result lands, and free memory without adding user-facing delay. Chat serving, the post said, never had to do this. Agents do.

That is the second-order bill hiding under the 1 trillion token line. If each user turn fans into tool calls, the cluster’s problem is occupancy, not the sticker on GLM-5.3 Flash. ScitiX is trying to sell the occupancy layer, with RadixArk as the team that already knows how ugly those jobs get on NVIDIA metal.

On Sep. 10, ScitiX said an 80-GPU GLM-5.2 job ran for 24 hours at a 99.4% warm KV cache hit rate, with no L3 timeouts. The Aug. 16 statement still labels the 1 trillion token figure, the one-second TTFT, and the uptime line as illustrative, and not as service-level guarantees.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending