AI
Approaching.AI Builds a Token Factory Under China’s Labs
Approaching.AI raised over 1 billion yuan in 2026 to mill tokens for GLM and Kimi rather than sell another model catalog.
Approaching.AI said on July 13, 2026 that 2026 fundraising had passed 1 billion yuan. The Beijing firm, 趋境科技, already mills tokens for Zhipu GLM and Kimi on a plant it calls ATaaS.
A May Pre-A of hundreds of millions of yuan had named that plant. The later check, led by Henan industrial capital, treats token output as a utility sitting under the model labs, not as another public model catalog.
Approaching.AI Already Supplies GLM and Kimi
Most Chinese inference talk still scores vendors on how many models they host. Approaching.AI sells the opposite product. It keeps a short list of high-volume models and tries to turn each GPU hour into tokens that arrive on time, in a usable shape, with function calls that do not flake.
At the May round the company said its ATaaS platform was already handling nearly 1 trillion tokens a day for enterprise customers that include Zhipu GLM and Moonshot’s Kimi. By July it said a leading trillion-parameter model on that plant had pushed daily high-quality token output past the trillion mark on a steady basis. Those are company operating claims, not a third-party audit, and they describe two different loads: the whole platform in May, then one flagship model in July.
Time to first token, tokens per second, stable structured output, reliable function calls, and service quality under heavy concurrency are the yardsticks it wants buyers to use. Founder and chief executive Ai Zhiyuan, a Tsinghua computer science Ph.D., frames tokens as a production input that ties model skill, system speed, uptime, and cost together. The pitch is closer to a mill than to an API directory.
More Than 1 Billion Yuan Landed in Half a Year
The May 20, 2026 Pre-A was jointly led by Xinglian Capital and Huakong, with Honghui Capital, Tianhao Energy, Shangshi Capital, Tianjin Ren’ai Hongsheng, and Hangzhou Fucheng following. GL Ventures, already on the cap table, added more. The company said it would spend that check on compute reserves and the inference system under ATaaS.
On July 13 it came back with a Series A led by Henan Investment Group’s Huirong Fund. Zhenzhi Capital, Shangshi Capital, Xinglian Capital, Shanghai Guofang Innovation, Honghui Fund, Huakong Fund, and Hangzhou Fucheng oversubscribed as existing holders. Corporate records list the operating company as Beijing Qujing Technology Co., Ltd., formed in December 2023 in Haidian, with Ai as legal representative.
THE 2026 CHECKS
- December 2023: Beijing Qujing Technology is formed in Haidian, later branded Approaching.AI.
- May 20, 2026: Pre-A of hundreds of millions of yuan, jointly led by Xinglian Capital and Huakong, with GL Ventures adding on.
- July 13, 2026: Series A led by Henan Investment Group Huirong Fund; the company says 2026 funds now exceed 1 billion yuan.
Earlier 2026 capital, including a Parallel Tech check described around the Spring Festival window, sits under those two public rounds. The firm did not print a Series A ticket size of its own, only the half-year stack above 1 billion yuan, so that cumulative figure is the one that counts here.
What Token as a Service Is Built to Deliver
Public Model-as-a-Service boards in China are still dominated by cloud catalogs. IDC figures put ByteDance’s Volcano Engine above 40 percent of that MaaS market by token calls in 2025, with Alibaba Cloud near 30 percent and SiliconFlow fourth. SiliconFlow told the market its own daily throughput rose from 47.8 billion tokens in December 2024 to 578.5 billion in April, then closed a Series B of more than 2 billion yuan at a $7.7 billion valuation.
Approaching.AI is not trying to climb that same leaderboard. It sells production capacity to the labs that already own the models, which is why GLM and Kimi show up as customers rather than as rows in a public catalog. Compared with a MaaS menu of hundreds of weights, the company says it will keep “fewer models, in-depth optimization,” because model count is a weak proxy for whether a call finishes a job.
ATAAS QUALITY BARS
- First token: Time to first token has to stay stable when many jobs hit at once, not only in a quiet demo.
- Output rate: The firm cites 30-50 TPS as the high-speed band it will hold while keeping cost in check.
- Structured replies: JSON-shaped output and function calls have to land in a form a production system can parse.
- Predictable quality: Heterogeneous scheduling, cross-cluster cache sharing, path isolation, elastic scale, and quality monitors are the plumbing it lists for that promise.
Zhang Yang, chairman of Huakong Fund, put the same shift in investor language: huge token demand is rearranging the compute chain, and the plants that can ship high-quality tokens at scale become the layer the rest of the industry leans on. That is a bet on conversion efficiency, not on who lists the longest model card.
Hot Experts Stay on the GPU
The open-source fame around this team still points at a different picture. Bookmark posts sell the KTransformers hybrid CPU-GPU framework as a way to park DeepSeek-class Mixture-of-Experts models on a single 24 GB card. The project’s own notes are blunter. DeepSeek-R1 and V3 on that path want the 24 GB of video memory plus about 382 GB of system DRAM, with speedups listed at 3 to 28 times against earlier local setups. A lone graphics card is not the bill of materials.
Latency is the first objection that path draws, and it is fair. Moving cold experts into CPU memory saves VRAM; it does not turn a workstation into a rack of H100s. The commercial product is the same conversion idea aimed the other way, at labs that already run production traffic and want more tokens from the boxes they already rent.
TWO OPEN STACKS, TWO JOBS
| Project | Job | Published proof point |
|---|---|---|
| KTransformers | CPU-GPU hybrid inference and fine-tunes for large MoE models | More than 19,000 GitHub stars, 1,600 forks, Apache-2.0, SOSP 2025 paper |
| Mooncake | KV-cache-centric disaggregated serving for Kimi | FAST 2025 Best Paper; 75% more real requests; up to 525% throughput in simulation |
GitHub’s own serving example for DeepSeek-R1-0528 in FP8 on eight L20 GPUs plus a Xeon Gold 6454S lists 227.85 tokens per second total and 87.58 output tokens per second at eight-way concurrency. Fine-tune notes put DeepSeek-V3 on four RTX 4090s with about 80 GB of GPU memory at 3.7 iterations per second, and 6 to 12 times the throughput of a ZeRO-Offload baseline in the workloads they measured. Those figures live in the repo, not in the fundraising deck.
The same group, with Tsinghua, Kimi, 9#AISoft, Alibaba Cloud, and Ant Group, built the Mooncake KV-cache serving architecture that splits prefill from decode and parks spare CPU, DRAM, and SSD as a shared cache pool. The open Mooncake Transfer Engine code is now wired into vLLM, SGLang, and TensorRT-LLM paths, and Mooncake joined the PyTorch ecosystem on February 12, 2026. A desktop demo and a Kimi serving plant share a lab. They are not the same product.
The Tsinghua Names on the Cap Table
Approaching.AI is a Tsinghua High-Performance Computing Institute transfer, with related research valued and injected as equity. Academician Zheng Weimin sits as chief scientific adviser. Professor Wu Yongwei is chief scientist. Associate professor Zhang Mingxing, a co-founder, leads inference architecture work and the open projects above. The SOSP 2025 hybrid inference paper lists Ai Zhiyuan, Wu Yongwei, and Zhang Mingxing among the authors, which is the cleanest public link between the company and the code.
Chairman Ren Xuyang is an early Baidu hand who later helped stand up iQiyi, Yidian Zixun, Haizhi, and News Break. President Wu Wenjie holds a finance doctorate and a CFA and has run industrial and capital deals. Chief technology officer Chen Xianglin appears on the same KTransformers author list. Yang Ke, a Tsinghua doctorate on the technical staff, is named as a core Mooncake contributor. The company also says it ships into SGLang, vLLM, and NVIDIA Dynamo, and that KTransformers is a first-day engine for GLM, Kimi, MiniMax, and Qwen releases.
May materials put KTransformers above 17,000 GitHub stars. Repo pages now show more than 19,000. That climb is the public trail. The quieter move is Tsinghua treating the same work as shareable industrial property.
As large models move fully into production systems, a high-quality AI token service that stays stable, answers fast, and keeps cost under control has become a core need for enterprises putting AI to work at scale. Approaching.AI will keep to a path of fewer models and deeper optimization, and this round will speed the large-scale rollout of high-quality AI token plants. We will work with industrial investors to put domestic prefill-decode heterogeneous schemes into commercial use, and help domestic chips run as normal kit in high-standard AI production.
Ai Zhiyuan, founder and CEO, Approaching.AI, Series A remarks
That last clause is the industrial tell. The open-source story is hybrid CPUs and leftover RAM. The paid story is getting Chinese accelerators through a production SLO.
Who Put Industrial Capital Behind Token Output
Xinglian Capital is an early-stage fund aimed at the large-model chain. Partner Li Wenjue said the firm was buying ATaaS as a system, and the speed at which Tsinghua work had been turned into live traffic, with token demand from agents as the next surge. Huakong’s Zhang Yang went further and called the company a “token factory” for a new computing layer, and said the fund would help later financing and a listing path.
Henan Investment Group’s Huirong Fund is a different kind of shareholder. A fund official said stable high-quality token supply had become core kit for the industry’s next stage, and that conversion efficiency would be the fight, then pointed to the group’s green-energy and compute plants as the reason it could sit next to Approaching.AI. That is provincial industrial money looking for a offtake story, not a seed fund hunting a model demo.
WHO WROTE THE CHECKS
- May lead: Xinglian Capital and Huakong jointly led the Pre-A of hundreds of millions of yuan.
- May follow: Honghui, Tianhao Energy, Shangshi, Tianjin Ren’ai Hongsheng, and Hangzhou Fucheng joined; GL Ventures added to an existing stake.
- July lead: Henan Investment Group Huirong Fund led the Series A.
- July follow: Zhenzhi, Shangshi, Xinglian, Shanghai Guofang Innovation, Honghui, Huakong, and Hangzhou Fucheng oversubscribed.
GL Ventures did not appear on the Series A follow list the company published in July, after adding in May. The industrial lead did. If token mills become regional plants, that mix is the point of the round.
With the rapid rise of domestic large models and a full burst of application demand, massive token demand is reshaping the computing power chain. Huakong Fund firmly believes that AI infra that can stably supply high-quality tokens at scale will become key kit for the industry, with wide market room and high investment value.
Zhang Yang, chairman, Huakong Fund
Output Per Chip Jumped After Spring Festival
Since the 2026 Spring Festival, Approaching.AI said, average token output per unit of compute rose more than three times, and total high-quality token capacity rose more than 30 times. It says that jump came from system work in live traffic, not from stacking more boxes or listing more models, and that it now has a repeatable loop for designing, building, running, and operating token plants.
COMPANY PRODUCTION CLAIMS
- Platform load: Nearly 1 trillion tokens a day at the May Pre-A, across enterprise customers including GLM and Kimi.
- Flagship load: Daily high-quality tokens for one leading trillion-parameter model held above 1 trillion by July.
- Per-box gain: More than 3 times the token output per unit of compute after the 2026 Spring Festival.
- Plant gain: More than 30 times total high-quality token capacity over the same stretch.
The July money is earmarked for more high-quality token capacity, an ATaaS upgrade, and putting domestic heterogeneous compute into core production, including plants aimed at leading models, internet platforms, and regional industry clusters. Technical notes from that round list domestic prefill-decode heterogeneous pairing, high-performance heterogeneous KV-cache conversion, and pooled heterogeneous compute as methods already in use.
None of those operating figures has been published as an independent benchmark. The open-source trail is easier to inspect, and it still needs a fat memory machine when the marketing image is a single 24 GB card. What the cap table shows is simpler. In half a year, Tsinghua inference work went from a Pre-A of hundreds of millions of yuan to a company-stated stack above 1 billion yuan, with a provincial energy-and-compute fund in the lead, while GLM and Kimi were already taking tokens off the line.
Disclaimer: This article is news reporting on a private company’s disclosed fundraising and product claims and is for information only. It is not investment advice, a solicitation, or a recommendation to buy or sell any security, or to invest in Approaching.AI or any fund named here. Readers should consult a licensed financial adviser or securities professional before acting on any funding, valuation, or capacity figure. Amounts, customer names, and production claims reflect company statements and primary technical pages as dated in this piece and may change with later rounds or later operating reports.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
