AI
Tencent Hy3 Free Push Turns Agent Traffic Into the Real AI Prize
Hy3 leads global agent tool-call volume at roughly 33 million daily while WorkBuddy runs free worldwide through August 31 at rock-bottom API rates.
Tencent integrated its Hy3 large language model into the international edition of WorkBuddy on August 5 and made the agentic workspace free for every global user through the end of the month. The move pairs a mid-tier open model with rock-bottom pricing and already-leading agent traffic, testing whether everyday task completion beats raw intelligence scores.
Hy3 sits at roughly 33 million autonomous tool calls per day and an 8.9% share of global agentic volume on OpenRouter. That metric tracks research, spreadsheets, code and multi-step workflows more than chat. The company is treating AI as a utility that ships inside products people already open.
WorkBuddy Goes Global Without a Paywall
Tencent’s August 5 announcement opened Hy3 across WorkBuddy, the design studio Miora and Tencent Cloud TokenHub. WorkBuddy users worldwide get the model at no charge until free worldwide until 31 August 2026 Pacific Time. After that the API starts at $0.1288 per million input tokens and $0.5336 per million output tokens.
WorkBuddy launched internationally in May and functions as an out-of-the-box agent workspace. Users issue natural-language instructions; the system handles research, presentations, spreadsheets and workplace messages. Concurrent multi-agent runs and more than 100 built-in skills let one prompt drive complex flows. Remote management works through Discord, Slack or Telegram.
- Hy3: 295B total parameters, 21B active, Mixture-of-Experts, 256K context
- Internal Tencent workplace tests: task success above 90%, average completion time cut 34% versus prior Hy generation
- Also live in Miora for multi-agent creative pipelines and TokenHub for multi-model routing
The free window removes the usual pilot friction. Teams can route live work through the same workspace they already open for messages and files, then measure whether multi-step flows finish without standing up separate agent stacks.
Poshu Yeung, Senior Vice President of Tencent Cloud and head of its international business, framed the release as moving organizations “beyond AI experimentation” into practical scaled deployment with enterprise controls.
Because the model also appears in Miora and TokenHub on the same day, a single integration path covers creative pipelines, workplace agents and multi-model routing. That breadth matters for firms that want one vendor relationship rather than three separate tool contracts.
How Hy3 Seized the Tool-Call Rankings
One week after the July 6 official launch, Hy3 invocations jumped more than 68-fold over its predecessor and hit No. 1 on OpenRouter’s global tool-call volume board. Recent tallies still place it near the top of usage charts even as other Chinese models crowd the leaderboard.
Tool-call volume measures autonomous agent executions, not casual prompts. On that score Hy3 processes roughly 33 million calls daily for an 8.9% share of tracked agentic traffic. Token volume also ranks high, but the agent metric is the sharper signal for real workflows.
| Model / Source | Daily Tool Calls / Share | Input / Output per 1M tokens |
|---|---|---|
| Hy3 (Tencent) | ~33M / 8.9% | $0.1288 / $0.5336 |
| Typical Western flagship range | Lower agent share | Often 5-20× higher |
| OpenRouter listing | Multiple providers | Matches Tencent Cloud figures |
Developers on X have noted the broader pattern: Chinese models now occupy most of the top OpenRouter usage slots by tokens and calls. Cost per real task can run an order of magnitude or more below closed Western options while speed stays competitive for ordinary work.
The ranking path itself is short and steep. Launch on July 6, a 68-fold jump within a week, then sustained near-top placement even as rivals arrive. That sequence shows how quickly volume follows when a model is cheap enough and stable enough for production agent loops.
Mid-Tier Scores, Agent-First Design
On the independent Artificial Analysis Intelligence Index, Hy3 lands around 41-42, solidly mid-pack and behind both leading Chinese systems and frontier Western models. Tencent does not claim benchmark supremacy.
The architecture is a hybrid fast-and-slow-thinking MoE with long-context reliability and multi-step tool stability as explicit targets. Public benchmarks understate product usefulness, the team argues. A blind evaluation with 270 working experts gave Hy3 2.67 out of 4, edging GLM-5.1, with the biggest edge in frontend, data and CI/CD tasks.
Internal fixes focused on tool-call error recovery, output format stability, hallucination cuts (from 12.5% to 5.4% in one internal set) and multi-turn intent tracking. Variance across agent scaffolds such as CodeBuddy, Cline and KiloCode stayed within 4% on SWE-Bench Verified. The model is tuned for production agent loops more than academic leaderboards.
Those tuning choices map directly onto the traffic mix OpenRouter records. Research, spreadsheets, code and multi-step workflows punish format drift and weak recovery. A model that holds variance inside 4% across scaffolds and halves hallucination rates on internal sets is built for that load, even when the Intelligence Index keeps it mid-pack.
- Intelligence Index: roughly 41-42, mid-pack
- Blind expert score: 2.67 out of 4 across 270 reviewers
- Hallucination rate in one internal set: cut from 12.5% to 5.4%
- Scaffold variance on SWE-Bench Verified: within 4%
Open Weights Rewrite the Access Rules
Hy3 ships under the commercially permissive Apache 2.0 license. Full weights sit on Hugging Face, ModelScope, GitCode and CNB. Developers can download, modify and self-host without a proprietary cloud lock-in or ongoing platform tax. The Apache 2.0 weights on GitHub repository also includes finetune and RL post-training recipes.
That distribution choice collides with export-control and sovereign-cloud strategies built around closed stacks. Once weights are public, traditional containment grows harder. Hugging Face data already shows Chinese models accounting for Chinese models 41% of downloads, ahead of the U.S. share near 36.5% in earlier tallies.
Western enterprises still flag data-sovereignty, audit and compliance worries. Pending U.S. procurement limits on Chinese-origin models could wall off government and highly regulated buyers. The open release still gives startups, researchers and non-U.S. teams a free high-volume option that closed APIs cannot match on price.
Safety and evaluation questions remain live. Separate testing episodes around open-model ecosystems, including documented open-model safety testing incidents, show how quickly public weights travel once released.
Self-hosting plus the published finetune and RL recipes lets teams adapt the 295B-parameter MoE to private data without routing every token through a third-party endpoint. For buyers outside regulated U.S. channels, that path turns the download share lead into deployable capacity.
Enterprise Reality Check Meets Utility Pricing
Many companies have learned that chatbots alone do not transform operations. The 95% failure rate often cited for enterprise AI pilots points to a different bottleneck: completing reliable workflows at acceptable cost, not generating clever answers.
WorkBuddy and Hy3 target exactly that gap. Natural-language instructions become multi-step actions without each firm standing up its own agent infrastructure. Tencent’s internal numbers (above 90% success, 34% faster) are company-run, yet they align with the product pitch of long-context, multi-agent workplace execution.
Hy3 marks a significant step in how Tencent Cloud International is helping organisations move beyond AI experimentation and put intelligence to work in practical, trusted and scalable ways.
Poshu Yeung said that in the August 5 release, tying the model to enterprise knowledge bases, agent runtimes and secure infrastructure across 66 availability zones in 23 regions.
API rates starting at $0.1288 per million input tokens undercut many Western flagships by roughly an order of magnitude. For unprofitable startups and high-volume internal tools, that difference decides whether an agent runs at all. Crowd observations on X describe teams cutting monthly inference spend 60-85% after switching to Chinese open or low-price models for routine agent work.
Utility pricing changes the pilot math. When input tokens start at $0.1288 per million and output at $0.5336, a failed experiment costs little enough that teams can keep iterating. The same 95% failure pattern looks different when each attempt is cheap and the workspace already handles remote management through Discord, Slack or Telegram.
The Launch Timeline Shows Volume First
The product sequence compresses distribution, measurement and free access into a few months. Each step builds on the last without waiting for a new intelligence-index peak.
- May: WorkBuddy launches internationally as an out-of-the-box agent workspace.
- July 6: Hy3 receives its official launch.
- One week later: invocations jump more than 68-fold and reach No. 1 on OpenRouter tool-call volume.
- August 5: Hy3 opens across WorkBuddy, Miora and TokenHub, free for global WorkBuddy users.
- 31 August: the free worldwide window ends and standard API rates apply.
That order puts working agents in front of users before the paid meter starts. Traffic data from OpenRouter already reflects the middle stretch of the timeline. The free month tests whether the same users stay once prices appear.
Internal workplace results (above 90% task success, 34% faster completion) and the 256K context window support the bet that ordinary multi-step jobs will keep running after the trial. The timeline does not depend on leapfrogging frontier scores first.
Price Gaps Steer Routine Workflow Spend
When Western flagship rates run 5-20× higher on the same token unit, the gap shows up first in high-volume, lower-stakes work. Spreadsheets, research digests, frontend fixes and CI/CD loops are exactly the tasks where Hy3 posted its clearest edge in the 270-expert blind review.
Crowd reports of 60-85% monthly inference savings after switches to Chinese open or low-price models describe the same pressure from the buyer side. A startup that cannot fund premium APIs can still run concurrent multi-agent flows inside WorkBuddy, then move to TokenHub routing or Apache 2.0 self-hosting once the free period closes.
Regulated buyers and government procurement remain constrained by sovereignty and audit rules. Everyone else faces a simpler comparison: does the agent finish the job at a cost the budget absorbs. Mid-tier Intelligence Index scores matter less in that frame than stable tool calls, long context and sub-dollar million-token rates.
Scarcity Bets Face a Different Demand Curve
The Stargate consortium of OpenAI, SoftBank and Oracle has talked of up to $500 billion in U.S. AI infrastructure. The thesis is straightforward: control chips, power and cloud and you control the intelligence layer. Independent estimates already warn of high-bandwidth memory strain and regional grid pressure.
Stanford’s 2026 AI Index finds the U.S.-China performance gap closed to a 2.7% lead for the top American model as of March 2026, with the two countries trading first place multiple times since early 2025. Capability convergence makes pure benchmark leads less decisive for most commercial loads.
Hy3 will not drive frontier scientific discovery or the highest-stakes autonomous systems soon. Its value sits in the large middle of workplace automation where reliability, context length, cost and ease of deployment decide adoption. A free browser workspace plus open weights plus sub-dollar APIs is a direct bid for that middle.
Chinese models already dominate large slices of measured open usage. If free trials convert even a fraction of global WorkBuddy users into paid TokenHub or self-hosted traffic after August 31, the volume gap can widen further. Western closed models retain advantages in regulated verticals, brand trust and absolute peak capability. Everyday agent traffic is drifting toward the cheapest reliable option that finishes the job.
Tencent is not trying to out-spend Stargate on compute. It is trying to put more working agents on whatever compute already exists. That second-order pressure on pricing and distribution will shape who actually captures enterprise workflow spend long before the next intelligence-index leap.
-
AI1 month agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI1 month agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
AI2 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING2 months agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO2 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
APPS2 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
NEWS2 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
AI2 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
