Claude 3.5 Sonnet through Cubence ranks in the top 3 on the hvoy.ai Claude Speed leaderboard — that’s the headline for anyone who’s tired of waiting 30 seconds for a Sonnet response on other relay stations. Cubence isn’t trying to be a one-stop model hub. It targets exactly one use case: fast Claude access from China, no VPN required. The model list is intentionally narrow — just GPT-4o and Claude 3.5 Sonnet — and the architecture is built around speed, not breadth.
Pricing & Payment
There’s no free trial. You pay upfront via Alipay or WeChat Pay, and pricing is premium — comparable to PackyCode. The trade-off is that you’re paying for “almost no dilution” on a max group setup. Other relay stations pool users aggressively, which kills response speed during peak hours. Cubence keeps group sizes small.
| Plan | Price | Notes |
|---|---|---|
| Pay-as-you-go | Premium (similar to PackyCode) | No free trial; recharge via Alipay/WeChat Pay |
| Min recharge | Not specified | Likely flexible top-up |
| No promo code is available at this time. |
Speed & Uptime
The 98% uptime is decent but not market-leading — some competitors hit 99.5%+. What matters more is the per-request latency. Cubence monitors its own performance actively, and the Claude Speed leaderboard ranking confirms low response times for Sonnet requests. The 100,000 max token context is standard for Claude 3.5 Sonnet. If you’re working with long documents or multi-turn agent loops, you’ll hit that ceiling less often than the 32K or 64K limits on cheaper relays.
API Compatibility
The platform uses a standard reverse-proxy setup. You call the API endpoint with your OpenAI-compatible client (LangChain, openai-python, NextChat, etc.) and point it at Cubence’s relay URL. No SDKs, no custom wrappers — just swap the base URL and API key. If you’ve used other relay stations like API2D or OhMyGPT, the integration process is identical. The difference is in the backend routing, not the client interface.
Pros & Cons
Pros
- Top 3 on Claude Speed leaderboard — verified low latency
- Max group configuration with minimal request dilution
- Active 24/7 monitoring; you’ll know immediately if something breaks
- Fair reverse-proxy pricing on some model groups Cons
- Premium pricing — not for budget-conscious developers
- No free trial; you must recharge before testing
- Limited to 2 models — no Gemini, DeepSeek, or open-source options
Verdict
Cubence is for one kind of developer: someone who needs Claude 3.5 Sonnet to respond fast and is willing to pay for it. If you’re running a customer-facing chatbot, an agent loop, or a real-time coding assistant where a 5-second delay breaks the UX, Cubence delivers where cheaper relays choke. If you need model diversity — DeepSeek for cheap reasoning, Gemini for long-context vision, or GPT-4o-mini for batch processing — look elsewhere. Cubence doesn’t pretend to be a general-purpose relay. For Claude speed specifically, this is one of the best options from China right now. Just factor in the cost.
FAQ
Can I use Cubence with my existing OpenAI-compatible client?
Yes. Replace the base URL in your client with Cubence’s relay endpoint and set the API key. Works with LangChain, openai-python, NextChat, and any tool that supports custom OpenAI API endpoints.
Does Cubence support streaming responses?
Streaming is supported on both GPT-4o and Claude 3.5 Sonnet. Given the speed ranking, streaming latency is lower than most alternatives.
Is there a minimum recharge amount?
Not specified in the platform data. Expect a flexible top-up model similar to other Chinese relay stations — likely starting at 10–50 CNY.
How does Cubence compare to free relay stations?
Cubence has no free tier. Free stations typically dilute groups heavily (100+ users per group) and cap rate limits. Cubence keeps groups small and charges premium pricing for consistent low-latency responses.
What happens if the Claude API is down on the upstream side?
Cubence’s active monitoring means they detect upstream outages quickly. The 98% uptime figure accounts for both relay and upstream failures. If Claude itself is down, no relay can help.
Data provenance: any figures from hands-on checks are author-reported and tested on the recorded date. Independent verification is unavailable unless a source is linked.