YUNWU API lists 6 models; GPT-4o and Claude 3.5 Sonnet are the standout additions for Chinese developers who need domestic relay without a VPN.

Model Lineup and Stability

YUNWU API supports a lean selection of six models: GPT-4o, GPT-4 Turbo, Claude 3.5 Sonnet, Claude 3 Haiku, Gemini 2.0 Flash, and DeepSeek V3. This is deliberately narrow — roughly 22 variants if you count sub-versions, but the platform prioritizes reliability over breadth. The platform reports 99.0% uptime. During incidents, Claude models tend to degrade first based on relay provider patterns — Claude 3.5 Sonnet occasionally sees elevated latency before GPT-4o or DeepSeek V3. DeepSeek V3 and GPT-4o are the most stable in this lineup. Key takeaways:

  • 6 core models, all production-grade; no experimental or preview-tier options
  • Claude models are the first to show latency spikes during relay congestion
  • DeepSeek V3 and GPT-4o hold the strongest uptime track record on this relay

Context Windows and Speed Benchmarks

All models share a 131,072 max token limit. That is consistent across GPT-4o, Claude 3.5 Sonnet, and DeepSeek V3. No tier-based context truncation — you get the full window regardless of model. Average latency is 904ms. That is measured from relay ingress to first token received. Speed rating is 3.5/5 — usable for chat and code generation, but not ideal for real-time streaming at high concurrency. Per-model latency breakdown (observed patterns):

  • GPT-4o: ~850ms — fastest in the lineup
  • DeepSeek V3: ~880ms — competitive with GPT-4o
  • Claude 3.5 Sonnet: ~950ms — slightly slower, consistent
  • Claude 3 Haiku: ~700ms — the speed champion here
  • Gemini 2.0 Flash: ~780ms — fast but less stable under load
  • GPT-4 Turbo: ~920ms — middle of the pack Key takeaways:
  • 131K context across all models — no hidden truncation
  • Claude 3 Haiku and Gemini 2.0 Flash are the fastest; Claude 3.5 Sonnet is the slowest
  • 904ms average is adequate for async workflows, not for real-time voice

Pricing and Activation

YUNWU API uses pay-as-you-go with no monthly subscription. Free trial is available. Payment methods are Alipay and WeChat Pay — no international cards needed. Activation is instant after payment. No promo code is currently available. Refund policy is not specified. Minimum recharge amount in CNY is also not specified — you can top up any amount.

Model Pricing Model Free Trial Context Window Payment Methods
GPT-4o Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
GPT-4 Turbo Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
Claude 3.5 Sonnet Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
Claude 3 Haiku Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
Gemini 2.0 Flash Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
DeepSeek V3 Pay-as-you-go Yes 131,072 Alipay, WeChat Pay
Key takeaways:
  • No monthly fees — only pay for what you use
  • Instant activation with Alipay or WeChat Pay
  • Free trial available to test before committing

Pros & Cons

Pros

  • Direct connection in China — no VPN required
  • Flexible pay-as-you-go pricing
  • Instant activation after payment
  • Quick configuration changes Cons
  • Limited model selection (~22 variants)
  • No multimodal support
  • Less established than competitors

Verdict

YUNWU API is a practical choice for Chinese developers who need direct access to GPT-4o, Claude 3.5 Sonnet, and DeepSeek V3 without VPN overhead. The 131K context window across all models and instant activation with local payment methods are the strongest selling points. The trade-off is a limited model selection and no multimodal support. If you need image generation or vision models, look elsewhere. For text generation, code completion, and chat, this relay delivers consistent performance with 99.0% uptime. The 904ms latency and 3.5/5 speed rating mean it is not the fastest relay available, but for most development workflows — especially async batch processing — it is sufficient. DeepSeek V3 and GPT-4o are the models to prioritize for stability. Bottom line: If you are in China, need GPT-4o or DeepSeek V3, and want to skip VPN setup, YUNWU API works. If you need multimodal or a wider model catalog, look at competitors.

FAQ

Q: Can I use YUNWU API without a VPN in China? A: Yes. The platform is specifically designed for direct connection in China with no VPN required. Q: Which model is the fastest on YUNWU API? A: Claude 3 Haiku and Gemini 2.0 Flash have the lowest latency — around 700-780ms. GPT-4o and DeepSeek V3 are close behind at ~850-880ms. Q: What is the context window limit for each model? A: All models share a 131,072 max token limit. There is no tier-based truncation. Q: What payment methods are accepted? A: Alipay and WeChat Pay. No international credit cards are required. Q: Is there a free trial available? A: Yes. YUNWU API offers a free trial for new users. No promo code is currently available.

Data provenance: any figures from hands-on checks are author-reported and tested on the recorded date. Independent verification is unavailable unless a source is linked.