Comparativa real de cuotas de tokens de modelos de IA

Cuotas reales, calculadas a partir de pruebas de la comunidad y de la redacción, no topes oficiales de catálogo.

There is no single "best" plan — it depends on your schedule. For night-time heavy use, Alibaba's Qwen 3.8 Max Token Plan (~1.5B tokens/month for ¥39) is unmatched. For all-day coding, Kimi K3 Allegretto (¥159/mo, ~950M tokens) is the steadier pick.

Última verificación:

Tabla comparativa de cuotas, uno al lado del otro

ModeloMínimo/mesCuota mensualTokens/¥Ventana 5hDescuento nocturnoIdeal paraDetalles
Kimi K3 (Allegretto)Moonshot AI¥159~950M~6M~40MNoneAll-day coding & agentsDetalles →
Qwen 3.8 Max Token PlanAlibaba Cloud¥39~1.5B (night)~38M (night)700 credits / 5h80% off (night)Night-owl heavy usersDetalles →
GLM 5.2Zhipu AI

Los datos se calculan a la inversa a partir del uso real de usuarios de la comunidad, no de topes oficiales fijos. Los proveedores pueden ajustar la cuota dinámicamente. Haz clic en una cabecera de columna para ordenar.

¿Qué plan te encaja?

Occasional Q&A / light use
Start free — Qoder or the Kimi free tier. Upgrade only when you hit limits.

Cómo estimamos la cuota

What "quota back-calculation" means

Official pages rarely publish a hard token cap. We infer it: Total quota ≈ tokens already used ÷ usage-percentage shown. E.g. 15M tokens at 37.3% of the 5h bar implies ~40M for that window. We cross-check across the 5h, 7-day and monthly bars.

Why no official fixed cap is published

Caps are dynamic: they shift with model version, server load, and business strategy. A published number becomes a promise and a target for abuse, so providers keep them fuzzy on purpose. Treat every figure here as a snapshot, not a contract.

How cache hit-rate changes real consumption

In coding/agent workflows the same context is sent again and again. One user logged a 93.7% cache-hit rate — meaning only ~6% of input was billed at full price. The quota you "feel" is therefore far larger than the raw token count suggests.

Subscription vs pay-as-you-go API

A subscription trades flexibility for a flat price within rolling windows. Pay-as-you-go API bills every token but never rate-limits you. If your usage is steady and predictable, subscriptions win; if it's spiky or you need concurrency, API wins.

Claves del sector

Why Chinese models push subscriptions over pure API

Subscriptions lock in monthly revenue and reduce billing friction for consumers who distrust pay-per-token unpredictability. They also let providers smooth demand with quota windows instead of raw price spikes.

The compute-scheduling logic behind night discounts

Inference clusters are over-provisioned for daytime peaks and idle at night. A steep night discount (Qwen's 0.2 折) simply prices the idle capacity — it costs the provider almost nothing extra and fills otherwise-wasted GPUs.

Where the gap between "price" and "felt value" comes from

Three hidden levers decide what you actually get: cache hit-rate, the input/output ratio of your workload, and time-of-day discounts. Two users on the same plan can see 50× different real throughput — which is exactly why we back-calculate instead of trusting the sticker.

Fichas rápidas de modelos

Kimi K3 (Allegretto)Best all-day coding pick

The most balanced choice for daily coders: ~950M tokens/month, a ~40M 5h window, and 20× Code quota. Available all day, no night-only catch.

Desde/mes
¥159
Mensual
~950M
Ventana 5h
~40M
Ver la prueba completa →
Qwen 3.8 Max Token PlanBest value — at night

Unbeatable value IF you run heavy jobs at night (22:00–08:00): ~1.5B tokens/month for ¥39. Daytime use drops to ~300M, and without the promo discount to ~30M.

Desde/mes
¥39
Mensual
~1.5B (night)
Ventana 5h
700 credits / 5h
Ver la prueba completa →

Preguntas frecuentes

Which is better value, Kimi K3 or Qwen 3.8 Max?

It depends on your hours. Qwen 3.8 Max is far cheaper per token IF you work at night (~38M tokens/yuan vs ~6M for Kimi) — but daytime Qwen drops to ~300M tokens/month. Kimi K3 gives a steady ~950M tokens/month all day with no time restriction. Night-owls: Qwen. All-day coders: Kimi.

Which Chinese AI plan is the best value overall?

For pure tokens-per-yuan, Qwen 3.8 Max at night is unmatched. For dependability and a coding-focused toolchain, Kimi K3 Allegretto is the strongest all-rounder. For trying things out, start with Qoder's free Qwen 3.8 Max quota before paying anything.

What happens when I hit the quota limit?

Usually a prompt tells you the quota is exhausted; you wait for the rolling window to release (5h windows free up hour by hour, 7-day windows day by day) or upgrade. Some providers throttle quality or speed rather than hard-blocking.

Which runs out first — the 5-hour or the monthly quota?

For heavy, bursty sessions (agent coding, big refactors) the 5-hour bar is almost always the bottleneck — you can exhaust it well before the monthly bar moves. For steady trickle usage, the monthly bar is what you'll watch.

Are there free Chinese large models I can use?

Yes. Qoder offers a free tier running Qwen 3.8 Max, and Kimi has a free (limited) tier. They're the best way to gauge your real usage before paying for a subscription.

Lecturas relacionadas