Confronto reale quote token dei modelli AI

Quote reali, ricalcolate da test della community e delle redazioni — non i tetti massimi ufficiali.

There is no single "best" plan — it depends on your schedule. For night-time heavy use, Alibaba's Qwen 3.8 Max Token Plan (~1.5B tokens/month for ¥39) is unmatched. For all-day coding, Kimi K3 Allegretto (¥159/mo, ~950M tokens) is the steadier pick.

Ultima verifica:

Confronto quote fianco a fianco

ModelloMinimo /mese↕Quota mensile↕Token / ¥↕Finestra 5hSconto notturnoIdeale perDettagli
Kimi K3 (Allegretto)Moonshot AI¥159~950M~6M~40MNoneAll-day coding & agentsDettagli →
Qwen 3.8 Max Token PlanAlibaba Cloud¥39~1.5B (night)~38M (night)700 credits / 5h80% off (night)Night-owl heavy usersDettagli →
GLM 5.2Zhipu AI———————

I dati sono ricalcolati dall'uso reale degli utenti della community, non tetti fissi ufficiali. I provider possono regolare la quota dinamicamente. Clicca su un'intestazione di colonna per ordinare.

Quale piano fa per te?

Occasional Q&A / light use
→
Start free — Qoder or the Kimi free tier. Upgrade only when you hit limits.

Come stimiamo la quota

What "quota back-calculation" means

Official pages rarely publish a hard token cap. We infer it: Total quota ≈ tokens already used ÷ usage-percentage shown. E.g. 15M tokens at 37.3% of the 5h bar implies ~40M for that window. We cross-check across the 5h, 7-day and monthly bars.

Why no official fixed cap is published

Caps are dynamic: they shift with model version, server load, and business strategy. A published number becomes a promise and a target for abuse, so providers keep them fuzzy on purpose. Treat every figure here as a snapshot, not a contract.

How cache hit-rate changes real consumption

In coding/agent workflows the same context is sent again and again. One user logged a 93.7% cache-hit rate — meaning only ~6% of input was billed at full price. The quota you "feel" is therefore far larger than the raw token count suggests.

Subscription vs pay-as-you-go API

A subscription trades flexibility for a flat price within rolling windows. Pay-as-you-go API bills every token but never rate-limits you. If your usage is steady and predictable, subscriptions win; if it's spiky or you need concurrency, API wins.

Insight di settore

Why Chinese models push subscriptions over pure API

Subscriptions lock in monthly revenue and reduce billing friction for consumers who distrust pay-per-token unpredictability. They also let providers smooth demand with quota windows instead of raw price spikes.

The compute-scheduling logic behind night discounts

Inference clusters are over-provisioned for daytime peaks and idle at night. A steep night discount (Qwen's 0.2 折) simply prices the idle capacity — it costs the provider almost nothing extra and fills otherwise-wasted GPUs.

Where the gap between "price" and "felt value" comes from

Three hidden levers decide what you actually get: cache hit-rate, the input/output ratio of your workload, and time-of-day discounts. Two users on the same plan can see 50× different real throughput — which is exactly why we back-calculate instead of trusting the sticker.

Schede rapide dei modelli

Kimi K3 (Allegretto)Best all-day coding pick

The most balanced choice for daily coders: ~950M tokens/month, a ~40M 5h window, and 20× Code quota. Available all day, no night-only catch.

Da /mese
Â¥159
Mensile
~950M
Finestra 5h
~40M
Leggi il test completo →
Qwen 3.8 Max Token PlanBest value — at night

Unbeatable value IF you run heavy jobs at night (22:00–08:00): ~1.5B tokens/month for ¥39. Daytime use drops to ~300M, and without the promo discount to ~30M.

Da /mese
Â¥39
Mensile
~1.5B (night)
Finestra 5h
700 credits / 5h
Leggi il test completo →

FAQ

Which is better value, Kimi K3 or Qwen 3.8 Max?

It depends on your hours. Qwen 3.8 Max is far cheaper per token IF you work at night (~38M tokens/yuan vs ~6M for Kimi) — but daytime Qwen drops to ~300M tokens/month. Kimi K3 gives a steady ~950M tokens/month all day with no time restriction. Night-owls: Qwen. All-day coders: Kimi.

Which Chinese AI plan is the best value overall?

For pure tokens-per-yuan, Qwen 3.8 Max at night is unmatched. For dependability and a coding-focused toolchain, Kimi K3 Allegretto is the strongest all-rounder. For trying things out, start with Qoder's free Qwen 3.8 Max quota before paying anything.

What happens when I hit the quota limit?

Usually a prompt tells you the quota is exhausted; you wait for the rolling window to release (5h windows free up hour by hour, 7-day windows day by day) or upgrade. Some providers throttle quality or speed rather than hard-blocking.

Which runs out first — the 5-hour or the monthly quota?

For heavy, bursty sessions (agent coding, big refactors) the 5-hour bar is almost always the bottleneck — you can exhaust it well before the monthly bar moves. For steady trickle usage, the monthly bar is what you'll watch.

Are there free Chinese large models I can use?

Yes. Qoder offers a free tier running Qwen 3.8 Max, and Kimi has a free (limited) tier. They're the best way to gauge your real usage before paying for a subscription.

Lettura correlata