#tHow Many Tokens?

News and pricing updates

Auto-generated from each provider's pricing page. Pricing changes, model launches, and tokenizer-accuracy upgrades, in reverse chronological order. The leaderboard at the top updates with every build.

Live ranking

AI cost leaderboard, ranked cheapest first

Cost to run a 1,000-token-input / 200-token-output prompt 1,000 times, across the eight cheapest non-deprecated models we track. Updates automatically with every pricing snapshot.

#ModelPer 1M calls
#1 GPT-5 Nano $130.00
#2 GPT-4.1 Nano $180.00
#3 Gemini 2.5 Flash-Lite $180.00
#4 DeepSeek V4 Flash $196.00
#5 Llama 3.1 8B $216.00
#6 Qwen3.5 9B $220.00
#7 GPT-4o mini $270.00
#8 GPT-5.6 Luna $440.00

full pricing changelog · try your own prompt

Models added

Added: claude-mythos-5

Added. claude-mythos-5 shares Claude Fable 5's specs and pricing — $10 / $50 with a full 1M-token context window — and appears on Anthropic's published pricing table. It is not generally available: access is limited to approved customers in Project Glasswing, for defensive cybersecurity workflows, with no self-serve sign-up. The entry is included for rate comparison and labelled accordingly.

source

Models added

Added: gemini-3-7-flash

Added. Google's newest and most capable Flash tier, at $0.75 / $3.75 with a 1,048,576-token context window. Introductory pricing runs through 2026-12-31 and increases on 2027-01-01, so this rate has a known expiry.

source

Models added

Added: gemini-3-6-flash

Added. The Flash tier directly below Gemini 3.7 Flash, at the same rate and the same 1,048,576-token context window. Same introductory pricing window through 2026-12-31.

source

Models added

Added: gemini-3-5-flash

Added. The most expensive Flash tier Google currently lists, at double the rate of the newer 3.6 and 3.7 Flash models above it. Output is 6x input. 1,048,576-token context window.

source

Models added

Added: gemini-3-5-flash-lite

Added. This is the model whose rates were misread against the Gemini 3.1 Flash-Lite row on 2026-08-12 and correctly held back by the swing guard. It now has its own entry, which closes that ambiguity: 3.5 Flash-Lite is $0.30 / $2.50 and 3.1 Flash-Lite is $0.25 / $1.50.

source

Models added

Added: deepseek-v4-pro-0813

Added. A dated V4 Pro release that Together AI lists alongside the undated DeepSeek V4 Pro at its own rate and its own context length — 1,048,576 tokens against 512,000. Both are live listings, so this is a second entry rather than a change to the existing one.

source

Update

gemini-2-5-pro: contextWindow

Correction of a stale stored value, not a capacity cut. Google's own spec page for gemini-2.5-pro publishes an input token limit of 1,048,576; the 2,000,000 we held came from a larger window that was announced but never shipped. A 48% reduction, inside the plausible-swing threshold and read from a single-model spec page rather than a multi-row table, so applied. This reorders the /longest-context-window/ ranking, where Gemini 2.5 Pro previously stood alone at the top. Rates unchanged at $1.25 / $10.00.

source

Update

gemini-lineup: contextWindow

Precision correction across five entries — Gemini 3.1 Pro, Gemini 3 Flash, Gemini 3.1 Flash-Lite, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite. We had recorded the rounded 1,000,000; every Gemini spec page publishes the exact figure as 1,048,576 tokens. A 4.9% adjustment, the same rounding fix applied to Llama 3.3 70B on 2026-08-13. No rate changes.

source

Update

qwen-3-8-2-4t: verify

Not added — no context length published. Together AI now lists two Qwen tiers above Qwen3.7-Plus: Qwen3.8-2.4T-A95B at $2.50 / $6.25 and Qwen3.7-Max at $1.25 / $3.75. Both show a blank context length in Together AI's model catalogue, so neither can be given a complete entry yet. Flagged for manual review; revisit once Alibaba or Together AI publishes the window.

source

Update

qwen-3-coder-480b: verify

Possible deprecation, verify. Qwen3 Coder 480B and DeepSeek R1 no longer appear on Together AI's pricing page or in its serverless model catalogue. One absent fetch is not proof of removal, so both entries are left active and unchanged pending a human check.

source

Update

gemini-3-flash: verify

No change. Gemini 3 Flash dropped off Google's pricing table this run, but the model catalogue still lists gemini-3-flash-preview as available and its spec page still resolves, so the entry stays active at $0.50 / $3.00. Its context window was corrected to 1,048,576 along with the rest of the Gemini lineup.

source

Update

together-catalogue: verify

Scope note. Together AI's pricing page carries a number of listings outside the families we track — among them Muse Glimmer 30B, Inkling, Cogito v2.1 671B, NVIDIA Nemotron 3 Ultra, Gemma 4 31B, LFM2.5-8B-A1B, gpt-oss-120B and MiniMax M2.7. None were added. Expanding the tracked set is a deliberate editorial decision rather than an automatic one, so it is left for a human.

source

Update

openai-lineup: verify

No change. Every OpenAI entry re-confirmed at its stored rate, including the full GPT-5.6 line (Sol $5 / $30, Terra $2 / $12, Luna $0.20 / $1.20), GPT-5.5 and GPT-5.5 Pro, the GPT-5.4 tiers, GPT-5.3, GPT-5.2, GPT-5.1, GPT-5, GPT-4.1, o3, o4-mini and GPT-4o.

source

Update

anthropic-lineup: verify

No change. Claude Opus 5 ($5 / $25), Opus 5 fast mode ($10 / $50), Sonnet 5 ($2 / $10), Haiku 4.5 ($1 / $5) and Fable 5 ($10 / $50) all re-confirmed. Sonnet 5's $2 / $10 remains the standard rate — the increase to $3 / $15 once scheduled for 2026-09-01 is still cancelled.

source

Models added

Added: kimi-k3

Added. moonshotai/Kimi-K3, Moonshot AI's flagship on Together AI and the first Kimi generation to carry a full 1M-token context window, against the 256K of the K2 line. Moonshot AI is a new provider for us. Its tokenizer is a Moonshot BPE we do not implement, so counts use the Llama-family BPE heuristic as a proxy and the entry is labelled approximate — the same treatment given to the GLM entries.

source

Models added

Added: kimi-k2-7-code

Added. moonshotai/Kimi-K2.7-Code, Moonshot AI's coding-specialised tier, at roughly a third the rate of Kimi K3 with a 256K context window (262,144 tokens). Tokenizer proxied as with Kimi K3.

source

Models added

Added: kimi-k2-6

Added. moonshotai/Kimi-K2.6, the general-purpose model of Moonshot AI's K2 generation, sharing the 256K context window of K2.7 Code but priced above it in both directions. Tokenizer proxied as with Kimi K3.

source

Models added

Added: minimax-m3

Added. MiniMaxAI/MiniMax-M3, pairing a 512K context window (524,288 tokens) with one of the lowest rates of any long-context open-weight model we track. MiniMax is a new provider for us, and its tokenizer is likewise proxied by the Llama-family heuristic, so the entry is labelled approximate.

source

Update

llama-3-3-70b: contextWindow

Precision correction, not a capacity change. We had recorded the rounded 128,000; Together AI publishes the exact figure as 131,072 tokens, which is how we already store the Qwen entries. A 2.4% adjustment, well inside the plausible-swing threshold, so applied. Rates are unchanged at $1.04 in both directions.

source

Update

gemini-3-1-flash-lite: verify

No change — this closes the review opened on 2026-08-12. Google's pricing page now returns $0.25 input / $1.50 output for Gemini 3.1 Flash-Lite, matching our stored rate exactly, and separately lists Gemini 3.5 Flash-Lite at $0.30 / $2.50. That confirms yesterday's reading of $0.30 / $2.50 against this row was a misread of the adjacent model rather than a real increase, and the swing guard was right to hold it. Every other Gemini entry re-confirmed unchanged.

source

Update

openai-lineup: verify

No change — this closes the possible-deprecation review opened on 2026-08-02. That run's fetch returned a partial lineup and omitted GPT-5.3, o3-pro and GPT-4 Turbo; this run returned the full table and all three are present at their stored rates ($1.75/$14, $20/$80 and $10/$30), so no deprecation was warranted. All thirty OpenAI entries re-confirmed exactly, including the GPT-5.6 Luna cut applied by hand on 2026-08-11.

source

Update

gpt-5-6-cyber: deferred

Deferred — no change applied. OpenAI's pricing table has added a security-focused tier, gpt-5.6-cyber at $12.50 input / $75 output (cached $1.25), alongside a gpt-5.5-cyber at the same rate. Both are the most expensive OpenAI models on the page. Prices and model ids are clear, but neither the pricing page nor the models reference publishes a context window for either, while it does publish 1.05M for Sol, Terra and Luna. Since each entry feeds the /longest-context-window/ ranking, they stay out until the figure is published.

source

Update

together-untracked-models: deferred

Still deferred — no change applied. Three models on Together AI's pricing page have clear rates but no published context length in the serverless-models reference: Qwen3.8-2.4T-A95B ($2.50/$6.25, new to this run and the largest Qwen tier listed), Qwen3.7-Max ($1.25/$3.75, unchanged and deferred since 2026-08-02) and MiniMax M2.7 ($0.30/$1.20, priced identically to the M3 added this run). Its sibling MiniMax M3 was added because its context length is published. Also untracked by choice: Qwen3 235B and the small legacy tiers Qwen2.5 7B Instruct Turbo and Llama 3 8B Instruct Lite.

source

Update

claude-mythos-5: deferred

Still deferred — no change applied, blocker unchanged. Claude Mythos 5 remains listed at $10 / $50, matching Fable 5, and the long-context section still confirms the full 1M-token window at standard pricing. It is still limited availability behind a waitlist, and the page still names it inconsistently (Claude Mythos 5 in the pricing table, Claude Mythos Preview elsewhere) without publishing an API model id. Adding it would mean inventing the string users copy into their code. Every other Anthropic entry re-confirmed unchanged, including the Sonnet 5 rate of $2 / $10 now recorded as permanent.

source

Update

gemini-3-5-and-3-6-flash: deferred

Still deferred for a fifth run — no change applied, blocker unchanged. Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50) all hold their prices, and Google still publishes no input token limit for any of them. A fourth model, Gemini 3.5 Live Translate ($3.50/$21.00), is new to this run and has the same gap.

source

Update

claude-sonnet: updated

No rate change — $2 input / $10 output is unchanged. What changed is that the increase we had been warning about is cancelled. Anthropic now states the $2/$10 rate announced as introductory pricing through 2026-08-31 is the standard price, and the scheduled move to $3/$15 on 2026-09-01 will not occur. The model note has been rewritten, since it previously told readers to budget for a 50% increase in under three weeks.

source

Models added

Added: glm-5-2

Added, resolving the deferral opened on 2026-08-02. Zhipu AI’s newest flagship on Together AI, priced identically to the GLM-5.1 we already carry but with a 512K-token context window against GLM-5.1’s 128K. The earlier deferral was blocked on the missing API string and context window; both come from Together’s serverless-models reference, which lists zai-org/GLM-5.2 at 512,000 tokens. Tokenizer remains the Llama-family BPE proxy used for GLM-5.1 (approx-10pct).

source

Models added

Added: qwen-3-7-plus

Added, resolving the deferral opened on 2026-08-02. Alibaba’s newest general-purpose tier on Together AI, Qwen/Qwen3.7-Plus at a 1M-token context window. Cheaper in both directions than the Qwen3.6-Plus it succeeds. The sibling Qwen3.7-Max stays deferred — see its own entry.

source

Models added

Added: qwen-3-6-plus

Added, resolving the deferral opened on 2026-08-02. Qwen/Qwen3.6-Plus, the first Qwen generation on Together AI to carry a full 1M-token context window. Already superseded on price by Qwen3.7-Plus, but still listed and callable.

source

Models added

Added: qwen-3-5-9b

Added. Qwen/Qwen3.5-9B, the small dense model of the Qwen3.5 generation, carrying the same 256K native context (262,144 tokens) as its 397B sibling at a fraction of the rate.

source

Models added

Added: deepseek-v4-flash

Added. deepseek-ai/DeepSeek-V4-Flash-0731, the small fast tier of DeepSeek’s V4 generation, at roughly one twelfth the rate of V4 Pro. Notable for a 1M-token context window — twice that of the larger V4 Pro.

source

Update

gemini-3-1-flash-lite: updated

Model id only — no rate change. Gemini 3.1 Flash-Lite has moved from preview to generally available and the id has dropped its -preview suffix. Callers pinned to the old preview id should update.

source

Update

gemini-3-1-flash-lite: flagged

Flagged for manual review — no change applied, and the stored rate is very likely correct. A read of Google’s pricing page returned $0.30 input / $2.50 output for this row against our stored $0.25 / $1.50. The output figure is a 67% jump, past the 60% threshold at which an automated read is not trusted. It is almost certainly a row misread rather than a real increase: $0.30 / $2.50 are the rates recorded for the separate Gemini 3.5 Flash-Lite model, which sits adjacent in the same table and was seen at exactly those figures in the 2026-08-02 run. The stored $0.25 / $1.50 stands.

source

Update

qwen-3-7-max: deferred

Still deferred — no change applied. Qwen3.7-Max is listed at $1.25 / $3.75, unchanged from the 2026-08-02 read, and the API string Qwen/Qwen3.7-Max is now confirmed. The remaining blocker is the context length, which Together’s serverless-models reference leaves blank for this model while publishing it for every other Qwen3.x tier. Since each entry feeds the /longest-context-window/ ranking, it stays out until the figure is published. Its siblings Qwen3.7-Plus and Qwen3.6-Plus were added this run now that their context lengths are available.

source

Update

claude-mythos-5: deferred

Deferred — no change applied. Anthropic’s pricing table has added Claude Mythos 5 at $10 / $50, matching Fable 5, and the long-context section confirms it carries the full 1M-token context window at standard pricing. Two things hold it back: access is limited availability behind a waitlist rather than general release, and the page names it inconsistently (Claude Mythos 5 in the pricing table, Claude Mythos Preview in the long-context section) without publishing an API model id anywhere. Adding it would mean inventing the string users copy into their code. It stays out until the id is published.

source

Update

gemini-3-5-and-3-6-flash: deferred

Still deferred for a fourth run — no change applied, and the blocker is unchanged. Gemini 3.6 Flash ($1.50/$7.50) and Gemini 3.5 Flash ($1.50/$9.00) are both confirmed Stable with ids gemini-3.6-flash and gemini-3.5-flash, but Google still publishes no input token limit for either, on the pricing page, the models overview, or the per-model pages (which returned 404 on the slug patterns tried). Adding them would mean inventing the one figure /longest-context-window/ is built on.

source

Update

openai-lineup: verify

No changes. Every OpenAI rate that appeared in this fetch matched our stored values exactly, including the GPT-5.6 Luna cut applied by hand on 2026-08-11 ($0.20 / $1.20), which the page now confirms — closing out the review opened on 2026-08-02. Sol ($5/$30), Terra ($2/$12), GPT-5.5, 5.4, 5.4-mini, 5.2, 5.1, 5, 4.1, o3, o4-mini and 4o all re-confirmed. The fetch returned a partial lineup — roughly a dozen models against the thirty we track, omitting the Pro and Nano tiers and GPT-5.3 — so no deprecation conclusions were drawn from the gaps.

source

Price change

gpt-5-6-luna input price 1 → 0.2

HUMAN-VERIFIED, now applied. The 80% cut flagged on 2026-08-02 (see the flagged entry below) was confirmed against OpenAI's pricing page and corroborated by multiple third-party trackers — OpenAI cut Luna 80% on 2026-07-30, from $1.00/$6.00 (cached $0.10) to $0.20/$1.20 (cached $0.02). The swing exceeded the automated 60% guard, so it was correctly held for manual review; this entry records the verified application. Luna now leads the /cheapest-ai-model/ ranking.

source

Price change

gpt-5-6-luna output price 6 → 1.2

See the input change above — same 80% reduction (cached input also cut from $0.10 to $0.02).

source

Price change

gpt-5-6-terra input price 2.5 → 2

OpenAI cut GPT-5.6 Terra by 20% in both directions one week after launch. Cached input drops in step, from $0.25 to $0.20 per 1M. Terra now undercuts GPT-5.4 ($2.50/$15) on price while carrying roughly 2.6x the context window, which makes the older model hard to justify for new work. The other two GPT-5.6 tiers were not cut in the same proportion: Sol held at $5/$30, and the figure shown for Luna was too large a swing to apply automatically (see the Luna entry below).

source

Price change

gpt-5-6-terra output price 15 → 12

See the input change above — same 20% reduction applied to output.

source

Update

gpt-5-6-luna: flagged

FLAGGED FOR MANUAL REVIEW — no change applied. OpenAI's pricing page listed GPT-5.6 Luna at $0.20 input / $1.20 output (cached $0.02) against our stored $1.00 / $6.00 (cached $0.10). That is an 80% reduction, well past the 60% threshold at which we stop trusting an automated read and require a human to confirm. Two things argue it may be genuine rather than a parse error: it is an exact 5x cut applied consistently across input, output and cached input, and its sibling Terra was independently confirmed cut 20% in the same fetch. But an 80% swing is also exactly what a misread pricing table looks like, and Luna feeds the /cheapest-ai-model/ ranking, where a wrong figure would reorder the page. The stored rate stands at $1.00 / $6.00 until verified by hand.

source

Update

gemini-3-5-and-3-6-flash: deferred

STILL DEFERRED for a third run — no change applied. Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50, new to this run) all have confirmed, stable prices, and the models overview now lists all three as Stable with model ids gemini-3.6-flash, gemini-3.5-flash and gemini-3.5-flash-lite. The blocker is unchanged: Google still does not publish an input token limit for any of them, on either the models overview or the Vertex AI reference. Since every model entry feeds the /longest-context-window/ ranking, adding them would mean inventing the one figure that page is built on. They stay out until Google publishes the limit.

source

Update

together-new-models: deferred

DEFERRED — no change applied. Together's pricing page has picked up a batch of models we do not yet track, most notably GLM-5.2 ($1.40/$4.40, priced identically to the GLM-5.1 we already carry) and a Qwen3.7 line (Plus at $0.32/$1.28, Max at $1.25/$3.75), alongside Kimi K3, MiniMax M3 and Qwen3.6-Plus. Prices are clear, but the pricing table alone gives neither a context window nor an exact API model string, and Together's per-model pages for GLM-5.2 returned 404 on the slug patterns tried. Deferred rather than guessed. Existing entries were all re-confirmed against this fetch: Llama 3.3 70B ($1.04/$1.04), DeepSeek V4 Pro ($1.74/$3.48), GLM-5.1 ($1.40/$4.40) and Qwen3.5 397B ($0.60/$3.60) are unchanged.

source

Update

openai-absent-models: verify

POSSIBLE DEPRECATION, VERIFY — no change applied, all entries left active. GPT-5.3 ($1.75/$14), o3-pro ($20/$80) and GPT-4 Turbo ($10/$30) did not appear in this run's OpenAI pricing fetch. A single absent fetch is not evidence of removal — pricing pages routinely drop older tiers into collapsed or secondary tables — so per policy the entries stay live and unchanged. If they are still missing on the next two runs, that is worth a manual check against OpenAI's deprecations page. Every other OpenAI model we track was re-confirmed at its stored rate.

source

Models added

Added: qwen-3-5-397b

Added Qwen3.5 397B at $0.60 input / $3.60 output per 1M tokens (cached input $0.35) with a 262,144-token native context window. This clears one of the models deferred earlier today: the price was confirmed on Together's pricing page, and the context window and exact API id (Qwen/Qwen3.5-397B-A17B) have now been sourced from Together's own model page, so no figure is being guessed. A 397B-parameter MoE with 17B active, it leads the Qwen3.5 generation that has displaced Qwen3 Coder 480B and Qwen 2.5 72B on Together's main pricing table. Note that it breaks the older Qwen convention of charging the same rate in both directions — output is 6x input here.

source

Update

gemini-new-models: deferred

STILL DEFERRED — no change applied. Re-checked the Gemini 3.6 Flash ($1.50/$7.50) and Gemini 3.5 Flash ($1.50/$9.00) entries held back earlier today. Prices re-confirmed on the pricing page, but the context windows are still unpublished: the Gemini models overview lists both as Stable without token limits, and the Vertex AI model reference does not state them either. Held back for a second run rather than shipping a guessed context window that would feed the /longest-context-window/ ranking. Separately, Gemini 3 Flash — which did not appear in this run's pricing fetch — is confirmed still available on the models overview under Preview status, not in the deprecated section, so its entry stays active and unchanged at $0.50/$3.00.

source

Update

claude-opus: updated

Anthropic shipped Claude Opus 5. The claude-opus entry now resolves to Opus 5 (apiId: claude-opus-5), following the same convention used for the 4.7 to 4.8 repoint. Standard-mode pricing is unchanged at $5 input / $25 output per 1M tokens — Anthropic has now held the Opus rate flat across 4.5, 4.6, 4.7, 4.8 and 5. Opus 4.8 and 4.7 remain available at the same rate.

source

Update

claude-opus-fast: updated

Fast mode (still a research preview) now covers Claude Opus 5 as well as Opus 4.8, at an unchanged $10 input / $50 output per 1M tokens. Fast-mode pricing applies across the full context window including requests over 200k input tokens. Claude API first-party only — not on Bedrock, Google Cloud, or the Batch API, and not available on Opus 4.7 or 4.6.

source

Price change

claude-sonnet input price 3 → 2

The claude-sonnet entry now resolves to Claude Sonnet 5 (apiId: claude-sonnet-5). Sonnet 5 is on INTRODUCTORY pricing of $2 input / $10 output through 2026-08-31; standard pricing of $3 / $15 takes effect 2026-09-01. This is a scheduled, dated increase published by Anthropic, not a promotion of unknown duration — anyone budgeting past August should plan on $3 / $15. Sonnet 4.6 remains available at $3 / $15.

source

Price change

claude-sonnet output price 15 → 10

See the input change above — Sonnet 5 introductory pricing, reverting to $15 output on 2026-09-01.

source

Update

claude-family: contextWindow

Corrected the context window for claude-opus, claude-opus-fast and claude-sonnet from 200K to 1M. Anthropic's pricing docs now state that Claude 4.6 and later models include the full 1M-token context window at standard pricing, with no long-context surcharge — a 900k-token request bills at the same per-token rate as a 9k one. Prompt caching and batch discounts apply at standard rates across the whole window. Claude Haiku 4.5 is below the 4.6 cutoff and stays at 200K.

source

Models added

Added: claude-fable-5

Added Claude Fable 5 at $10 input / $50 output per 1M tokens with a 1M-token context window — a new tier priced above Opus. Note that Fable 5 uses the newer Claude 4.7+ tokenizer, which produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier, so the real cost gap versus older Claude models is wider than the headline rate. Claude Mythos 5 is priced identically but was NOT added: it is limited-availability (via anthropic.com/glasswing) rather than generally purchasable.

source

Models added

Added: openai-5-6

Added the GPT-5.6 line: Sol ($5/$30, cached $0.50), Terra ($2.50/$15, cached $0.25) and Luna ($1/$6, cached $0.10). All three carry a 1.05M-token context window, up from 400K across the GPT-5.0 to 5.5 generations. Sol and Terra match the headline rates of GPT-5.5 and GPT-5.4 respectively while offering ~2.6x the context, so there is little reason to stay on the older models at the same price.

source

Price change

llama-3-3-70b input price 0.88 → 1.04

Together.ai raised Llama 3.3 70B from $0.88 to $1.04 per 1M tokens, an 18% increase. Within the plausible-swing threshold and confirmed on the live pricing page, so applied.

source

Price change

llama-3-3-70b output price 0.88 → 1.04

See the input change above — Together prices Llama 3.3 70B identically in both directions.

source

Models added

Added: deepseek-v4-pro

Added DeepSeek V4 Pro at $1.74 input / $3.48 output per 1M tokens (cached input $0.20) with a 512K-token context window — the largest context of any open-weight model we track, and roughly 4x the 128K that the V3 generation offered. It has replaced V3.1 as the DeepSeek entry on Together's main pricing page.

source

Update

gemini-new-models: deferred

NOT ADDED PENDING VERIFICATION. Three new Gemini models appeared with confirmed prices — Gemini 3.6 Flash ($1.50/$7.50), Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.5 Flash-Lite ($0.30/$2.50) — but neither the pricing page nor the models overview page publishes their context windows or exact API model ids, and the per-model doc pages 404. Rather than guess a context window (which would feed the /longest-context-window/ ranking directly), these are held back until the figures can be sourced. All six existing Gemini entries were re-verified against the live page and are unchanged.

source

Update

together-oss-new: deferred

NOT ADDED PENDING VERIFICATION. Together listed several new models whose context windows could not be sourced (their model pages 404): GLM-5.2 ($1.40/$4.40, identical to GLM-5.1), Qwen3.7-Plus ($0.32/$1.28), Qwen3.6-Plus ($0.50/$3.00), Qwen3.5-397B-A17B ($0.60/$3.60) and Qwen3.5 9B ($0.17/$0.25). Prices recorded here for the next run; entries held back rather than shipped with a guessed context window.

source

Update

openai-legacy: possible-deprecation

POSSIBLE DEPRECATION, VERIFY — no change applied. The GPT-5.0 through 5.2 tiers, the GPT-4.1 family, the o3/o4 reasoning series and the GPT-4o family did not appear in this run's fetch of either the pricing page or the models page, which now foreground the GPT-5.6 line. That is suggestive but not proof of removal — the fetch may simply have captured the flagship section. All entries left active per the never-remove-on-one-fetch rule. Worth a manual check of the OpenAI deprecations page before flipping any flags.

source

Update

together-oss-missing: possible-deprecation

POSSIBLE DEPRECATION, VERIFY — no change applied. DeepSeek V3.1, DeepSeek R1, Qwen3 Coder 480B and Qwen 2.5 72B are no longer visible on Together's serverless pricing table, which now leads with DeepSeek V4 Pro and the Qwen3.5+ generation. These entries were already flagged as indicative pricing and are left active. Anyone billing against them should confirm their provider's current rate.

source

Update

claude-opus: updated

Anthropic released Claude Opus 4.8 on 2026-05-28. Standard mode pricing unchanged at $5 input / $25 output per 1M tokens. The claude-opus entry now resolves to 4.8 (apiId: claude-opus-4-8). Inherits the Opus 4.7 tokenizer behavior (up to 35% more tokens than legacy Claude models for the same text). Anthropic describes 4.8 as 'a modest but tangible improvement' with gains in agentic coding, reasoning, knowledge work, and honesty.

source

Models added

Added: claude-opus-fast

Added Claude Opus 4.8 Fast Mode as a separate entry: $10 input / $50 output per 1M tokens, producing tokens at ~2.5x normal speed. Anthropic dropped the fast-tier price 3x vs Opus 4.7 (was $30/$150). Useful for latency-sensitive workloads where Opus quality is needed but the standard tier's throughput is the bottleneck.

source

Accuracy upgrade

llama-family tokenizer: approx-3pct → exact

Shipped real BPE tokenization for the Llama family via llama-tokenizer-js (lazy-loaded ~2MB chunk on first Llama count). All llama-* models now labeled 'exact' instead of '≈±3%'. Mistral / Qwen / DeepSeek / GLM still use heuristic (now character-class-aware — buckets text into ASCII / digit / CJK / whitespace and applies per-class ratios — more accurate than the prior constant-ratio version, still labeled ≈±3%).

source

Models added

Added: oss

Added 5 OSS models: Llama 3.3 70B ($0.88/$0.88, current Together flagship Meta), DeepSeek V3.1 ($0.60/$1.70 Together listing), DeepSeek R1 ($3/$7 reasoning), Qwen3 Coder 480B ($2/$2 current Alibaba coding flagship), GLM-5.1 ($1.40/$4.40 new Zhipu provider). Llama 3.1 entries kept but flagged 'no longer on Together's main page — verify provider'. Llama 4 NOT added — not on Together's current pricing page.

source

Models added

Added: openai

Added 21 OpenAI models: full GPT-5 family (5, 5.1, 5.2, 5.2 Pro, 5.3, 5.4, 5.4 Mini, 5.4 Nano, 5.4 Pro, 5.5, 5.5 Pro, mini, nano, Pro tiers), GPT-4.1 family (4.1, mini, nano), and o-series reasoning (o3, o3-mini, o3-pro, o4-mini).

source

Models added

Added: google

Added Gemini 3 family: 3.1 Pro Preview ($2/$12), 3 Flash Preview ($0.50/$3), 3.1 Flash-Lite Preview ($0.25/$1.50). Added Gemini 2.5 Flash-Lite GA ($0.10/$0.40).

source

Price change

gemini-2-5-flash input price 0.075 → 0.3

Correction — Gemini 2.5 Flash priced higher than initial entry. Live Google pricing page lists $0.30 input / $2.50 output.

source

Price change

gemini-2-5-flash output price 0.3 → 2.5

Correction — see input change above.

source

Price change

claude-opus input price 15 → 5

Correction — initial Opus 4.7 entry was based on prior-generation Opus pricing. Anthropic prices Opus 4.7 at $5 input / $25 output.

source

Price change

claude-opus output price 75 → 25

Correction — see input change above.

source

Price change

claude-haiku input price 0.8 → 1

Correction — Haiku 4.5 priced higher than initial entry suggested.

source

Price change

claude-haiku output price 4 → 5

Correction — see input change above.

source

Update

all: initial

Initial pricing snapshot — 15 models across 7 providers.

source