AI Model Cost per Task 2026: Frontier vs Open-Weight Benchmarks
The short answer: in 2026 a single AI task costs anywhere from well under a cent to a few dollars, and the spread is almost entirely model choice. Published reference points (Artificial Analysis, Aug 2026) put Muse Spark 1.2 at ~$0.40 per task and Claude Opus 5 at ~$2.34 per task — a ~6x spread on the same class of work. This page benchmarks per-task cost across the frontier and open-weight models agencies actually route to, with methodology and dated sources.
Capacity context: Anthropic's reported $45 billion Nscale commitment and its SpaceX Colossus 1 compute deal signal real Claude capacity growth — but also pricing pressure ahead of the record IPO. For the full breakdown, see Anthropic's $45B Nscale compute deal: what it means for Claude capacity and pricing.
Published per-task reference points (Artificial Analysis, Aug 2026)
| Model | Reference cost per task | Source |
|---|---|---|
| Meta Muse Spark 1.2 | ~$0.40 | Artificial Analysis (Aug 2026) |
| Claude Opus 5 | ~$2.34 | Artificial Analysis (Aug 2026) |
| DeepSeek V4 Flash (pre-hike) | ~$0.03 | Artificial Analysis (Aug 2026; pre-Aug 16 rates) |
These are vendor-independent benchmark figures, not our own measurements. Note the DeepSeek row is the pre-hike reference — after the Aug 16, 2026 increase, the per-task cost for DeepSeek V4 depends heavily on peak vs off-peak and cache hits.
Current per-1M rates (Aug 28, 2026) — the inputs for your own math
| Model | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|
| Qwen 3.8 Flash (Alibaba) | $0.15 | $0.47 | Open-weight 125B MoE, ~6B active (~95% sparsity), multimodal; qwen-community-1.0 license; cache read $0.016; 1M context (YaRN); verified Aug 27–28, 2026 |
| GPT-5.6 Luna | $0.20 | $1.20 | −80% Jul 30, 2026 (from $1/$6); 13.8x usage in the Jul 27 – Aug 14 discount window |
| DeepSeek V4-Flash (off-peak) | $0.22 | $0.66 | Peak $0.44/$1.32; cache-hit input $0.007 off-peak / $0.014 peak |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | Peak $1.32/$3.96; cache-hit input $0.022 off-peak / $0.044 peak |
| Gemini 3.8 Flash (intro) | $0.75 | $3.75 | Released Sept 2, 2026 — current Gemini Flash; intro through 2026-12-31, then $1.50/$7.50; 1M context |
| Gemini 3.7 Flash (intro — previous version) | $0.75 | $3.75 | Intro through 2026-12-31, then $1.50/$7.50; 1M context; superseded by 3.8 Flash Sept 2, 2026 |
| Meta Muse Glimmer (hosted, Together AI) | $0.35 | $1.50 | Local self-host = electricity after hardware amortized |
| GPT-5.6 Terra | $2.00 | $12.00 | −20% Jul 30, 2026 (from $2.50/$15); 5.6x usage in the discount window |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 500K context; AA Intelligence Index 61; cache hit $0.50 |
| Qwen 3.8 Max (open weights) | $2.00 | $6.00 | Hosted API list price; open weights live Aug 12, 2026 |
| Claude Sonnet 5 | $2.00 | $10.00 | Permanent as of Aug 10, 2026; Sept 1 increase cancelled |
| Kimi K3 (Moonshot) | $3.00 | $15.00 | Open weights; $20M/12mo commercial-use threshold |
| GPT-5.6 Sol | $4.00 | $20.00 | Official OpenAI promo rate (Aug 21 – Nov 21, 2026); cached input $0.40; previously $5/$30 |
| GPT-6 Astra | $10.00 | $50.00 | NEW Sept 3, 2026 — OpenAI flagship; $1 cached input / $12.50 cache writes; prompts >272K input reprice the full request to $20/$75 per 1M; 1.05M context |
OpenAI's discount experiment (Jul 27 – Aug 14, 2026): when Terra and Luna were discounted 50% via OpenRouter, daily Terra token usage rose 5.6x and Luna 13.8x, while Sol at list price rose only 1.1x (control); ~1/3 of users who tried a discounted model kept using it after expiry. Terra/Luna went from 0.7% to 7.8% of OpenRouter tokens. For the agency-pricing breakdown, see OpenAI's discount experiment: what 13.8x usage growth means for agency pricing.
Cost per task on a typical 10K-in / 2K-out workload
Illustrative math (input tokens ÷ 1M × input price + output tokens ÷ 1M × output price), not a benchmark claim — your real token counts will differ:
| Model | Input cost | Output cost | Total per task |
|---|---|---|---|
| Qwen 3.8 Flash | $0.0015 | $0.0009 | ~$0.0024 |
| DeepSeek V4-Flash (off-peak) | $0.0022 | $0.0013 | ~$0.004 |
| GPT-5.6 Luna | $0.002 | $0.0024 | ~$0.004 |
| Gemini 3.8 Flash (intro) | $0.0075 | $0.0075 | ~$0.015 |
| Gemini 3.7 Flash (intro — previous version) | $0.0075 | $0.0075 | ~$0.015 |
| Muse Glimmer (hosted) | $0.0035 | $0.0030 | ~$0.007 |
| Grok 4.6 | $0.020 | $0.012 | ~$0.032 |
| Qwen 3.8 Max | $0.020 | $0.012 | ~$0.032 |
| Claude Sonnet 5 | $0.020 | $0.020 | ~$0.040 |
| GPT-5.6 Terra | $0.020 | $0.024 | ~$0.044 |
| GPT-5.6 Sol | $0.040 | $0.040 | ~$0.080 |
| GPT-6 Astra | $0.100 | $0.100 | ~$0.200 — 2.5x Sol on this workload; over 272K input the full request reprices to $20/$75 per 1M |
At 10K tasks/month, that spread is $24/mo (Qwen 3.8 Flash) to $2,000/mo (GPT-6 Astra) on identical volume — GPT-6 Astra (Sept 3, 2026) at $10/$50 per 1M bills ~$0.20/task vs GPT-5.6 Sol's ~$0.08 (Sol's Aug 21 promo cut $5/$30 → $4/$20 dropped Sol from $1,100 to $800/mo on this volume). This is why model routing is a margin lever, not a footnote. The demand side of those prices is firming up too: Anthropic's reported $65B run rate signals a vendor with real pricing power behind Claude's per-token rates.
What the 2026 data says about routing
- Small structured tasks belong on cheap models. Published reference (AA): Muse Spark 1.2 ≈ $0.40/task vs Claude Opus 5 ≈ $2.34 — a 6x spread for the same class of work. Our 10K/2K math shows Qwen 3.8 Flash (~$0.0024/task), DeepSeek off-peak, Gemini 3.8 Flash intro (same rate as 3.7 Flash), and hosted Muse Glimmer under ~2 cents per task.
- Frontier costs cluster at the top. GPT-5.6 Sol at $4/$20 (official promo through Nov 21, 2026) is the highest per-task on the board; Grok 4.6 at $2/$6 matches Qwen 3.8 Max and undercuts GPT-5.6 Sol by ~2.5x on the illustrative task while matching its AA Intelligence Index of 61.
- DeepSeek's hike changes the "cheapest stack" answer. Post-Aug 16, DeepSeek V4-Flash off-peak ($0.22/$0.66) is still cheap on cache-miss input but output is 2.4x the old rate; peak output ($1.32) is now above Gemini 3.8 Flash intro (the 3.7 Flash rate it inherited) and hosted Muse Glimmer.
- Qwen 3.8 Flash resets the cheap-hosted tier. Released Aug 26, 2026 at $0.15/$0.47 per 1M on QwenCloud/OpenRouter ($0.016 cache reads), the open-weight 125B MoE activating ~6B per token (~95% sparsity) lands at ~$0.0024 on the 10K/2K task — below DeepSeek V4-Flash off-peak and GPT-5.6 Luna (~$0.004 each) and roughly a quarter of DeepSeek V4 Pro off-peak (~$0.0106).
- Sustained volume flips the answer to self-host. Open weights (Muse Glimmer 30B local, DeepSeek V4 MIT weights, Qwen 3.8 Max, Qwen 3.8 Flash-Next FP8 at 172.78 GiB, Kimi K3) run at electricity-cost marginal inference once hardware is amortized.
Routing decides which model reasons; it does not decide how the work is metered. Point the same routing table at a voice interface and the meter splits in two — duration billed by the second, reasoning billed by the token — so one call runs on two cost curves at once and a per-task figure stops being the whole answer. The two-meter voice agent cost model (voice seconds plus backend reasoning tokens) prices a cheap-reasoner preset and an expensive-reasoner preset against an identical $0.05-per-minute voice meter, so the pair is what you quote rather than the model alone.
How to compute your agency's real cost per task
- Instrument a pilot: log input/output tokens per task type for 100–500 real tasks.
- Multiply by your model's per-1M rate (use the table above; apply cache hits and off-peak where eligible).
- Sum per task type, then add retry/failure multipliers — real agent workflows rarely run clean the first time (levelsio reported ~$0.06 per request and "a few dollars per task" baselines, with failure modes turning into $500 loops and $2,000 overnight bills).
- Re-run quarterly — 2026 pricing moves weekly (Gemini 3.8 Flash Sept 2, GPT-5.6 Sol Aug 21, DeepSeek Aug 16, Gemini 3.7 Flash Aug 13, Grok 4.6 Aug 12, Claude Sonnet 5 Aug 10).
Per-task math has one blind spot worth naming: it prices work, not time. A voice deployment bills two meters at once — realtime audio transport per connected minute, and the backend reasoning per turn — so a per-task figure understates the run. The voice agent two-meter cost model takes your calls per day and call length, bills the WebRTC session by the second with the 15-second init floor applied as a transport toggle, and returns cost per call, per day and per month with the month convention printed.
Model the blended cost of your exact task mix
Open the AI Agency Pricing Calculator →Setup fees, retainers, model strategy (DeepSeek, Gemini, Grok 4.6, Qwen, Sonnet 5, local), and margin — current 2026 rates.
Frequently asked questions
How much does an AI task cost in 2026?
It depends on the model and the task size. Published reference points (Artificial Analysis, Aug 2026): Muse Spark 1.2 ≈ $0.40 per task, Claude Opus 5 ≈ $2.34 per task, DeepSeek V4 Flash ≈ $0.03 per task at pre-hike rates. A 10K-token-in / 2K-token-out task on current 2026 rates ranges from ~$0.0024 (Qwen 3.8 Flash at its official $0.15/$0.47 rate) to ~$0.08 (GPT-5.6 Sol at its official $4/$20 promo rate), with frontier models like Claude Opus 5 and GPT-5.6 Sol at the high end for heavy tasks.
Which AI model is cheapest per task in 2026?
Among hosted APIs, Qwen 3.8 Flash ($0.15/$0.47 per 1M, ~$0.0024 per 10K/2K task) is now the cheapest per task, with DeepSeek V4-Flash off-peak ($0.22/$0.66 after the Aug 16, 2026 increase) and Gemini 3.8 Flash intro pricing ($0.75/$3.75 through 2026-12-31, same rate as the 3.7 Flash it superseded) close behind for small work; Meta Muse Spark 1.2 at ~$0.40/task remains the published small-task reference. For sustained volume, self-hosted open weights (Meta Muse Glimmer local, DeepSeek V4 MIT weights) undercut every hosted API once hardware is amortized.
What is the cost per task for Grok 4.6 vs GPT-5.6 Sol?
On a 10K-token-in / 2K-token-out task, Grok 4.6 ($2/$6 per 1M) costs about $0.032 vs GPT-5.6 Sol (official $4/$20 per 1M promo through Nov 21, 2026) at about $0.08 — roughly 2.5x cheaper per task on that workload. Grok 4.6 also has a 500K context window and an AA Intelligence Index of 61, matching GPT-5.6 Sol max.
How much does GPT-6 Astra cost per task?
On a 10K-in / 2K-out task, GPT-6 Astra (launched Sept 3, 2026 at OpenAI's official $10/$50 per 1M, $1 cached input) costs about $0.20 — 2.5x GPT-5.6 Sol's ~$0.08 at Sol's promotional $4/$20. The cost cliff is long context: any prompt over 272K input tokens reprices the FULL request at $20 input / $75 output per 1M, so a 300K-in / 5K-out run bills roughly $6.375 vs about $3.25 at sub-272K rates (a ~96% jump). See GPT-6 Astra API Pricing for the full rate card.
How do I calculate my agency's cost per task?
Cost per task = (input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price). Measure real token counts per task type in a pilot, then multiply by the model's per-1M rate. Use cache hits and off-peak scheduling to cut input cost, and route small structured tasks to a cheap model while keeping frontier models for complex work.
Sources
- OpenAI developer docs — GPT-6 Astra model + pricing (verified Sept 3, 2026): developers.openai.com/api/docs/models/gpt-6-astra · developers.openai.com/api/docs/pricing
- Artificial Analysis, model pages + cost-per-task estimates (Aug 2026): artificialanalysis.ai/models
- DeepSeek official pricing page (live verified Aug 16, 2026 — new peak/off-peak rates): api-docs.deepseek.com/quick_start/pricing
- Google, "Introducing Gemini 3.8 Flash and 3.8 Flash Cyber" (Sept 2, 2026): blog.google · x.com/Google · Google, "Introducing Gemini 3.7 Flash" (Aug 13, 2026, previous version): blog.google
- SpaceXAI, "Grok 4.6" (Aug 12, 2026): x.ai/news/grok-4-6
- Anthropic, "Claude Sonnet 5" + pricing docs (Aug 10, 2026): anthropic.com
- Together AI, Muse Glimmer hosted pricing (verified Aug 12, 2026): together.ai/pricing
- OpenRouter Blog, "GPT 5.6 Discounts & Jevons Paradox" (Aug 25, 2026): openrouter.ai
- OpenAI API pricing (platform docs, verified Aug 28, 2026 — Terra $2/$12, Luna $0.20/$1.20, Sol $4/$20 promo): platform.openai.com/docs/pricing
- TLDR AI newsletter (Aug 28, 2026): tldr.tech
- QwenCloud, "Qwen3.8-Flash" official model & pricing page (verified Aug 27, 2026 — $0.15/$0.47/$0.016 per 1M): qwencloud.com/models/qwen3.8-flash
- OpenRouter, "Qwen3.8 Flash" API page + public models API (live listing, fetched Aug 28, 2026): openrouter.ai/qwen/qwen3.8-flash
- Hugging Face, "Qwen3.8-Flash-Next" model card (125B total / 6B active, multimodal MoE, qwen-community-1.0; released Aug 26, 2026): huggingface.co/Qwen/Qwen3.8-Flash-Next
- OrcaRouter, "Qwen 3.8 Flash release" (Alibaba’s Aug 28 announcement — ~95% sparsity, official QwenCloud rates): orcarouter.ai
- Hacker News / levelsio cost-per-task baseline (Aug 5, 2026): news.ycombinator.com/item?id=45931825
Accuracy note: GPT-6 Astra per-token pricing ($10/$50 per 1M, $1 cached input; >272K input reprices the full request to $20/$75) is OpenAI's official rate, verified Sept 3, 2026 from OpenAI's model + pricing docs; the ~$0.20 per 10K/2K task and ~$6.375 per 300K-in/5K-out run are our illustrative arithmetic on those official rates, not benchmark claims. The three per-task reference points (Muse Spark $0.40, Claude Opus 5 $2.34, DeepSeek V4 Flash $0.03) are Artificial Analysis figures from Aug 2026 and are attributed as such; the DeepSeek row predates the Aug 16, 2026 increase. The 10K/2K per-task column is our illustrative arithmetic on published per-1M rates — labeled as such, not a benchmark claim. GPT-5.6 Sol per-token pricing is OpenAI's official promotional rate ($4/$20 per 1M, cached input $0.40, valid Aug 21 – Nov 21, 2026; previously $5/$30). GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20) are OpenAI's official Jul 30, 2026 list prices; the usage multiples (Luna 13.8x, Terra 5.6x, Sol 1.1x control) and ~1/3 retention are OpenRouter's measured discount-window data (Jul 27 – Aug 14, 2026), attributed as such (research brief t_5c8841ca). Per-1M rates are as of Aug 28, 2026 and move frequently — re-verify before quoting. Qwen 3.8 Flash pricing is the official QwenCloud rate confirmed Aug 27, 2026, cross-checked on OpenRouter (fetched Aug 28); 125B total / 6B active and ~95% sparsity per the Hugging Face model card and Alibaba’s Aug 28 announcement (6/125 = 95.2% derived); Qwen benchmarks vendor-reported and unreproduced as of Aug 28. No hands-on model testing was performed for this page; all figures trace to named sources.