← Blog
PricingClaude APIGPT APIComparison

Claude Sonnet vs Opus vs Haiku vs GPT: API Cost Compared

(prices in the text follow the live price list)

Short answer

For most API work the sensible default is Claude Sonnet 5.5: as of 2026-10-10 it costs $0.77 / $3.82 per million input / output tokens on APIVAI, against $2.00 / $10.00 at Anthropic's list price. Step up to Claude Opus 5.5 ($1.54 / $7.65) for long-running agentic coding and hard reasoning, and down to Claude Haiku 4.5 ($0.38 / $1.92) for high-volume classification, extraction and routing; Claude Fable 5.1 is the most capable and most expensive Claude tier and pays off only when Opus falls short. On the OpenAI side, GPT-6.1 Sol ($0.30 / $1.52) is the mid-priced option and GPT-6 Luna ($0.0144 / $0.0768) the cheapest model on the list. One APIVAI key works for all of them, so you can pick a model per task instead of committing to one.

What is the difference between Haiku, Sonnet, Opus and Fable?

Anthropic sells Claude in four tiers. The descriptions below follow Anthropic's models overview and its guide to choosing a model:

  • Haiku: built for high-volume, latency-sensitive tasks such as classification, extraction and routing. It is the fastest and cheapest tier and also works well as a sub-agent that handles bulk steps for a larger model.
  • Sonnet: described by Anthropic as the best combination of speed and intelligence. It covers everyday code generation, data analysis, content creation and agentic tool use.
  • Opus: built for long-running agentic coding and knowledge work, such as multi-hour coding agents and large refactors. It is slower than Sonnet. When you are unsure, Anthropic's own docs suggest starting with Claude Opus 5.5.
  • Fable: for demanding reasoning and long-horizon agentic work, such as agent sessions that run for hours and multistep research. It is the slowest and most expensive tier.

Claude Sonnet 5.5, Claude Opus 5.5 and Claude Fable 5.1 have a 1M-token context window and up to 128K output tokens; Claude Haiku 4.5 has a 200K context window. On APIVAI thinking is on by default and thinking tokens are billed as output tokens, so a short visible answer can still use a fair amount of output. If replies come back empty or cut off, raise max_tokens to 4096 or more.

How do GPT models compare?

OpenAI's current models page lists three tiers: Astra ("our most capable model for the most demanding work"), Sol (GPT-6.1 Sol, "near-Astra performance for complex work at a lower cost") and Luna (GPT-6 Luna, "our most efficient model for focused, high-volume tasks"). By official list price, Astra sits next to Fable, Sol next to Sonnet and Luna below Haiku.

APIVAI also offers earlier GPT releases: GPT-6 Sol, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna and GPT-5.5. OpenAI's current models page does not describe them, so this guide compares them only by price and context. One practical difference: Claude models work in both the Anthropic and the OpenAI API format on APIVAI, while GPT models use the OpenAI format (/v1/chat/completions or /v1/responses).

How much does each model cost per million tokens?

Prices per million tokens as of 2026-10-10. The pricing page shows every model, and the model list has a page for each one.

ModelAPIVAI input / outputOfficial input / outputAPIVAI cache readContext
Claude Fable 5.1$4.42 / $22.08$10.00 / $50.00$0.111M
Claude Opus 5.5$1.54 / $7.65$4.00 / $20.00$0.07681M
Claude Sonnet 5.5$0.77 / $3.82$2.00 / $10.00$0.03841M
Claude Sonnet 4.6$1.15 / $5.74$3.00 / $15.00$0.121M
Claude Haiku 4.5$0.38 / $1.92$1.00 / $5.00$0.0384200K
GPT-6 Astra$1.52 / $7.60$10.00 / $50.00$0.151M
GPT-6.1 Sol$0.30 / $1.52$2.00 / $10.00$0.01441M
GPT-6 Sol$0.30 / $1.52$2.00 / $10.00$0.03041M
GPT-5.5$0.77 / $4.56$5.00 / $30.00$0.07681M
GPT-6 Luna$0.0144 / $0.0768$0.10 / $0.50$0.00161M

For every model, output costs several times as much as input, so long answers and long thinking drive the bill more than long prompts. Cache reads cost a small fraction of normal input, which is why caching matters so much for agents (example 1 below). The Claude API pricing comparison lists every Claude version, including cache-write prices.

Why do several Opus models cost the same?

Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7 and Claude Opus 4.6 share one price: $1.92 / $9.57 on APIVAI ($5.00 / $25.00 official). The newest Opus, Claude Opus 5.5, is actually cheaper at $1.54 / $7.65. The same pattern holds for Sonnet: Claude Sonnet 5.5 costs less than Claude Sonnet 4.6, and Claude Sonnet 5 has the same input and output price as Sonnet 5.5 but a higher cache-read price. Claude Fable 5 and Claude Fable 5.1 share input and output prices, and 5.1 has the cheaper cache reads.

At an equal or lower price, the newer model is usually the better choice. Keep an older version only when your prompts were tuned for it and your own tests show it does better, or when a tool pins a specific ID. Anthropic lists a few breaking changes for Opus 5.5 (for example forced tool use and turning thinking off), so test before switching a production integration. Each older model has its own page: Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, Fable 5.

What does a coding agent session cost with caching?

Assumptions: a two-hour session in a coding agent such as Claude Code or Cline makes 40 requests. Each request resends about 25,000 tokens of context (system prompt, tool definitions, open files, earlier turns) that is read from the cache, so the session reads 1,000,000 cached tokens. On top of that it sends 60,000 new input tokens (new files, tool results) and receives 30,000 output tokens, thinking included.

ModelCache reads (1M tokens)New input + outputSame session without cache (APIVAI)Without cache (official)
Claude Sonnet 5.5$0.0384$0.161$0.931$2.42
Claude Sonnet 4.6$0.12$0.241$1.39$3.63
Claude Opus 5.5$0.0768$0.322$1.86$4.84
Claude Fable 5.1$0.11$0.928$5.35$12.10
GPT-6.1 Sol$0.0144$0.0636$0.364$2.42

With caching, the session costs the first two columns added together. The third and fourth columns show what the same tokens would cost if nothing were cached. Cache writes on Claude cost a little more than normal input ($0.96 against $0.77 per million on Sonnet 5.5), so the real total is slightly higher than the sum. Caching only works when the tool resends an identical prefix; most coding agents do. To turn a session price into a monthly budget, and to compare it with subscription editors, see Claude Code vs Cursor: cost comparison.

What does a chatbot with 1,000 conversations a day cost?

Assumptions: a support bot handles 1,000 conversations a day, 30 days a month. A conversation averages 2,500 input tokens (instructions, a few knowledge snippets, the history) and 350 output tokens including thinking.

ModelPer day (APIVAI)Per month (APIVAI)Per month (official)
GPT-6 Luna$0.0629$1.89$12.75
Claude Haiku 4.5$1.62$48.66$128
GPT-6.1 Sol$1.28$38.46$255
Claude Sonnet 5.5$3.26$97.86$255

For routine questions a Haiku- or Luna-class model is usually enough, and many teams route only the hard tickets to Sonnet. Caching the fixed instructions lowers these numbers further.

What does summarizing 2,000 documents cost?

Assumptions: a one-off job turns 2,000 reports of about 15,000 tokens each into 800-token summaries. Input dominates here, so the input price matters most.

ModelTotal (APIVAI)Total (official)
GPT-6 Luna$0.555$3.80
Claude Haiku 4.5$14.47$38.00
GPT-6.1 Sol$11.43$76.00
Claude Sonnet 5.5$29.21$76.00
Claude Opus 5.5$58.44$152

Run a sample of 20 documents on two or three models and compare the summaries before you run the whole batch; the cheapest model that passes your check is the right one. APIVAI runs these as normal requests and does not offer a Batch API. Anthropic's own Batch API is billed below its standard rates for jobs that can wait, so for very large offline jobs compare that price too.

Which model should I use for each task?

TaskStart withStep up or down
Everyday coding in an agent or IDEClaude Sonnet 5.5Claude Opus 5.5 for large refactors and long autonomous runs
Hardest reasoning, hours-long agent runsClaude Opus 5.5Claude Fable 5.1 or GPT-6 Astra when Opus falls short
Writing and contentClaude Sonnet 5.5 or GPT-6.1 SolClaude Opus 5.5 for long, structured pieces

A good habit is to test your own prompts on two neighbouring tiers. If the cheaper one passes, keep it; if it needs many retries, the stronger model is often cheaper overall. Pick model IDs from GET /v1/models or the pricing page rather than copying old names, because the list changes over time.

FAQ

Is Opus better than Sonnet for coding?

For long autonomous runs, large refactors and hard debugging, Opus is the stronger tier. For everyday edits and short agent tasks Sonnet is usually enough, at $0.77 / $3.82 against $1.54 / $7.65 for Opus, so many developers default to Sonnet and switch to Opus per task.

When is Haiku enough instead of Sonnet?

Haiku is built for high-volume, simple tasks: classification, extraction, routing, short support answers and sub-agent steps. If your outputs need multi-step reasoning or careful code, test Sonnet.

Is Claude or GPT cheaper through the API?

It depends on the tier. As of 2026-10-10 on APIVAI, GPT-6 Luna costs $0.0144 / $0.0768 against $0.38 / $1.92 for Claude Haiku 4.5, and GPT-6.1 Sol costs $0.30 / $1.52 against $0.77 / $3.82 for Claude Sonnet 5.5. Compare quality on your own prompts, not only the price.

Do thinking tokens cost extra?

They are billed as output tokens at the model's normal output price. Thinking is on by default on APIVAI, so budget for some output beyond the visible answer.

Can I use several models with one key?

Yes. One APIVAI key works for every Claude and GPT model; you change only the model ID in each request.

Create an account, top up from $10 and try each model with one key.

Ready to start?

Get your API key in 30 seconds. Pay as you go for Claude and GPT.

Get Started