Short answer
The cheapest way to run Claude Code, Cursor or your own scripts on Claude and GPT models is to pay per token at a lower rate and pick the smallest model that does the job. APIVAI is a pay-as-you-go API gateway: you point your tools at https://api.apivai.com and, as of 2026-10-06, pay 58–66% less than Anthropic's list prices for Claude models and up to 86% less than OpenAI's for GPT models. A typical Claude Code message on Claude Sonnet 4.6 costs about $0.012 instead of $0.0345. There is no subscription; top-ups start at $10 by card or crypto, or ¥20 via Alipay / WeChat Pay.
What drives the cost of an AI API?
Every request is billed by tokens, in three kinds:
- Input tokens: everything you send, including the system prompt, the conversation so far, open files and tool definitions. Coding agents send a lot of input, because each turn repeats the context.
- Output tokens: what the model writes back. Output is priced several times higher than input per token, so long answers and large code rewrites add up quickly.
- Cache tokens: when the same prefix (a long system prompt, a file, earlier turns) is sent again, it can be written to a prompt cache once and read back on later turns at a much lower price than fresh input.
The other lever is the model. Within one provider, the most capable model can cost several times more per token than the everyday one. On APIVAI, Fable is the most capable and most expensive tier, Opus is for complex coding and analysis, Sonnet handles everyday coding at the best value, and Haiku is the fast, cheap option. On the OpenAI side, GPT-5.5 and the GPT-6 family are the main models, and the Luna models are the smallest and cheapest.
Official vs APIVAI prices for the main models
| Model | Official input / output (per 1M tokens) | APIVAI input / output | Below list |
|---|---|---|---|
| Claude Fable 5.1 | $10.00 / $50.00 | $4.21 / $21.02 | 58% |
| Claude Opus 5.5 | $4.00 / $20.00 | $1.39 / $6.96 | 65% |
| Claude Sonnet 5.5 | $2.00 / $10.00 | $0.69 / $3.47 | 66% |
| Claude Sonnet 4.6 | $3.00 / $15.00 | $1.04 / $5.22 | 65% |
| Claude Haiku 4.5 | $1.00 / $5.00 | $0.35 / $1.74 | 65% |
| GPT-6 Sol | $2.00 / $10.00 | $0.30 / $1.52 | 85% |
| GPT-5.5 | $5.00 / $30.00 | $0.77 / $4.56 | 85% |
| GPT-6 Luna | $0.10 / $0.50 | $0.0144 / $0.0768 | 86% |
Prices as of 2026-10-06. Official prices are the providers' published list prices (Anthropic, OpenAI). The pricing page always shows today's numbers for every model, including cache prices. For a deeper look at Claude pricing across providers, see the Claude API pricing comparison.
How much does Claude Code cost per message and per month?
A Claude Code chat turn is roughly 4,000 input tokens and 1,500 output tokens. Agentic turns that read many files are larger, so treat these as a floor:
| Usage | Model | Official API | APIVAI |
|---|---|---|---|
| One message | Claude Sonnet 4.6 | $0.0345 | $0.012 |
| One message | Claude Opus 5.5 | $0.046 | $0.016 |
| One message | Claude Haiku 4.5 | $0.0115 | $0.00401 |
| 20 messages a day for a month (600) | Claude Sonnet 4.6 | $20.70 | $7.19 |
| 20 messages a day for a month (600) | Claude Opus 5.5 | $27.60 | $9.60 |
| 10M input + 2M output tokens a month | Claude Sonnet 4.6 | $60.00 | $20.84 |
| 10M input + 2M output tokens a month | GPT-5.5 | $110 | $16.82 |
Prices as of 2026-10-06; see /pricing for current rates. These numbers leave out prompt caching, which lowers both columns when the same context is sent again. If you are deciding between tools, the Claude Code vs Cursor cost comparison works through more scenarios.
How to set up Claude Code
Claude Code speaks the Anthropic Messages format, so its base URL has no /v1. Set two environment variables and start it.
Mac/Linux:
ANTHROPIC_AUTH_TOKEN="your-apivai-key" \ ANTHROPIC_BASE_URL="https://api.apivai.com" \ claude
Windows PowerShell:
$env:ANTHROPIC_AUTH_TOKEN="your-apivai-key" $env:ANTHROPIC_BASE_URL="https://api.apivai.com" claude
If ANTHROPIC_API_KEY is also set in your environment, it takes priority and requests will fail with 401, so unset it first. Inside Claude Code, use /model to switch between Sonnet, Opus and Haiku. The full walkthrough is in the Claude Code setup guide.
How to set up Cursor
- Open Cursor Settings > Models.
- Turn on Override OpenAI Base URL and enter
https://api.apivai.com/v1. - Paste your APIVAI key as the OpenAI API key.
- Add the model IDs you want as custom models, for example
claude-sonnet-4-6orgpt-5.5.
More detail is in the Cursor setup guide.
How to set up Codex CLI
Codex CLI (0.121 and later) reads providers from ~/.codex/config.toml. Add:
model = "gpt-5.5" model_provider = "apivai" [model_providers.apivai] name = "APIVAI" base_url = "https://api.apivai.com/v1" env_key = "OPENAI_API_KEY" wire_api = "responses"
Then set OPENAI_API_KEY to your APIVAI key and run codex.
How to call it from Python (OpenAI SDK)
Any OpenAI-compatible client works by changing the base URL. Claude and GPT model IDs both work here:
from openai import OpenAI
client = OpenAI(
api_key="your-apivai-key",
base_url="https://api.apivai.com/v1"
)
response = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)How to call it from Python (Anthropic SDK)
The Anthropic SDK uses the base URL without /v1:
import anthropic
client = anthropic.Anthropic(
api_key="your-apivai-key",
base_url="https://api.apivai.com"
)
message = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello!"}]
)
print(message.content[0].text)Endpoint and parameter details are in the docs.
How to keep spend under control
- Give each key its own budget. Create separate keys on the Dashboard for each tool, project or teammate and set a budget on each. When a key reaches its budget, its requests stop with a 401 instead of draining your balance.
- Pick the model per task. Use Claude Sonnet 5.5 ($0.69 / $3.47 per 1M) for most coding, switch to Claude Opus 5.5 only for hard debugging or architecture work, and use Claude Haiku 4.5 ($0.35 / $1.74) for quick edits, summaries and simple scripts.
- Use prompt caching. On Claude Sonnet 5.5, cached reads cost $0.0688 per 1M tokens and cache writes $0.86, compared with $0.69 for fresh input. Claude Code uses caching on its own; in your own Anthropic SDK code, mark long, stable prefixes with
cache_control. - Keep context short. Start new sessions for unrelated tasks, and avoid pasting whole files when a function will do. Every repeated token is billed as input or cache.
- Check the Dashboard. It lists every request with its model, tokens and cost, so you can see which tool or habit is using the budget.
Limits worth knowing
- Each key allows 60 requests per minute by default; ask support if you need more.
- Supported endpoints are Messages, token counting, Chat Completions, Responses and model listing, with streaming, tool calling and images. The Batch API, Files API, embeddings, image generation, audio and fine-tuning are not offered; use the official API for those.
- You pay for input, output and cache tokens from a prepaid balance. There is no monthly fee or minimum.
- Top up from $10 by card or crypto, or ¥20 by Alipay / WeChat Pay. Consumed credits are non-refundable.
- APIVAI does not store the content of requests or responses; only usage metadata is kept for billing.
FAQ
Is APIVAI cheaper than a Claude Pro or Max subscription?
It depends on how much you use. Subscriptions charge a fixed monthly fee with usage limits (Claude pricing); pay-as-you-go suits uneven or moderate usage, and you never hit a plan cap. Compare your monthly token volume with the examples above.
Which base URL do I use?
Use https://api.apivai.com for Anthropic-format tools (Claude Code, Anthropic SDK) and https://api.apivai.com/v1 for OpenAI-format tools (Cursor, Codex CLI, OpenAI SDK). A 404 usually means the /v1 is missing or doubled.
Can one key use both Claude and GPT models?
Yes. The same key works in both formats; just change the model ID.
Does Claude Code work exactly as with the official API?
Yes, Claude Code talks to the same Messages endpoint, with streaming and tool calls. You only change the two environment variables.
How do I stop a runaway script from spending my balance?
Give the script its own key with a small budget. When the budget is used up, that key stops working while your other keys continue.
What if I need more than 60 requests per minute?
Contact support by email at w69787200@gmail.com or on Telegram @e3655 and ask for a higher limit.
Create an account and get your first key in a few minutes.