Short answer
To use Open WebUI with APIVAI, go to Settings > Admin > Connections, add a connection to the Manage OpenAI API Connections list, set the URL to https://api.apivai.com/v1 and paste your APIVAI key into API Key. Open WebUI reads the model list from /v1/models, so Claude and GPT models appear in the model selector with no extra setup, and one key covers both. As of 2026-10-07, Claude Sonnet 4.6 costs $1.04 / $5.22 per million input / output tokens through APIVAI, against the official $3.00 / $15.00. Chat works right away; document search (RAG) needs an embedding model from elsewhere, because APIVAI does not provide embeddings.
What is Open WebUI?
Open WebUI is a self-hosted chat interface, usually run in Docker, with user accounts, chat history, model presets and document chat. It does not ship its own cloud models. Instead it connects to any server that speaks the OpenAI Chat Completions protocol, which is exactly what APIVAI offers at https://api.apivai.com/v1.
This guide follows the Open WebUI documentation as of October 2026. If a label in your version looks slightly different, the fields are still the same: a URL, an API key and an optional list of model IDs.
Which connection type should you use?
Open WebUI talks to cloud models through OpenAI connections. It has no field for a custom Anthropic Messages endpoint; its built-in Anthropic support only recognises Anthropic's own URLs. So the right choice is the OpenAI format:
- URL:
https://api.apivai.com/v1(with/v1, exactly once) - Models: Claude and GPT both work in this format, so one connection serves every model on the pricing page
APIVAI's Anthropic-format address (https://api.apivai.com, no /v1) is for tools that use the Anthropic SDK, such as Claude Code. You do not need it in Open WebUI. The OpenAI-compatible API page explains the format in more detail.
How do you get an APIVAI key?
- Create an account with your email.
- Top up your balance from $10 by card, crypto, Alipay or WeChat Pay. There is no subscription; you pay per token.
- On the Dashboard, create a key and copy it. You can give the key its own budget limit, which is useful when many people share one Open WebUI instance.
How do you add APIVAI to Open WebUI step by step?
- Sign in to Open WebUI with an admin account.
- Open Settings > Admin > Connections and find the Manage OpenAI API Connections list.
- Click + Add Connection.
- Fill in URL and API Key with the values from the table below. Leave Connection Type on External.
- Click Verify Connection beside the URL field. Open WebUI calls
/modelswith your key and should report Server connection verified. - Optional: open Advanced. In Model IDs, type the IDs you want to offer (for example
claude-sonnet-4-6) and click + after each one. If you leave it empty, every model your key can see is listed. Prefix ID adds a label such asapivai.in front of each model name; Open WebUI removes it again before the request is sent. Leave Provider on Default. - Click Save. The models now appear in the model selector at the top of a new chat.
| Field | Value |
|---|---|
| Connection Type | External |
| URL | https://api.apivai.com/v1 |
| API Key | Your APIVAI key |
| Model IDs (Advanced) | Empty for all models, or e.g. claude-sonnet-4-6, claude-haiku-4-5, gpt-5.5 |
| Prefix ID (Advanced) | Optional, e.g. apivai |
| Provider (Advanced) | Default |
How do you set it with Docker environment variables?
If you deploy with Docker, you can pass the same values as environment variables when the container is created:
docker run -d -p 3000:8080 \ -v open-webui:/app/backend/data \ -e WEBUI_SECRET_KEY=your-secret-key \ -e OPENAI_API_BASE_URL=https://api.apivai.com/v1 \ -e OPENAI_API_KEY=your-apivai-key \ --name open-webui --restart always \ ghcr.io/open-webui/open-webui:main
If you do not run Ollama, add -e ENABLE_OLLAMA_API=False so Open WebUI stops trying to reach it. To limit the model list from the environment, use OPENAI_API_CONFIGS, a JSON object keyed by connection index with fields such as model_ids and prefix_id.
One catch: these variables are ConfigVar values. With the default ENABLE_PERSISTENT_CONFIG=True, whatever is saved in the database through the Admin panel wins over the environment. If you change OPENAI_API_BASE_URL later and nothing happens, edit the connection in Settings > Admin > Connections instead.
Which model should you choose?
Prices per million tokens as of 2026-10-07. The model list changes over time, so pick IDs from GET https://api.apivai.com/v1/models or the pricing page rather than copying old names.
| Model | Input | Output | Good for |
|---|---|---|---|
| Claude Sonnet 4.6 | $1.04 | $5.22 | Default chat model for most teams |
| Claude Sonnet 5.5 | $0.69 | $3.47 | Newer Sonnet, strong at code and long documents |
| Claude Opus 5.5 | $1.39 | $6.96 | The hardest questions |
| Claude Haiku 4.5 | $0.35 | $1.74 | Quick answers and background tasks |
| GPT-5.5 | $0.77 | $4.56 | GPT alternative |
| GPT-6 Luna | $0.0144 | $0.0768 | Cheapest option |
If you run Open WebUI for a team, use the Model IDs list to offer only the models you want to pay for. Hiding Opus from everyday users is the simplest way to keep a shared bill predictable.
Set a cheap task model
Open WebUI makes small background calls to write chat titles, generate tags, suggest follow-up questions and autocomplete prompts. By default they run on the model of the current chat, so a chat with an expensive model also pays for its titles at that price. In Settings > Admin > Interface, set External Task Model to claude-haiku-4-5 or gpt-6-luna.
Background calls get 1,000 output tokens unless you set more, and thinking tokens count against that budget. If titles come back empty or cut off, open Task Model Parameters > Configure in the same section and set max_tokens to 4096.
What does it cost?
APIVAI bills per token from a prepaid balance. The real size of each request depends on how long the conversation is, because Open WebUI sends the chat history with every message.
Example 1: one person. Assume 40 messages a day for 30 days, each about 3,000 input tokens (question plus history) and 800 output tokens, on Claude Sonnet 4.6. That is 1,200 messages for $8.76, compared with $25.20 at official prices.
Example 2: a team of ten. Ten people with 40 messages each per workday over 22 workdays make 8,800 messages: $64.20 on Claude Sonnet 4.6 ($185 official), or $21.49 if everyone uses Claude Haiku 4.5. If the task model runs on Claude Haiku 4.5 with about two background calls per message of 1,500 input and 300 output tokens, it adds roughly $18.43.
The Dashboard lists every request with its tokens and cost, so after a few days you can replace these assumptions with your own numbers. For a wider price comparison, see Claude API pricing in 2026.
Do documents and knowledge bases work?
Chat with Claude and GPT works fully through APIVAI. Document search (RAG) is different: it needs an embedding model to index files, and APIVAI does not provide embeddings.
Out of the box, Open WebUI uses a local SentenceTransformers model for embeddings (RAG_EMBEDDING_ENGINE left empty), so uploaded documents keep working without any change. Just do not switch the embedding engine to OpenAI with APIVAI's URL. If you want hosted embeddings, configure a separate embedding provider with RAG_OPENAI_API_BASE_URL and RAG_OPENAI_API_KEY, or in Settings > Admin > Documents. The answers themselves still come from your APIVAI chat model. Image generation and speech features also need their own provider.
Troubleshooting
- Verify Connection fails with 401: the key is wrong or incomplete. Copy it again from the Dashboard, check that it has not been deleted, and check that the key's budget and your balance are not used up.
- 404: the URL is wrong. It must be
https://api.apivai.com/v1: nothttps://api.apivai.comwithout/v1, not/v1/v1, not a full path like/v1/chat/completions, and no trailing slash. - Model not found: an ID in Model IDs does not exist (often an old name or a display name). Compare it with
GET /v1/models. A model shown asapivai.claude-sonnet-4-6is fine; that is your Prefix ID. - Empty or cut-off replies: thinking is on by default and thinking tokens count as output. Set
max_tokensto 4096 or more under Advanced Params, either per chat in Chat Controls or for everyone on the model in Workspace > Models. - Environment variables seem ignored: a value saved in the Admin panel overrides them (see the Docker section above).
- 429 Too Many Requests: the default limit is 60 requests per minute per key, and background tasks count too. A busy team can give users separate keys or ask support to raise the limit.
FAQ
Can Open WebUI use Claude without the Anthropic format?
Yes. APIVAI serves Claude models in the OpenAI format at https://api.apivai.com/v1, so an ordinary OpenAI connection in Open WebUI is enough. You only need the Anthropic format for tools built on the Anthropic SDK; see the Claude Code setup guide.
Can I offer Claude and GPT models with one key?
Yes. One APIVAI key works for every model on the pricing page. Users switch models in the model selector at the top of a chat.
Does APIVAI store our conversations?
APIVAI does not log the content of requests and responses. Open WebUI keeps chat history in its own database on your server.
Can each user have their own key?
Open WebUI also has personal Direct Connections that users add in their own settings, using the same URL and an APIVAI key of their own. Alternatively, keep one admin connection and give that key a budget limit on the Dashboard.
Is there a free trial?
No. APIVAI is prepaid with no subscription; the minimum top-up is $10 and you pay only for tokens you use. Setup details for other tools are in the docs and in our LobeChat guide.
Create an account, top up from $10 and paste your key into Open WebUI.