Short answer
To use Dify with APIVAI, install the OpenAI-API-compatible model provider plugin, click Add Model on its card and enter Model Type LLM, an APIVAI model ID such as claude-sonnet-4-6 or gpt-5.5 as the Model Name, your APIVAI key as the API Key and https://api.apivai.com/v1 as the API Base URL. Add one entry per model, then choose the models in your apps or under Default Models. For Claude only, you can also use Dify's Anthropic plugin with API URL https://api.apivai.com and no /v1, because the plugin's Anthropic client adds /v1/messages itself. As of 2026-10-07, Claude Sonnet 4.6 costs $1.04 / $5.22 per million input / output tokens through APIVAI, against the official $3.00 / $15.00.
What is Dify, and how does it connect to models?
Dify is an open-source platform for building LLM apps: chatbots, agents, workflows and knowledge-base (RAG) apps. You can use Dify Cloud or self-host it with Docker.
Dify reaches models through model provider plugins installed from the Dify Marketplace. Each provider stores a key and an address for the whole workspace, and only workspace owners and admins can change them. The providers live on the Model Provider page: in current Dify Cloud it is under Integrations, in older self-hosted versions under Settings in the account menu. The field names below were checked in October 2026 against the Dify docs and the source of the official openai_api_compatible plugin (version 0.0.68) and anthropic plugin.
Which provider should you choose: OpenAI-API-compatible or Anthropic?
| Provider plugin | Address field | Models | Pick it when |
|---|---|---|---|
| OpenAI-API-compatible | API Base URL: https://api.apivai.com/v1 | Claude and GPT | You want one provider for every model and full control of model IDs |
| Anthropic | API URL: https://api.apivai.com | Claude only | You only use Claude and want the native format with its prompt caching |
Start with OpenAI-API-compatible. You type each model ID yourself, so new models work as soon as they appear in the APIVAI list. The general details of both formats are on the OpenAI-compatible API and Claude API proxy pages.
How do you add APIVAI to Dify step by step?
Step 1: Get an API key
- Sign up with your email
- Top up your balance, from $10, by card, crypto, Alipay or WeChat Pay
- Create a key on the Dashboard and copy it. Give it a budget limit if several people build apps in the same workspace.
Step 2: Install the provider plugin
- Open the Model Provider page.
- In the list of providers you can install (or on the Dify Marketplace), find OpenAI-API-compatible and click Install.
Step 3: Add a model
On the OpenAI-API-compatible card, click Add Model and fill in the form:
| Field | Value |
|---|---|
| Model Type | LLM |
| Model Name | claude-sonnet-4-6 (the exact APIVAI ID) |
| Model display name | optional, for example "Claude Sonnet 4.6 (APIVAI)" |
| API Key | your APIVAI key |
| API Base URL | https://api.apivai.com/v1 |
| model name for API endpoint | leave empty |
| Completion mode | Chat |
| Model context size | no larger than the context window on the model's page |
| Upper bound for max tokens | 8192 or more |
| Function Call Type | Tool Call |
| Stream function calling | Support |
| API Type | Chat Completions API (/chat/completions) |
Leave the other fields at their defaults. Older versions of the plugin label the address field API endpoint URL; the value is the same.
A few notes on these fields:
- Model Name is what Dify sends to the API, so it must be the exact ID. The display name is only for the interface.
- Model context size is the number Dify uses when it fits conversation history and retrieved knowledge into the prompt. A value below the model's real window is fine and keeps long chats cheaper.
- Upper bound for max tokens only sets how high the Max Tokens parameter can go in your apps. Thinking is on by default and its tokens count as output, so leave room above 4096.
- Function Call Type: Tool Call lets Agent apps use native tool calling, which APIVAI supports.
Click Save. Dify checks the credentials with a test request before it saves the model. Repeat Add Model for every model you want, for example gpt-5.5 and claude-haiku-4-5, with the same key and address.
Step 4: Use the models in your apps
- In an app, open the model selector, choose your model under OpenAI-API-compatible, and set Max Tokens to 4096 or more in its parameters.
- In a workflow, each LLM node has its own model, so cheap and strong models can sit in the same flow.
- Under Default Models on the Model Provider page, set the System Reasoning Model, the workspace's default model for general LLM tasks.
Alternative: the Anthropic plugin for Claude
- Install the Anthropic provider plugin and open its setup.
- Set API Key to your APIVAI key and API URL to:
https://api.apivai.com
Do not add /v1: the plugin passes this address to the Anthropic client, which appends /v1/messages. Leave Enable request metadata on Disabled.
- The plugin comes with a list of Claude models. Some IDs in it match APIVAI's, such as
claude-sonnet-4-6; others use different or dated names. If a model is missing or returns "model not found", use Add Model on the Anthropic card and enter the exact ID from the APIVAI list, such asclaude-opus-5-5.
Which model should you choose for Dify apps?
Prices per 1M input / output tokens as of 2026-10-07:
| Model | APIVAI | Official | Good for |
|---|---|---|---|
| Claude Haiku 4.5 | $0.35 / $1.74 | $1.00 / $5.00 | Question Classifier nodes, extraction, high-volume bots |
| Claude Sonnet 4.6 | $1.04 / $5.22 | $3.00 / $15.00 | Default for chatbots and agents |
| Claude Sonnet 5.5 | $0.69 / $3.47 | $2.00 / $10.00 | Agents with many tools, long documents |
| Claude Opus 5.5 | $1.39 / $6.96 | $4.00 / $20.00 | The hardest reasoning steps |
| GPT-5.5 | $0.77 / $4.56 | $5.00 / $30.00 | GPT alternative for chat and agents |
| GPT-6 Sol | $0.30 / $1.52 | $2.00 / $10.00 | Newer GPT for complex workflows |
| GPT-6 Luna | $0.0144 / $0.0768 | $0.10 / $0.50 | Cheapest option for simple nodes |
The model list changes over time. Take IDs from the pricing page or from the API:
curl https://api.apivai.com/v1/models -H "Authorization: Bearer YOUR_KEY"
What does a Dify app cost with APIVAI?
Two examples, with thinking tokens counted as output:
- Support chatbot: 300 messages a day on Claude Sonnet 4.6, each about 3,000 input tokens (system prompt, recent history and retrieved knowledge) and 600 output tokens. A month of 9,000 messages costs $56.27, against $162 at official prices. The same month on GPT-6 Luna would be $0.804.
- Routing workflow: a Question Classifier node on Claude Haiku 4.5 sorts 1,000 requests a day, about 800 input and 300 output tokens each. A month of 30,000 calls costs $24.06, against $69.00 at official prices.
Knowledge retrieval makes each request bigger, so the number of chunks you retrieve matters as much as the model. Your Dify logs and the APIVAI Dashboard both show tokens per request. For a wider price comparison, see Claude API pricing in 2026.
Can you use APIVAI for Dify Knowledge (RAG)?
Only for the answering part. A Dify knowledge base needs an Embedding Model to index and search documents, and APIVAI does not provide embeddings. Install a separate embedding provider and set it under Default Models → Embedding Model (a Rerank Model is optional and also comes from another provider). The LLM that writes the answer from the retrieved text can still be an APIVAI model.
Troubleshooting
- Credentials validation fails or 401: check that the key was copied in full, still exists and has balance or budget left.
- 404: the
/v1is wrong. API Base URL in OpenAI-API-compatible must behttps://api.apivai.com/v1, with/v1once. API URL in the Anthropic plugin must behttps://api.apivai.com; with/v1the requests go to/v1/v1/messages. - Model not found: Model Name must be the exact ID, such as
claude-sonnet-4-6, not a display name. Check it against/v1/models. - Empty or cut-off answers: raise Max Tokens in the app to 4096 or more, and raise Upper bound for max tokens in the model entry if the slider stops too low.
- The agent never calls tools: edit the model entry and set Function Call Type to Tool Call.
- Self-hosted Dify cannot connect: the Dify containers that run plugins need outbound internet access to
api.apivai.com. - 429 Too Many Requests: the default limit is 60 requests per minute per key. Batch workflows can hit it; slow them down or ask support for a higher limit.
FAQ
Do I need a separate provider for each model?
No. One OpenAI-API-compatible provider holds all your models; you add one entry per model ID with the same key and API Base URL.
Does this work on Dify Cloud and self-hosted Dify?
Yes. Both use the same provider plugins and fields. A self-hosted instance needs outbound access to api.apivai.com.
Can one workflow use both Claude and GPT?
Yes. Each LLM node chooses its own model, and one APIVAI key covers every model.
Can APIVAI provide the embeddings for my knowledge base?
No. APIVAI does not offer embeddings, so Dify Knowledge needs a separate embedding provider. Chat, agents and workflow LLM nodes work through APIVAI.
Does APIVAI store what my users type?
Request and response content is not logged. Only usage data such as model, tokens and cost is kept for billing. Keys and limits are explained in the docs. If you also automate with n8n, see n8n with Claude and GPT.
Create an account, top up from $10 and paste your key into Dify.