AI Providers
Codeep supports all major AI providers. You can configure multiple providers and switch between them at any time using /provider or /settings.
Supported Providers
| Provider | Env Variable | Default Model | Protocol |
|---|---|---|---|
Anthropic | ANTHROPIC_API_KEY | claude-opus-5-5 | Anthropic |
Custom (OpenAI-compatible) | none required | your model | OpenAI-compatible |
DeepSeek | DEEPSEEK_API_KEY | deepseek-flash | OpenAI-compatible |
Google AI | GOOGLE_API_KEY | gemini-3.1-pro-preview | OpenAI-compatible |
Grok (xAI) | XAI_API_KEY | grok-build-0.1 | OpenAI-compatible |
Kimi (Kimi Code subscription) | KIMI_CODE_API_KEY | kimi-for-coding (K2.8 Preview; K3 on eligible plans) | OpenAI-compatible |
Kimi API (pay-per-use) | MOONSHOT_API_KEY | kimi-k3 | OpenAI-compatible |
Kimi China | MOONSHOT_CN_API_KEY | kimi-k3 | OpenAI-compatible |
MiniMax (Coding Plan) | MINIMAX_API_KEY | MiniMax-M3 | OpenAI / Anthropic |
MiniMax API (pay-per-use) | MINIMAX_API_KEY | MiniMax-M3 | OpenAI-compatible |
MiniMax China | MINIMAX_CN_API_KEY | MiniMax-M3 | OpenAI / Anthropic |
ModelScope (live free catalog) | MODELSCOPE_API_KEY | dynamic (Qwen3.5-397B fallback) | OpenAI-compatible |
Ollama (local / remote) | none required | dynamic | OpenAI-compatible |
OpenAI | OPENAI_API_KEY | gpt-6-sol | OpenAI Responses |
OpenRouter (100+ models) | OPENROUTER_API_KEY | openrouter/auto | OpenAI-compatible |
Qwen (Coding Plan subscription) | BAILIAN_CODING_PLAN_API_KEY | qwen3.7-plus | OpenAI-compatible |
Qwen (Token Plan subscription) | BAILIAN_TOKEN_PLAN_API_KEY | qwen3.8-max | OpenAI-compatible |
Qwen API (pay-per-use) | DASHSCOPE_API_KEY | qwen3.8-max | OpenAI-compatible |
Qwen China (Coding Plan / API) | BAILIAN_CODING_PLAN_CN_API_KEY / DASHSCOPE_CN_API_KEY | qwen3.7-plus / qwen3.7-max | OpenAI-compatible |
Z.AI (GLM Coding Plan) | ZAI_API_KEY | glm-5.3 | OpenAI / Anthropic |
Z.AI API (pay-per-use) | ZAI_API_KEY | glm-5.3 | OpenAI-compatible |
Z.AI China (GLM Coding Plan) | ZAI_CN_API_KEY | glm-5.3 | OpenAI / Anthropic |
Z.AI China API (pay-per-use) | ZAI_CN_API_KEY | glm-5.3 | OpenAI-compatible |
/model. Codeep reads the catalog available to your token at that moment and falls back to Qwen3.5-397B only if the catalog request fails.Every model Codeep ships
Generated straight from the client catalogue, so it lists exactly what /model will offer — 97 models across 24 providers, as of [email protected]. Ids are the exact strings the API receives.
Z.AI (ZhipuAI) z.ai
Uses your Z.AI subscription — no per-token charges.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
glm-5.3 (default) | 1M | Included in plan | low · high · max |
glm-5.3-flash | 1M | Included in plan | low · high · max |
Z.AI API (pay-per-use) z.ai-api
Pay-per-use via Z.AI API key (zai.ai → API Keys).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
glm-5.3 (default) | 1M | $1.4 / $4.4 | low · high · max |
glm-5.3-flash | 1M | $0.15 / $0.5 | low · high · max |
glm-5.3-flashx | 1M | $0.37 / $1.25 | low · high · max |
glm-5.2 | 1M | $1.4 / $4.4 | high · max |
Z.AI China (ZhipuAI) z.ai-cn
Uses your ZhipuAI China subscription.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
glm-5.3 (default) | 1M | Included in plan | low · high · max |
glm-5.3-flash | 1M | Included in plan | low · high · max |
Z.AI China API (pay-per-use) z.ai-cn-api
Pay-per-use via ZhipuAI China API key.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
glm-5.3 (default) | 1M | $1.4 / $4.4 | low · high · max |
glm-5.3-flash | 1M | $0.15 / $0.5 | low · high · max |
glm-5.3-flashx | 1M | $0.37 / $1.25 | low · high · max |
glm-5.2 | 1M | $1.4 / $4.4 | high · max |
glm-5-turbo | 198K | $1.2 / $4 | — |
MiniMax minimax
Uses your MiniMax subscription — no per-token charges.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
MiniMax-M3 (default) | 1M | Included in plan | — |
MiniMax API (pay-per-use) minimax-api
Pay-per-use via MiniMax API key (minimaxi.com → API Keys).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
MiniMax-M3 (default) | 1M | $0.3 / $1.2 | — |
MiniMax China minimax-cn
Uses your MiniMax China subscription.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
MiniMax-M3 (default) | 1M | Included in plan | — |
DeepSeek deepseek
Pay-per-use via DeepSeek API key (platform.deepseek.com).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
deepseek-flash (default) | 1M | $0.3 / $1.2 | low · high · max |
deepseek-v4-pro | 1M | $1.32 / $3.96 | low · high · max |
Kimi (Moonshot) — Coding Plan kimi
Uses your Kimi Code subscription — no per-token charges. Key from kimi.com/code/console.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
kimi-for-coding (default) | 1M | Included in plan | low · high · max |
k3 | 1M | Included in plan | low · high · max |
k3-256k | 256K | Included in plan | low · high · max |
kimi-for-coding-highspeed | 256K | Included in plan | — |
Kimi (Moonshot) API (pay-per-use) kimi-api
Pay-per-use via Moonshot API key (platform.kimi.ai). Kimi K3 supports 1M context and graded reasoning.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
kimi-k3 (default) | 1M | $3 / $15 | low · high · max |
kimi-k2.7-code | 256K | $0.95 / $4 | — |
kimi-k2.7-code-highspeed | 256K | $1.9 / $8 | — |
kimi-k2.6 | 256K | $0.95 / $4 | — |
Kimi China (Moonshot) kimi-cn
Pay-per-use via Moonshot China API key (platform.moonshot.cn). Kimi K3 supports 1M context.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
kimi-k3 (default) | 1M | $3 / $15 | low · high · max |
kimi-k2.7-code | 256K | $0.95 / $4 | — |
kimi-k2.7-code-highspeed | 256K | $1.9 / $8 | — |
kimi-k2.6 | 256K | $0.95 / $4 | — |
Grok (xAI) grok
Pay-per-use via xAI API key (console.x.ai).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
grok-4.7 | 488K | $2 / $6 | low · medium · high · max |
grok-4.6 | 488K | $2 / $6 | low · medium · high · max |
grok-4.5 | 488K | $2 / $6 | low · medium · high |
grok-build-0.1 (default) | 250K | $1 / $2 | — |
grok-4.3 | 1M | $1.25 / $2.5 | low · medium · high |
Qwen (Alibaba) — Coding Plan qwen
Uses your Qwen Coding Plan — no per-token charges. sk-sp-… key from Model Studio. Interactive coding use only.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
qwen3.7-plus (default) | 1M | Included in plan | — |
qwen3.6-plus | 1M | Included in plan | — |
qwen3.5-plus | 1M | Included in plan | — |
Qwen (Alibaba) API (pay-per-use) qwen-api
Pay-per-use via Alibaba Model Studio key (DASHSCOPE_API_KEY).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
qwen3.8-max (default) | 1M | $2 / $6 | — |
qwen3.8-flash | 1M | $0.15 / $0.47 | — |
qwen3.7-max | 1M | $2.5 / $7.5 | — |
qwen3.7-plus | 1M | $0.4 / $1.6 | — |
qwen3.6-flash | 1M | $0.25 / $1.5 | — |
Qwen (Alibaba) — Token Plan qwen-token-plan
Uses monthly Token Plan credits. Requires a separate sk-sp-… Token Plan key.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
qwen3.8-max (default) | 1M | Included in plan | — |
qwen3.8-flash | 1M | Included in plan | — |
qwen3.7-max | 1M | Included in plan | — |
qwen3.7-plus | 1M | Included in plan | — |
qwen3.6-plus | 1M | Included in plan | — |
qwen3.6-flash | 1M | Included in plan | — |
Qwen China — Coding Plan qwen-cn
Uses your Qwen Coding Plan (China). sk-sp-… key from Bailian.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
qwen3.7-plus (default) | 1M | Included in plan | — |
qwen3.6-plus | 1M | Included in plan | — |
qwen3.5-plus | 1M | Included in plan | — |
Qwen China API (pay-per-use) qwen-cn-api
Pay-per-use via Alibaba Model Studio China key.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
qwen3.8-max | 1M | $2 / $6 | — |
qwen3.8-flash | 1M | $0.15 / $0.47 | — |
qwen3.7-max (default) | 1M | $2.5 / $7.5 | — |
qwen3.7-plus | 1M | $0.4 / $1.6 | — |
qwen3.6-flash | 1M | $0.25 / $1.5 | — |
ModelScope (free Qwen) modelscope
Fetches the live free catalog for your ModelScope token; availability and limits vary by account. Its full catalogue is fetched live — the models below are the offline fallback.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
Qwen/Qwen3.5-397B-A17B (default) | 256K | Included in plan | — |
OpenAI openai
Pay-per-use via OpenAI API key (platform.openai.com).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
gpt-6-astra | 1.1M | $10 / $50 | low · medium · high · max |
gpt-6-sol (default) | 1.1M | $2 / $10 | low · medium · high · max |
gpt-6-luna | 1.1M | $0.1 / $0.5 | low · medium · high · max |
gpt-5.6-sol | 1.1M | $4 / $20 | low · medium · high · max |
gpt-5.6-terra | 1.1M | $2 / $12 | low · medium · high · max |
gpt-5.6-luna | 1.1M | $0.2 / $1.2 | low · medium · high · max |
Anthropic anthropic
Pay-per-use via Anthropic API key (console.anthropic.com).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
claude-fable-5-1 | 1M | $10 / $50 | low · medium · high · max |
claude-opus-5-5 (default) | 1M | $4 / $20 | low · medium · high · max |
claude-fable-5 | 1M | $10 / $50 | low · medium · high · max |
claude-opus-5 | 1M | $5 / $25 | low · medium · high · max |
claude-sonnet-5-5 | 1M | $2 / $10 | low · medium · high · max |
claude-sonnet-5 | 1M | $2 / $10 | low · medium · high · max |
claude-haiku-4-5-20251001 | 195K | $1 / $5 | — |
Google AI google
Pay-per-use via Google AI API key (aistudio.google.com).
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
gemini-3.1-pro-preview (default) | 1M | $2 / $12 | low · medium · high |
gemini-3.8-flash | 1M | $0.75 / $3.75 | low · medium · high |
gemini-3.7-flash | 1M | $0.75 / $3.75 | low · medium · high |
gemini-3.6-flash | 1M | $0.75 / $3.75 | low · medium · high |
gemini-3.5-flash | 1M | $1.5 / $9 | low · medium · high |
gemini-3.5-flash-lite | 1M | $0.3 / $2.5 | low · medium · high |
OpenRouter openrouter
One key for 100+ models. Pay-per-use via openrouter.ai. Its full catalogue is fetched live — the models below are the offline fallback.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
openrouter/auto (default) | — | Not published | low · medium · high |
anthropic/claude-fable-5.1 | 1M | Not published | low · medium · high · max |
anthropic/claude-opus-5.5 | 1M | Not published | low · medium · high · max |
anthropic/claude-fable-5 | 1M | Not published | low · medium · high · max |
anthropic/claude-opus-5 | 1M | Not published | low · medium · high · max |
anthropic/claude-sonnet-5.5 | 1M | Not published | low · medium · high · max |
anthropic/claude-sonnet-5 | 1M | Not published | low · medium · high · max |
openai/gpt-6-astra | 1.1M | Not published | low · medium · high · max |
openai/gpt-6-sol | 1.1M | Not published | low · medium · high · max |
openai/gpt-6-luna | 1.1M | Not published | low · medium · high · max |
openai/gpt-5.6-sol | 1.1M | Not published | low · medium · high · max |
openai/gpt-5.6-luna | 1.1M | Not published | low · medium · high · max |
google/gemini-3.8-flash | 1M | Not published | low · medium · high |
google/gemini-3.7-flash | 1M | Not published | low · medium · high |
deepseek/deepseek-v4.1-flash | — | Not published | low · medium · high · max |
deepseek/deepseek-v4-pro-0813 | — | Not published | low · medium · high · max |
moonshotai/kimi-k3 | 1M | Not published | low · medium · high · max |
qwen/qwen3.8-max-0902 | — | Not published | low · medium · high · max |
x-ai/grok-4.7 | 488K | Not published | low · medium · high · max |
Ollama (local) ollama
Runs locally — no API key or account needed. Its full catalogue is fetched live — the models below are the offline fallback.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|---|---|---|
llama3.2 (default) | — | Not published | — |
Custom (OpenAI-compatible) custom
Point at any OpenAI-compatible server. Set the URL in /settings (Custom Base URL) or the OPENAI_BASE_URL env var, then pick your model with /model. Its full catalogue is fetched live — the models below are the offline fallback.
| Model id | Context | Input / Output per 1M | Thinking levels |
|---|
OpenRouter (Aggregator)
One API key, 100+ models from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Qwen, xAI and more. Useful when you want to experiment with multiple providers without managing separate keys, or want OpenRouter to auto-route to the best model for the task.
export OPENROUTER_API_KEY=sk-or-v1-...
codeep /provider openrouter
codeep /model # fetches the live 100+ model catalogue with pricingOr, save your key once on the dashboard under Provider keys and sync to any machine with codeep account sync.
Codeep uses the cost OpenRouter reports per call (in usage.cost) — your dashboard sees the same figures as your OpenRouter invoice, no local pricing lookup needed.
Routing preferences
OpenRouter lets you bias which upstream provider its router picks for a given model (latency, cost, geography, privacy). Configure with:
/openrouter # show current preferences
/openrouter prefer DeepInfra,Together # try these first in order
/openrouter ignore OpenAI # never route through these
/openrouter fallbacks on|off # allow fallback when preferred providers fail
/openrouter privacy strict|allow # strict = data_collection: deny
/openrouter clear # drop all preferencesopenrouter/auto as your model and OpenRouter chooses the best provider for each task. Combine with /openrouter prefer to bias the auto-router without locking it down.Ollama (Local AI)
Run AI models fully locally — or on a remote server — without an API key. Codeep fetches your installed models automatically via the Ollama API.
Local setup
Install Ollama, pull a model, and start the server:
ollama pull qwen2.5-coder:7b
ollama serveThen in Codeep, select the provider and pick your model:
/provider → select "ollama"
/model → pick from installed models (shows on-disk size)
/model browse → curated catalog of coding models; pick one to pull
/model rm <m> → remove a local model to reclaim diskNew to local models? /model browse lists recommended coding models (Qwen2.5 Coder, DeepSeek Coder V2, Llama 3.1, DeepSeek R1, …) with parameter sizes, rough VRAM, and an agent-mode hint — select one and Codeep pulls it for you.
Remote Ollama (different machine)
If Ollama runs on another machine (e.g. a home server, NAS, or Docker container), start it with OLLAMA_HOST=0.0.0.0 so it accepts external connections:
OLLAMA_HOST=0.0.0.0 ollama serveThen in Codeep, open /settings and update the Ollama URL field to point to your server:
http://192.168.1.100:11434/model after switching to the Ollama provider to see and select all models installed on your server. No manual configuration needed.Agent mode and model size
Codeep's agent mode requires a model capable of reliable tool use and instruction following. With Ollama, use at least a 7B parameter model for agent tasks — for example qwen2.5-coder:7b or llama3.1:8b.
/settings and use Codeep as a chat assistant instead.Native API (beta)
By default Codeep talks to Ollama through its OpenAI-compatible /v1 endpoint, which ignores a couple of Ollama-specific options. Turn on Ollama Native API (beta) in /settings to use the native /api/chat endpoint instead, which honors:
• num_ctx — the model uses its full context window instead of Ollama's small default (auto-detected from /api/show, or set ollamaNumCtx).
• keep_alive — keeps the model resident between turns, avoiding reload latency (ollamaKeepAlive, default 30m).
Custom (OpenAI-compatible endpoints)
Run your own model behind an OpenAI-compatible server — vLLM, LiteLLM, LM Studio, or text-generation-webui — and point Codeep straight at it. No commercial provider or API key required.
Pick Custom (OpenAI-compatible) in the welcome screen or with /provider, then set the endpoint under /settings → Custom Base URL (config key customBaseUrl). Use the full base, including /v1:
# ~/.codeep/config.json
{
"provider": "custom",
"customBaseUrl": "http://100.88.112.5:8000/v1",
"model": "qwen3-coder-30b"
}Then run /model to pick from the models your server advertises (fetched live from its /models endpoint). If your endpoint requires a key, set one with /login; otherwise leave it blank.
OPENAI_BASE_URL variable, so a proxy that serves gpt-* model names works with no config changes — just export OPENAI_BASE_URL and use the OpenAI provider.Switching providers
Use the /provider command to interactively switch between configured providers. The new provider takes effect immediately for the next message.
Using multiple providers
You can configure API keys for multiple providers. Codeep stores all keys securely. Use /login to add a key for any provider without changing the active one.