CodeepCodeep

AI Providers

Codeep supports all major AI providers. You can configure multiple providers and switch between them at any time using /provider or /settings.

★
Set your API key as an environment variable and Codeep will detect it automatically on startup, no configuration needed.

Supported Providers

ProviderEnv VariableDefault ModelProtocol
AnthropicANTHROPIC_API_KEYclaude-opus-5-5Anthropic
Custom (OpenAI-compatible)none requiredyour modelOpenAI-compatible
DeepSeekDEEPSEEK_API_KEYdeepseek-flashOpenAI-compatible
Google AIGOOGLE_API_KEYgemini-3.1-pro-previewOpenAI-compatible
Grok (xAI)XAI_API_KEYgrok-build-0.1OpenAI-compatible
Kimi (Kimi Code subscription)KIMI_CODE_API_KEYkimi-for-coding (K2.8 Preview; K3 on eligible plans)OpenAI-compatible
Kimi API (pay-per-use)MOONSHOT_API_KEYkimi-k3OpenAI-compatible
Kimi ChinaMOONSHOT_CN_API_KEYkimi-k3OpenAI-compatible
MiniMax (Coding Plan)MINIMAX_API_KEYMiniMax-M3OpenAI / Anthropic
MiniMax API (pay-per-use)MINIMAX_API_KEYMiniMax-M3OpenAI-compatible
MiniMax ChinaMINIMAX_CN_API_KEYMiniMax-M3OpenAI / Anthropic
ModelScope (live free catalog)MODELSCOPE_API_KEYdynamic (Qwen3.5-397B fallback)OpenAI-compatible
Ollama (local / remote)none requireddynamicOpenAI-compatible
OpenAIOPENAI_API_KEYgpt-6-solOpenAI Responses
OpenRouter (100+ models)OPENROUTER_API_KEYopenrouter/autoOpenAI-compatible
Qwen (Coding Plan subscription)BAILIAN_CODING_PLAN_API_KEYqwen3.7-plusOpenAI-compatible
Qwen (Token Plan subscription)BAILIAN_TOKEN_PLAN_API_KEYqwen3.8-maxOpenAI-compatible
Qwen API (pay-per-use)DASHSCOPE_API_KEYqwen3.8-maxOpenAI-compatible
Qwen China (Coding Plan / API)BAILIAN_CODING_PLAN_CN_API_KEY / DASHSCOPE_CN_API_KEYqwen3.7-plus / qwen3.7-maxOpenAI-compatible
Z.AI (GLM Coding Plan)ZAI_API_KEYglm-5.3OpenAI / Anthropic
Z.AI API (pay-per-use)ZAI_API_KEYglm-5.3OpenAI-compatible
Z.AI China (GLM Coding Plan)ZAI_CN_API_KEYglm-5.3OpenAI / Anthropic
Z.AI China API (pay-per-use)ZAI_CN_API_KEYglm-5.3OpenAI-compatible
ℹ
GLM-5.3-Flash is on every Z.AI entry, subscription and pay-per-use, in both regions: the same million-token window and tool calling as GLM-5.3 at roughly a ninth of the price — $0.15 in / $0.50 out per 1M against 5.3's $1.40 / $4.40 — which makes it the obvious default for high-volume work like CI fixes. It takes the same graded thinking control, and like 5.3 it cannot have thinking disabled. Both GLM Coding Plans now accept only GLM-5.3 and GLM-5.3-Flash; the faster GLM-5.3-FlashX is on pay-per-use only.
★
Subscription plans. Kimi (Kimi Code), Qwen (Coding Plan), Z.AI (GLM Coding Plan) and MiniMax all offer a flat-fee subscription you can drive directly — Codeep points at the plan's dedicated endpoint with your subscription key, so there are no per-token charges. Qwen Token Plan is a separate endpoint/key and consumes monthly plan credits. Each also has a separate pay-per-use API entry. (Grok adds Kimi/Qwen- style coding models via a pay-per-use xAI key; its SuperGrok subscription isn't API-drivable.)
★
Live ModelScope models. After selecting ModelScope, run /model. Codeep reads the catalog available to your token at that moment and falls back to Qwen3.5-397B only if the catalog request fails.

Every model Codeep ships

Generated straight from the client catalogue, so it lists exactly what /model will offer — 97 models across 24 providers, as of [email protected]. Ids are the exact strings the API receives.

ℹ
About the price column. Codeep quotes a rate only where the provider publishes one and the account is actually metered. Subscription plans read Included in plan — your tokens are real, the per-token dollar figure would not be. A model whose provider has published no rate reads Not published rather than borrowing a sibling's price.

Z.AI (ZhipuAI) z.ai

Uses your Z.AI subscription — no per-token charges.

Model idContextInput / Output per 1MThinking levels
glm-5.3 (default)1MIncluded in planlow · high · max
glm-5.3-flash1MIncluded in planlow · high · max

Z.AI API (pay-per-use) z.ai-api

Pay-per-use via Z.AI API key (zai.ai → API Keys).

Model idContextInput / Output per 1MThinking levels
glm-5.3 (default)1M$1.4 / $4.4low · high · max
glm-5.3-flash1M$0.15 / $0.5low · high · max
glm-5.3-flashx1M$0.37 / $1.25low · high · max
glm-5.21M$1.4 / $4.4high · max

Z.AI China (ZhipuAI) z.ai-cn

Uses your ZhipuAI China subscription.

Model idContextInput / Output per 1MThinking levels
glm-5.3 (default)1MIncluded in planlow · high · max
glm-5.3-flash1MIncluded in planlow · high · max

Z.AI China API (pay-per-use) z.ai-cn-api

Pay-per-use via ZhipuAI China API key.

Model idContextInput / Output per 1MThinking levels
glm-5.3 (default)1M$1.4 / $4.4low · high · max
glm-5.3-flash1M$0.15 / $0.5low · high · max
glm-5.3-flashx1M$0.37 / $1.25low · high · max
glm-5.21M$1.4 / $4.4high · max
glm-5-turbo198K$1.2 / $4—

MiniMax minimax

Uses your MiniMax subscription — no per-token charges.

Model idContextInput / Output per 1MThinking levels
MiniMax-M3 (default)1MIncluded in plan—

MiniMax API (pay-per-use) minimax-api

Pay-per-use via MiniMax API key (minimaxi.com → API Keys).

Model idContextInput / Output per 1MThinking levels
MiniMax-M3 (default)1M$0.3 / $1.2—

MiniMax China minimax-cn

Uses your MiniMax China subscription.

Model idContextInput / Output per 1MThinking levels
MiniMax-M3 (default)1MIncluded in plan—

DeepSeek deepseek

Pay-per-use via DeepSeek API key (platform.deepseek.com).

Model idContextInput / Output per 1MThinking levels
deepseek-flash (default)1M$0.3 / $1.2low · high · max
deepseek-v4-pro1M$1.32 / $3.96low · high · max

Kimi (Moonshot) — Coding Plan kimi

Uses your Kimi Code subscription — no per-token charges. Key from kimi.com/code/console.

Model idContextInput / Output per 1MThinking levels
kimi-for-coding (default)1MIncluded in planlow · high · max
k31MIncluded in planlow · high · max
k3-256k256KIncluded in planlow · high · max
kimi-for-coding-highspeed256KIncluded in plan—

Kimi (Moonshot) API (pay-per-use) kimi-api

Pay-per-use via Moonshot API key (platform.kimi.ai). Kimi K3 supports 1M context and graded reasoning.

Model idContextInput / Output per 1MThinking levels
kimi-k3 (default)1M$3 / $15low · high · max
kimi-k2.7-code256K$0.95 / $4—
kimi-k2.7-code-highspeed256K$1.9 / $8—
kimi-k2.6256K$0.95 / $4—

Kimi China (Moonshot) kimi-cn

Pay-per-use via Moonshot China API key (platform.moonshot.cn). Kimi K3 supports 1M context.

Model idContextInput / Output per 1MThinking levels
kimi-k3 (default)1M$3 / $15low · high · max
kimi-k2.7-code256K$0.95 / $4—
kimi-k2.7-code-highspeed256K$1.9 / $8—
kimi-k2.6256K$0.95 / $4—

Grok (xAI) grok

Pay-per-use via xAI API key (console.x.ai).

Model idContextInput / Output per 1MThinking levels
grok-4.7488K$2 / $6low · medium · high · max
grok-4.6488K$2 / $6low · medium · high · max
grok-4.5488K$2 / $6low · medium · high
grok-build-0.1 (default)250K$1 / $2—
grok-4.31M$1.25 / $2.5low · medium · high

Qwen (Alibaba) — Coding Plan qwen

Uses your Qwen Coding Plan — no per-token charges. sk-sp-… key from Model Studio. Interactive coding use only.

Model idContextInput / Output per 1MThinking levels
qwen3.7-plus (default)1MIncluded in plan—
qwen3.6-plus1MIncluded in plan—
qwen3.5-plus1MIncluded in plan—

Qwen (Alibaba) API (pay-per-use) qwen-api

Pay-per-use via Alibaba Model Studio key (DASHSCOPE_API_KEY).

Model idContextInput / Output per 1MThinking levels
qwen3.8-max (default)1M$2 / $6—
qwen3.8-flash1M$0.15 / $0.47—
qwen3.7-max1M$2.5 / $7.5—
qwen3.7-plus1M$0.4 / $1.6—
qwen3.6-flash1M$0.25 / $1.5—

Qwen (Alibaba) — Token Plan qwen-token-plan

Uses monthly Token Plan credits. Requires a separate sk-sp-… Token Plan key.

Model idContextInput / Output per 1MThinking levels
qwen3.8-max (default)1MIncluded in plan—
qwen3.8-flash1MIncluded in plan—
qwen3.7-max1MIncluded in plan—
qwen3.7-plus1MIncluded in plan—
qwen3.6-plus1MIncluded in plan—
qwen3.6-flash1MIncluded in plan—

Qwen China — Coding Plan qwen-cn

Uses your Qwen Coding Plan (China). sk-sp-… key from Bailian.

Model idContextInput / Output per 1MThinking levels
qwen3.7-plus (default)1MIncluded in plan—
qwen3.6-plus1MIncluded in plan—
qwen3.5-plus1MIncluded in plan—

Qwen China API (pay-per-use) qwen-cn-api

Pay-per-use via Alibaba Model Studio China key.

Model idContextInput / Output per 1MThinking levels
qwen3.8-max1M$2 / $6—
qwen3.8-flash1M$0.15 / $0.47—
qwen3.7-max (default)1M$2.5 / $7.5—
qwen3.7-plus1M$0.4 / $1.6—
qwen3.6-flash1M$0.25 / $1.5—

ModelScope (free Qwen) modelscope

Fetches the live free catalog for your ModelScope token; availability and limits vary by account. Its full catalogue is fetched live — the models below are the offline fallback.

Model idContextInput / Output per 1MThinking levels
Qwen/Qwen3.5-397B-A17B (default)256KIncluded in plan—

OpenAI openai

Pay-per-use via OpenAI API key (platform.openai.com).

Model idContextInput / Output per 1MThinking levels
gpt-6-astra1.1M$10 / $50low · medium · high · max
gpt-6-sol (default)1.1M$2 / $10low · medium · high · max
gpt-6-luna1.1M$0.1 / $0.5low · medium · high · max
gpt-5.6-sol1.1M$4 / $20low · medium · high · max
gpt-5.6-terra1.1M$2 / $12low · medium · high · max
gpt-5.6-luna1.1M$0.2 / $1.2low · medium · high · max

Anthropic anthropic

Pay-per-use via Anthropic API key (console.anthropic.com).

Model idContextInput / Output per 1MThinking levels
claude-fable-5-11M$10 / $50low · medium · high · max
claude-opus-5-5 (default)1M$4 / $20low · medium · high · max
claude-fable-51M$10 / $50low · medium · high · max
claude-opus-51M$5 / $25low · medium · high · max
claude-sonnet-5-51M$2 / $10low · medium · high · max
claude-sonnet-51M$2 / $10low · medium · high · max
claude-haiku-4-5-20251001195K$1 / $5—

Google AI google

Pay-per-use via Google AI API key (aistudio.google.com).

Model idContextInput / Output per 1MThinking levels
gemini-3.1-pro-preview (default)1M$2 / $12low · medium · high
gemini-3.8-flash1M$0.75 / $3.75low · medium · high
gemini-3.7-flash1M$0.75 / $3.75low · medium · high
gemini-3.6-flash1M$0.75 / $3.75low · medium · high
gemini-3.5-flash1M$1.5 / $9low · medium · high
gemini-3.5-flash-lite1M$0.3 / $2.5low · medium · high

OpenRouter openrouter

One key for 100+ models. Pay-per-use via openrouter.ai. Its full catalogue is fetched live — the models below are the offline fallback.

Model idContextInput / Output per 1MThinking levels
openrouter/auto (default)—Not publishedlow · medium · high
anthropic/claude-fable-5.11MNot publishedlow · medium · high · max
anthropic/claude-opus-5.51MNot publishedlow · medium · high · max
anthropic/claude-fable-51MNot publishedlow · medium · high · max
anthropic/claude-opus-51MNot publishedlow · medium · high · max
anthropic/claude-sonnet-5.51MNot publishedlow · medium · high · max
anthropic/claude-sonnet-51MNot publishedlow · medium · high · max
openai/gpt-6-astra1.1MNot publishedlow · medium · high · max
openai/gpt-6-sol1.1MNot publishedlow · medium · high · max
openai/gpt-6-luna1.1MNot publishedlow · medium · high · max
openai/gpt-5.6-sol1.1MNot publishedlow · medium · high · max
openai/gpt-5.6-luna1.1MNot publishedlow · medium · high · max
google/gemini-3.8-flash1MNot publishedlow · medium · high
google/gemini-3.7-flash1MNot publishedlow · medium · high
deepseek/deepseek-v4.1-flash—Not publishedlow · medium · high · max
deepseek/deepseek-v4-pro-0813—Not publishedlow · medium · high · max
moonshotai/kimi-k31MNot publishedlow · medium · high · max
qwen/qwen3.8-max-0902—Not publishedlow · medium · high · max
x-ai/grok-4.7488KNot publishedlow · medium · high · max

Ollama (local) ollama

Runs locally — no API key or account needed. Its full catalogue is fetched live — the models below are the offline fallback.

Model idContextInput / Output per 1MThinking levels
llama3.2 (default)—Not published—

Custom (OpenAI-compatible) custom

Point at any OpenAI-compatible server. Set the URL in /settings (Custom Base URL) or the OPENAI_BASE_URL env var, then pick your model with /model. Its full catalogue is fetched live — the models below are the offline fallback.

Model idContextInput / Output per 1MThinking levels

OpenRouter (Aggregator)

One API key, 100+ models from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Qwen, xAI and more. Useful when you want to experiment with multiple providers without managing separate keys, or want OpenRouter to auto-route to the best model for the task.

export OPENROUTER_API_KEY=sk-or-v1-...
codeep /provider openrouter
codeep /model                        # fetches the live 100+ model catalogue with pricing

Or, save your key once on the dashboard under Provider keys and sync to any machine with codeep account sync.

Codeep uses the cost OpenRouter reports per call (in usage.cost) — your dashboard sees the same figures as your OpenRouter invoice, no local pricing lookup needed.

Routing preferences

OpenRouter lets you bias which upstream provider its router picks for a given model (latency, cost, geography, privacy). Configure with:

/openrouter                      # show current preferences
/openrouter prefer DeepInfra,Together   # try these first in order
/openrouter ignore OpenAI               # never route through these
/openrouter fallbacks on|off            # allow fallback when preferred providers fail
/openrouter privacy strict|allow        # strict = data_collection: deny
/openrouter clear                       # drop all preferences
★
Pick openrouter/auto as your model and OpenRouter chooses the best provider for each task. Combine with /openrouter prefer to bias the auto-router without locking it down.

Ollama (Local AI)

Run AI models fully locally — or on a remote server — without an API key. Codeep fetches your installed models automatically via the Ollama API.

Local setup

Install Ollama, pull a model, and start the server:

ollama pull qwen2.5-coder:7b
ollama serve

Then in Codeep, select the provider and pick your model:

/provider     → select "ollama"
/model        → pick from installed models (shows on-disk size)
/model browse → curated catalog of coding models; pick one to pull
/model rm <m> → remove a local model to reclaim disk

New to local models? /model browse lists recommended coding models (Qwen2.5 Coder, DeepSeek Coder V2, Llama 3.1, DeepSeek R1, …) with parameter sizes, rough VRAM, and an agent-mode hint — select one and Codeep pulls it for you.

Remote Ollama (different machine)

If Ollama runs on another machine (e.g. a home server, NAS, or Docker container), start it with OLLAMA_HOST=0.0.0.0 so it accepts external connections:

OLLAMA_HOST=0.0.0.0 ollama serve

Then in Codeep, open /settings and update the Ollama URL field to point to your server:

http://192.168.1.100:11434
★
Use /model after switching to the Ollama provider to see and select all models installed on your server. No manual configuration needed.

Agent mode and model size

Codeep's agent mode requires a model capable of reliable tool use and instruction following. With Ollama, use at least a 7B parameter model for agent tasks — for example qwen2.5-coder:7b or llama3.1:8b.

⚠
Models smaller than 7B (1B–3B) often fail to produce correctly formatted tool calls and may behave unexpectedly in agent mode. For small models, set Agent Mode to Manual or Off in /settings and use Codeep as a chat assistant instead.

Native API (beta)

By default Codeep talks to Ollama through its OpenAI-compatible /v1 endpoint, which ignores a couple of Ollama-specific options. Turn on Ollama Native API (beta) in /settings to use the native /api/chat endpoint instead, which honors:

• num_ctx — the model uses its full context window instead of Ollama's small default (auto-detected from /api/show, or set ollamaNumCtx).
• keep_alive — keeps the model resident between turns, avoiding reload latency (ollamaKeepAlive, default 30m).

★
It's marked beta while it gets real-world coverage across more models and longer agent sessions — it's off by default, so nothing changes unless you opt in. Hit a problem? Please open an issue on GitHub — feedback decides when it becomes the default.

Custom (OpenAI-compatible endpoints)

Run your own model behind an OpenAI-compatible server — vLLM, LiteLLM, LM Studio, or text-generation-webui — and point Codeep straight at it. No commercial provider or API key required.

Pick Custom (OpenAI-compatible) in the welcome screen or with /provider, then set the endpoint under /settings → Custom Base URL (config key customBaseUrl). Use the full base, including /v1:

# ~/.codeep/config.json
{
  "provider": "custom",
  "customBaseUrl": "http://100.88.112.5:8000/v1",
  "model": "qwen3-coder-30b"
}

Then run /model to pick from the models your server advertises (fetched live from its /models endpoint). If your endpoint requires a key, set one with /login; otherwise leave it blank.

★
Prefer environment variables? The OpenAI provider honors the standard OPENAI_BASE_URL variable, so a proxy that serves gpt-* model names works with no config changes — just export OPENAI_BASE_URL and use the OpenAI provider.

Switching providers

Use the /provider command to interactively switch between configured providers. The new provider takes effect immediately for the next message.

Using multiple providers

You can configure API keys for multiple providers. Codeep stores all keys securely. Use /login to add a key for any provider without changing the active one.