Skip to content

Models

Every workspace picks its own model, and you can change it whenever you like. There are two ways to connect one.

If you already use a command-line coding assistant, Rookery can drive it directly and reuse the sign-in you already did. No key to paste.

Supported: Claude Code, Codex, OpenCode, Cursor and Gemini CLI. Rookery looks for them on your PATH and in ~/.local/bin.

Each workspace gets its own isolated configuration directory, so one workspace’s agents can never read another’s session or history. Your operator credentials are copied in per invocation.

Give Rookery a provider, a model name and a key, and it talks to the API directly in-process — no separate tool involved.

Hosted covers the frontier labs (OpenAI, Anthropic, Gemini, xAI, Mistral, DeepSeek, Moonshot, Z.AI), the routers (OpenRouter, Perplexity, OpenCode Zen and Go), the enterprise clouds (AWS Bedrock, Alibaba Cloud) and the open-weight inference clouds (Groq, Together, Fireworks, Cerebras, SambaNova, Nebius, DeepInfra, Hugging Face, GitHub Models, Ollama Cloud).

Self-hosted covers OpenAI-compatible servers on your own hardware: Ollama, vLLM, LM Studio, llama.cpp, LocalAI and Jan. These need no key — point Rookery at the address and nothing leaves your network.

There is also a Custom (OpenAI-compatible) option for anything else that speaks the same protocol.

Settings → Coder in the workspace. Pick the kind, then the provider and model. You can paste a key straight into the form and Rookery stores it as an encrypted secret for you.

The base URL is prefilled with the provider’s default and can be overridden — which is how you point at a local server on a non-standard port. An untouched prefill keeps following the default rather than freezing on today’s address.

The published image ships no coder tool and enforces it: a workspace running in the container must use a provider API. Attempting to configure a local coder fails with a message saying so, rather than trying to run a binary that is not there.

If youUse
Already have a coder tool set upThat tool
Want the fewest moving partsA provider API
Want nothing leaving your networkA self-hosted model
Run the containerA provider API

When a provider runs out of credit or rate-limits you, the run fails with a message saying which — “out of quota” and “try again shortly” are different problems and are reported differently. Token usage is recorded per run on the agent’s page.