User manual

Camel CLI - AI Providers and Local Models

The AI features of the Camel CLI — camel ask, camel explain, camel harden, camel overview --ai, and the AI panel of the Camel TUI — talk to a large language model (LLM). This page explains how the provider is chosen, and how to run a model on your own machine.

Choosing a provider

The provider is detected from the environment. The first match wins:

  1. ANTHROPIC_API_KEY → Anthropic API

  2. CLOUD_ML_REGION + ANTHROPIC_VERTEX_PROJECT_ID → Anthropic models on Google Vertex AI

  3. AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT → Azure OpenAI (uses the api-key header; optional AZURE_OPENAI_DEPLOYMENT_NAME and AZURE_OPENAI_API_VERSION)

  4. GEMINI_API_KEY → Google Gemini native API (generativelanguage.googleapis.com). With --api-type=gemini, GOOGLE_API_KEY is also accepted.

  5. OPENAI_API_KEY → OpenAI API (api.openai.com)

  6. LLM_API_KEY + optional LLM_BASE_URL (or OPENAI_BASE_URL) → any OpenAI-compatible API

  7. WATSONX_APIKEY + optional WATSONX_URL → IBM watsonx.ai

  8. Ollama started with camel infra run ollama → that Ollama

  9. Ollama at localhost:11434 → local Ollama

Override the detection with explicit options:

camel ask "check health" --api-type=anthropic --api-key=sk-...
camel ask "check health" --api-type=openai --model=gpt-4
camel ask "check health" --api-type=gemini --model=gemini-2.0-flash
camel ask "check health" --api-type=ollama --model=llama3.1
camel harden only detects Ollama (from camel infra or at localhost:11434). For another provider, give it the endpoint: --url=https://api.openai.com --api-type=openai --api-key=sk-…​ (any OpenAI-compatible API).

In the Camel TUI, press Ctrl+P in the AI panel to switch provider or model, or pin them in F2 → Settings (camel.tui.ai.provider, camel.tui.ai.model, camel.tui.ai.url). See Camel TUI AI Panel.

Running a model locally

Ollama

Install Ollama natively for the best performance — the native binary uses GPU acceleration (Metal on macOS, CUDA/ROCm on Linux). Running Ollama through Docker (camel infra run ollama) bypasses the GPU and makes inference much slower.

# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# Pull a model and start asking
ollama pull qwen3.6:35b-a3b
camel ask "what routes are running?"

Ollama at localhost:11434 is auto-detected — no environment variable needed. The CLI checks what models are available and auto-selects a suitable one.

Model requirements

camel ask and the TUI F8 panel rely on tool calling to inspect your running Camel process. Models smaller than ~14B do not reliably invoke tools and answer from training knowledge instead. Use at least a 14B model. Prefer a mixture-of-experts model such as qwen3.6:35b-a3b: with only 3B parameters active per token it processes the tool-heavy prompt many times faster than a dense 27B/32B model, so answers start in seconds instead of a minute.

Model RAM (Q4) Notes

qwen3.6:35b-a3b

~23 GB

Recommended: fastest prompt processing, needs 32 GB+

qwen2.5:14b

~9 GB

Minimum for 16 GB machines

qwen3.6:27b

~18 GB

Strong dense model, several times slower prompt processing

qwen2.5:32b

~20 GB

Good quality, slow prompt processing

hermes3:70b

~43 GB

Excellent tool calling, needs 64 GB+

llama3.3:70b

~43 GB

Best open model, needs 64 GB+

On Apple Silicon, all RAM is unified — a 64 GB M-series Mac can run llama3.3:70b comfortably alongside the OS and other dev tools. Use the default (GGUF) tags rather than the -mlx tags: the Ollama MLX engine cannot yet reuse the cached prompt for Qwen 3.x models, so every question re-processes the whole prompt.

Context window and keep-alive

Every Ollama request from the CLI (camel ask, camel explain, camel harden) and from the TUI asks Ollama to keep the model loaded for 30 minutes, so a follow-up question reuses the cached prompt instead of reloading the model, and for a context window (num_ctx) chosen once per model:

  1. OLLAMA_CONTEXT_LENGTH in the environment wins when set.

  2. If Ollama already has the model loaded, its window is adopted (raised to 32k when smaller), so the CLI never makes Ollama reload a model that another client is using. A request with a different num_ctx makes Ollama reload the model, which costs a cold start of tens of seconds and throws away the prompt cache of every client.

  3. Otherwise 64k when the model’s weights plus the KV cache of a 64k window fit the machine’s memory, judged from what ollama show reports about the model, else 32k. A window that does not fit makes Ollama move layers to the CPU, which slows generation far more than a smaller window costs.

The tool-calling prompt alone is about 4.5k tokens, which is why 32k is the floor: Ollama’s own default on machines with less than 24 GB is 4k and would truncate it. Keep all your Ollama clients on the same value; if you set OLLAMA_CONTEXT_LENGTH for the CLI, set it for ollama serve too. The TUI’s Ollama tab shows the window in use, and Working with a local Ollama model explains how the AI panel manages its history inside it.

OpenAI-compatible local servers

Many local LLM servers expose an OpenAI-compatible API. Use LLM_API_KEY and LLM_BASE_URL to point the CLI at any of them:

export LLM_API_KEY=any-value      # required but can be any non-empty string
export LLM_BASE_URL=http://localhost:1234   # your server's base URL
camel ask "what routes are running?"

OPENAI_BASE_URL is accepted as an alternative to LLM_BASE_URL (common in other tools).

The first model the server lists on /v1/models is used unless --model names another one; the model must support tool calling.

Common OpenAI-compatible servers:

Server Default port Notes

LM Studio

1234

GUI app, Mac/Windows/Linux

vLLM

8000

Production-grade, NVIDIA GPU

llama.cpp server

8080

Runs on CPU and GPU

GPT4All

4891

Desktop app

See also

  • AI Tools — camel ask, camel explain, camel overview --ai and camel harden

  • Camel TUI Local Models — the tool set a local model gets in the TUI, what a question costs, and the Ollama tab

  • Camel MCP Server — the other way round: your own AI assistant uses Camel tools