# Camel TUI Local Models

The [AI panel](camel-jbang-tui-ai.md) (**F8**) works with a model running on your own machine, through [Ollama](https://ollama.com) or any OpenAI-compatible server, so no API key is needed and nothing has to leave your machine. A local model spends its time differently from a hosted one, and the TUI shows you where.

See [Camel TUI](camel-jbang-tui.md) for getting started and the other pages.

## Key Features

-   [**Ollama**](#_using_ollama_local_no_api_key) — auto-detected at `localhost:11434`, with the models that work well
    
-   [**OpenAI-compatible servers**](#_using_an_openai_compatible_local_server) — LM Studio, vLLM, llama.cpp and others
    
-   [**A smaller tool set**](#_tool_set_for_local_models) — local models get the core tools, which roughly halves the prompt
    
-   [**What a question costs**](#_working_with_a_local_ollama_model) — load, prefill and decode, the prompt cache and the context window
    
-   [**Ollama tab**](#_ollama_tab) — the loaded model, tokens per second, cache hit and context fill, live
    

## Using Ollama (local, no API key)

Install Ollama natively for best performance — the native binary uses GPU acceleration (Metal on macOS, CUDA/ROCm on Linux):

```bash
# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# Pull a model — then open the TUI and press F8
ollama pull qwen3.6:35b-a3b
camel tui
```

Ollama at `localhost:11434` is auto-detected. No configuration needed.

> **Important**
> The F8 AI panel works by invoking built-in tools to inspect your running Camel process. Models smaller than ~14B do not reliably call tools and answer from training knowledge instead. Use at least a 14B model. Prefer a mixture-of-experts model such as `qwen3.6:35b-a3b`: with only 3B parameters active per token it processes the tool-heavy prompt many times faster than a dense 27B/32B model, so answers start in seconds instead of a minute.

**Models that work well** (tool-calling capable, ≥14B, default Q4\_K\_M quantization):

  
| Model | RAM | Notes |
| --- | --- | --- |
| `qwen3.6:35b-a3b` | ~23 GB | Recommended: fastest prompt processing, needs 32 GB+ |
| `qwen2.5:14b` | ~9 GB | Minimum for 16 GB machines |
| `qwen3.6:27b` | ~18 GB | Strong dense model, several times slower prompt processing |
| `qwen2.5:32b` | ~20 GB | Good quality, slow prompt processing |
| `hermes3:70b` | ~43 GB | Excellent tool calling, needs 64 GB+ |
| `llama3.3:70b` | ~43 GB | Best open model, needs 64 GB+ |
> **Note**
> `camel infra run ollama` runs Ollama in Docker and bypasses GPU acceleration, making inference significantly slower. Native install is preferred for development.

> **Note**
> On Apple Silicon, use the default (GGUF) tags rather than the `-mlx` tags. The Ollama MLX engine cannot yet reuse the cached prompt for Qwen 3.x models, so every question re-processes the whole prompt, while the default engine reuses it and only processes what is new.

## Using an OpenAI-compatible local server

Set `LLM_API_KEY` and `LLM_BASE_URL` to connect to any OpenAI-compatible server (LM Studio, vLLM, llama.cpp, GPT4All, …):

```bash
export LLM_API_KEY=any-value
export LLM_BASE_URL=http://localhost:1234
camel tui
```

`OPENAI_BASE_URL` is also accepted as an alternative to `LLM_BASE_URL`.

The panel uses the first model the server lists on `/v1/models`. To use another one, run `/model <name>` in the panel (`/model` alone lists what the server offers), set **AI Model** in **F2 → Settings**, or set `camel.tui.ai.model`. The model must support tool calling, otherwise the panel answers from training data instead of inspecting your integration. When a request fails, the panel shows the HTTP status and the server’s error message, for example a model that the server does not host.

## Tool set for local models

Every question sends the definitions of the `tui_*` tools the model may call, and a local model pays for each of them in prompt-processing time. The panel therefore sends only the core set of tools (state, tables, logs, errors, diagrams, topology, processor details, catalog docs, traces, spans, route control, sending messages, source files, infra services, navigation, log level and filters) to Ollama and to any provider on `localhost`, which roughly halves the prompt. Hosted providers get every tool, including the drawing, animation and automation tools. Use `/tools full` in the panel to send all tools to a local model too, `/tools core` to trim the set for a hosted one, pick **AI Tools** in **F2 → Settings**, or set `camel.tui.ai.tools` in `.camel-cli.properties`. Ollama requests also ask the server to keep the model loaded for 30 minutes and for a context window of 32k or 64k (see [Working with a local Ollama model](#_working_with_a_local_ollama_model); `OLLAMA_CONTEXT_LENGTH` overrides it), so follow-up questions reuse the cached prompt instead of reloading the model.

## Working with a local Ollama model

A local model is not a slower version of a hosted one; it spends its time differently, and the TUI shows you where. This section explains what a question costs with Ollama and which knobs matter, with figures measured on an Apple M4 Pro (64 GB) running `qwen3.6:35b-a3b`.

**How a question is spent.** Ollama answers in three phases, and every timing the TUI shows maps to one of them:

1.  **Load**: if the model is not in memory (first question, the keep-alive expired, or a request asked for a different context size) Ollama starts a runner and loads the weights: 10 to 20 seconds for a 22 GB model. The TUI asks Ollama to keep the model loaded for 30 minutes after each request.
    
2.  **Prefill**: the prompt (system prompt, tool definitions, conversation history, your question) is processed in one batch, at roughly 600 to 700 tokens per second when nothing is cached.
    
3.  **Decode**: the answer is generated token by token, at 50 to 60 tokens per second for this model.
    

The wait before anything appears, the time to first token, is load plus prefill. A cold first question therefore takes 20 seconds before the first word; a warm one under a second.

**One question, many requests.** The panel answers by calling the `tui_*` tools, and every tool call the model makes costs a new request that sends the whole prompt again. A simple question can take 3 to 13 requests. This is affordable only because Ollama caches the prompt prefix: the system prompt, the tool definitions and the history are identical from one step to the next, so each step prefills only the new tokens and takes about half a second. Across a session the cache hit is typically above 90%. The Ollama tab shows the request count per question as `×N` and the cache hit per question; a question with a high count and a short answer is the model exploring, which a smaller, sharper tool set reduces (see [Tool set for local models](#_tool_set_for_local_models)).

**The context window.** The TUI decides the window it asks Ollama for (`num_ctx`) once per model:

1.  `OLLAMA_CONTEXT_LENGTH` in the environment wins when set.
    
2.  If the model is already loaded, its window is adopted (raised to 32k when smaller), so the TUI never makes Ollama reload a model another client is using. A request with a different `num_ctx` costs a cold start and throws away everyone’s prompt cache.
    
3.  Otherwise 64k when the model’s weights plus the KV cache of a 64k window fit the machine’s memory, computed from what `ollama show` reports (layers, KV heads, head size), else 32k. Overshooting is worse than being conservative: Ollama then moves layers to the CPU and generation slows to a crawl.
    

The static prefix of system prompt and core tools is about 4.5k tokens, and each question with tool calls adds another 2k to 4k of history. The panel compacts the history once the prompt Ollama measured for the last request passes half the window, capped at half of 64k even when a larger window was adopted: every token of history is prefill time again after a cache loss (an idle unload after 30 minutes, another client, a restart), and 64k at a few hundred tokens per second is already well over a minute. Compacting rewrites the history, which invalidates Ollama’s prompt cache, so the request after a compaction prefills the whole prompt again: about 40 seconds for a 21k-token prompt. The panel therefore compacts local history rarely and then thoroughly, down to about a quarter of the window, and prints one line saying what it did and how long the next reply will take to start. `/compact` does the same on demand with the same line; `/context` shows the window, the compaction point and the last measured prompt; the panel title shows the fill as `ctx 34%`.

**Choosing the model.** With no `camel.tui.ai.model` set the panel takes `llama3.2` when it is installed, otherwise the first installed model; run `/model <name>` in the panel or set **AI Model** in **F2 → Settings** to pin one. The model must support tool calling and should have at least 14B parameters; a mixture-of-experts model such as `qwen3.6:35b-a3b` prefills several times faster than a dense model of similar quality, which is what matters for a tool-heavy prompt. **F2 → Run Doctor** shows whether Ollama was found, which models are installed and whether they are large enough.

**Where to look.** The [Ollama tab](#_ollama_tab) is the instrument for all of the above: tokens per second live and per request, time to first token with cold starts marked, cache hit, how full the context window is and how it grows per question, GPU and process load, and one line per question with the requests it took, the question being answered right now on top with `working` as its reason and its time counting up, and a footer with the average per question. In the AI panel, `/context` prints what the next request will cost, `/usage` and **Ctrl+U** the session totals per question.

**Remote and containerised Ollama.** Everything above applies to an Ollama on another host or inside `camel infra run ollama` as well, with two differences: the container runs without GPU acceleration, and the live runner state and host load on the Ollama tab need the server on the same machine.

For the wider picture see the blog posts [We had a frontier AI coach a small local model through Camel](/blog/2026/09/camel-local-model-benchmark/) on what a local model can do with Camel and what was changed to help it, and [Observe Your Camel AI Routes with GenAI OpenTelemetry](/blog/2026/09/camel-genai-observability-jbang/) on observing routes that call Ollama.

## Ollama tab

The Ollama tab (under More, in the **AI** group) shows how the model served by a local [Ollama](https://ollama.com) is performing, in the spirit of an LLM dashboard. It works with or without a running integration: it finds Ollama at `localhost:11434`, at the address of `camel infra run ollama`, or at the endpoint the AI panel (**F8**) is using. The tab is listed only while an Ollama server answers; the TUI checks every ten seconds, so it appears shortly after `ollama serve` starts.

-   **Model** — the loaded model with its family, parameters, quantization, layers, experts (and how many are active per token for a mixture-of-experts model), how much of it sits in GPU memory, the allocated context length and when Ollama will unload it. With no model loaded, the installed models are listed instead.
    
-   **Throughput** — decode and prefill tokens per second, **live** while the model is generating and otherwise from the last request; time to first token and load time (a load of a second or more is a cold start); session averages and a sparkline of the decode rate.
    
-   **Context** — how full the context window is, the share of the prompt served from Ollama’s cache, whether the model is working or idle, the speculative decoding method in use, and a per-turn trend of how much of the window each AI panel prompt filled, with the session peak and the compactions seen.
    
-   **Host** — GPU utilization and memory (Apple silicon through `ioreg`, NVIDIA through `nvidia-smi`), and CPU and memory of the Ollama server and its model runner.
    
-   **Requests** — one line per question asked in the AI panel (a question with tool calls costs one request per step; **Enter** unfolds the steps) with the question text, prompt and generated tokens, cache hit, the share of the context window reached, prefill and decode tokens per second, time to first token, the time you waited and the stop reason. Calls made by Camel routes are listed as their own lines.
    

Two kinds of requests appear in the log. Questions asked in the AI panel with Ollama as the provider come with the timings Ollama reports for each request (prompt evaluation, generation, model load, total). Calls made by Camel routes through `camel-langchain4j-chat`, `camel-openai` or `camel-spring-ai-chat` appear when the integration runs with GenAI observability (`--observe`, or `--dep=camel:ai-observability`), tagged with the route id; Camel records tokens and duration for those, not the phase split.

The live figures, the context panel and the host panel need the model runner on the same machine: Ollama starts a `llama-server` process per loaded model and the tab reads its slot state a few times a second. Against a remote or containerised Ollama the tab keeps the model, per-request and session data and says which panels are unavailable.

Press **r** to reset the request log and the session totals, **F5** to refresh immediately. The same data is available to AI agents through the `tui_get_ollama` MCP tool. For what the figures mean for the AI panel and which knobs to turn, see [Working with a local Ollama model](#_working_with_a_local_ollama_model).