OAICA

A terminal AI CLI — a fork of Ollama's client — for running models you already have. Point it at a local Ollama or llama.cpp server, a self-hosted oaica serve, or any OpenAI-compatible endpoint. OAICA runs on your own machine: it hosts no models and needs no account.

Install

macOS / Linux

curl -fsSL https://oaica.com/install.sh | bash

Detects your architecture automatically.

Windows (PowerShell)

irm https://oaica.com/install.ps1 | iex

Installs to %LOCALAPPDATA%\Programs\OAICA and adds oaica to your session PATH. Set $env:OAICA_INSTALL_DIR first for a custom location.

Direct download

The installer is a convenience wrapper around these archives; they are hosted here on Cloudflare Pages and are byte-identical to the GitHub release.

PlatformArchive
Windows x64oaica-windows-amd64.zip
macOS Apple siliconoaica-darwin-arm64.zip
macOS Inteloaica-darwin-amd64.zip
Linux x64oaica-linux-amd64.tgz · .tar.zst
Linux arm64oaica-linux-arm64.tgz · .tar.zst

Checksums: SHA256SUMS · VERSION.txt. The installer verifies the archive against SHA256SUMS before extracting. Layout inside every archive is bin/oaica (bin/oaica.exe on Windows).

Upgrade

Same command as install — it overwrites the existing copy cleanly, no separate upgrade command needed:

curl -fsSL https://oaica.com/install.sh | bash          # macOS/Linux
irm https://oaica.com/install.ps1 | iex                 # Windows

Uninstall

curl -fsSL https://oaica.com/install.sh | OAICA_UNINSTALL=1 bash          # macOS/Linux
$env:OAICA_UNINSTALL=1; irm https://oaica.com/install.ps1 | iex           # Windows

Removes the binary and its lib directory. Doesn't touch any shell profile export you added yourself.

Point OAICA at a model

OAICA speaks the OpenAI-compatible wire to whatever endpoint you give it. By default it talks to a local server on 127.0.0.1:11434 — the Ollama/llama.cpp default — so the common case needs no configuration at all.

Local (default)

ollama serve                 # or: oaica serve <model>
oaica run llama3.2           # chat with a model your local server serves

Any model your local Ollama daemon answers for works, including pulled models and :cloud aliases. Self-hosted oaica serve models are addressed as name:local.

Any other endpoint

Set OAICA_HOST to an OpenAI-compatible base URL — a box on your network, a tunnel, or a hosted provider you already pay for. If that endpoint needs a key, set OAICA_API_KEY; alternatively put the key in the URL as userinfo (https://key@host), and OAICA promotes it to the Authorization header so it never rides along in URLs.

export OAICA_HOST=https://<your-endpoint>
export OAICA_API_KEY=<your-key>          # macOS/Linux
$env:OAICA_HOST="https://<your-endpoint>"  # Windows PowerShell
$env:OAICA_API_KEY="<your-key>"

Add those to your shell profile (~/.bashrc, ~/.zshrc) or Windows profile to persist them. Third-party cloud providers are also selectable per-launch by their own env key (e.g. OPENROUTER_API_KEY, ANTHROPIC_API_KEY); run oaica launch claude to pick from a searchable menu.

Run a model

oaica run <model>                          # interactive chat
oaica run <model> "explain this error: ..."  # one-shot
oaica launch claude --model <model>        # run Claude Code / opencode / codex on it
oaica launch claude                        # or pick from a searchable menu

Run oaica run with an unknown name to print the models your configured endpoint offers.

Plan with one model, execute with another. Claude Code's tiers (Opus for plan mode, Sonnet for execution) can point at different backends — a cloud model and a local one, two different remotes, anything:

oaica launch claude --model <cloud-model> -- --sonnet-model <local-model>
# inside Claude Code:  /model opusplan   → plans on the cloud model, executes on the local one

Any name works for either side: a remote from ~/.oaica/remotes.json, a running oaica serve (name:local), or anything your local Ollama daemon answers for. Plain launches send every main-conversation turn to the primary; the second model is used by /model opusplan, /model sonnet and subagents. Details: docs/CLAUDE_TIERS.md in the repo.

Failover, compaction headroom, and the wizard (v0.5.0+). A plain interactive oaica launch claude walks a wizard — when saved plans exist it first offers to reuse the plan last launched from this directory, then the tiers: primary model, Sonnet/execution tier, Haiku tier (Claude Code's background work: titles, topic detection), a compaction model (only models with a probed context window at least as large as the primary's are offered), and a route policy — then offers to save the whole setup as a named plan (Enter at the save prompt overwrites the last-used plan name, blank skips; oaica plan set ... --oversize ... --route-policy ... does the same with flags, and oaica config set sonnet-model|haiku-model <model> sets the two secondary tiers for every later launch, where a flag or a plan still wins). Route policies (local-first default, remote-first, auto, local-only, remote-only) drive what happens when the chosen backend fails: a health circuit breaker (3 consecutive failures → 90s open, 30s probes) fails the conversation over to the other leg without ever re-routing mid-stream, auto additionally escalates a failing session to its strongest healthy leg, and --oversize <model> catches the auto-compaction call when the prompt exceeds the primary's context. Every response carries X-Oaica-Route naming the leg that served it, and oaica doctor checks all legs read-only (exit 1 on any failure). Details: docs/CLAUDE_TIERS.md.

Inside the interactive session

CommandWhat it does
/model <name>Switch the active model for the rest of the session
/model listList models currently available
/lora add <name>Activate a LoRA adapter globally on its backend model
/lora remove <name>Deactivate it globally
/lora listList configured adapters
/lora use <name> [name2 ...]Use adapter(s) for this session only — per-request, doesn't touch other users
/lora stack <name>Add one more adapter to the current session's active set (compose multiple)
/lora offClear this session's per-request adapter(s)

Two different LoRA mechanisms — pick the right one

/lora add/remove is a global toggle: it flips the adapter on/off for every concurrent caller of that model, everywhere. Use this only when you actually want a shared, server-wide change (e.g. you administer the box).

/lora use/stack/off is per-request: the adapter choice rides along in your own request only, isolated from every other user hitting the same model at the same time. This is what almost everyone wants. Verified concurrently: 3 simultaneous sessions, one using a LoRA, two not — correct/independent output on all 3, zero cross-talk.

oaica run fitness_en
> /lora list
Configured LoRA adapters:
  fitness_en    (model: fitness_en, slot: 0)
  nutrition_en  (model: fitness_en, slot: 1)
> /lora use fitness_en
Using LoRA(s) [fitness_en] for this session only (per-request — doesn't affect other users)
> What should I eat before a morning workout?
...
> /lora stack nutrition_en
Stacked LoRA(s) [nutrition_en] for this session only (per-request — doesn't affect other users)
> What should I eat before a morning workout?
...(blended answer — both adapters' influence, not just the last one)
> /lora off
Per-request LoRA disabled for this session.

Stacked LoRA composes multiple adapters in one request (e.g. a fitness adapter + a nutrition adapter together) — real weighted composition, verified to produce a blended response distinct from either adapter alone. All stacked adapters must share the same backend model (they need to already be loaded together on one server via multiple --lora flags at launch — this is a private/local-deployment feature, not something the CLI provisions on the fly). Adapters can be toggled/scaled/stacked freely at runtime; loading a brand-new adapter file not already on the server requires a restart — that's a real llama.cpp limitation, not ours.

Build a website from a sentence

oaica site new ./clinic --prompt "Landing page for a family dental clinic in Johor Bahru: same-day appointments, kids' dentistry, prices in MYR"
oaica site edit ./clinic --prompt "hero should mention we open 7 days"
oaica site preview ./clinic                       # sandboxed local preview
oaica site deploy ./clinic --project my-clinic     # publish to Cloudflare Pages (needs wrangler)

Plans the page, writes each section separately, sanitizes everything (no scripts, no inline styles), assembles one index.html. Edits regenerate only the section you name. Static output — host it anywhere.

Calling your endpoint directly (curl, any language)

curl -X POST "$OAICA_HOST/v1/chat/completions" \
  -H "Authorization: Bearer $OAICA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"<model>","messages":[{"role":"user","content":"..."}],"max_tokens":256}'

Standard OpenAI-compatible shape, so any client that speaks it works against the same endpoint. GET /v1/models lists what the endpoint currently serves.

Known limitations