A terminal AI CLI — a fork of Ollama's client — for running models you already have. Point it at a local Ollama or llama.cpp server, a self-hosted oaica serve, or any OpenAI-compatible endpoint. OAICA runs on your own machine: it hosts no models and needs no account.
macOS / Linux
curl -fsSL https://oaica.com/install.sh | bash
Detects your architecture automatically.
Windows (PowerShell)
irm https://oaica.com/install.ps1 | iex
Installs to %LOCALAPPDATA%\Programs\OAICA and adds oaica to your session PATH. Set $env:OAICA_INSTALL_DIR first for a custom location.
The installer is a convenience wrapper around these archives; they are hosted here on Cloudflare Pages and are byte-identical to the GitHub release.
| Platform | Archive |
|---|---|
| Windows x64 | oaica-windows-amd64.zip |
| macOS Apple silicon | oaica-darwin-arm64.zip |
| macOS Intel | oaica-darwin-amd64.zip |
| Linux x64 | oaica-linux-amd64.tgz · .tar.zst |
| Linux arm64 | oaica-linux-arm64.tgz · .tar.zst |
Checksums: SHA256SUMS · VERSION.txt. The installer verifies the archive against SHA256SUMS before extracting. Layout inside every archive is bin/oaica (bin/oaica.exe on Windows).
Same command as install — it overwrites the existing copy cleanly, no separate upgrade command needed:
curl -fsSL https://oaica.com/install.sh | bash # macOS/Linux
irm https://oaica.com/install.ps1 | iex # Windows
curl -fsSL https://oaica.com/install.sh | OAICA_UNINSTALL=1 bash # macOS/Linux
$env:OAICA_UNINSTALL=1; irm https://oaica.com/install.ps1 | iex # Windows
Removes the binary and its lib directory. Doesn't touch any shell profile export you added yourself.
OAICA speaks the OpenAI-compatible wire to whatever endpoint you give it. By default it talks to a local server on 127.0.0.1:11434 — the Ollama/llama.cpp default — so the common case needs no configuration at all.
ollama serve # or: oaica serve <model>
oaica run llama3.2 # chat with a model your local server serves
Any model your local Ollama daemon answers for works, including pulled models and :cloud aliases. Self-hosted oaica serve models are addressed as name:local.
Set OAICA_HOST to an OpenAI-compatible base URL — a box on your network, a tunnel, or a hosted provider you already pay for. If that endpoint needs a key, set OAICA_API_KEY; alternatively put the key in the URL as userinfo (https://key@host), and OAICA promotes it to the Authorization header so it never rides along in URLs.
export OAICA_HOST=https://<your-endpoint>
export OAICA_API_KEY=<your-key> # macOS/Linux
$env:OAICA_HOST="https://<your-endpoint>" # Windows PowerShell
$env:OAICA_API_KEY="<your-key>"
Add those to your shell profile (~/.bashrc, ~/.zshrc) or Windows profile to persist them. Third-party cloud providers are also selectable per-launch by their own env key (e.g. OPENROUTER_API_KEY, ANTHROPIC_API_KEY); run oaica launch claude to pick from a searchable menu.
oaica run <model> # interactive chat
oaica run <model> "explain this error: ..." # one-shot
oaica launch claude --model <model> # run Claude Code / opencode / codex on it
oaica launch claude # or pick from a searchable menu
Run oaica run with an unknown name to print the models your configured endpoint offers.
Plan with one model, execute with another. Claude Code's tiers (Opus for plan mode, Sonnet for execution) can point at different backends — a cloud model and a local one, two different remotes, anything:
oaica launch claude --model <cloud-model> -- --sonnet-model <local-model>
# inside Claude Code: /model opusplan → plans on the cloud model, executes on the local one
Any name works for either side: a remote from ~/.oaica/remotes.json, a running oaica serve (name:local), or anything your local Ollama daemon answers for. Plain launches send every main-conversation turn to the primary; the second model is used by /model opusplan, /model sonnet and subagents. Details: docs/CLAUDE_TIERS.md in the repo.
Failover, compaction headroom, and the wizard (v0.5.0+). A plain interactive oaica launch claude walks a wizard — when saved plans exist it first offers to reuse the plan last launched from this directory, then the tiers: primary model, Sonnet/execution tier, Haiku tier (Claude Code's background work: titles, topic detection), a compaction model (only models with a probed context window at least as large as the primary's are offered), and a route policy — then offers to save the whole setup as a named plan (Enter at the save prompt overwrites the last-used plan name, blank skips; oaica plan set ... --oversize ... --route-policy ... does the same with flags, and oaica config set sonnet-model|haiku-model <model> sets the two secondary tiers for every later launch, where a flag or a plan still wins). Route policies (local-first default, remote-first, auto, local-only, remote-only) drive what happens when the chosen backend fails: a health circuit breaker (3 consecutive failures → 90s open, 30s probes) fails the conversation over to the other leg without ever re-routing mid-stream, auto additionally escalates a failing session to its strongest healthy leg, and --oversize <model> catches the auto-compaction call when the prompt exceeds the primary's context. Every response carries X-Oaica-Route naming the leg that served it, and oaica doctor checks all legs read-only (exit 1 on any failure). Details: docs/CLAUDE_TIERS.md.
| Command | What it does |
|---|---|
/model <name> | Switch the active model for the rest of the session |
/model list | List models currently available |
/lora add <name> | Activate a LoRA adapter globally on its backend model |
/lora remove <name> | Deactivate it globally |
/lora list | List configured adapters |
/lora use <name> [name2 ...] | Use adapter(s) for this session only — per-request, doesn't touch other users |
/lora stack <name> | Add one more adapter to the current session's active set (compose multiple) |
/lora off | Clear this session's per-request adapter(s) |
/lora add/remove is a global toggle: it flips the adapter on/off for every concurrent caller of that model, everywhere. Use this only when you actually want a shared, server-wide change (e.g. you administer the box).
/lora use/stack/off is per-request: the adapter choice rides along in your own request only, isolated from every other user hitting the same model at the same time. This is what almost everyone wants. Verified concurrently: 3 simultaneous sessions, one using a LoRA, two not — correct/independent output on all 3, zero cross-talk.
oaica run fitness_en
> /lora list
Configured LoRA adapters:
fitness_en (model: fitness_en, slot: 0)
nutrition_en (model: fitness_en, slot: 1)
> /lora use fitness_en
Using LoRA(s) [fitness_en] for this session only (per-request — doesn't affect other users)
> What should I eat before a morning workout?
...
> /lora stack nutrition_en
Stacked LoRA(s) [nutrition_en] for this session only (per-request — doesn't affect other users)
> What should I eat before a morning workout?
...(blended answer — both adapters' influence, not just the last one)
> /lora off
Per-request LoRA disabled for this session.
Stacked LoRA composes multiple adapters in one request (e.g. a fitness adapter + a nutrition adapter together) — real weighted composition, verified to produce a blended response distinct from either adapter alone. All stacked adapters must share the same backend model (they need to already be loaded together on one server via multiple --lora flags at launch — this is a private/local-deployment feature, not something the CLI provisions on the fly). Adapters can be toggled/scaled/stacked freely at runtime; loading a brand-new adapter file not already on the server requires a restart — that's a real llama.cpp limitation, not ours.
oaica site new ./clinic --prompt "Landing page for a family dental clinic in Johor Bahru: same-day appointments, kids' dentistry, prices in MYR"
oaica site edit ./clinic --prompt "hero should mention we open 7 days"
oaica site preview ./clinic # sandboxed local preview
oaica site deploy ./clinic --project my-clinic # publish to Cloudflare Pages (needs wrangler)
Plans the page, writes each section separately, sanitizes everything (no scripts, no inline styles), assembles one index.html. Edits regenerate only the section you name. Static output — host it anywhere.
curl -X POST "$OAICA_HOST/v1/chat/completions" \
-H "Authorization: Bearer $OAICA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"<model>","messages":[{"role":"user","content":"..."}],"max_tokens":256}'
Standard OpenAI-compatible shape, so any client that speaks it works against the same endpoint. GET /v1/models lists what the endpoint currently serves.
oaica doctor to check every configured leg read-only.