Copilocal is a tool designed to streamline the process of launching GitHub Copilot CLI with your own local large language models (LLMs). It integrates with popular LLM providers such as Ollama, Foundry Local, LM Studio, and LiteLLM, allowing you to easily discover and select pre-installed models. Once selected, Copilocal ensures the provider's OpenAI-compatible server is running, sets up the necessary BYOK (Bring Your Own Key/Model) environment variables, and launches GitHub Copilot CLI with your chosen model.
Key Features:
Local Model Discovery: Automatically detects models installed in Ollama, Foundry Local, LM Studio, and LiteLLM.
Terminal-Based UI: Provides an intuitive arrow-key menu for selecting models.
Secure Environment Setup: Configures BYOK environment variables exclusively for the Copilot child process without affecting your shell.
Provider Installation Support: Offers to install missing providers (e.g., Ollama, Foundry Local) if they are not already installed.
Offline Mode: Enables air-gapped operation by setting COPILOT_OFFLINE=true to disable internet access and telemetry.
Audience & Benefit:
Ideal for developers and AI enthusiasts who want to use local LLMs with GitHub Copilot CLI. It empowers users to leverage self-hosted models, reducing reliance on third-party services while maintaining the functionality of GitHub Copilot. This tool is particularly useful for those seeking cost savings, enhanced privacy, or offline capabilities.
Copilocal can be installed via winget and is available for Windows x64/ARM64 and macOS arm64/x64 platforms. It operates independently of GitHub, Microsoft, or other providers, serving as a community-driven solution to enhance the versatility of GitHub Copilot CLI.
README
copilocal
Pick a local LLM from an arrow-key terminal menu and launch GitHub Copilot CLI against it.
copilocal discovers the models you already have in Ollama, Foundry Local,
LM Studio, and LiteLLM, lets you choose one in an arrow-key menu, makes sure that provider's
OpenAI-compatible server is running, then starts copilot with the right BYOK
environment variables — set only on the Copilot child process, never persisted to
your shell.
If a provider isn't installed, copilocal offers to install it (checkbox opt-in, with
links to each tool's docs so you can decide).
> ⚠️ Not affiliated with GitHub or Microsoft. copilocal is an independent,
> community-built tool. It is not affiliated with, endorsed by, or sponsored by
> GitHub, Microsoft, OpenAI, Ollama, or LM Studio. It simply launches the official
> GitHub Copilot CLI that you install separately. See
> Disclaimer & trademarks.
(startup animation: icon flies in left→right, reveals the wordmark, then settles on the right)
██████╗ ██████╗ ██████╗ ██╗██╗ ██████╗ ██████╗ █████╗ ██╗ ╭─════════════─╮
██╔════╝██╔═══██╗██╔══██╗██║██║ ██╔═══██╗██╔════╝██╔══██╗██║ │ ┄┄┄┄┄┄┄┄┄┄ ├─●
██║ ██║ ██║██████╔╝██║██║ ██║ ██║██║ ███████║██║ ●─┤ >_ │
██║ ██║ ██║██╔═══╝ ██║██║ ██║ ██║██║ ██╔══██║██║ │ ┄┄┄┄┄┄┄┄┄┄ ├─●
╚██████╗╚██████╔╝██║ ██║███████╗╚██████╔╝╚██████╗██║ ██║███████╗ ╰─════════════─╯
╚═════╝ ╚═════╝ ╚═╝ ╚═╝╚══════╝ ╚═════╝ ╚═════╝╚═╝ ╚═╝╚══════╝
Pick a local model · launch GitHub Copilot CLI against it
Discovering local models…
✓ Ollama — 2 models
✓ Foundry Local — 2 models
✓ LM Studio — 1 model
Select a local model (↑/↓, Enter to launch):
Ollama
> Ollama qwen2.5-coder:7b
Ollama llama3.2:3b
Foundry Local
Foundry Local qwen2.5-coder-7b-generic-gpu
Foundry Local phi-4-mini
LM Studio
LM Studio qwen3-0.6b
⚙ Configure launch options
✖ Quit
# 1. Install copilocal
winget install Gjlumsden.Copilocal
# 2. Prerequisites (install separately):
# - GitHub Copilot CLI on PATH: copilot --version
# - Optional now: local runtime + model (you can also install a runtime from inside copilocal on launch)
# e.g. Ollama:
ollama pull qwen2.5-coder:7b
# 3. Ollama only — give Copilot's prompt room (64k–128k if memory allows):
setx OLLAMA_CONTEXT_LENGTH 131072 # then restart Ollama
# 4. Launch
copilocal
No local runtime installed yet? Run copilocal and choose ⚙ Install / manage providers
to install local runtimes (Ollama / Foundry Local / LM Studio), or enable LiteLLM.
When setting up LiteLLM, copilocal can import all currently discovered local-provider
models into the LiteLLM config in one step.
Pick a model with ↑/↓ and Enter, then choose how to run it:
Launch GitHub Copilot CLI (current behavior)
Chat-only mode (local model chat loop; no tool execution)
Back to picker
For Copilot launch mode, copilocal starts the provider if needed, warms the model up,
sets BYOK env vars, and launches copilot. When Copilot exits you can pick a different
model to continue the same session, or Exit.
Why
GitHub Copilot CLI supports BYOK (bring-your-own-key/model) via environment variables,
but that's one model per session, set by hand. copilocal turns it into a picker over
all your local runtimes — no proxy, no config files, no shell pollution.
Requirements
GitHub Copilot CLI (copilot on PATH) — https://github.com/github/copilot-cli
(on Windows copilocal can install it for you at startup via winget; on macOS install it
yourself, e.g. with Homebrew)
At least one provider path: Ollama, Foundry Local, LM Studio, or LiteLLM
(copilocal can install/manage local runtimes and LiteLLM runtime flows from the UI)
Foundry Local note: copilocal expects the newer preview CLI surface (0.10+:
foundry cache list -o json, foundry server ...). If winget install Microsoft.FoundryLocal gives you an older 0.8.x service-based CLI, use copilocal's
installer flow or the cli-preview-* GitHub release instead.
Windows x64 / ARM64, or macOS arm64 / x64 — a single self-contained binary; no .NET runtime required
Validated with
copilocal talks to each runtime's CLI/REST surface, which shifts over time. The current
behaviour is verified against these versions (Windows 11 x64, June 2026):
Component
Version
Ollama
0.30.6
LM Studio
0.4.16
Foundry Local (CLI)
0.10.0
GitHub Copilot CLI
1.0.62
.NET SDK (build)
10.0.300
Newer releases usually work too; if discovery or launch misbehaves after a runtime update,
please open an issue noting the version.
Install
winget (Windows)
winget install Gjlumsden.Copilocal
Manual (Windows / macOS)
Download the binary for your platform from Releases
and put it on your PATH:
> On macOS, copilocal discovers and launches your local providers and Copilot CLI, but
> the one-click install flows (winget / Add-AppxPackage) are Windows-only — install
> Copilot CLI and the runtimes yourself there.
Usage
copilocal # interactive picker -> choose Copilot launch or chat-only mode
copilocal -- --resume # everything after -- is forwarded to copilot
copilocal --name "my feature" # name the managed session (resume it later by name)
copilocal --pick 1 # non-interactive: pick model #1
copilocal --pick 1 --offline # non-interactive: pick #1 and run air-gapped
copilocal --dry-run # show what it would set, don't launch
> Continue with a different model: in the interactive flow copilocal assigns a
> stable --session-id to the Copilot session it launches. When Copilot exits, it shows
> the captured resume id/name (and the copilot --resume= command to reopen it
> later), then drops you back to the model picker. Pick another local model to continue
> the same conversation with it, or choose Exit. This is skipped when you drive
> sessions yourself (e.g. --resume, --continue, --session-id).
> Air-gapped mode: after you pick a model, copilocal asks whether to run
> air-gapped (default No). Choosing yes sets COPILOT_OFFLINE=true, so Copilot CLI
> never contacts GitHub's servers and disables telemetry. Pass --offline to default
> the prompt to yes (or to enable it directly on the --pick path).
Chat-only mode
Choose Chat-only mode after selecting a model to stay inside copilocal and chat
directly with that provider endpoint (no GitHub Copilot CLI launch, no tool execution).
Commands inside chat:
/help show commands
/clear reset conversation history
/multi compose multiline input (finish with /send, cancel with /cancel)
/exit return to model picker
Ctrl+C also returns to model picker (same behavior as /exit). Note: while a model
response is in flight (the Thinking… spinner), Ctrl+C is honored only after the
request returns or the request timeout elapses — the HTTP call itself is not cancelled mid-flight.
slash command autocomplete: type / then Enter to pick a command, or use unique prefixes like /h
bottom-right token tracker showing active model, cumulative tokens, and last-turn token usage (when provider returns usage fields)
assistant replies wrap to terminal width and render common markdown (headings, emphasis, inline code, links) plus markdown tables as native terminal tables
interactive menus (picker/options/install) use the terminal's alternate screen so each is a
distinct, cleared page; chat mode and the GitHub Copilot CLI launch run on the normal buffer
so the terminal's native scrollback works and the launched CLI's own UI renders cleanly
How models are discovered
Provider
Discovery command / source
Endpoint
Ollama
ollama list
http://localhost:11434/v1
Foundry Local
foundry cache list -o json
http://127.0.0.1:/v1 (resolved at runtime)
LM Studio
lms ls --json
http://localhost:1234/v1
LiteLLM
GET /v1/models
configurable (default http://localhost:4000/v1)
copilocal sets these on the childcopilot process only:
Environment variable
Value / when set
COPILOT_PROVIDER_BASE_URL
the chosen provider's OpenAI base URL
COPILOT_PROVIDER_TYPE
openai
COPILOT_MODEL
the chosen model id
COPILOT_PROVIDER_API_KEY
local for local runtimes; LiteLLM uses your configured key/env var
COPILOT_PROVIDER_WIRE_API
responses when a reasoning model needs the Responses API
COPILOT_OFFLINE
true in air-gapped mode
COPILOT_PROVIDER_MAX_PROMPT_TOKENS
auto-derived from model context, or your launch-options override
COPILOT_PROVIDER_MAX_OUTPUT_TOKENS
auto-derived from model context, or your launch-options override
> Models must support tool calling and streaming to work well with Copilot CLI.
> Before launch copilocal runs preflight guards: it flags models that don't advertise
> tool calling, checks the model's context window is big enough for Copilot's large
> prompt, and—at launch—probes the chosen model to catch ones that emit tool calls as plain
> text instead of native tool_calls (e.g. Ollama's qwen2.5-coder), which silently breaks
> Copilot's agentic loop. Interactively you can override a warning; a non-interactive
> --pick is blocked so it fails fast with a clear reason.
Reasoning models & context windows
A few gotchas copilocal now handles for you:
Reasoning models (e.g. gpt-oss) reply in a reasoning field and leave content
empty on the chat/completions wire — Copilot then loops and Ollama returns
400 invalid message content type: . When the warm-up detects a reasoning model
and the endpoint exposes /v1/responses (Ollama, LM Studio do), copilocal switches it
to the OpenAI Responses wire API (COPILOT_PROVIDER_WIRE_API=responses).
Context too small for Copilot's prompt. Copilot's system prompt + tools run to 20k+
tokens, so a small window truncates it (blank replies / loops / 400). copilocal reads each
provider's effective context and warns below 16384 tokens:
Ollama auto-sizes context from available VRAM when OLLAMA_CONTEXT_LENGTH is unset
(for example 4k under 24 GiB VRAM, 32k at 24–48 GiB, 256k at 48+ GiB). For Copilot, set
an explicit large value such as 64000 or 131072 if memory allows, then restart Ollama.
Foundry Local bakes the context into each compiled variant. On the validated
qwen2.5-coder variants with Foundry CLI 0.10.0, the …-openvino-npu (NPU) build
reported 4224 tokens (overflows as input_ids size … exceeds max length) and couldn't
be widened, while GPU/CPU variants of the same model reported 32768. Treat those as
observed model/version values, not universal guarantees. Use a non-NPU variant, e.g.
foundry model download -generic-gpu (or -generic-cpu / -openvino-gpu), or
foundry model run --device GPU.
LM Studio loads each model at the context chosen in its load dialog, and the
default custom length (often 8192) is too small for Copilot. When you load the model,
set Context Length to the model maximum (turn off the custom limit / slide it to max)
so the loaded window fits Copilot's prompt. Otherwise you'll see
n_keep (…) >= n_ctx (8192) … load the model with a larger context length.
Recommended models
Models need tool calling (for Copilot's agentic loop) and ideally fit your VRAM.
Small, fast, non-reasoning coders are the safest start; reasoning models work too
(copilocal routes them via the Responses API automatically).
Use
Ollama
Foundry Local
LM Studio
Best small coder
qwen2.5-coder:7b
qwen2.5-coder-7b
qwen2.5-coder-7b-instruct
Lighter / faster
qwen2.5-coder:3b, llama3.2:3b
qwen2.5-coder-1.5b, phi-4-mini
llama-3.2-3b-instruct
Tiny (quick tests)
llama3.2:1b
qwen2.5-coder-0.5b
qwen3-0.6b
Recent / agentic
granite4:3b, qwen3:4b
phi-4-mini, qwen2.5-7b
granite-4.0-h-tiny, phi-4-mini-instruct
> Tags change fast — check the latest live: Ollama
> ollama.com/library?sort=newest,
> Foundry with foundry model list, and LM Studio's in-app Discover catalog.
> NPU note: Foundry Local's *-openvino-npu variants run on a supported Intel/Qualcomm
> NPU, freeing the GPU/CPU — but the validated qwen2.5-coder NPU variants on Foundry CLI
> 0.10.0 reported a 4224-token context, too small for Copilot CLI's prompt. The validated
> GPU/CPU variants of the same model reported 32768 tokens. Treat these as observed values,
> not guarantees; for copilocal, prefer a GPU/CPU variant (-generic-gpu / -generic-cpu /
> -openvino-gpu) when the NPU variant warns on context.
Example: tuning to your machine
There's no single "best" model — it depends on your VRAM and RAM. A model's
weights must fit in VRAM + RAM; what fits in VRAM alone runs fastest; and a
MoE model (high total params, few active per token) lets you punch above your
VRAM if you have plenty of RAM — because memory is set by total params while speed
is set by active params.
As an example only, on a Ryzen 9 5900X · 128 GB RAM · Radeon RX 6800 XT (16 GB VRAM):
Fits in VRAM (fastest): a 7–14B dense model at Q4 — e.g. qwen2.5-coder:14b
(~9 GB) as an everyday coder, with headroom for context.
Lean / snappy:qwen2.5-coder:7b, llama3.2:3b.
More capability via RAM: a low-active MoE such as qwen3-coder:30b (~3B active)
— its ~18 GB of weights spill into the ample 128 GB RAM while only ~3B params compute
per token, so it stays usable.
Context: 16 GB VRAM comfortably handles OLLAMA_CONTEXT_LENGTH=32768; 131072
is heavy (large KV cache).
Avoid: big dense 24B+ models — they overflow 16 GB VRAM and crawl.
Configure launch options
The menu's ⚙ Configure launch options item opens a page to set the flags copilocal
passes to copilot on every launch. Choices are saved to ~/.copilocal/config.json and
applied automatically. Toggle (multi-select), pick a reasoning effort, and add any
extra raw args. Available toggles include:
MCP/skills: disable built-in MCP (--disable-builtin-mcps), disable user MCP servers
(--disable-mcp-server= for each in ~/.copilot/mcp-config.json), disable skills
(--excluded-tools=skill)
Reasoning effort:--reasoning-effort = none / low / medium / high / xhigh / max
Plus a free-text extra args field for any other copilot flag. The page shows a
preview of the resulting launch command. Disabling MCP servers and skills shrinks
Copilot's prompt a lot — useful for local models with limited context.
Token budget
Unknown local models aren't in Copilot's catalog, so Copilot falls back to generic
COPILOT_PROVIDER_MAX_PROMPT_TOKENS / MAX_OUTPUT_TOKENS defaults that can overshoot the
model's real context and truncate to empty/garbled output. copilocal auto-derives them
from the model's actual context — Ollama via OLLAMA_CONTEXT_LENGTH, LM Studio via its
loaded context (/api/v1/models) — reserving room for the reply
(prompt ≈ context − output − buffer, output ≈ context/4, capped at 8192). Override
either with the Max prompt tokens / Max output tokens fields in Configure launch
options (blank = auto).
Installing providers
If a runtime is missing, copilocal shows it in the menu. Choosing Install / manage
providers opens a checkbox list (space to toggle) with a docs link for each:
Ollama — winget (Ollama.Ollama)
LM Studio — winget (ElementLabs.LMStudio)
Foundry Local — latest CLI MSIX from microsoft/Foundry-Local releases (matches your CPU architecture)
LiteLLM — choose runtime mode in UI:
setup mode can either install local LiteLLM runtime or skip install and configure an existing LiteLLM instance
if local install fails, copilocal shows the failure reason and can immediately switch to existing-instance setup
start failures now include the concrete reason (e.g., missing Docker or missing LiteLLM key)
start now waits until LiteLLM is actually reachable (/v1/models) before reporting success
if LiteLLM is enabled on a local endpoint and discovery can't resolve it at startup, copilocal attempts one automatic LiteLLM start, then re-discovers models
docker: scaffolds a compose stack (LiteLLM + DB + UI) and manages start/stop/status
python: installs litellm[proxy] (uv/pip path), uses a local SQLite-backed config, and manages start/stop/status
after a successful LiteLLM start, copilocal prints the login key + endpoint + clickable UI link so you can sign in and manage models/config
you can print the clickable UI link any time via Manage LiteLLM runtime → Show runtime status
key hint: keys are normalized to sk-... format; docker UI login uses admin + the same LiteLLM key
setup prompt: optionally add all discovered local-provider models into LiteLLM config
manage action: add any missing local-provider models later (no duplicate entries)
LiteLLM troubleshooting
UI says invalid credentials
For local docker mode, sign in with:
username: admin
password: your LiteLLM key (normalized to sk-...)
Credentials are written to ~/.copilocal/litellm/.env as UI_USERNAME / UI_PASSWORD.
Copilot launch gets 401 on LiteLLM model
Make sure the same LiteLLM key is configured in copilocal (Set endpoint + auth).
copilocal now uses that key for:
LiteLLM /v1/models discovery
launch auth (COPILOT_PROVIDER_API_KEY)
Changed key or auth settings?
Run Manage LiteLLM runtime → Stop LiteLLM runtime, then Start LiteLLM runtime to rewrite .env.
LiteLLM can’t call local routed models (500 connection error)
In docker mode, local provider routes must use container-reachable hostnames.
copilocal now rewrites api_base loopback routes (localhost / 127.0.0.1) to
host.docker.internal on LiteLLM start, and when syncing local models into config.
Need a full clean reset?
Use Manage LiteLLM runtime → Reset LiteLLM local setup. This:
runs docker compose down (if present)
removes ~/.copilocal/litellm/
resets LiteLLM settings in ~/.copilocal/config.json to defaults (LiteLLM disabled, default base URL/env var, docker mode)
Where settings are stored
Global app settings: ~/.copilocal/config.json
LiteLLM local runtime files (compose/config/.env/pid): ~/.copilocal/litellm/
For Foundry Local, that installer intentionally uses the compatible cli-preview-* release
instead of assuming the Microsoft.FoundryLocal winget package has the same CLI surface.
Build from source
git clone https://github.com/garylumsden/copilocal.git
cd copilocal
dotnet run # debug
dotnet publish -c Release -r win-x64 # single AOT exe -> bin/Release/net10.0/win-x64/publish/
dotnet publish -c Release -r win-arm64
Requires the .NET 10 SDK. Output is a self-contained, Native-AOT single executable.
> Native AOT on Windows needs the Visual Studio C++ toolchain (Desktop development
> with C++) for the linker; dotnet run/dotnet build work without it.
Run the test suite:
dotnet test tests/Copilocal.Tests/Copilocal.Tests.csproj
Project layout
Source is grouped by responsibility, one namespace per folder:
Folder / namespace
What's in it
Copilocal (root)
Program — entry point and interactive orchestration
Dependency direction stays one-way: Infrastructure ← Providers ← Launch ← Ui, with Cli
standalone and Program wiring everything together.
Contributing
Issues and PRs are welcome. Please keep changes small and focused, run
dotnet build -c Debug and dotnet test before opening a PR (both must be green), and
match the existing style (interface-backed IProcessRunner/IHttpGateway for testability,
hand-built JSON helpers so the build stays Native-AOT friendly). Report bugs and request
features on the issue tracker.
copilocal is an independent, unofficial project and is not affiliated with, endorsed
by, or sponsored by GitHub, Microsoft, OpenAI, Ollama, or LM Studio. It launches the
official GitHub Copilot CLI that you install and authenticate separately, and it
talks to local runtimes you install yourself.
"GitHub", "GitHub Copilot", "Microsoft", "Foundry Local", "Ollama", and "LM Studio" are
trademarks of their respective owners and are used here for identification only. Use of
GitHub Copilot remains subject to GitHub's terms; bringing your own model does not change
your obligations under those terms.