Local coding

Component map (what you need)

Local coding is five layers. Missing any required layer = “nothing works.” Optional RAG only if you need repo retrieval beyond the model context.

  1. 1

    1. IDE / Agent client

    VS Code / Cursor extensions (Continue, Cline), JetBrains plugins (Continue and similar), plus CLI agents like Aider. Point them at a local OpenAI-compatible endpoint.

    The host is the IDE or terminal—the plugin is not the model. Browse Agent / Toolchain boards for repos.

  2. 2 Optional

    2. RAG / repo index

    Most clients ship an index (Continue/Cline codebase, Aider repo map). Standalone layer: AnythingLLM, RAGFlow, or LlamaIndex; vector stores Chroma, Qdrant, or FAISS; run embeddings locally.

    Skip if short-context complete is enough; required for large monorepos with small local models. This is not the inference runtime.

  3. 3

    3. Local runtime (OpenAI-compatible server)

    Ollama, llama.cpp server, MLX, vLLM, or SGLang—exposes localhost chat/completions.

    This is the “stack” decision on this page. Community stars below measure popularity of this layer.

  4. 4

    4. Model weights on disk

    GGUF / MLX / safetensors from the Models allowlist—size locks your hardware band.

    API-only MoE coders never land here.

  5. 5

    5. Hardware (VRAM / unified memory)

    NVIDIA CUDA, Apple unified + Metal/MLX, or limited AMD/Intel paths—see Hardware.

    Weights fit ≠ context fit (KV tax).

Quick “what am I missing?”

  • Have an IDE but no localhost runtime → install Ollama / llama.cpp / MLX first.
  • Runtime up but empty models → download an allowlisted coder GGUF (Models page).
  • Model loads but agent is blind in a big repo → add RAG / repo index (optional layer).
  • OOM or tiny context → Hardware band is wrong, or KV budget is too high.

Decision tree

Desktop NVIDIA

Ollama to start → llama.cpp when you need exact quants and context → vLLM / SGLang only on a dedicated machine.

Apple Silicon

Prefer MLX for native unified memory. Use llama.cpp or Ollama when you need cross-platform GGUF workflows.

IDE wiring

Point a VS Code / Cursor extension (Continue, Cline), a JetBrains plugin (Continue and similar), or Aider at localhost. Browse Agent and Toolchain boards for open repos—this page does not invent a second leaderboard.

Agent · Toolchain

Community signal — runtime popularity (GitHub OSS)

Stars from Trends segments inference / serving / llm_lib. This is popularity of layer 3 (local runtime)—not SWE skill, not IDE quality, not model quality.

Use it to sanity-check “is this runtime still alive?”, not to pick a coding model.

Repository Trends segment Stars
ggml-org/llama.cpp inference 129,951
earendil-works/pi inference 110,617
BerriAI/litellm inference 59,917
sgl-project/sglang inference 36,654
microsoft/onnxruntime inference 21,961
huggingface/transformers llm_lib 166,848
fighting41love/funNLP llm_lib 83,583
unslothai/unsloth llm_lib 77,072
ComposioHQ/awesome-claude-skills llm_lib 75,939
asgeirtj/system_prompts_leaks llm_lib 68,644
open-webui/open-webui serving 153,613
Mintplex-Labs/anything-llm serving 66,618
janhq/jan serving 44,717
lm-sys/FastChat serving 39,556
chatchat-space/Langchain-Chatchat serving 38,664

← Local coding hub

Assumptions and sources

Hardware bands are classes with USD reference ranges (not street quotes). VRAM estimates use curated GGUF footprints plus size-class KV factors. Coding scores are hand-curated SWE/Aider/LiveCode/HumanEval references—not harness re-runs. API job cost uses OpenRouter primary listing when present.