A local-first Electron IDE that pairs a Monaco editor with OpenAI-compatible chat, multi-step orchestration, and a live model terminal — tools run on your machine, not in someone else’s cloud sandbox.

LocalCode is a desktop LLM coding environment built with Electron and TypeScript. It puts an editor, streaming chat, workflow runner, and live terminal output in one workspace that talks to any OpenAI-compatible endpoint — Ollama, LM Studio, vLLM, or a remote API.

The pitch is simple: keep the model wherever you want, keep the files and shell on your box, and give the agent a real tool runtime with scopes, approvals, and optional local vector memory.

Repo: github.com/java-shell/LocalCode

What it’s for

Surface What you get
Editor Monaco tabs, dirty-state tracking, save-all, disk change detection, workspace folder picker + tree
Chat OpenAI-compatible /v1/chat/completions, streaming, model discovery, per-tool enablement
Workflows Basic linear plans or an advanced stage graph with conditional links and model-chosen handoffs
Terminal Live stdout/stderr for tool-driven commands, stdin, interrupt via (ctrl+c)

It’s aimed at people who want Cursor-ish agent loops without shipping their whole tree to a hosted runner — or who just want a single app that can drive a local long-context model through a coding pipeline.

Core concepts

Electron shell

The app lives under llm-ide/:

Path Role
src/main.ts Main process, IPC, window lifecycle, chat/orch memory glue
src/preload.ts Context-isolated bridge into the renderer
src/backend/ LLM client, orchestration engine, plan schema/storage, tool runtime, TurboVec, ntfy
renderer/ Editor, chat, orchestration UI (including the advanced graph canvas), diagnostics, terminal

Renderer runs with contextIsolation: true and nodeIntegration: false. System access goes through whitelisted preload methods.

Chat + tool loop

  1. You send a prompt; LocalCode builds system context (TurboVec recall, .ai-rules, step prompt).
  2. The model streams tokens and may request tools.
  3. Each tool runs locally in the selected scope; results stream back into the conversation.
  4. Rounds continue up to maxToolTurns (default 40, hard cap 200).
  5. Final answer lands in the chat panel; optional ntfy events fire when you’re done / stuck / errored.
flowchart LR
  user([You]) --> chat[Chat / Orchestration UI]
  chat --> ctx[Build system context]
  ctx --> mem[TurboVec + GraphRAG recall]
  ctx --> rules[.ai-rules]
  ctx --> llm[OpenAI-compatible endpoint]
  llm -->|tool calls| tools[Tool runtime]
  tools -->|workspace / root scope| fs[Files · git · shell · web]
  tools -->|query_user| youAsk[Dialog + toast]
  tools --> llm
  llm --> chat

Tool runtime

Tools are scoped to workspace (default — the folder you opened) or root (the LocalCode app tree). Paths resolve relative to that scope; anything that tries to escape gets an approval dialog.

Category Tools Notes
Read list_dir, file_search, read_file, grep_search, git_status, git_diff Inspect without mutating
Mutation create_file, write_file, replace_in_file, insert_in_file, create_directory, move_path, delete_path Scoped; destructive ops need explicit enablement
Git + shell git_add, git_commit, run_terminal_command Commits and commands show in the Model Terminal
Web browser_fetch, web_search Pull docs / search when allowed
Memory list_namespaces, query_memory Talk to the TurboVec store
Interaction query_user Pause for text, yes/no, or checkbox answers

query_user surfaces an info toast (with a little blip) and a dialog in chat so the model can ask before trampling something important.

TurboVec + GraphRAG memory

Memory is not a separate “GraphRAG database.” It is a TurboVec vector index (Python TurboQuantIndex via turbovec_bridge.py) plus a GraphRAG expansion pass over metadata links implemented in Node (turboVecMemory.ts).

On disk, each namespace is roughly:

Artifact Role
memory.tvec Compressed vector index (bridge / TurboVec)
memory.jsonl One JSON record per memory unit
namespace.config.json Embedding dimension + schema version
graphrag.stats.json Persisted expansion diagnostics

Workspace namespaces live under Electron userData/turbovec-memory/; optional global namespaces under turbovec-global-memory/. Chat/orch default names look like chat:<workspaceDir> and orchestration:<workspaceDir>, with an optional global namespace merged in for multi-namespace retrieve.

Memory units (what actually gets stored)

A unit is a JSONL row: { id, text, source, createdAt, metadata }.

Typical writers today:

Source label When Text budget (approx.)
chat:user After a successful chat turn 4k chars
chat:assistant Same turn’s answer 6k chars
chat:compaction-summary / chat:compaction-tool-state When context-compaction v2 fired up to ~7k / ~4.5k
orchestration:<stepName> Step/final output when the memory adapter stores 7k
Orchestration handoff capsule Stored alongside step output (same batch) 7k

Every store() batch gets a shared metadata.turnId (turn_<counter>_<timestamp>) so the units written together become GraphRAG neighbors. Optional metadata keys relatedIds / neighborIds / adjacentIds define explicit edges when present (ingest tools / custom archives can set them; the main chat path usually does not).

Units are embedded through the LLM’s embeddings endpoint when available (with a hash fallback), deduped against recent neighbors, and appended with monotonic integer IDs.

GraphRAG retrieval (how units become context)

flowchart LR
  q[Query text] --> emb[Embed query]
  emb --> seed[TurboVec ANN seed hits]
  seed --> bfs[hybridExpand BFS]
  bfs --> links[Neighbor link types]
  links --> topk[Rank + truncate to k]
  topk --> inject[Inject into system prompt]
  1. Embed the query (namespace embedding dim from config / index).
  2. Bridge query returns up to vectorK seed hits (default ≈ 2 × k, capped).
  3. hybridExpand walks neighbors for up to maxHops (default 1, hard cap 3).
  4. Rank = seed similarity, decayed by link weight per hop; return top k (default 6, chat/orch often request 8).
  5. Injection is further capped by adaptive limits from the model context window (baseline ~8k → 4 items / 1k chars each / 3.5k total; scales up, capped).

Neighbor link types (with relative weights):

Link How it connects Weight
turn Same metadata.turnId (batch co-stored units) 1.00
source Same source string (e.g. all chat:user) 0.94
explicit IDs listed in relatedIds / neighborIds / adjacentIds 0.88
sequential Positional id ± 1 in the namespace 0.82

Multi-namespace retrieve weights each namespace’s seed budget by record count, min-max normalizes scores per namespace, then merges. Diagnostics expose per-namespace GraphRAG counters (retrievals, vectorSeedHits, expandedHits, linkHits.*).

Injection order (chat): base/system sections → recalled TurboVec block (plus optional compaction continuity capsules) → .ai-rules → user prompt. Orchestration stages retrieve with originalPrompt + stepInput as the query.

Export/import uses *.turbovec-namespace.json archives; namespaces can be locked against accidental writes.

Orchestration patterns

There are two execution modes in orchestrationEngine.ts, selected by the plan shape:

Mode Plan fields UI
Basic steps[] linear pipeline Orchestration tab
Advanced advanced.enabled + advanced.stages[] + advanced.links[] Orchestration Advanced canvas (draggable stage cards)

Built-in Import Preset plans (full_coding_pipeline, quick_fix, code_review in schema.ts) are basic linear pipelines. Advanced graphs are authored on the canvas (or imported JSON) — research templates like the DPC branching plan in markdown-reference/DPC_ADVANCED_ORCHESTRATION.md show the fuller pattern.

Basic: linear steps

Stages run in steps order. Each step has its own model, tools, retries, and passOutputToNext (default chain; only false opts out). Output handoffs are truncated (~12k) so the next step doesn’t drown.

Optional outer loop:

  • loopUntilComplete + maxLoops (1–20, default 3)
  • Primary stop: a stage with allowCompletionTool: true calls complete_orchestration_task
  • Fallback stop: output still contains completionMarker (default TASK_COMPLETE) for older plans

By default only the terminal successful output is stored to TurboVec (orchestration-final). Set storeIntermediateOutputs: true to persist every successful step (noisier namespaces).

Preset shapes:

  1. Full Coding Pipeline — Context Scan → Planner → Implementer → Validator → Reviewer
  2. Quick Fix — Analyze → Implement
  3. Code Review — Scan → Review → Suggestions

Advanced: stage graph

Advanced mode is a queue + fan-in walker, not a single linked list:

  1. Start at entryStageId (or first stage).
  2. Merge pending inputs for a stage (per-source budget inside a 12k fan-in cap).
  3. Run the stage (same tool loop as basic).
  4. Choose next hops:
    • Model routing: if the stage has outbound links, it receives a call_downstream_stage tool listing those direct neighbors. Calling it enqueues exactly that stage (supervisor / router pattern).
    • Link conditions: otherwise follow matching links.
  5. Stop on complete_orchestration_task, or when the queue drains / maxStageExecutions is hit (default 100, cap 300). Per-stage runs are also capped (ceil(maxStageExecutions / stageCount)) so a feedback loop can’t monopolize the budget.

Link condition values:

Condition Fires when
always Always (still filtered by matchPattern if set)
success Stage succeeded
error Stage failed
error:<tag> Failed and detail text contains <tag>

Optional matchPattern is a case-insensitive regex (or substring fallback) tested against the handoff text / error detail — used for router token branching (ROUTE_JAVA, PASS / FAIL, etc.).

Patterns you can actually build:

Pattern How
Linear pipeline Basic steps, or advanced stages linked with success / always
Router → specialists Entry stage emits route tokens; outbound links use success + matchPattern, and/or call_downstream_stage
Error recovery branch error (or error:timeout) links into a recovery stage, then back toward QA
QA gate Success links split on matchPattern (PASS → finisher, FAIL → recovery)
Outer loop-until-complete Same loopUntilComplete / completion tool / marker as basic, wrapping whole graph iterations
Fan-in Multiple upstream links into one stage; inputs concatenate with per-source truncation

Note: plan schema validation currently requires each link’s target to appear later in advanced.stages than its source (array order). The runtime walker can follow backward edges if they are present, but validated/saved plans are expected to keep a topological stage order — place recovery/finisher stages accordingly when you want feedback-style graphs.

flowchart TB
  entry[Entry / router stage] -->|success + matchPattern| implA[Specialist A]
  entry -->|success + matchPattern| implB[Specialist B]
  entry -->|call_downstream_stage| implA
  implA -->|success| qa[QA stage]
  implB -->|success| qa
  implA -->|error| recovery[Recovery stage]
  implB -->|error| recovery
  recovery --> qa
  qa -->|success + PASS| done[Finisher + complete_orchestration_task]
  qa -->|success + FAIL| recovery

.ai-rules

Drop a .ai-rules/ folder of markdown at the workspace root. On every request LocalCode concatenates README.md first, then the rest alphabetically (capped around 8k characters) and injects it as system guidance. No folder? Silently skipped.

External notifications (ntfy)

Optional pushes from the main process (so they still fire when the window is minimized) via ntfy — no account, no API key, no new dependency.

Event When
chat.completed Chat finished successfully
chat.input_needed Model called query_user
chat.error Chat or orchestration failed
orchestration.completed Plan finished

Model output stays off the wire unless you opt into excerpts. Notifications are not model-callable tools, so a prompt-injected agent can’t retarget your topic. Details: NTFY.md in the repo.

How to run it

Requirements: Node 18+ (20 recommended), npm 9+, a desktop OS.

cd llm-ide
npm install
npm run build
npm start

Dev loop:

cd llm-ide
npm run dev

Point chat at an OpenAI-compatible endpoint, e.g. Ollama:

Field Example
Endpoint http://localhost:11434/v1/chat/completions
Model llama3
API key Only if your provider needs one

Model discovery hits GET /v1/models derived from that chat URL. Chat presets cover local Ollama (coding / 32K), a generic OpenAI-compatible template, and a read-only assistant.

Linux AppImage / Windows builds are available via npm run dist:linux and npm run dist:win in llm-ide.

Security notes

  • Preload bridge is the only renderer → OS path.
  • Filesystem tools enforce scope; outside-scope mutation asks for approval.
  • delete_path and other destructive tools stay off until you enable them for that step.
  • ntfy topic is treated as a secret (redacted in logs) and never exposed as a tool.

Architecture at a glance

flowchart TB
  subgraph ui [Renderer]
    editor[Monaco editor]
    chatUi[Chat panel]
    orchUi[Orchestration + Advanced canvas]
    term[Model Terminal]
  end

  subgraph main [Electron main]
    ipc[IPC handlers]
    llmClient[llmClient]
    orch[orchestrationEngine]
    tools[toolRuntime]
    tv[TurboVecMemory]
    bridge[turbovec_bridge.py]
    notify[ntfy notify]
  end

  subgraph external [Outside the app]
    endpoint[OpenAI-compatible LLM + embeddings]
    workspace[Workspace files]
    ntfyServer[ntfy server]
  end

  editor --> ipc
  chatUi --> ipc
  orchUi --> ipc
  ipc --> llmClient
  ipc --> orch
  orch --> llmClient
  llmClient --> endpoint
  llmClient --> tools
  orch --> tools
  tools --> workspace
  tools --> term
  ipc --> tv
  orch --> tv
  tv --> bridge
  tv --> endpoint
  ipc --> notify
  notify --> ntfyServer

Where to look in the code

Area Start here
LLM + tool turns llm-ide/src/backend/llmClient.ts
Tool implementations / scopes llm-ide/src/backend/toolRuntime.ts
Plan execution (basic + advanced) llm-ide/src/backend/orchestrationEngine.ts
Plan schema + presets llm-ide/src/backend/schema.ts
TurboVec + GraphRAG llm-ide/src/backend/turboVecMemory.ts, turbovec_bridge.py
Chat/orch memory wiring llm-ide/src/main.ts
Shared plan/stage types llm-ide/src/shared/types.ts
Advanced graph UI llm-ide/renderer/app.js (orchestration-advanced)
Example advanced SWE graph markdown-reference/DPC_ADVANCED_ORCHESTRATION.md
ntfy llm-ide/src/backend/notify.ts, NTFY.md

Current limitations / honest caveats

  • You’re bringing your own model endpoint — quality and tool-calling fidelity vary a lot by provider.
  • GraphRAG’s richest edges (turn, optional relatedIds) only help when metadata is populated; chat/orch stores assign turnId per batch, but explicit cross-links are mostly for custom ingest.
  • Built-in presets are linear; branching/recovery graphs are an Advanced-canvas (or JSON) exercise.
  • Schema’s “links must point later in stages” rule constrains how you order feedback loops when validating plans.
  • TurboVec needs the Python bridge / venv; treat namespace export as your backup story.
  • This is an actively evolving desktop app — pin a commit if you need a frozen toolchain.