A local-first Electron IDE that pairs a Monaco editor with OpenAI-compatible chat, multi-step orchestration, and a live model terminal — tools run on your machine, not in someone else’s cloud sandbox.
LocalCode is a desktop LLM coding environment built with Electron and TypeScript. It puts an editor, streaming chat, workflow runner, and live terminal output in one workspace that talks to any OpenAI-compatible endpoint — Ollama, LM Studio, vLLM, or a remote API.
The pitch is simple: keep the model wherever you want, keep the files and shell on your box, and give the agent a real tool runtime with scopes, approvals, and optional local vector memory.
Repo: github.com/java-shell/LocalCode
What it’s for
| Surface | What you get |
|---|---|
| Editor | Monaco tabs, dirty-state tracking, save-all, disk change detection, workspace folder picker + tree |
| Chat | OpenAI-compatible /v1/chat/completions, streaming, model discovery, per-tool enablement |
| Workflows | Basic linear plans or an advanced stage graph with conditional links and model-chosen handoffs |
| Terminal | Live stdout/stderr for tool-driven commands, stdin, interrupt via (ctrl+c) |
It’s aimed at people who want Cursor-ish agent loops without shipping their whole tree to a hosted runner — or who just want a single app that can drive a local long-context model through a coding pipeline.
Core concepts
Electron shell
The app lives under llm-ide/:
| Path | Role |
|---|---|
src/main.ts |
Main process, IPC, window lifecycle, chat/orch memory glue |
src/preload.ts |
Context-isolated bridge into the renderer |
src/backend/ |
LLM client, orchestration engine, plan schema/storage, tool runtime, TurboVec, ntfy |
renderer/ |
Editor, chat, orchestration UI (including the advanced graph canvas), diagnostics, terminal |
Renderer runs with contextIsolation: true and nodeIntegration: false. System access goes through whitelisted preload methods.
Chat + tool loop
- You send a prompt; LocalCode builds system context (TurboVec recall,
.ai-rules, step prompt). - The model streams tokens and may request tools.
- Each tool runs locally in the selected scope; results stream back into the conversation.
- Rounds continue up to
maxToolTurns(default 40, hard cap 200). - Final answer lands in the chat panel; optional ntfy events fire when you’re done / stuck / errored.
flowchart LR
user([You]) --> chat[Chat / Orchestration UI]
chat --> ctx[Build system context]
ctx --> mem[TurboVec + GraphRAG recall]
ctx --> rules[.ai-rules]
ctx --> llm[OpenAI-compatible endpoint]
llm -->|tool calls| tools[Tool runtime]
tools -->|workspace / root scope| fs[Files · git · shell · web]
tools -->|query_user| youAsk[Dialog + toast]
tools --> llm
llm --> chat
Tool runtime
Tools are scoped to workspace (default — the folder you opened) or root (the LocalCode app tree). Paths resolve relative to that scope; anything that tries to escape gets an approval dialog.
| Category | Tools | Notes |
|---|---|---|
| Read | list_dir, file_search, read_file, grep_search, git_status, git_diff |
Inspect without mutating |
| Mutation | create_file, write_file, replace_in_file, insert_in_file, create_directory, move_path, delete_path |
Scoped; destructive ops need explicit enablement |
| Git + shell | git_add, git_commit, run_terminal_command |
Commits and commands show in the Model Terminal |
| Web | browser_fetch, web_search |
Pull docs / search when allowed |
| Memory | list_namespaces, query_memory |
Talk to the TurboVec store |
| Interaction | query_user |
Pause for text, yes/no, or checkbox answers |
query_user surfaces an info toast (with a little blip) and a dialog in chat so the model can ask before trampling something important.
TurboVec + GraphRAG memory
Memory is not a separate “GraphRAG database.” It is a TurboVec vector index (Python TurboQuantIndex via turbovec_bridge.py) plus a GraphRAG expansion pass over metadata links implemented in Node (turboVecMemory.ts).
On disk, each namespace is roughly:
| Artifact | Role |
|---|---|
memory.tvec |
Compressed vector index (bridge / TurboVec) |
memory.jsonl |
One JSON record per memory unit |
namespace.config.json |
Embedding dimension + schema version |
graphrag.stats.json |
Persisted expansion diagnostics |
Workspace namespaces live under Electron userData/turbovec-memory/; optional global namespaces under turbovec-global-memory/. Chat/orch default names look like chat:<workspaceDir> and orchestration:<workspaceDir>, with an optional global namespace merged in for multi-namespace retrieve.
Memory units (what actually gets stored)
A unit is a JSONL row: { id, text, source, createdAt, metadata }.
Typical writers today:
| Source label | When | Text budget (approx.) |
|---|---|---|
chat:user |
After a successful chat turn | 4k chars |
chat:assistant |
Same turn’s answer | 6k chars |
chat:compaction-summary / chat:compaction-tool-state |
When context-compaction v2 fired | up to ~7k / ~4.5k |
orchestration:<stepName> |
Step/final output when the memory adapter stores | 7k |
| Orchestration handoff capsule | Stored alongside step output (same batch) | 7k |
Every store() batch gets a shared metadata.turnId (turn_<counter>_<timestamp>) so the units written together become GraphRAG neighbors. Optional metadata keys relatedIds / neighborIds / adjacentIds define explicit edges when present (ingest tools / custom archives can set them; the main chat path usually does not).
Units are embedded through the LLM’s embeddings endpoint when available (with a hash fallback), deduped against recent neighbors, and appended with monotonic integer IDs.
GraphRAG retrieval (how units become context)
flowchart LR
q[Query text] --> emb[Embed query]
emb --> seed[TurboVec ANN seed hits]
seed --> bfs[hybridExpand BFS]
bfs --> links[Neighbor link types]
links --> topk[Rank + truncate to k]
topk --> inject[Inject into system prompt]
- Embed the query (namespace embedding dim from config / index).
- Bridge
queryreturns up tovectorKseed hits (default ≈2 × k, capped). hybridExpandwalks neighbors for up tomaxHops(default 1, hard cap 3).- Rank = seed similarity, decayed by link weight per hop; return top
k(default 6, chat/orch often request 8). - Injection is further capped by adaptive limits from the model context window (baseline ~8k → 4 items / 1k chars each / 3.5k total; scales up, capped).
Neighbor link types (with relative weights):
| Link | How it connects | Weight |
|---|---|---|
| turn | Same metadata.turnId (batch co-stored units) |
1.00 |
| source | Same source string (e.g. all chat:user) |
0.94 |
| explicit | IDs listed in relatedIds / neighborIds / adjacentIds |
0.88 |
| sequential | Positional id ± 1 in the namespace |
0.82 |
Multi-namespace retrieve weights each namespace’s seed budget by record count, min-max normalizes scores per namespace, then merges. Diagnostics expose per-namespace GraphRAG counters (retrievals, vectorSeedHits, expandedHits, linkHits.*).
Injection order (chat): base/system sections → recalled TurboVec block (plus optional compaction continuity capsules) → .ai-rules → user prompt. Orchestration stages retrieve with originalPrompt + stepInput as the query.
Export/import uses *.turbovec-namespace.json archives; namespaces can be locked against accidental writes.
Orchestration patterns
There are two execution modes in orchestrationEngine.ts, selected by the plan shape:
| Mode | Plan fields | UI |
|---|---|---|
| Basic | steps[] linear pipeline |
Orchestration tab |
| Advanced | advanced.enabled + advanced.stages[] + advanced.links[] |
Orchestration Advanced canvas (draggable stage cards) |
Built-in Import Preset plans (full_coding_pipeline, quick_fix, code_review in schema.ts) are basic linear pipelines. Advanced graphs are authored on the canvas (or imported JSON) — research templates like the DPC branching plan in markdown-reference/DPC_ADVANCED_ORCHESTRATION.md show the fuller pattern.
Basic: linear steps
Stages run in steps order. Each step has its own model, tools, retries, and passOutputToNext (default chain; only false opts out). Output handoffs are truncated (~12k) so the next step doesn’t drown.
Optional outer loop:
loopUntilComplete+maxLoops(1–20, default 3)- Primary stop: a stage with
allowCompletionTool: truecallscomplete_orchestration_task - Fallback stop: output still contains
completionMarker(defaultTASK_COMPLETE) for older plans
By default only the terminal successful output is stored to TurboVec (orchestration-final). Set storeIntermediateOutputs: true to persist every successful step (noisier namespaces).
Preset shapes:
- Full Coding Pipeline — Context Scan → Planner → Implementer → Validator → Reviewer
- Quick Fix — Analyze → Implement
- Code Review — Scan → Review → Suggestions
Advanced: stage graph
Advanced mode is a queue + fan-in walker, not a single linked list:
- Start at
entryStageId(or first stage). - Merge pending inputs for a stage (per-source budget inside a 12k fan-in cap).
- Run the stage (same tool loop as basic).
- Choose next hops:
- Model routing: if the stage has outbound links, it receives a
call_downstream_stagetool listing those direct neighbors. Calling it enqueues exactly that stage (supervisor / router pattern). - Link conditions: otherwise follow matching links.
- Model routing: if the stage has outbound links, it receives a
- Stop on
complete_orchestration_task, or when the queue drains /maxStageExecutionsis hit (default 100, cap 300). Per-stage runs are also capped (ceil(maxStageExecutions / stageCount)) so a feedback loop can’t monopolize the budget.
Link condition values:
| Condition | Fires when |
|---|---|
always |
Always (still filtered by matchPattern if set) |
success |
Stage succeeded |
error |
Stage failed |
error:<tag> |
Failed and detail text contains <tag> |
Optional matchPattern is a case-insensitive regex (or substring fallback) tested against the handoff text / error detail — used for router token branching (ROUTE_JAVA, PASS / FAIL, etc.).
Patterns you can actually build:
| Pattern | How |
|---|---|
| Linear pipeline | Basic steps, or advanced stages linked with success / always |
| Router → specialists | Entry stage emits route tokens; outbound links use success + matchPattern, and/or call_downstream_stage |
| Error recovery branch | error (or error:timeout) links into a recovery stage, then back toward QA |
| QA gate | Success links split on matchPattern (PASS → finisher, FAIL → recovery) |
| Outer loop-until-complete | Same loopUntilComplete / completion tool / marker as basic, wrapping whole graph iterations |
| Fan-in | Multiple upstream links into one stage; inputs concatenate with per-source truncation |
Note: plan schema validation currently requires each link’s target to appear later in advanced.stages than its source (array order). The runtime walker can follow backward edges if they are present, but validated/saved plans are expected to keep a topological stage order — place recovery/finisher stages accordingly when you want feedback-style graphs.
flowchart TB
entry[Entry / router stage] -->|success + matchPattern| implA[Specialist A]
entry -->|success + matchPattern| implB[Specialist B]
entry -->|call_downstream_stage| implA
implA -->|success| qa[QA stage]
implB -->|success| qa
implA -->|error| recovery[Recovery stage]
implB -->|error| recovery
recovery --> qa
qa -->|success + PASS| done[Finisher + complete_orchestration_task]
qa -->|success + FAIL| recovery
.ai-rules
Drop a .ai-rules/ folder of markdown at the workspace root. On every request LocalCode concatenates README.md first, then the rest alphabetically (capped around 8k characters) and injects it as system guidance. No folder? Silently skipped.
External notifications (ntfy)
Optional pushes from the main process (so they still fire when the window is minimized) via ntfy — no account, no API key, no new dependency.
| Event | When |
|---|---|
chat.completed |
Chat finished successfully |
chat.input_needed |
Model called query_user |
chat.error |
Chat or orchestration failed |
orchestration.completed |
Plan finished |
Model output stays off the wire unless you opt into excerpts. Notifications are not model-callable tools, so a prompt-injected agent can’t retarget your topic. Details: NTFY.md in the repo.
How to run it
Requirements: Node 18+ (20 recommended), npm 9+, a desktop OS.
cd llm-ide
npm install
npm run build
npm start
Dev loop:
cd llm-ide
npm run dev
Point chat at an OpenAI-compatible endpoint, e.g. Ollama:
| Field | Example |
|---|---|
| Endpoint | http://localhost:11434/v1/chat/completions |
| Model | llama3 |
| API key | Only if your provider needs one |
Model discovery hits GET /v1/models derived from that chat URL. Chat presets cover local Ollama (coding / 32K), a generic OpenAI-compatible template, and a read-only assistant.
Linux AppImage / Windows builds are available via npm run dist:linux and npm run dist:win in llm-ide.
Security notes
- Preload bridge is the only renderer → OS path.
- Filesystem tools enforce scope; outside-scope mutation asks for approval.
delete_pathand other destructive tools stay off until you enable them for that step.- ntfy topic is treated as a secret (redacted in logs) and never exposed as a tool.
Architecture at a glance
flowchart TB
subgraph ui [Renderer]
editor[Monaco editor]
chatUi[Chat panel]
orchUi[Orchestration + Advanced canvas]
term[Model Terminal]
end
subgraph main [Electron main]
ipc[IPC handlers]
llmClient[llmClient]
orch[orchestrationEngine]
tools[toolRuntime]
tv[TurboVecMemory]
bridge[turbovec_bridge.py]
notify[ntfy notify]
end
subgraph external [Outside the app]
endpoint[OpenAI-compatible LLM + embeddings]
workspace[Workspace files]
ntfyServer[ntfy server]
end
editor --> ipc
chatUi --> ipc
orchUi --> ipc
ipc --> llmClient
ipc --> orch
orch --> llmClient
llmClient --> endpoint
llmClient --> tools
orch --> tools
tools --> workspace
tools --> term
ipc --> tv
orch --> tv
tv --> bridge
tv --> endpoint
ipc --> notify
notify --> ntfyServer
Where to look in the code
| Area | Start here |
|---|---|
| LLM + tool turns | llm-ide/src/backend/llmClient.ts |
| Tool implementations / scopes | llm-ide/src/backend/toolRuntime.ts |
| Plan execution (basic + advanced) | llm-ide/src/backend/orchestrationEngine.ts |
| Plan schema + presets | llm-ide/src/backend/schema.ts |
| TurboVec + GraphRAG | llm-ide/src/backend/turboVecMemory.ts, turbovec_bridge.py |
| Chat/orch memory wiring | llm-ide/src/main.ts |
| Shared plan/stage types | llm-ide/src/shared/types.ts |
| Advanced graph UI | llm-ide/renderer/app.js (orchestration-advanced) |
| Example advanced SWE graph | markdown-reference/DPC_ADVANCED_ORCHESTRATION.md |
| ntfy | llm-ide/src/backend/notify.ts, NTFY.md |
Current limitations / honest caveats
- You’re bringing your own model endpoint — quality and tool-calling fidelity vary a lot by provider.
- GraphRAG’s richest edges (
turn, optionalrelatedIds) only help when metadata is populated; chat/orch stores assignturnIdper batch, but explicit cross-links are mostly for custom ingest. - Built-in presets are linear; branching/recovery graphs are an Advanced-canvas (or JSON) exercise.
- Schema’s “links must point later in
stages” rule constrains how you order feedback loops when validating plans. - TurboVec needs the Python bridge / venv; treat namespace export as your backup story.
- This is an actively evolving desktop app — pin a commit if you need a frozen toolchain.