Skip to main content
Compiling is the step that turns raw source files into a structured, interlinked wiki. When you run llmwiki compile, the pipeline runs in two phases: Phase 1 reads every changed source in sources/ and asks the LLM to extract all concepts, entities, and topics it finds. Phase 2 takes those extracted concepts and generates typed wiki pages - one Markdown file per concept, with YAML frontmatter, paragraph-level citations back to source line ranges, and [[wikilinks]] connecting related concepts. Splitting into two phases matters: by completing all extraction before any pages are written, the compiler can merge concepts that appear across multiple sources into a single page, catch extraction failures before anything is committed to disk, and resolve cross-references that would be impossible to wire up in a single pass.

Basic usage

Run from your project root. If sources/ is empty or doesn’t exist yet, compile will tell you to run llmwiki ingest first.

Flags

Nested source folders

By default, compile reads only top-level .md files in sources/. To compile your existing folder structure directly, add sources to .llmwiki/config.json:
Keep any existing review or connector settings in the same file. Then run llmwiki compile. No mirror directory or preprocessing step is needed. sources/records/notes.md keeps the identity records/notes.md in compilation state, citations and source previews. It is distinct from sources/notes.md. Use / in source IDs on every platform, including Windows. Exclusions are literal paths relative to sources/, not glob patterns. inbox excludes that path and its descendants, but not inbox-old. Leading ./ and trailing / are optional. Empty, absolute and parent-traversing paths are invalid; malformed settings stop the command rather than silently changing the selection. Dot-prefixed files and directories are not automatically ignored: exclude .trash explicitly if needed. Symlinks are not compiled or followed.
Excluding an already-compiled source removes its contribution on the next ordinary compile; it is not a temporary pause. The original source file stays untouched. Pages with no remaining contributors become orphaned, and shared pages are rebuilt from selected contributors. Selecting that source again compiles it as new and can incur model costs. Turning recursion off has the same effect for previously compiled nested files.
Run an ordinary llmwiki compile after changing selection. refresh --stale does not report unchanged, deselected owners as stale. llmwiki status --json lists their pending contribution retirement as deleted; that status does not mean the source file will be removed. The refresh command is scoped to stale owners and is not a substitute for applying all selection changes. Existing review and retry behavior still applies to generated pages. Source listings in the viewer, SDK and MCP use the same selection. An explicit re-ingest can still update an existing excluded source by its identity; with recursion enabled, moving an ingested file into a nested folder does not create a second copy on re-ingest. New ingests are still written at the top level.

Project instructions

Add editorial guidance without replacing the compiler’s built-in instructions:
Use the file for writing style, terminology, concept-selection preferences and level of detail. Instructions apply to extraction and page generation. They are advisory model guidance, not an enforcement mechanism; built-in citation, source-grounding and output requirements remain in the prompt. The path is relative to your working directory, or absolute. Explicit symlinks are allowed. Missing, unreadable, non-regular, invalid UTF-8 or oversized files fail before compilation. Empty or whitespace-only files mean no extra policy. No instruction file is discovered automatically: SOUL.md, AGENTS.md and CLAUDE.md are not read unless you explicitly select one. Neither project configuration nor schema files supply this policy. Pass the option on every compile that should use it. Changing its trimmed text or omitting the option regenerates affected pages and replaces obsolete review candidates; moving an identical file does not. The existing promptModifiers records the policy digest, not its text or path. The text is sent to your selected model provider, so do not include secrets. SDK callers can supply the same policy through compile({ systemPolicy }). watch, quickstart, refresh --stale and the MCP compile tool do not accept this option. They compile without the policy; using them after an instructed compile can regenerate pages without that guidance. Use the explicit compile command above when you need the policy preserved. After pages are written and interlinks resolved, compile repairs wikilinks whose target does not exist but unambiguously names a page that does. Generation tends to link a concept by its short canonical name ([[Argo CD]]) while the page carries a longer descriptive title (argo-cd-image-update-ownership-model); the link is repointed to it and the displayed text is left exactly as written. A link whose slug prefixes two pages, or none, is not touched. Guessing between two candidates would be wrong half the time, and a link to a concept your wiki genuinely lacks is worth keeping — llmwiki lint still reports it as broken-wikilink, which is how you find the pages your sources justify but extraction never created. Pending review candidates count toward ambiguity. A pending-only match stays unchanged until approval makes its target live; approval retries the repair pass. Repair preserves inline code, fenced examples, indented code, and frontmatter. It reads only confined wiki pages and warns when an escaping symlink or special file is dropped. The pass uses no model calls.

Incremental behaviour

Compile is hash-based and incremental. Each source file’s content is fingerprinted on ingest; on subsequent compiles, only sources whose content hash has changed are re-processed through the LLM. When sources and generation settings are unchanged, compile skips extraction and page generation: it makes no chat-model calls and does not regenerate those pages. Embedding reconciliation still runs as described below. When a changed or new source shares concepts with existing sources, compile also includes those contributors when rebuilding pages. For unchanged contributors, it can reuse successful extraction metadata stored with their live ownership in .llmwiki/state.json, avoiding another extraction call while still reading their current source text for page generation. Index growth alone does not reassign an unchanged source’s concepts. Only current or previous concepts of freshly extracted sources, plus explicit reconciliation targets, are regenerated. Reused contributors supply evidence to those pages without causing their other pages to regenerate. Their existing ownership records remain intact. A page shared by many sources still receives all contributing evidence within the configured prompt budget. Reuse requires matching source bytes, extraction prompt and tool schema, instructions, language, provider and model, and llmwiki’s prompt version, plus matching published concept ownership. Missing or invalid metadata falls back to extraction; older projects acquire it as sources compile. Deletion reconciliation, pending retries and --review use fresh extraction. Held or rejected concept assignments are not stored as reusable metadata. To force fresh extraction for a source, remove its entry from .llmwiki/state.json before compiling; deleting only its extraction field disables reuse the next time that source is scheduled, but does not itself schedule work. Embedding reuse depends on page content and the active embedding configuration, independently of source compilation state. Within a single run, both phases call the LLM in parallel: concept extraction across changed sources and page generation across merged concepts both run under one shared concurrency cap. The default is 5 simultaneous calls. On a cold start or a large refresh, raise it with --concurrency or LLMWIKI_COMPILE_CONCURRENCY to cut wall-clock; lower it if your provider rate-limits the run. The same flag is available on llmwiki refresh, llmwiki watch, and llmwiki quickstart.
Very high concurrency on a first compile can slightly reduce cross-reference quality. Page generation reads a handful of sibling pages for [[wikilink]] context, so with many pages generating at once, more are written before their neighbours exist. Page content is never corrupted — only the initial interlinking may be less complete. An ordinary re-compile does not backfill it, because unchanged pages are skipped; to get richer sibling-page context, run at a lower concurrency, or force a recompile of the affected sources (edit them, or remove their entries from .llmwiki/state.json). The default of 5 is a safe balance; reserve large values for big batch or CI runs where raw speed matters most.
Compile sends page and chunk embeddings to the provider in batches rather than one request at a time, which cuts compile time on cold starts and large refreshes. The embeddings summary reports how many requests this took, for example Embeddings: 12/12 pages, 84/84 chunks embedded (3 batched requests). Tune the request size with LLMWIKI_EMBED_BATCH_SIZE, or set LLMWIKI_EMBED_STRICT to fail the run when an embedding provider is broken or misconfigured.
The summary only prints when embeddings actually run. With no embedding backend configured (for example, the default anthropic provider without VOYAGE_API_KEY), compile prints Skipped embeddings update instead and semantic search is not built.

Embedding reconciliation

Every non-review compile with embeddings enabled scans eligible live pages and the embedding store, even when no sources changed and even if you have never disabled embeddings. It backfills missing vectors and refreshes stale vectors without regenerating source-derived pages. A healthy index with unchanged embedding configuration requires no embedding-provider calls or store writes. Changing the effective embedding model, provider, or endpoint can require a full re-embed of the eligible wiki on an otherwise unchanged compile. This sends page and chunk text to the embedding provider and can incur provider costs. Set LLMWIKI_EMBEDDINGS=off to skip embedding writes, including this scan. After removing the setting, run llmwiki compile to reconcile the index. Existing vectors remain available to retrieval while refreshes are disabled. Failed reconciliation uses the same five-attempt limit as changed-page refreshes. Attempts are bound to the page’s content: after five failures for the same content the page is quarantined, and compiles stop retrying it until its content changes. Content changed while refreshes were disabled is picked up by the next enabled compile. With LLMWIKI_EMBED_STRICT=1, an embedding-generation failure also fails an otherwise unchanged compile. See the embedding environment reference for failure handling and quarantine recovery.

What compile produces

Small embedding stores use JSON. When the store exceeds 64 MiB, compile saves the already-generated vectors to binary storage automatically. You can also opt in with LLMWIKI_BINARY_EMBEDDINGS=true. Later writes remain binary even without the flag. See storage limits and downgrade guidance. After a successful compile, your project contains:
Each page in wiki/concepts/ has YAML frontmatter with the page title, summary, kind (concept, entity, comparison, or overview), contributing source filenames, confidence, and provenance state. Paragraphs carry ^[source.md] citation markers; specific claims can pin to line ranges like ^[source.md:42-58].

Example compile output

llmwiki refresh --stale

If sources have changed since the last compile, affected pages are marked stale. Rather than running a full recompile, you can target only those pages:
This recompiles only the sources that own stale pages and cleans up pages whose sources were all deleted (orphaned pages). Unrelated new sources that haven’t been compiled yet are deliberately skipped - use llmwiki compile to bring those in separately.

--dry-run

Prints the plan - which pages would be recompiled, which sources are involved, how many orphaned pages would be cleaned up - without making any LLM calls or writing any files. Use this to verify the plan before committing.

llmwiki watch

Watches sources/ for file changes and automatically triggers a compile whenever a source is added or modified. Useful during active research sessions where you’re ingesting and reviewing pages in the same working session.

Review mode and the review policy

By default, compile writes pages directly to wiki/. Two mechanisms let you route pages through a review queue instead: --review flag (all-or-nothing): All generated pages are written to .llmwiki/candidates/ as JSON records. Nothing lands in wiki/ until you explicitly approve each candidate:
Review policy (selective): Add a .llmwiki/config.json to hold only pages that trip specific risk conditions - low confidence, contradictions, schema violations, or broken provenance - while writing the rest live. See Review Policy for the full configuration reference.
llmwiki compile sends source content to your configured LLM provider. Make sure your provider credentials are set before running compile. For Anthropic, set ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN. For other providers, see the Providers reference.
For production wikis, consider running llmwiki compile --review to inspect all generated pages before they go live. Once you’re confident in your source quality and prompt configuration, you can switch to a targeted review policy that holds only low-confidence or contradicted pages automatically.
For a deeper look at how the two-phase pipeline works, see How It Works.

Long sources and output limits

To allow longer completions for concept extraction and page generation:
The default is 4096 output tokens per call. Choose a limit supported by your model; higher limits can increase cost and latency. Invalid settings fail before an LLM request, and llmwiki does not automatically increase the budget or retry with a larger one. The claude-agent and codex-agent providers manage their own output limits and ignore this setting. This is not an input-context limit. LLMWIKI_PROMPT_BUDGET_CHARS limits combined source text during page generation, not the complete source passed to concept extraction. For textbooks, split sources into coherent chapters or sections and inspect the stored files under sources/ to check what ingestion retained. Extraction asks for 3–8 concepts per source, not an exhaustive textbook index; a larger output budget alone does not guarantee more concepts or valid pages. Changing the output budget does not invalidate successfully compiled sources. Failed sources remain retryable; unchanged successful sources are still skipped. To regenerate an existing source, edit its content before compiling again.

Prompt modifiers and recompilation

A prompt modifier is a setting that changes what the page prompt asks for without changing llmwiki’s own prompt wording. Today that is the output language, set by --lang or LLMWIKI_OUTPUT_LANG, and the Sources-section preference, set by --no-sources-section or LLMWIKI_SOURCES_SECTION=off. The Sources flag only disables the request. To restore the default, omit the flag and unset LLMWIKI_SOURCES_SECTION (or set it to on). Setting or clearing this preference regenerates affected pages, including already-compiled pages. It changes the instruction sent to the model; it does not strip a heading from the response or guarantee that a model will never include one. Compile is incremental: it normally skips a source whose bytes have not changed. A modifier is an input to the page prompt exactly as the source text is, so llmwiki records which modifiers each compile ran under and regenerates the affected pages when that selection changes:
Because those pages are regenerated by the model, this costs a full compile of every affected page. Clearing a modifier invalidates the same pages that setting it did. The instruction-file policy participates in the same change detection through its content digest; changing or omitting --instructions also invalidates pages. Each generated page also records the modifiers it was produced under in its promptModifiers frontmatter, which the JSON export surfaces per page.