Skip to main content
Once you’ve compiled a wiki, llmwiki query lets you ask natural language questions and get grounded, cited answers drawn from your compiled pages. Unlike sending a question directly to an LLM, query retrieves the most relevant content from your own wiki first and uses that as the grounding context - so answers cite specific pages and trace back to your original sources. The query pipeline runs in two steps: first it identifies the most relevant pages or chunks for your question using hybrid retrieval, then it generates a grounded answer using only the content those pages contain. If the wiki doesn’t have enough information to answer your question, the model says so.

Basic usage

Run from your project root after compiling. The answer streams to the terminal as it’s generated.

Flags

Answer citations

After the answer finishes streaming, query reports unique recognized body wikilinks in first-occurrence order:
  • Resolved means the target matches a retained concept or query page. The report shows its qualified identity, such as concepts/alpha. Resolution checks exact concept names, exact query names, concept aliases, then query aliases.
  • Pending means no retained page resolves the target, but an admitted review candidate has that exact eventual concept/query slug. It does not mean the page is published. Typed candidates and candidate aliases do not qualify.
  • Broken means neither retained resolution nor an eligible pending candidate matches. Typed-only pages do not resolve through the current bare-wikilink contract.
Targets are normalized using the viewer’s wikilink rules: [[Alpha|label]] reports alpha, and repeated targets appear once. Code, escaped markers, and wikilinks inside Markdown link text are excluded. Generated title and summary frontmatter are excluded; the original streamed answer stays unchanged. If collection succeeds and the answer has no recognized links, the CLI prints Answer citations: none recognized (not a factual-support assessment). Neither a resolved target nor an empty report verifies that the answer is factually supported. An empty answer retains the existing no-match message. The report uses a fresh, unlocked snapshot collected after generation and before saving. Files may change afterward. The report adds no model request and is not permission to publish: requested saves receive a separate check under the project lock. Without --save, query leaves page and review-candidate bytes unchanged; existing query activity logging still runs. Citation collection reads retained-page frontmatter incrementally, stopping at the closing fence for complete headers, including normal headers within 64 KiB. Small reads avoid loading retained body bytes; they are not a speed optimization and can require many filesystem calls. Late or missing closing fences have no fixed byte bound and may require reading the whole file. Review-candidate JSON is still read in full. Malformed candidates are skipped with the existing warnings. If advisory citation collection fails, the CLI prints Answer citations: unavailable with the error diagnostic on stderr and omits answerCitations from the structured result. The generated answer remains available. A requested save retries validation under the lock: it may succeed if that fresh check succeeds, or refuse publication if validation remains unavailable. An unavailable report is never a successful empty report.

Publication and review policy

Publishing replaces wiki/queries/<slug>.md, so the three validated operations resolve links as they will be after that write: the page being replaced still resolves by its filename, but aliases declared only on the old page no longer count. An answer whose link resolved only through such an alias is refused as broken. These checks resolve recognized links; they do not verify factual support. No recognized links is accepted and is not evidence that the answer is correct. Typed candidates retain their existing typed policy. An imported query candidate without validated-answer metadata follows generic policy. Connector pins, trust planning, and profile restrictions still apply. Citation refusals display the generated answer and a diagnostic, then exit nonzero. They do not write a query page, review candidate, index, or embeddings. Direct saves and no-save queries keep their existing activity logging; reviewed staging appends no log.md entry. Filesystem failures updating log.md warn and do not fail the query. Actual generation, page/candidate write, or other uncaught persistence failures still fail the operation; SDK calls reject rather than returning an answer-preserving citation refusal for those failures. Direct saving requires the default profile: a non-default profile produces the existing warning and exit zero, with a structured profile-disabled refusal. Reviewed staging is allowed in profile-enabled projects because the review candidate is the trust-routed destination; approval routes the write through the planner under the profile in force at approval time, and a planner refusal keeps the candidate and writes nothing. The authoritative check reloads current retained pages and pending candidates while holding the project lock. It uses the canonical answer body and does not reuse the advisory report or stored proposal observations. This serializes cooperating writers only: arbitrary external edits can race the check, and later changes can break published links. A currently resolving alias may point to different content than when the answer was generated; resolution does not verify semantic equivalence. Validation adds no provider request, although ordinary generation and enabled embedding refresh still use providers.
Staging prints a candidate ID and leaves live pages, indexes, and embeddings unchanged. It also records a closed precondition on the target page: its current digest, or that it is absent. Approval requires all recognized links to resolve now and the precondition to still hold; if the page changed, appeared, or was deleted after staging, approval refuses with a restaging instruction and keeps both the page and the candidate. If approval refuses, the candidate stays available for a retry after its targets are published. If the body was edited, regenerate and restage the answer, then reject the old candidate by ID. See validated-answer review and recovery for digest limits, sibling proposals, and approval failure recovery.

How retrieval works

llmwiki query uses a two-step hybrid retrieval strategy:
  1. Chunk-level semantic search. The embedding store (.llmwiki/embeddings.bin, or embeddings.json for small stores) holds vectors for indexed wiki chunks. Your question is embedded and cosine similarity narrows the full corpus down to a top-K set of the most semantically relevant chunks.
  2. BM25 reranking. The top-K chunks are reranked using BM25 (lexical scoring) to surface chunks where your question’s exact terms appear prominently. This corrects for cases where semantic similarity alone would pick conceptually adjacent but not directly relevant content.
The final evidence pack is the full content of the selected pages, loaded and passed to the LLM as grounding context. When no embedding store is present, the command falls back to asking the LLM to pick pages from a list of the live pages’ titles and summaries (pages excluded from search are never listed). It also falls back when the embedding call fails, for example with a provider that cannot embed, and prints an embedding-degraded warning after the answer. Only pages whose content reached the LLM count as grounding. If a selected page can’t be read, it is left out of the answer’s grounding and the activity log, and a warning names it. Excerpts are sent only for pages in that grounding. If you need wikilink-graph expansion - pulling in pages directly referenced by the top results - use llmwiki context instead. context packages an extended evidence pack with graph neighbors and citation metadata, designed for agent pipelines that build their own prompts.

Example session

llmwiki context "<prompt>" [--json]

llmwiki context and llmwiki query both perform hybrid retrieval, but they produce different outputs:
  • llmwiki context packages the evidence - primary pages, semantic chunks, graph neighbors, citations, per-page freshness warnings, and suggested next actions - into a structured evidence pack without generating an answer. The pack is designed to be fed directly to an LLM agent or MCP tool as a token-budgeted context window.
  • llmwiki query generates a grounded answer using that same retrieved evidence.
Use llmwiki context when you want to build your own prompt around the retrieved evidence, or when you’re integrating wiki retrieval into an agent pipeline. Use llmwiki query when you want a ready-to-read answer. --json emits the same v1 JSON envelope as the MCP get_context_pack tool - stable format suitable for agent consumption.

--include-sources

The --include-sources flag is opt-in and appends raw text windows from the ingested source files alongside the compiled wiki page content. Source windows are path-confined to the sources/ directory - the flag cannot be used to read files outside that boundary. Only enable source windows for agents and pipelines you trust with the full text of your ingested sources.

Compounding queries with --save

When llmwiki query "<question>" --save passes the profile and fresh citation checks, the answer is written to wiki/queries/<slug>.md, added to the index, and scheduled for the ordinary embedding refresh unless disabled. It then becomes context for future queries. A refusal leaves the answer displayed without publishing it:
Over time, a wiki that accumulates saved query answers becomes progressively richer - later questions can build on earlier ones, and the evidence pack grows to include synthesized knowledge alongside raw source-derived pages.

Example with --debug

llmwiki query detects both embedding formats automatically. An existing binary file takes precedence; if it is unavailable, query warns rather than using an older JSON snapshot. If no usable embedding store is present, query falls back to index-based selection. Compile your wiki first to build the embedding store. See storage and recovery behavior.
For using query and context via an AI agent, see MCP Agent Integration.