llmwiki query lets you ask natural language questions and get grounded, cited answers drawn from your compiled pages. Unlike sending a question directly to an LLM, query retrieves the most relevant content from your own wiki first and uses that as the grounding context - so answers cite specific pages and trace back to your original sources.
The query pipeline runs in two steps: first it identifies the most relevant pages or chunks for your question using hybrid retrieval, then it generates a grounded answer using only the content those pages contain. If the wiki doesn’t have enough information to answer your question, the model says so.
Basic usage
Flags
Answer citations
After the answer finishes streaming,query reports unique recognized body wikilinks in first-occurrence order:
- Resolved means the target matches a retained concept or query page. The report shows its qualified identity, such as
concepts/alpha. Resolution checks exact concept names, exact query names, concept aliases, then query aliases. - Pending means no retained page resolves the target, but an admitted review candidate has that exact eventual concept/query slug. It does not mean the page is published. Typed candidates and candidate aliases do not qualify.
- Broken means neither retained resolution nor an eligible pending candidate matches. Typed-only pages do not resolve through the current bare-wikilink contract.
[[Alpha|label]] reports alpha, and repeated targets appear once. Code, escaped markers, and wikilinks inside Markdown link text are excluded. Generated title and summary frontmatter are excluded; the original streamed answer stays unchanged.
If collection succeeds and the answer has no recognized links, the CLI prints Answer citations: none recognized (not a factual-support assessment). Neither a resolved target nor an empty report verifies that the answer is factually supported. An empty answer retains the existing no-match message.
The report uses a fresh, unlocked snapshot collected after generation and before saving. Files may change afterward. The report adds no model request and is not permission to publish: requested saves receive a separate check under the project lock. Without --save, query leaves page and review-candidate bytes unchanged; existing query activity logging still runs.
Citation collection reads retained-page frontmatter incrementally, stopping at the closing fence for complete headers, including normal headers within 64 KiB. Small reads avoid loading retained body bytes; they are not a speed optimization and can require many filesystem calls. Late or missing closing fences have no fixed byte bound and may require reading the whole file. Review-candidate JSON is still read in full. Malformed candidates are skipped with the existing warnings.
If advisory citation collection fails, the CLI prints Answer citations: unavailable with the error diagnostic on stderr and omits answerCitations from the structured result. The generated answer remains available. A requested save retries validation under the lock: it may succeed if that fresh check succeeds, or refuse publication if validation remains unavailable. An unavailable report is never a successful empty report.
Publication and review policy
Publishing replaces
wiki/queries/<slug>.md, so the three validated operations resolve links as they will be after that write: the page being replaced still resolves by its filename, but aliases declared only on the old page no longer count. An answer whose link resolved only through such an alias is refused as broken. These checks resolve recognized links; they do not verify factual support. No recognized links is accepted and is not evidence that the answer is correct. Typed candidates retain their existing typed policy. An imported query candidate without validated-answer metadata follows generic policy. Connector pins, trust planning, and profile restrictions still apply.
Citation refusals display the generated answer and a diagnostic, then exit nonzero. They do not write a query page, review candidate, index, or embeddings. Direct saves and no-save queries keep their existing activity logging; reviewed staging appends no log.md entry. Filesystem failures updating log.md warn and do not fail the query. Actual generation, page/candidate write, or other uncaught persistence failures still fail the operation; SDK calls reject rather than returning an answer-preserving citation refusal for those failures.
Direct saving requires the default profile: a non-default profile produces the existing warning and exit zero, with a structured profile-disabled refusal. Reviewed staging is allowed in profile-enabled projects because the review candidate is the trust-routed destination; approval routes the write through the planner under the profile in force at approval time, and a planner refusal keeps the candidate and writes nothing.
The authoritative check reloads current retained pages and pending candidates while holding the project lock. It uses the canonical answer body and does not reuse the advisory report or stored proposal observations. This serializes cooperating writers only: arbitrary external edits can race the check, and later changes can break published links. A currently resolving alias may point to different content than when the answer was generated; resolution does not verify semantic equivalence. Validation adds no provider request, although ordinary generation and enabled embedding refresh still use providers.
Stage an answer and recover pending links
How retrieval works
llmwiki query uses a two-step hybrid retrieval strategy:
-
Chunk-level semantic search. The embedding store (
.llmwiki/embeddings.bin, orembeddings.jsonfor small stores) holds vectors for indexed wiki chunks. Your question is embedded and cosine similarity narrows the full corpus down to a top-K set of the most semantically relevant chunks. - BM25 reranking. The top-K chunks are reranked using BM25 (lexical scoring) to surface chunks where your question’s exact terms appear prominently. This corrects for cases where semantic similarity alone would pick conceptually adjacent but not directly relevant content.
embedding-degraded warning after the answer.
Only pages whose content reached the LLM count as grounding. If a selected page can’t be read, it is left out of the answer’s grounding and the activity log, and a warning names it. Excerpts are sent only for pages in that grounding.
If you need wikilink-graph expansion - pulling in pages directly referenced by the top results - use llmwiki context instead. context packages an extended evidence pack with graph neighbors and citation metadata, designed for agent pipelines that build their own prompts.
Example session
llmwiki context "<prompt>" [--json]
llmwiki context and llmwiki query both perform hybrid retrieval, but they produce different outputs:
llmwiki contextpackages the evidence - primary pages, semantic chunks, graph neighbors, citations, per-page freshness warnings, and suggested next actions - into a structured evidence pack without generating an answer. The pack is designed to be fed directly to an LLM agent or MCP tool as a token-budgeted context window.llmwiki querygenerates a grounded answer using that same retrieved evidence.
llmwiki context when you want to build your own prompt around the retrieved evidence, or when you’re integrating wiki retrieval into an agent pipeline. Use llmwiki query when you want a ready-to-read answer.
--json emits the same v1 JSON envelope as the MCP get_context_pack tool - stable format suitable for agent consumption.
--include-sources
--include-sources flag is opt-in and appends raw text windows from the ingested source files alongside the compiled wiki page content. Source windows are path-confined to the sources/ directory - the flag cannot be used to read files outside that boundary. Only enable source windows for agents and pipelines you trust with the full text of your ingested sources.
Compounding queries with --save
When llmwiki query "<question>" --save passes the profile and fresh citation checks, the answer is written to wiki/queries/<slug>.md, added to the index, and scheduled for the ordinary embedding refresh unless disabled. It then becomes context for future queries. A refusal leaves the answer displayed without publishing it:
Example with --debug
llmwiki query detects both embedding formats automatically. An existing binary file takes precedence; if it is unavailable, query warns rather than using an older JSON snapshot. If no usable embedding store is present, query falls back to index-based selection. Compile your wiki first to build the embedding store. See storage and recovery behavior.query and context via an AI agent, see MCP Agent Integration.