> ## Documentation Index
> Fetch the complete documentation index at: https://llmwiki.atomicstrata.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# llmwiki query - Ask Questions Against Your Compiled Wiki

> llmwiki query answers questions using hybrid semantic search and BM25 reranking over compiled wiki pages. Use --save to persist answers as new pages.

Once you've compiled a wiki, `llmwiki query` lets you ask natural language questions and get grounded, cited answers drawn from your compiled pages. Unlike sending a question directly to an LLM, `query` retrieves the most relevant content from your own wiki first and uses that as the grounding context - so answers cite specific pages and trace back to your original sources.

The query pipeline runs in two steps: first it identifies the most relevant pages or chunks for your question using hybrid retrieval, then it generates a grounded answer using only the content those pages contain. If the wiki doesn't have enough information to answer your question, the model says so.

## Basic usage

```bash theme={null}
llmwiki query "<question>"
```

Run from your project root after compiling. The answer streams to the terminal as it's generated.

```bash theme={null}
llmwiki query "What is multi-head attention and how does it differ from self-attention?"
```

## Flags

| Flag | Description |
| - | - |
| `--save` | Request publication in `wiki/queries/` after fresh citation validation. On success, refresh the index and embeddings (unless embedding refresh is disabled). |
| `--review` | With `--save`, stage a validated-answer review candidate instead of publishing. Requires `--save`; an invalid pair fails before provider checks or requests. |
| `--lang <code>` | Generate the answer in the specified language (e.g. `zh-CN`, `ja`, `Japanese`). Wins over `LLMWIKI_OUTPUT_LANG`. |
| `--debug` | Print a retrieval detail snapshot after the answer, showing which pages and chunks were selected, their scores, and whether BM25 reranking changed the order. |

## Answer citations

After the answer finishes streaming, `query` reports unique recognized body wikilinks in first-occurrence order:

```text theme={null}
Answer citations: 1 resolved, 1 pending, 1 broken
resolved: alpha -> concepts/alpha
pending: beta
broken: missing
```

* **Resolved** means the target matches a retained concept or query page. The report shows its qualified identity, such as `concepts/alpha`. Resolution checks exact concept names, exact query names, concept aliases, then query aliases.
* **Pending** means no retained page resolves the target, but an admitted review candidate has that exact eventual concept/query slug. It does not mean the page is published. Typed candidates and candidate aliases do not qualify.
* **Broken** means neither retained resolution nor an eligible pending candidate matches. Typed-only pages do not resolve through the current bare-wikilink contract.

Targets are normalized using the viewer's wikilink rules: `[[Alpha|label]]` reports `alpha`, and repeated targets appear once. Code, escaped markers, and wikilinks inside Markdown link text are excluded. Generated title and summary frontmatter are excluded; the original streamed answer stays unchanged.

If collection succeeds and the answer has no recognized links, the CLI prints `Answer citations: none recognized (not a factual-support assessment)`. Neither a resolved target nor an empty report verifies that the answer is factually supported. An empty answer retains the existing no-match message.

The report uses a fresh, unlocked snapshot collected after generation and before saving. Files may change afterward. The report adds no model request and is not permission to publish: requested saves receive a separate check under the project lock. Without `--save`, query leaves page and review-candidate bytes unchanged; existing query activity logging still runs.

Citation collection reads retained-page frontmatter incrementally, stopping at the closing fence for complete headers, including normal headers within 64 KiB. Small reads avoid loading retained body bytes; they are not a speed optimization and can require many filesystem calls. Late or missing closing fences have no fixed byte bound and may require reading the whole file. Review-candidate JSON is still read in full. Malformed candidates are skipped with the existing warnings.

If advisory citation collection fails, the CLI prints `Answer citations: unavailable` with the error diagnostic on stderr and omits `answerCitations` from the structured result. The generated answer remains available. A requested save retries validation under the lock: it may succeed if that fresh check succeeds, or refuse publication if validation remains unavailable. An unavailable report is never a successful empty report.

## Publication and review policy

| Operation | Resolved-only or no recognized links | Pending links | Broken links | Validation unavailable |
| - | - | - | - | - |
| Query without saving | Return answer and report | Return answer and report | Return answer and report | Return answer, omit report, warn |
| `query --save` | Publish | Refuse publication | Refuse publication | Refuse publication, preserve answer |
| `query --save --review` | Stage only | Stage only | Refuse staging | Refuse staging, preserve answer |
| Validated-answer approval | Publish | Refuse, retain candidate | Refuse, retain candidate | Refuse, retain candidate |
| Generic default concept/query approval | Existing approval policy | Existing approval policy | Refuse genuinely broken, non-repairable links | Warn and continue existing approval policy |

Publishing replaces `wiki/queries/<slug>.md`, so the three validated operations resolve links as they will be after that write: the page being replaced still resolves by its filename, but aliases declared only on the old page no longer count. An answer whose link resolved only through such an alias is refused as broken. These checks resolve recognized links; they do not verify factual support. No recognized links is accepted and is not evidence that the answer is correct. Typed candidates retain their existing typed policy. An imported query candidate without validated-answer metadata follows generic policy. Connector pins, trust planning, and profile restrictions still apply.

Citation refusals display the generated answer and a diagnostic, then exit nonzero. They do not write a query page, review candidate, index, or embeddings. Direct saves and no-save queries keep their existing activity logging; reviewed staging appends no `log.md` entry. Filesystem failures updating `log.md` warn and do not fail the query. Actual generation, page/candidate write, or other uncaught persistence failures still fail the operation; SDK calls reject rather than returning an answer-preserving citation refusal for those failures.

Direct saving requires the default profile: a non-default profile produces the existing warning and exit zero, with a structured `profile-disabled` refusal. Reviewed staging is allowed in profile-enabled projects because the review candidate is the trust-routed destination; approval routes the write through the planner under the profile in force at approval time, and a planner refusal keeps the candidate and writes nothing.

The authoritative check reloads current retained pages and pending candidates while holding the project lock. It uses the canonical answer body and does not reuse the advisory report or stored proposal observations. This serializes cooperating writers only: arbitrary external edits can race the check, and later changes can break published links. A currently resolving alias may point to different content than when the answer was generated; resolution does not verify semantic equivalence. Validation adds no provider request, although ordinary generation and enabled embedding refresh still use providers.

### Stage an answer and recover pending links

```bash theme={null}
llmwiki query "How does attention work?" --save --review
llmwiki review list
llmwiki review show <answer-candidate-id>
# Inspect and approve each required pending target first:
llmwiki review approve <target-candidate-id>
llmwiki review approve <answer-candidate-id>
```

Staging prints a candidate ID and leaves live pages, indexes, and embeddings unchanged. It also records a closed precondition on the target page: its current digest, or that it is absent. Approval requires all recognized links to resolve now and the precondition to still hold; if the page changed, appeared, or was deleted after staging, approval refuses with a restaging instruction and keeps both the page and the candidate. If approval refuses, the candidate stays available for a retry after its targets are published. If the body was edited, regenerate and restage the answer, then reject the old candidate by ID. See [validated-answer review and recovery](/cli/review#validated-answer-candidates) for digest limits, sibling proposals, and approval failure recovery.

## How retrieval works

`llmwiki query` uses a two-step hybrid retrieval strategy:

1. **Chunk-level semantic search.** The embedding store (`.llmwiki/embeddings.bin`, or `embeddings.json` for small stores) holds vectors for indexed wiki chunks. Your question is embedded and cosine similarity narrows the full corpus down to a top-K set of the most semantically relevant chunks.

2. **BM25 reranking.** The top-K chunks are reranked using BM25 (lexical scoring) to surface chunks where your question's exact terms appear prominently. This corrects for cases where semantic similarity alone would pick conceptually adjacent but not directly relevant content.

The final evidence pack is the full content of the selected pages, loaded and passed to the LLM as grounding context. When no embedding store is present, the command falls back to asking the LLM to pick pages from a list of the live pages' titles and summaries (pages excluded from search are never listed). It also falls back when the embedding call fails, for example with a provider that cannot embed, and prints an `embedding-degraded` warning after the answer.

Only pages whose content reached the LLM count as grounding. If a selected page can't be read, it is left out of the answer's grounding and the activity log, and a warning names it. Excerpts are sent only for pages in that grounding.

If you need wikilink-graph expansion - pulling in pages directly referenced by the top results - use [`llmwiki context`](#llmwiki-context-prompt---json) instead. `context` packages an extended evidence pack with graph neighbors and citation metadata, designed for agent pipelines that build their own prompts.

### Example session

```bash theme={null}
$ llmwiki query "What is the attention mechanism in transformers?"
* Selecting relevant pages…
  i Reasoning: Selected 3 page(s) from 8 reranked chunks: self-attention#2 (0.891), multi-head-attention#0 (0.847), transformer#1 (0.812)
  * Selected 3 page(s): self-attention, multi-head-attention, transformer
* Generating answer…

The attention mechanism allows a model to weigh the importance of different
positions in the input sequence when encoding each token. [[self-attention]]
describes the core operation: for each token, the model computes query, key,
and value vectors, then produces a weighted sum of the values where weights
are determined by the dot-product similarity between queries and keys.

[[multi-head-attention]] extends this by running multiple attention heads in
parallel, each learning different relationship patterns, and concatenating
the results.

→ Tip: use --save to add this answer to your wiki
```

## `llmwiki context "<prompt>" [--json]`

```bash theme={null}
llmwiki context "<prompt>"
llmwiki context "<prompt>" --json
```

`llmwiki context` and `llmwiki query` both perform hybrid retrieval, but they produce different outputs:

* **`llmwiki context`** packages the evidence - primary pages, semantic chunks, graph neighbors, citations, per-page freshness warnings, and suggested next actions - into a structured evidence pack without generating an answer. The pack is designed to be fed directly to an LLM agent or MCP tool as a token-budgeted context window.
* **`llmwiki query`** generates a grounded answer using that same retrieved evidence.

Use `llmwiki context` when you want to build your own prompt around the retrieved evidence, or when you're integrating wiki retrieval into an agent pipeline. Use `llmwiki query` when you want a ready-to-read answer.

`--json` emits the same `v1` JSON envelope as the MCP `get_context_pack` tool - stable format suitable for agent consumption.

### `--include-sources`

```bash theme={null}
llmwiki context "<prompt>" --include-sources
```

The `--include-sources` flag is opt-in and appends raw text windows from the ingested source files alongside the compiled wiki page content. Source windows are path-confined to the `sources/` directory - the flag cannot be used to read files outside that boundary. Only enable source windows for agents and pipelines you trust with the full text of your ingested sources.

## Compounding queries with `--save`

When `llmwiki query "<question>" --save` passes the profile and fresh citation checks, the answer is written to `wiki/queries/<slug>.md`, added to the index, and scheduled for the ordinary embedding refresh unless disabled. It then becomes context for future queries. A refusal leaves the answer displayed without publishing it:

```bash theme={null}
# First query: answer drawn from concept pages only
llmwiki query "What are the training techniques used for large language models?"

# Save it
llmwiki query "What are the training techniques used for large language models?" --save

# Future queries can now incorporate this answer as context
llmwiki query "How does RLHF relate to the training techniques we documented?"
```

Over time, a wiki that accumulates saved query answers becomes progressively richer - later questions can build on earlier ones, and the evidence pack grows to include synthesized knowledge alongside raw source-derived pages.

### Example with `--debug`

```bash theme={null}
$ llmwiki query "How does positional encoding work?" --debug

[answer streams here…]

--- Retrieval debug ---
  i Source: chunk-level; reranked: yes
  • positional-encoding (best chunk score 0.923)
  • transformer (best chunk score 0.741)
  · positional-encoding#1 score=0.923 :: Positional encodings are added to the input embeddings at the bottom of the encoder and decoder stacks…
  · positional-encoding#0 score=0.889 :: The model uses sine and cosine functions of different frequencies…
  · transformer#2 score=0.741 :: The transformer architecture relies on positional information…
```

<Note>
  `llmwiki query` detects both embedding formats automatically. An existing binary file takes precedence; if it is unavailable, query warns rather than using an older JSON snapshot. If no usable embedding store is present, query falls back to index-based selection. Compile your wiki first to build the embedding store. See [storage and recovery behavior](/configuration/environment-variables#embedding-storage).
</Note>

For using `query` and `context` via an AI agent, see [MCP Agent Integration](/guides/mcp-agent-integration).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.