Skip to main content
llmwiki reads configuration from environment variables and an optional .env file placed in your project directory (alongside sources/ and wiki/). Variables set in your shell take precedence over those in .env, which in turn take precedence over the Claude Code settings fallback (~/.claude/settings.json → env block) and built-in defaults. You don’t need to set everything - only the variables relevant to your chosen provider are required. Most projects need only two or three exports before running llmwiki compile.

Provider selection

LLMWIKI_EMBEDDINGS controls embedding writes, not retrieval. While it is disabled, write paths do not read or change the embedding store (binary index or legacy .llmwiki/embeddings.json), .llmwiki/pending-embeddings.json, or .llmwiki/quarantined-embeddings.json. An existing pending marker may therefore remain visible in lint and status. Query and search commands can still read an existing embedding store, and they keep their normal embedding-provider validation. After you re-enable refreshes, run llmwiki compile. The compile reconciles missing and stale vectors even when no source changed. It does not regenerate source-derived pages in that case, and a healthy store is not rewritten. When you set LLMWIKI_EMBEDDING_PROVIDER, the named provider’s own credential is required: VOYAGE_API_KEY for anthropic or claude-agent, OPENAI_API_KEY (or OPENAI_EMBEDDINGS_API_KEY) for openai, ORCAROUTER_API_KEY for orcarouter, and no key for ollama. Setting OPENAI_EMBEDDINGS_BASE_URL marks a self-hosted OpenAI-compatible endpoint that needs no key, so OPENAI_API_KEY is not required in that case. This check only runs when LLMWIKI_EMBEDDING_PROVIDER is set - leaving it unset keeps today’s behavior unchanged. The check runs at startup, before any compile or query work begins, so a misspelled provider name or a missing key fails immediately instead of part-way through a run. Compile skips embedding-only checks when refreshes are disabled, but still validates the provider used to generate pages. codex-agent is the exception to the default embedding-provider value. Because Codex cannot embed, commands that consume or produce embeddings require LLMWIKI_EMBEDDING_PROVIDER to be set explicitly. A missing or unusable embedding backend fails actionably rather than silently degrading. Compile does not require one when LLMWIKI_EMBEDDINGS disables refreshes. The keyless combination with embeddings enabled is codex-agent chat plus ollama embeddings. When codex-agent is selected explicitly through --provider or a shell-exported LLMWIKI_PROVIDER, llmwiki does not open the project .env file, because that file may contain API keys. Export the Codex model, embedding provider, and embedding-backend settings in the shell for that run. Other providers retain the existing shell-over-.env precedence.
If you point OPENAI_EMBEDDINGS_BASE_URL at a host you do not operate, set OPENAI_EMBEDDINGS_API_KEY as well. Without it the embeddings client reuses OPENAI_API_KEY, which sends your cloud OpenAI key to that host. llmwiki warns when this happens, except for endpoints on localhost.

Changing the embedding backend rebuilds the index

The embedding index records which provider, model, and endpoint produced its vectors. Changing the effective LLMWIKI_EMBEDDING_PROVIDER, LLMWIKI_EMBEDDING_MODEL, or embedding endpoint invalidates the index. The next non-review llmwiki compile with refreshes enabled can re-embed the entire eligible wiki, even when no source changed. This incurs embedding-provider requests and potential costs without regenerating source-derived pages. Quarantined pages remain excluded until you explicitly re-queue them; see retry recovery. This matters because vectors from different backends are not comparable even when the model name matches - a local server answering to text-embedding-3-small does not produce the same vectors as cloud OpenAI, and nomic-embed-text served by Ollama does not match the same model served over an OpenAI-compatible endpoint. Until the index is rebuilt, llmwiki query reports the index as outdated and falls back to lexical ranking rather than ranking against a mixed index. Moving between anthropic and claude-agent does not rebuild the index. Both send embeddings to Voyage with the same model, so their vectors are interchangeable. An index built before llmwiki recorded the endpoint carries only its model name, which cannot distinguish those backends. Such an index is kept as-is while you run without LLMWIKI_EMBEDDING_PROVIDER or an endpoint override, so upgrading does not re-embed your project. With either override active, the next llmwiki compile rebuilds it once and records the full configuration from then on.

Anthropic

Either ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN satisfies authentication - you do not need both. If neither is set in your environment or .env, llmwiki attempts to read these values from ~/.claude/settings.json.

OpenAI

OpenAI-compatible request options

These settings apply to chat requests built by the shared OpenAI-compatible client, including OpenAI, Copilot, MiniMax, Atlas Cloud, and Ollama’s completion and streaming paths. They do not affect embeddings, Anthropic, the local agent providers, or Ollama’s native structured-output requests. Invalid values fail on the first request attempt, before network access, without retry backoff. Valid values are checked against the installed SDK vocabulary; individual models can support a smaller subset. In particular, older GPT-5 models do not support none.
Reasoning models reject max_tokens outright — the request fails with Unsupported parameter: 'max_tokens' is not supported with this model. llmwiki picks the right field from the model id, so LLMWIKI_OPENAI_TOKEN_PARAM is only needed when the id itself does not reveal the model family, which is common behind OpenAI-compatible gateways.

OpenAI Codex CLI

Authentication is managed only by the installed Codex CLI (codex login). llmwiki does not read Codex auth files or any OpenAI API key for chat. The CLI does not expose a maxTokens control, so llmwiki cannot honor that internal hint for this provider. Each request has a fixed 10-minute timeout and 1 MiB process/final-output caps. The minimum verified compatible Codex CLI version is 0.152.1. If an older binary rejects a required secure flag, llmwiki reports it as incompatible; update with npm install -g @openai/codex@latest.

Ollama


GitHub Copilot


Atlas Cloud

Model names are namespaced by publisher (e.g. qwen/qwen3.5-35b-a3b, the default). An override set through LLMWIKI_MODEL must be a model Atlas Cloud lists as supporting tools, because compile extracts concepts through a tool call.

OrcaRouter


Voyage (embeddings)


Embeddings

llmwiki compile embeds pages and chunks in provider-native batches. These variables are optional.
Batching controls requests independently of storage. Small stores keep .llmwiki/embeddings.json; stores exceeding its 64 MiB limit automatically use .llmwiki/embeddings.bin without repeating embedding requests.
Enabled non-review compiles scan eligible pages and the store for missing or stale embeddings even when their source delta is empty. This also applies when you have never used LLMWIKI_EMBEDDINGS. An unchanged, healthy store makes no embedding-provider calls and is not rewritten. See compile reconciliation. LLMWIKI_EMBED_STRICT applies to automatically discovered work too: if embedding generation fails, an otherwise unchanged compile exits non-zero after recording retry state. Without strict mode, it warns and continues. Disabling refreshes with LLMWIKI_EMBEDDINGS skips generation entirely, including in strict mode.

Embedding retries and quarantine

Explicit changes and automatically discovered missing or stale vectors share a durable retry budget in .llmwiki/pending-embeddings.json. Each budget is bound to the page’s embeddable content (its title, summary and body chunks): an attempt is charged against the content actually sent to the embedding provider, recorded before the request is made. Only real failures are charged:
  • Pages in the request that failed are charged one attempt.
  • Pages sent earlier in the same run that were not at fault are not charged.
  • Pages the run never reached are not charged.
  • A run that fails before any provider request charges nobody, but is still reported.
  • If every request succeeds but the results cannot be saved, the pages sent are charged, so paid work that can never be saved stops.
After five failed attempts for the same content, llmwiki warns and moves the page into .llmwiki/quarantined-embeddings.json. Later refreshes skip it while its content is unchanged, allowing other pages to embed. A quarantined page is retried automatically, with a fresh budget, as soon as its content changes. That includes changes made while refreshes were disabled, or outside llmwiki. Re-saving identical content does not reset the budget. These files use bounded, root-confined reads and atomic writes, and llmwiki re-reads them after writing. Neither file is read or changed while refreshes are disabled. Each retry file is limited to 5,000 entries and 512 KiB, which fits 5,000 typical page ids with their content hashes. A refresh only attempts pages whose budgets fit in the pending file; additional pages wait for a later compile. Existing entries are never evicted to make room for new work. If the quarantine file is full, exhausted entries stay in the pending file but are not retried. If both files fill with exhausted entries, new embedding work pauses until you recover those entries. The limits do not cause retries to restart. Deferred work is reported with a page count. No provider request is made for a page whose budget could not be recorded. Strict embedding mode reports this as a failure after settling any admitted work, so successful refreshes are not charged another retry. Fix marker storage errors or free capacity, then compile again. An unchanged compile can continue the deferred work when space is available. For a valid v3 store using the same embedding backend, deferred or quarantined pages keep their previous page and passage vectors, with the original hashes and timestamps. These cached vectors may be stale until refresh succeeds. Deleted pages, pages no longer eligible for embeddings, invalid vectors, and vectors from another backend are not retained. Legacy-store upgrades do not preserve stale vectors for deferred pages; those pages wait for a successful refresh. Entries from earlier versions. Quarantined entries written before retries were bound to content are re-queued once by the first compile with refreshes enabled. Each such page is bound to its current content and gets up to five additional retry rounds; one round can issue more than one provider request. llmwiki prints a one-time notice with the count, and llmwiki status reports these entries until then. After that, they follow the rules above. Resetting after fixing a provider. Unchanged content stays quarantined even after the provider is fixed, so reset its retry entries explicitly:
  1. Make sure no compile or refresh is running.
  2. Remove the page’s entries from both .llmwiki/quarantined-embeddings.json and .llmwiki/pending-embeddings.json. To reset every page, remove both files.
  3. Run llmwiki compile with refreshes enabled.
Removing both files also releases exhausted entries retained in pending at capacity. Resetting restarts retry budgets and may incur costs. Changing only embedding configuration does not clear quarantine.

Embedding storage

Binary storage keeps metadata and Float32 vectors in one atomically replaced file. It retains chunk text and the existing logical index version. Float32 introduces small rounding differences; it does not change the embedding model. The limits are 256 MiB of metadata, 512 MiB for the complete binary file, and 100,000 combined page and chunk records. Higher-dimensional vectors and longer chunk text use more space, so the record limit is not a promise that every such corpus fits. The index is still loaded into memory for retrieval. Once embeddings.bin exists, readers and writers use it without the flag. Any older JSON file is left untouched as a historical backup, not a fallback. An unreadable or corrupt binary index produces embedding-store-unavailable and degraded retrieval rather than silently reading that older snapshot. The context command uses its existing embedding-store-missing warning code for this condition, meaning no usable store, not necessarily no file on disk. An enabled non-review llmwiki compile discovers the missing vectors even when no sources changed. Rebuilding an unavailable index embeds the live eligible corpus again, except quarantined pages, and can incur provider costs. Unsetting the flag does not convert binary back to JSON. Before downgrading to a release without binary support, back up and move both derived index files aside, then rebuild with that release. This only works for a corpus that fits the older JSON limit; larger corpora require a binary-capable release. Do not delete just the binary file and reuse stale JSON. Keep your source files, wiki pages and state.json intact.

Timeouts

Timeout resolution for Ollama: explicit constructor option → OLLAMA_TIMEOUT_MS → LLMWIKI_REQUEST_TIMEOUT_MS → built-in 30-minute default. Non-numeric, zero, or negative values are silently ignored and the next source in the chain is used.

Compile


Profiles, workflows, and connectors

Connector etiquette such as contactEmail, minRequestIntervalMs, and allowedHosts lives in .llmwiki/config.json. Local connector config can only tighten the first-party registry policy. It cannot activate connectors or add new hosts.

Gateway-specific request fields

LLMWIKI_OPENAI_EXTRA_BODY accepts a JSON object of additional chat-completion body fields. It applies to OpenAI-compatible providers, including streaming and tool calls through the OpenAI client, but never to embeddings, agent-backed providers, or Ollama’s native structured-output requests. Blank means no extensions. For an endpoint that requires thinking disabled for forced tools:
Check your endpoint’s documentation before using vendor-specific fields. For example, DeepSeek documents its thinking toggle. Fields are added at the top level, not nested under extra_body, and nested objects are not merged. Unsupported extensions can still produce provider errors. Malformed JSON, non-object values, and overrides of model, messages, tools, tool_choice, stream, max_tokens, max_completion_tokens, or reasoning_effort fail locally without retry backoff. Use the dedicated token and reasoning settings for those limits. Required tool selection remains in place because extraction depends on structured arguments, not optional prose.

Output


Debug


Example .env file

Place this file in your project root (the same directory that contains sources/ and wiki/):
Except for explicitly selected codex-agent runs, the .env file is read from your project root at startup. It is a convenience for projects where you don’t want to export variables in your shell every session. Shell environment variables always take precedence over .env values, so you can override any .env setting with a one-off export without editing the file.