.env file placed in your project directory (alongside sources/ and wiki/). Variables set in your shell take precedence over those in .env, which in turn take precedence over the Claude Code settings fallback (~/.claude/settings.json → env block) and built-in defaults.
You don’t need to set everything - only the variables relevant to your chosen provider are required. Most projects need only two or three exports before running llmwiki compile.
Provider selection
LLMWIKI_EMBEDDINGS controls embedding writes, not retrieval. While it is
disabled, write paths do not read or change the embedding store (binary index
or legacy .llmwiki/embeddings.json), .llmwiki/pending-embeddings.json, or
.llmwiki/quarantined-embeddings.json. An existing pending marker may therefore
remain visible in lint and status. Query and search commands can still read
an existing embedding store, and they keep their normal embedding-provider
validation.
After you re-enable refreshes, run llmwiki compile. The compile reconciles
missing and stale vectors even when no source changed. It does not regenerate
source-derived pages in that case, and a healthy store is not rewritten.
When you set LLMWIKI_EMBEDDING_PROVIDER, the named provider’s own credential is required: VOYAGE_API_KEY for anthropic or claude-agent, OPENAI_API_KEY (or OPENAI_EMBEDDINGS_API_KEY) for openai, ORCAROUTER_API_KEY for orcarouter, and no key for ollama. Setting OPENAI_EMBEDDINGS_BASE_URL marks a self-hosted OpenAI-compatible endpoint that needs no key, so OPENAI_API_KEY is not required in that case. This check only runs when LLMWIKI_EMBEDDING_PROVIDER is set - leaving it unset keeps today’s behavior unchanged.
The check runs at startup, before any compile or query work begins, so a misspelled provider name or a missing key fails immediately instead of part-way through a run. Compile skips embedding-only checks when refreshes are disabled, but still validates the provider used to generate pages.
codex-agent is the exception to the default embedding-provider value. Because
Codex cannot embed, commands that consume or produce embeddings require
LLMWIKI_EMBEDDING_PROVIDER to be set explicitly. A missing or unusable
embedding backend fails actionably rather than silently degrading. Compile does
not require one when LLMWIKI_EMBEDDINGS disables refreshes. The keyless
combination with embeddings enabled is codex-agent chat plus ollama
embeddings.
When codex-agent is selected explicitly through --provider or a
shell-exported LLMWIKI_PROVIDER, llmwiki does not open the project .env
file, because that file may contain API keys. Export the Codex model, embedding
provider, and embedding-backend settings in the shell for that run. Other
providers retain the existing shell-over-.env precedence.
Changing the embedding backend rebuilds the index
The embedding index records which provider, model, and endpoint produced its vectors. Changing the effectiveLLMWIKI_EMBEDDING_PROVIDER, LLMWIKI_EMBEDDING_MODEL, or embedding endpoint invalidates the index. The next non-review llmwiki compile with refreshes enabled can re-embed the entire eligible wiki, even when no source changed. This incurs embedding-provider requests and potential costs without regenerating source-derived pages. Quarantined pages remain excluded until you explicitly re-queue them; see retry recovery.
This matters because vectors from different backends are not comparable even when the model name matches - a local server answering to text-embedding-3-small does not produce the same vectors as cloud OpenAI, and nomic-embed-text served by Ollama does not match the same model served over an OpenAI-compatible endpoint. Until the index is rebuilt, llmwiki query reports the index as outdated and falls back to lexical ranking rather than ranking against a mixed index.
Moving between anthropic and claude-agent does not rebuild the index. Both send embeddings to Voyage with the same model, so their vectors are interchangeable.
An index built before llmwiki recorded the endpoint carries only its model name, which cannot distinguish those backends. Such an index is kept as-is while you run without LLMWIKI_EMBEDDING_PROVIDER or an endpoint override, so upgrading does not re-embed your project. With either override active, the next llmwiki compile rebuilds it once and records the full configuration from then on.
Anthropic
Either
ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN satisfies authentication - you do not need both. If neither is set in your environment or .env, llmwiki attempts to read these values from ~/.claude/settings.json.
OpenAI
OpenAI-compatible request options
These settings apply to chat requests built by the shared OpenAI-compatible client, including OpenAI, Copilot, MiniMax, Atlas Cloud, and Ollama’s completion and streaming paths. They do not affect embeddings, Anthropic, the local agent providers, or Ollama’s native structured-output requests.
Invalid values fail on the first request attempt, before network access, without
retry backoff. Valid values are checked against the installed SDK vocabulary;
individual models can support a smaller subset. In particular, older GPT-5
models do not support
none.
Reasoning models reject
max_tokens outright — the request fails with Unsupported parameter: 'max_tokens' is not supported with this model. llmwiki picks the right field from the model id, so LLMWIKI_OPENAI_TOKEN_PARAM is only needed when the id itself does not reveal the model family, which is common behind OpenAI-compatible gateways.OpenAI Codex CLI
Authentication is managed only by the installed Codex CLI (
codex login).
llmwiki does not read Codex auth files or any OpenAI API key for chat. The CLI
does not expose a maxTokens control, so llmwiki cannot honor that internal
hint for this provider. Each request has a fixed 10-minute timeout and 1 MiB
process/final-output caps.
The minimum verified compatible Codex CLI version is 0.152.1. If an older
binary rejects a required secure flag, llmwiki reports it as incompatible;
update with npm install -g @openai/codex@latest.
Ollama
GitHub Copilot
Atlas Cloud
Model names are namespaced by publisher (e.g.
qwen/qwen3.5-35b-a3b, the default). An override set through LLMWIKI_MODEL must be a model Atlas Cloud lists as supporting tools, because compile extracts concepts through a tool call.
OrcaRouter
Voyage (embeddings)
Embeddings
llmwiki compile embeds pages and chunks in provider-native batches. These variables are optional.
Batching controls requests independently of storage. Small stores keep
.llmwiki/embeddings.json; stores exceeding its 64 MiB limit automatically use .llmwiki/embeddings.bin without repeating embedding requests.LLMWIKI_EMBEDDINGS. An unchanged, healthy store makes no
embedding-provider calls and is not rewritten. See compile reconciliation.
LLMWIKI_EMBED_STRICT applies to automatically discovered work too: if embedding
generation fails, an otherwise unchanged compile exits non-zero after recording
retry state. Without strict mode, it warns and continues. Disabling refreshes
with LLMWIKI_EMBEDDINGS skips generation entirely, including in strict mode.
Embedding retries and quarantine
Explicit changes and automatically discovered missing or stale vectors share a durable retry budget in.llmwiki/pending-embeddings.json. Each budget is bound to
the page’s embeddable content (its title, summary and body chunks): an attempt is
charged against the content actually sent to the embedding provider, recorded
before the request is made.
Only real failures are charged:
- Pages in the request that failed are charged one attempt.
- Pages sent earlier in the same run that were not at fault are not charged.
- Pages the run never reached are not charged.
- A run that fails before any provider request charges nobody, but is still reported.
- If every request succeeds but the results cannot be saved, the pages sent are charged, so paid work that can never be saved stops.
.llmwiki/quarantined-embeddings.json. Later refreshes skip it while its
content is unchanged, allowing other pages to embed.
A quarantined page is retried automatically, with a fresh budget, as soon as its
content changes. That includes changes made while refreshes were disabled, or
outside llmwiki. Re-saving identical content does not reset the budget. These files
use bounded, root-confined reads and atomic writes, and llmwiki re-reads them after
writing. Neither file is read or changed while refreshes are disabled.
Each retry file is limited to 5,000 entries and 512 KiB, which fits 5,000 typical
page ids with their content hashes. A refresh only attempts pages whose budgets fit
in the pending file; additional pages wait for a later compile. Existing entries are
never evicted to make room for new work. If the quarantine file is full, exhausted
entries stay in the pending file but are not retried. If both files fill with
exhausted entries, new embedding work pauses until you recover those entries. The
limits do not cause retries to restart.
Deferred work is reported with a page count. No provider request is made for a
page whose budget could not be recorded. Strict embedding mode reports this as
a failure after settling any admitted work, so successful refreshes are not
charged another retry. Fix marker storage errors or free capacity, then compile
again. An unchanged compile can continue the deferred work when space is available.
For a valid v3 store using the same embedding backend, deferred or quarantined
pages keep their previous page and passage vectors, with the original hashes and
timestamps. These cached vectors may be stale until refresh succeeds. Deleted
pages, pages no longer eligible for embeddings, invalid vectors, and vectors from another
backend are not retained. Legacy-store upgrades do not preserve stale vectors
for deferred pages; those pages wait for a successful refresh.
Entries from earlier versions. Quarantined entries written before retries were
bound to content are re-queued once by the first compile with refreshes enabled.
Each such page is bound to its current content and gets up to five additional retry
rounds; one round can issue more than one provider request. llmwiki prints a
one-time notice with the count, and llmwiki status reports these entries until
then. After that, they follow the rules above.
Resetting after fixing a provider. Unchanged content stays quarantined even
after the provider is fixed, so reset its retry entries explicitly:
- Make sure no compile or refresh is running.
- Remove the page’s entries from both
.llmwiki/quarantined-embeddings.jsonand.llmwiki/pending-embeddings.json. To reset every page, remove both files. - Run
llmwiki compilewith refreshes enabled.
Embedding storage
Binary storage keeps metadata and Float32 vectors in one atomically replaced file. It retains chunk text and the existing logical index version. Float32 introduces small rounding differences; it does not change the embedding model. The limits are 256 MiB of metadata, 512 MiB for the complete binary file, and 100,000 combined page and chunk records. Higher-dimensional vectors and longer chunk text use more space, so the record limit is not a promise that every such corpus fits. The index is still loaded into memory for retrieval. Onceembeddings.bin exists, readers and writers use it without the flag.
Any older JSON file is left untouched as a historical backup, not a fallback.
An unreadable or corrupt binary index produces embedding-store-unavailable
and degraded retrieval rather than silently reading that older snapshot.
The context command uses its existing embedding-store-missing warning code
for this condition, meaning no usable store, not necessarily no file on disk.
An enabled non-review llmwiki compile discovers the missing vectors even when
no sources changed. Rebuilding an unavailable index embeds the live eligible
corpus again, except quarantined pages, and can incur provider costs.
Unsetting the flag does not convert binary back to JSON. Before downgrading to
a release without binary support, back up and move both derived index files
aside, then rebuild with that release. This only works for a corpus that fits
the older JSON limit; larger corpora require a binary-capable release. Do not
delete just the binary file and reuse stale JSON. Keep your source files, wiki
pages and state.json intact.
Timeouts
Timeout resolution for Ollama: explicit constructor option →
OLLAMA_TIMEOUT_MS → LLMWIKI_REQUEST_TIMEOUT_MS → built-in 30-minute default.
Non-numeric, zero, or negative values are silently ignored and the next source in the chain is used.
Compile
Profiles, workflows, and connectors
Connector etiquette such as
contactEmail, minRequestIntervalMs, and
allowedHosts lives in .llmwiki/config.json. Local connector config can only
tighten the first-party registry policy. It cannot activate connectors or add
new hosts.
Gateway-specific request fields
LLMWIKI_OPENAI_EXTRA_BODY accepts a JSON object of additional chat-completion
body fields. It applies to OpenAI-compatible providers, including streaming and
tool calls through the OpenAI client, but never to embeddings, agent-backed
providers, or Ollama’s native structured-output requests. Blank means no
extensions. For an endpoint that requires thinking disabled for forced tools:
extra_body, and nested
objects are not merged. Unsupported extensions can still produce provider errors.
Malformed JSON, non-object values, and overrides of model, messages, tools,
tool_choice, stream, max_tokens, max_completion_tokens, or
reasoning_effort fail locally without retry backoff. Use the dedicated token
and reasoning settings for those limits. Required tool selection remains in
place because extraction depends on structured arguments, not optional prose.
Output
Debug
Example .env file
Place this file in your project root (the same directory that contains sources/ and wiki/):
Except for explicitly selected
codex-agent runs, the .env file is read from your project root at startup. It is a convenience for projects where you don’t want to export variables in your shell every session. Shell environment variables always take precedence over .env values, so you can override any .env setting with a one-off export without editing the file.