Skip to main content

Configuration Reference

OKT loads configuration from a layered YAML file. The defaults are embedded in the binary and auto-written to <binary_dir>/configs/config.default.yaml on first run when no on-disk copy exists.

Override values by creating configs/config.local.yaml (gitignored) alongside the binary. The load order is:

  1. configs/config.default.yaml (embedded fallback)
  2. configs/config.local.yaml (local overrides)
  3. Environment variables (highest priority)

You can also point to a custom config path with --config <file|dir>.

Example config

A minimal config.local.yaml for a typical local setup:

server:
port: 8080

databases:
default:
host: localhost
port: 5432
user: okt
password: okt_dev
name: okt
ssl_mode: disable
tasks:
host: localhost
port: 5432
user: okt
password: okt_dev
name: okt_tasks
ssl_mode: disable

system:
database: default

task:
database: tasks
job_timeout: 4h
queues:
retrieve_source: 20
source_decomposition: 100
embed_facts: 50
deduplicate_facts: 50
extract_concepts: 100

auth:
jwt_secret: "change-me-in-production"
token_ttl: 24h

providers:
search:
provider: "serper"
serper:
api_key: "your-serper-key"
decomposition:
fact_extraction:
provider: "openrouter"
model: "google/gemma-4-31b-it"
concept_extraction:
enabled: true
provider: "openrouter"
model: "google/gemma-4-31b-it"
embedding:
provider: "openrouter"
model: "google/gemini-embedding-2"
dimensions: 3072
qdrant:
host: localhost
port: 6334
ai:
openrouter:
api_key: "your-openrouter-key"

bootstrap:
default_repository: true
default_admin: false

Environment variable overrides

Many config values can be set via environment variables. Env vars take precedence over YAML. The API reads these at startup:

Env varConfig pathDescription
SERPER_API_KEYproviders.search.serper.api_keySerper web search key
OPENROUTER_API_KEYproviders.ai.openrouter.api_keyOpenRouter LLM key
OLLAMA_API_KEYproviders.ai.ollama_cloud.api_keyOllama Cloud key
OLLAMA_BASE_URLproviders.ai.ollama.base_urlLocal Ollama endpoint
OPENALEX_EMAILproviders.search.openalex.emailOpenAlex polite-pool email
UNPAYWALL_EMAILproviders.resolution.unpaywall.emailUnpaywall DOI email
OKT_FETCH_IMPERSONATEproviders.resolution.tls.impersonateTLS impersonation profile
FLARESOLVERR_URLproviders.resolution.flaresolverr.urlFlareSolverr endpoint
FLARESOLVERR_ENDPOINTSproviders.resolution.flaresolverr.endpointsComma-separated endpoints
FLARESOLVERR_MAX_CONCURRENCYproviders.resolution.flaresolverr.max_concurrencyMax in-flight calls
QDRANT_HOSTproviders.qdrant.hostQdrant host
QDRANT_PORTproviders.qdrant.portQdrant gRPC port
QDRANT_API_KEYproviders.qdrant.api_keyQdrant API key
REGISTRY_URLproviders.registry.urlKnowledge Registry URL
REGISTRY_AUTH_MODEproviders.registry.auth_modeRegistry auth mode
REGISTRY_API_KEYproviders.registry.api_keyRegistry write key
REGISTRY_READ_API_KEYproviders.registry.read_api_keyRegistry read key
AUTH_JWTSECRETauth.jwt_secretJWT signing secret
OKT_OAUTH_ISSUERoauth.issuerOAuth 2.1 issuer URL
OKT_BOOTSTRAP_AUTO_PROMOTEbootstrap.auto_promote_first_userFirst-user autopromotion toggle (true/false/1/0)
OKT_BOOTSTRAP_DEFAULT_ADMIN_EMAILbootstrapExplicit admin email (see bootstrap section)
OKT_BOOTSTRAP_DEFAULT_ADMIN_PASSWORDbootstrapExplicit admin password
OKT_BOOTSTRAP_DEFAULT_ADMIN_DISPLAY_NAMEbootstrapExplicit admin display name
OKT_AUDIT_RETENTION_DAYSaudit.retention_daysAudit log retention in days

Sections

server

HTTP server settings.

server:
port: 8080 # Listen port
read_timeout: 15s # HTTP read timeout
write_timeout: 15s # HTTP write timeout

databases

Named Postgres database registry. default is always required. Each entry connects to a Postgres instance and carries the full okt_system + okt_repository schema.

databases:
default:
host: localhost
port: 5432
user: okt
password: okt_dev
name: okt
ssl_mode: disable # disable | require | verify-ca | verify-full
max_conns: 200 # Connection pool size
tasks: # Dedicated River task queue database
host: localhost
port: 5432
user: okt
password: okt_dev
name: okt_tasks
ssl_mode: disable
max_conns: 200
# Per-tenant databases (optional):
# iso_8f3a:
# host: tenant-pg.internal
# port: 5432
# ...

system

Which named database carries the system tables (users, sessions, casbin, repositories). Empty falls back to default.

system:
database: default

isolation

Per-repository data isolation controls.

isolation:
default_database: default # Where new repos land by default
allowed_databases: [] # Picker allow-list (empty = closed)
# - iso_8f3a # Add names to open the picker

task

River task queue configuration. This is a top-level key — do not nest it under another section.

task:
database: tasks # Named database for River
job_timeout: 4h # Wall-clock cap per job (0s = unlimited)
heartbeat_interval: 1m # Worker heartbeat cadence
heartbeat_timeout: 10m # Stale heartbeat threshold
rescue_on_startup: true # Re-queue orphaned jobs on boot
refresh_concept_relations_interval: 10m
queues: # Worker counts per job kind
retrieve_source: 20 # Capped by FlareSolverr pool
source_decomposition: 100 # LLM-heavy
embed_facts: 50
deduplicate_facts: 50
extract_concepts: 100 # LLM-heavy
refine_concepts: 100
embed_concepts: 50
cleanup_facts: 50
summarize_concepts: 100 # LLM-heavy
synthesize_concept: 100 # LLM-heavy
refresh_concept_relations: 50
annotate_report: 50
migrate_context: 50
contribute_source: 50
pull_all_from_registry: 50

auth

JWT session authentication.

auth:
jwt_secret: "change-me-in-production" # HMAC secret for JWT signing
token_ttl: 24h # Session token lifetime

oauth

OAuth 2.1 authorization server (for MCP clients). Access tokens are HS256 JWTs signed with auth.jwt_secret.

oauth:
issuer: "" # MUST override in production (e.g. "https://okt.example.com")
access_token_ttl: 15m # Short-lived access JWT
refresh_token_ttl: 720h # 30 days
auth_code_ttl: 10m # Single-use authorization codes

api_keys

Personal API keys (personal access tokens). Users create these from the Profile page to authenticate the REST surface without a browser session — useful for CI, scripts, and the CLI. Each key is scoped to a subset of (object, action) permission pairs and optionally to a single repository; the key never grants anything its owner couldn't already do via RBAC. Keys are opaque, stored as sha256(hex) hashes; the raw okt_-prefixed token is shown once at creation.

api_keys:
max_per_user: 20 # Cap on active keys per user (0 = built-in default 20)
default_ttl: 0 # Default lifetime in days when omitted (0 = no expiry)
max_ttl: 2160h # 90d upper bound on any key's lifetime (0 = disable cap)

In-repo fact/concept search behavior (lexical vs hybrid fusion with the Qdrant embedding index). Separate from providers.search (external search providers). When hybrid is disabled, or when Qdrant or the embedding provider is not configured at boot, the search endpoints fall back to lexical-only.

search:
hybrid:
enabled: true # Master switch (false = lexical-only)
rrf_k: 60 # Reciprocal Rank Fusion damping constant
min_score: 0.0 # Qdrant cosine similarity floor (0.0 = accept all)
over_fetch_multiplier: 3 # x limit each channel fetches before fusion

concepts

Concept-relations read surface (REST /concepts/{id}/relations and the MCP getRelatedConcepts tool). Classifies each relation by shared_fact_count into three bands: insignificant (< min_shared_fact_count, hidden by default), weak (between the two thresholds), stable (>= stable_shared_fact_count).

concepts:
min_shared_fact_count: 3 # Below this = insignificant (hidden unless show_insignificant=true)
stable_shared_fact_count: 6 # At/above this = stable
min_concept_fact_count: 5 # Floor for concept list/search (show_small=true to lower)

providers.search

Web and academic search providers.

providers:
promptset_default: "" # Global default promptset hash (empty = built-in default)
search:
provider: "serper" # Default search provider
serper:
api_key: "" # https://serper.dev
openalex:
email: "" # Polite-pool email for higher rate limits
registry:
per_page: 20 # Registry search page size (0 = default)
timeout: 15s # Registry search timeout

providers.resolution

Source fetching and resolution chain. Providers self-disable when their config is empty.

providers:
resolution:
fetch:
enabled: true
user_agent: "Mozilla/5.0 ..."
timeout: 60s
retry:
max_attempts: 3 # Including first attempt
base_delay: 2s
max_delay: 15s
retry_403_max_attempts: 2 # 403 retry cap (0 = 403 permanent)
unpaywall:
email: "" # https://unpaywall.org — email is the API key
tls:
impersonate: "" # e.g. "chrome_133" — empty disables
timeout: 30s
flaresolverr:
url: "" # Single endpoint (env: FLARESOLVERR_URL)
endpoints: [] # Multi-endpoint pool
timeout: 60s
max_concurrency: 0 # 0 = no cap; set to ~number of containers
retry:
max_attempts: 2 # 1 disables; sidecar transient-retry budget
base_delay: 5s
max_delay: 30s
retry_403_max_attempts: 1
host_overrides: # Static host → provider-id map (defaults pin Reddit)
"reddit.com": "flaresolverr"
"www.reddit.com": "flaresolverr"
chain: "" # Comma-separated provider order override
auto_skip: # Learned (host,provider) skip list
enabled: true
min_sample: 100 # Min total attempts before auto-skip
failure_threshold: 0.85 # Skip when failures/total >= threshold
cooldown: 24h # How long a skip stays active

providers.decomposition

Source decomposition (fact extraction, concept extraction, image extraction).

providers:
decomposition:
chunking:
chunk_size: 2000 # Characters per chunk
chunk_overlap: 200
fact_extraction:
provider: "openrouter"
model: "google/gemma-4-31b-it"
concurrency: 4 # Parallel chunks per source
image_extraction:
enabled: true
provider: "ollama_cloud"
model: "gemma4:31b-cloud" # Must be multimodal/vision
max_image_bytes: 5242880 # 5 MB
max_images_per_source: 20
concurrency: 4
concept_extraction:
enabled: true
provider: "openrouter"
model: "google/gemma-4-31b-it"
fact_batch_size: 10 # Facts per LLM call
concurrency: 4

providers.embedding

Vector embedding provider. The model and dimensions must match the Qdrant collection config.

providers:
embedding:
provider: "openrouter"
model: "google/gemini-embedding-2"
dimensions: 3072
# Free local alternative (requires Qdrant collection rebuild):
# provider: "ollama"
# model: "qwen3-embedding"
# dimensions: 1024

providers.qdrant

Vector store connection.

providers:
qdrant:
host: localhost
port: 6334 # gRPC port (REST is port-1)
api_key: "" # Empty when Qdrant has no auth
collection: "okt_facts"
concept_collection: "okt_concepts"
allow_recreate: false # true in dev to rebuild collections

providers.dedup

Cross-source deduplication thresholds.

providers:
dedup:
threshold: 0.94 # Cosine similarity above which facts are duplicates
catchup_max_age: 168h # Reap stuck facts older than this

providers.reports

Report auto-annotation (autocitation).

providers:
reports:
enabled: true
similarity_threshold: 0.7 # Lower than dedup — we want "related", not "duplicate"
max_facts_per_sentence: 5
min_sentence_runes: 40
posture_classifier:
enabled: true # Labels: related | supports | contradicts
provider: "openrouter"
model: "google/gemma-4-31b-it"
batch_size: 8
max_concurrent: 4
max_tokens: 800

providers.storage

File storage for source assets (images, PDF bodies).

providers:
storage:
backend: filesystem # Only "filesystem" implemented today
filesystem:
root: var/source_assets
# s3: # Reserved for future use
# bucket: ""
# region: ""
# endpoint: ""

providers.registry

Knowledge Registry connection (pre-decomposed source cache).

providers:
registry:
url: "https://registry.openktree.com" # Leave empty to disable
auth_mode: "" # none | bearer | hmac
api_key: "" # Write key
read_api_key: "" # Read key (falls back to api_key)
allowed_models: [] # ["*"] to allow all
registries: [] # Multi-registry list (see below)

providers.registries

Multi-registry list (optional). When set, each repository can pick which registry it uses (cache lookup on fetch + remote browse/pull) from its Settings page. The legacy single registry: block above is the default entry (id default) when this list is empty; you don't need to duplicate it here.

Each entry needs a unique id (used as the per-repo selector) and a url.

providers:
registries:
- id: public
url: "https://registry.example.org"
auth_mode: bearer
api_key: "write-secret"
read_api_key: "read-secret"
allowed_models: ["*"]
- id: internal
url: "http://internal-registry:8081"
allowed_models: ["openai/gpt-4o"]

providers.summarization

Concept summarization (incremental summary slices).

providers:
summarization:
enabled: false # Off by default — enable after bootstrapping facts
provider: "openrouter"
model: "google/gemma-4-31b-it"
batch_size: 20 # Facts per summary slice
max_concepts_per_run: 40
lock_staleness: 2h
max_tokens: 600 # ~450 words per slice

providers.refinement

Concept refinement (resolve unresolved concept candidates).

providers:
refinement:
enabled: false # Off by default
provider: "openrouter"
model: "google/gemma-4-31b-it"
max_candidates_per_run: 40
prune_threshold: 5
max_tokens: 400
max_concurrency: 5

providers.synthesis

Concept synthesis ("definitions" — the authoritative crystallized knowledge).

providers:
synthesis:
enabled: true
provider: "openrouter"
model: "deepseek/deepseek-v4-flash:turbo"
image_picker_model: "google/gemma-4-31b-it"
max_tokens: 10000
thinking_level: "low" # low | medium | high
max_images: 10
max_image_candidates: 50
max_related_concepts: 10
max_related_syntheses: 3

providers.ai

LLM provider connections and model catalog.

providers:
ai:
ollama:
base_url: "" # Local Ollama endpoint (e.g. "http://localhost:11434")
ollama_cloud:
api_key: "" # Ollama Cloud API key
openrouter:
api_key: ""
embed_batch_size: 32 # Inputs per embedding POST
rate_limit_wait_timeout: 1h # How long a Chat/Embed call blocks for a rate-limiter token
models: # Catalog with rate limits and cost
- id: "google/gemma-4-31b-it"
provider: "openrouter"
input_cost_per_1m: 0.12
output_cost_per_1m: 0.4
rate_limit_rpm: 500
- id: "google/gemini-embedding-2"
provider: "openrouter"
input_cost_per_1m: 0.0
output_cost_per_1m: 0.0
rate_limit_rpm: 500

providers.llm_retry

LLM retry budget shared by every AI-provider Chat/Embed call (decomposition, summarization, refinement, synthesis). The defaults let a single LLM call ride out a transient provider degradation.

providers:
llm_retry:
max_attempts: 4 # Total attempts including the first
base_delay: 2s # Backoff for attempt 2; grows exponentially
max_delay: 30s # Cap on per-attempt backoff
per_call_timeout: 5m # Per-attempt ctx timeout

per_call_timeout must be less than or equal to the smallest worker LLM timeout; the worker timeout must exceed max_attempts × per_call_timeout + backoffs.

bootstrap

One-time startup data creation. Each step is idempotent and only acts when its target table is empty.

bootstrap:
default_repository: true # Create a starter repository on first boot
auto_promote_first_user: true # First /auth/register on empty users → sysadmin
default_admin: false # Seed a sysadmin (credentials via env vars)
# OKT_BOOTSTRAP_DEFAULT_ADMIN_EMAIL: [email protected]
# OKT_BOOTSTRAP_DEFAULT_ADMIN_PASSWORD: change-me
# OKT_BOOTSTRAP_DEFAULT_ADMIN_DISPLAY_NAME: "Default Admin"
FieldEnv varDefaultPurpose
auto_promote_first_userOKT_BOOTSTRAP_AUTO_PROMOTEtrueWhen true, the first successful POST /api/v1/auth/register on an empty users table grants the sysadmin role on the system domain to that user. Smooth out-of-the-box path for docker compose up + register. Turn off for public deployments so an attacker cannot become sysadmin by registering first. The env var accepts true/false/1/0 (case-insensitive).
default_adminOKT_BOOTSTRAP_DEFAULT_ADMIN_*falseWhen true, seeds a sysadmin at boot from the OKT_BOOTSTRAP_DEFAULT_ADMIN_{EMAIL,PASSWORD,DISPLAY_NAME} env vars when the users table is empty. Skipped if any env var is missing or the users table is non-empty.

When both auto_promote_first_user and default_admin are enabled, default_admin wins: it runs at boot before any Register call, so the users table is non-empty by the time autopromote's CountUsers() == 1 guard would fire. A log line is emitted when autopromote fires so an operator notices if it fires unexpectedly on a public deployment.

repository_presets

The repository "types" offered by the create-repository UI. Each preset bundles a provider set and context allow-list.

repository_presets:
- id: general
label: General
description: "All providers enabled, full context vocabulary."
providers:
search: ["serper", "openalex"]
resolution: ["fetch", "unpaywall", "tls", "flaresolverr"]
contexts: ["all"]
- id: scientific
label: Scientific
description: "Academic search + open-access resolution."
providers:
search: ["openalex"]
resolution: ["fetch", "unpaywall", "tls"]
contexts: ["Biomolecule", "drug", "gene", "protein", ...]
- id: enterprise
label: Enterprise
description: "Web search + plain fetch, custom contexts."
providers:
search: ["serper"]
resolution: ["fetch"]
contexts: []
custom_contexts: ["Product", "Application", "Role"]

default_repository_preset: general

audit

Audit-log retention. The daily audit_cleanup periodic job deletes okt_system.permission_audit rows older than retention_days. The default of 30 matches the common compliance window; raise it for longer history, lower it to bound table growth.

audit:
retention_days: 30 # Env override: OKT_AUDIT_RETENTION_DAYS