Installation
No module named 'fleet_rlm'
Sync the workspace with all extras:
Fleet RLM requires exactly DSPy 3.3.1
Fleet certifies exactly dspy==3.3.1 and fails closed on any other installed version before starting any resource or binding a listener. --help stays reachable. Resync from the lock to restore the certified version:
Python version mismatch
fleet-rlm targets Python 3.13. Install a matching interpreter through uv if needed:Policy and profile
default_profile not set
Fleet refuses to start without an explicit [config] default_profile in config/fleet.toml. The shipped policy declares one profile:
/profiles command edits this key interactively for the next restart.
fleet cli rejects a non-daytona profile
fleet cli accepts only profiles whose runtime environment is daytona. Selecting a different environment fails before database preflight, MLflow startup, or backend spawning. Switch to a Daytona profile or use fleet-rlm serve-api for backend-only setups.
Unknown TOML key error at startup
Fleet fails startup on unknown keys, missing profiles, invalid variable references, and absent TOML. Common culprits:mlflow.trace_content_mode— removed. Delete the key.rlm.recursion_max_depth— removed. The native child boundary is a fixed invariant. Delete the key.
Configuration
Daytona configuration missing
Every profile requires FLEET_DAYTONA_API_KEY. Set it in .env or export it:
Provider credentials missing
Only the environment variables named by the selected profile are read. The shipped policy names the Databricks Unity AI Gateway Chat Completions endpoint:model, api_key_env, and base_url_env in config/fleet.toml, then set the variables named there (for example, FLEET_OPENAI_API_KEY and FLEET_OPENAI_BASE_URL=https://api.openai.com/v1).
Backend startup
Connection refused at 127.0.0.1:8000
Check what is holding the port and rerun on a different port:
Backend refuses to bind a non-loopback host
Launchers default to127.0.0.1 and reject 0.0.0.0, LAN addresses, and hostnames other than localhost. Pass --allow-non-loopback-bind to opt in when running behind a proxy.
Database preflight fails
fleet cli verifies the configured database is at the canonical Alembic head. Recover with:
Daytona
Diagnose before a real Turn
Run the doctor to validate settings, database head, provider auth, Volume visibility, mount scoping, and interpreter execution:finally, and prints only bounded categories and corrective actions.
Missing Daytona base snapshot
Recursive delegations bootstrap from a reusable snapshot. If it is missing, rebuild it:Turn SSE stream
Transport 200 but the stream ends with error
The Turn stream begins immediately, so a healthy transport status does not imply a successful Turn. When claim or preparation fails, the stream closes with error + finish chunks carrying one of these messages:
Session not foundA Turn is already runningIdempotency key input mismatchInvalid Skill selectionTurn preparation timed outTurn is unavailableInvalid request
start chunk metadata, not from a response header.
Cancelled Turns show no reasoning or output
By design. Cancellation ends the live stream with a single terminalabort chunk. The persisted tombstone carries only observed usage, a cancelled data-status part, and the closed text Turn cancelled — never reasoning, code, output, or Tool evidence.
Turn fails with an adapter parse error
Fleet automatically re-asks the provider up to 2 more times with corrective feedback when a response cannot be parsed. If the Turn still fails withAdapterParseError, the provider returned empty or malformed output on every attempt. Check provider health and the MLflow trace for the raw responses.
MLflow tracing
Traces not appearing in the local server
fleet cli starts or reuses a local MLflow server on 127.0.0.1:5001; local tracing is the committed default. Backend-only commands (fleet web, fleet-rlm serve-api) require the tracking server to be started separately.
Managed MLflow inputs missing
A profile that declaresmlflow.tracking_uri = "databricks" with the managed *_env references requires FLEET_MLFLOW_EXPERIMENT_NAME, FLEET_MLFLOW_TRACE_CATALOG, FLEET_MLFLOW_TRACE_SCHEMA, FLEET_MLFLOW_TRACE_TABLE_PREFIX, and FLEET_MLFLOW_TRACING_SQL_WAREHOUSE_ID. Startup fails without them.
Workspace Memory
Workspace Memory degraded warnings in the logs
Workspace Memory preparation is fail-soft, so a warning does not fail the Turn. Read the five bounded fields to decide what to fix next:
category=provider_unavailable— the Volume or mounted agent is unreachable. Checkuv run fleet doctor daytonafor Volume visibility and mount scoping.category=invariant_violation— the durable store contains duplicate or invalid rows. Repair or dedupememory/MEMORIES.mdinside the Workspace Volume.category=corrupt_record_set— a mounted-agent Memory payload violated its response shape. Investigate the agent or reset the affected records.category=legacy_migration— the legacy rootMEMORIES.md→memory/MEMORIES.mdsequence failed. Remove any non-regular file at the legacy path.category=search_failureornormalization— the Turn already fell back to the recency-only digest; no action is needed unless it repeats.category=unexpected_internal— file a bug and includecause_type.
Still stuck?
- File an issue: github.com/qredence/fleet-rlm/issues
- Read the source —
src/fleet_rlm/api/mounts the routes andsrc/fleet_rlm/chat/turn_coordinator.pyowns the Turn lifecycle. The architecture page lists the reading order.