Skip to main content

Installation

No module named 'fleet_rlm'

Sync the workspace with all extras:
Verify the entrypoints:

Fleet RLM requires exactly DSPy 3.3.1

Fleet certifies exactly dspy==3.3.1 and fails closed on any other installed version before starting any resource or binding a listener. --help stays reachable. Resync from the lock to restore the certified version:

Python version mismatch

fleet-rlm targets Python 3.13. Install a matching interpreter through uv if needed:

Policy and profile

default_profile not set

Fleet refuses to start without an explicit [config] default_profile in config/fleet.toml. The shipped policy declares one profile:
The pi-tui /profiles command edits this key interactively for the next restart.

fleet cli rejects a non-daytona profile

fleet cli accepts only profiles whose runtime environment is daytona. Selecting a different environment fails before database preflight, MLflow startup, or backend spawning. Switch to a Daytona profile or use fleet-rlm serve-api for backend-only setups.

Unknown TOML key error at startup

Fleet fails startup on unknown keys, missing profiles, invalid variable references, and absent TOML. Common culprits:
  • mlflow.trace_content_mode — removed. Delete the key.
  • rlm.recursion_max_depth — removed. The native child boundary is a fixed invariant. Delete the key.

Configuration

Daytona configuration missing

Every profile requires FLEET_DAYTONA_API_KEY. Set it in .env or export it:

Provider credentials missing

Only the environment variables named by the selected profile are read. The shipped policy names the Databricks Unity AI Gateway Chat Completions endpoint:
To use OpenAI or another OpenAI-compatible gateway, rewrite the selected profile’s model, api_key_env, and base_url_env in config/fleet.toml, then set the variables named there (for example, FLEET_OPENAI_API_KEY and FLEET_OPENAI_BASE_URL=https://api.openai.com/v1).

Backend startup

Connection refused at 127.0.0.1:8000

Check what is holding the port and rerun on a different port:

Backend refuses to bind a non-loopback host

Launchers default to 127.0.0.1 and reject 0.0.0.0, LAN addresses, and hostnames other than localhost. Pass --allow-non-loopback-bind to opt in when running behind a proxy.

Database preflight fails

fleet cli verifies the configured database is at the canonical Alembic head. Recover with:

Daytona

Diagnose before a real Turn

Run the doctor to validate settings, database head, provider auth, Volume visibility, mount scoping, and interpreter execution:
It creates one uniquely labelled disposable Sandbox, deletes it in finally, and prints only bounded categories and corrective actions.

Missing Daytona base snapshot

Recursive delegations bootstrap from a reusable snapshot. If it is missing, rebuild it:

Turn SSE stream

Transport 200 but the stream ends with error

The Turn stream begins immediately, so a healthy transport status does not imply a successful Turn. When claim or preparation fails, the stream closes with error + finish chunks carrying one of these messages:
  • Session not found
  • A Turn is already running
  • Idempotency key input mismatch
  • Invalid Skill selection
  • Turn preparation timed out
  • Turn is unavailable
  • Invalid request
Read the Run id from the start chunk metadata, not from a response header.

Cancelled Turns show no reasoning or output

By design. Cancellation ends the live stream with a single terminal abort chunk. The persisted tombstone carries only observed usage, a cancelled data-status part, and the closed text Turn cancelled — never reasoning, code, output, or Tool evidence.

Turn fails with an adapter parse error

Fleet automatically re-asks the provider up to 2 more times with corrective feedback when a response cannot be parsed. If the Turn still fails with AdapterParseError, the provider returned empty or malformed output on every attempt. Check provider health and the MLflow trace for the raw responses.

MLflow tracing

Traces not appearing in the local server

fleet cli starts or reuses a local MLflow server on 127.0.0.1:5001; local tracing is the committed default. Backend-only commands (fleet web, fleet-rlm serve-api) require the tracking server to be started separately.

Managed MLflow inputs missing

A profile that declares mlflow.tracking_uri = "databricks" with the managed *_env references requires FLEET_MLFLOW_EXPERIMENT_NAME, FLEET_MLFLOW_TRACE_CATALOG, FLEET_MLFLOW_TRACE_SCHEMA, FLEET_MLFLOW_TRACE_TABLE_PREFIX, and FLEET_MLFLOW_TRACING_SQL_WAREHOUSE_ID. Startup fails without them.

Workspace Memory

Workspace Memory degraded warnings in the logs

Workspace Memory preparation is fail-soft, so a warning does not fail the Turn. Read the five bounded fields to decide what to fix next:
  • category=provider_unavailable — the Volume or mounted agent is unreachable. Check uv run fleet doctor daytona for Volume visibility and mount scoping.
  • category=invariant_violation — the durable store contains duplicate or invalid rows. Repair or dedupe memory/MEMORIES.md inside the Workspace Volume.
  • category=corrupt_record_set — a mounted-agent Memory payload violated its response shape. Investigate the agent or reset the affected records.
  • category=legacy_migration — the legacy root MEMORIES.md → memory/MEMORIES.md sequence failed. Remove any non-regular file at the legacy path.
  • category=search_failure or normalization — the Turn already fell back to the recency-only digest; no action is needed unless it repeats.
  • category=unexpected_internal — file a bug and include cause_type.
See the Workspace Memory degradation section for the full field contract.

Still stuck?

Last modified on September 4, 2026