Skip to main content
Fleet starts from the required, committed config/fleet.toml policy file. [config] default_profile inside that file selects the active profile. The shipped policy declares exactly one profile, daytona-recursive, and the whole policy lives in [defaults]. Policy is strict, resolved once at process startup, and takes effect only after restart. The TOML file contains no secret values. It declares environment-variable names for Root/Sub API keys, the database URL, the Daytona API key, and managed MLflow destinations. Fleet reads only those named values from the process environment or repository .env (process values win). Fleet ignores other FLEET_* variables, including model, RLM, endpoint, runtime, and MLflow settings, unless the selected profile explicitly names them. Fleet does not consult FLEET_CONFIG_PROFILE. Unknown TOML keys, absent profiles, missing TOML, and invalid variable references fail startup.

Runtime prerequisites

The committed policy calls an OpenAI-compatible Chat Completions endpoint and routes Root and Sub through the Databricks Unity AI Gateway MLflow endpoint. It requires DATABRICKS_TOKEN, FLEET_LLM_BASE_URL, and FLEET_DAYTONA_API_KEY. To use OpenAI or another compatible provider instead, update the selected profile’s model, api_key_env, and base_url_env entries in config/fleet.toml; the base URL is typically the provider’s /v1 root. Profiles are explicit and do not fall back to each other. Daytona startup never applies migrations; use uv run python scripts/db_init.py or Alembic directly. The generated profile matrix shows the provider, token, recursion, and environment contract derived from config/fleet.toml.

Policy structure

config/fleet.toml deep-merges [defaults] into the selected [profiles.<name>]. It centralizes:
  • Application identity.
  • The runtime variant selector, runtime timeouts, leases, liveness, and the credentialed-command live switch.
  • Root/Sub model ids, provider-service routing, endpoint, token limit, per-role request timeout, temperature, cache, retries, and secret-variable references.
  • RLM limits, the wrap-up reserve, and host verbosity.
  • Storage limits and the database variable reference.
  • Daytona API-key/Volume/Snapshot policy.
  • MLflow tracking policy.
  • Fleet/DSPy logger level.
storage.max_upload_bytes bounds uploads and workspace files, storage.max_url_bytes bounds fetched public URL sources, and storage.max_artifact_bytes bounds artifact bodies.

Root and Sub models

Both roles run databricks-deepseek-v4-flash-0731 through the OpenAI-compatible Chat Completions format, using the DATABRICKS_TOKEN and FLEET_LLM_BASE_URL references. FLEET_LLM_BASE_URL must be the Databricks Unity AI Gateway /ai-gateway/mlflow/v1 base; the client appends /chat/completions. The role ceiling is a deliberately bounded max_tokens = 16384; Fleet’s character-level output caps bound retained output independently. Each role also carries its own provider request timeout: llm.root.timeout_seconds = 300 and llm.sub.timeout_seconds = 90. LM caching is disabled for both roles. The shipped Root and Sub roles set num_retries = 1. This is a committed runtime policy choice; custom profiles that omit the field inherit the shipped default of 1. The typed settings default of 3 applies only when both the defaults and the selected profile omit the field. Model ids may use an explicit provider/model prefix. For an OpenAI-compatible base URL, bare ids are normalized with the openai/ prefix before constructing dspy.LM.

Runtime variant

runtime.variant selects the execution architecture. Its default and only implemented value is legacy. Fleet rejects native and capsule at startup, and the settings editor offers only implemented choices. Policies that omit the key keep the legacy behavior; the committed policy names it explicitly. runtime.environment = "daytona" selects the provider environment independently. It does not select an execution architecture.
config/fleet.toml

Live commands

runtime.live_enabled defaults to true for explicitly invoked provider, Daytona, and Prime Oolong commands. Set it to false in the selected TOML policy to fail closed before those commands construct provider or Daytona clients. This policy replaces the old FLEET_LIVE=1 shell switch; invoking a live command remains an explicit operator action, and the required credentials are still validated.

MLflow tracing

When tracing is enabled, mlflow.async_logging keeps trace export off the Turn critical path and mlflow.trace_sampling_ratio controls the fraction of Turns sent to MLflow. The committed default is asynchronous export with a 1.0 sampling ratio. Trace payloads retain bounded, readable prompts, reasoning, generated code, tool payloads, and responses. mlflow.trace_content_max_chars bounds each readable field and defaults to 10000 characters. The trace export boundary still protects credentials, connection strings, private paths, and system-prompt dumps.
The mlflow.trace_content_mode setting is removed. fleet.toml files that still set trace_content_mode = "safe" fail validation with an unknown-key error; delete the key. Trace content is now always readable (bounded by mlflow.trace_content_max_chars).
The committed default routes traces to the local fleet-rlm experiment at http://127.0.0.1:5001; the supervised fleet cli command starts or reuses that server. Databricks-hosted tracing remains available for local policy: declare mlflow.tracking_uri = "databricks" together with the experiment_name_env, trace_catalog_env, trace_schema_env, trace_table_prefix_env, and tracing_sql_warehouse_id_env references in a profile. The loader resolves those names the same way. Fleet enables MLflow DSPy inference autologging for the selected experiment; compile and evaluator traces remain disabled for live Turn observability.

PostHog product analytics

The optional [posthog] policy section controls fail-soft PostHog product analytics. The shipped [defaults.posthog] policy enables analytics against the EU ingestion host and stays disabled whenever the named token variable is absent. Analytics never block startup, and re-init during the FastAPI lifespan is idempotent. Every event shares one stable per-installation distinct_id persisted under the storage data root at <data_root>/analytics-instance-id. The deterministic local user id is never used as a PostHog identity, so multiple installations remain distinct. PostHog exception autocapture is disabled; Turn failures are captured through sanitized failure messages only. The client emits these events from the corresponding HTTP routes: The Settings API exposes posthog.enabled, posthog.project_token_env, and posthog.host for editing through the loopback /api/settings surface. Changes apply after the next Fleet restart. Example policy:
config/fleet.toml

RLM bounds and the wrap-up reserve

The native RLM policy fields map directly to DSPy 3.3.x. max_iters bounds Root/child action iterations. max_llm_calls bounds prompts sent through native llm_query and llm_query_batched tools; each batched prompt counts. max_output_chars bounds each REPL output when DSPy renders native history for the next action; it is not a total-history limit. The shipped policy deliberately lowers the effective Root values to 12, 32, and 6000 (the generic DSPy fallback values are 20, 50, and 10000). The child values remain 8, 12, and 4000. The tightened Root budget leaves room for deliberate verification without allowing an unproductive long tail. rlm.wrap_up_seconds reserves a final-answer window before the Turn deadline. When the remaining time inside a Turn drops to this reserve, Fleet directs the Root to submit its final answer instead of starting new actions. The committed default is 300 seconds.

Turn budget

Each Turn owns one shared, atomic budget. Provider attempts, retries, adapter repairs, Tool calls, recursive children, retained execution output, and finalization all draw from it, so parallel children cannot double-count against the same allowance. When a Turn exhausts a budget dimension, Fleet stops admitting that kind of work and moves the Root toward finalization. These budget controls are separate from the DSPy-mapped max_iters, max_llm_calls, and max_output_chars fields above. The Turn deadline (runtime.turn_timeout_seconds) and the recursive call and concurrency limits below settle against the same shared budget.
config/fleet.toml

Recursive RLM

The [rlm] recursion settings bound the native rlm_query(prompt=prompt) child harness: The native recursive-child boundary is a fixed product invariant (RLM_NATIVE_CHILD_DEPTH = 1), not an editable policy value. Policies that still set rlm.recursion_max_depth fail validation; delete the key. Under daytona-recursive, each child receives a fresh, dedicated Daytona Sandbox, ordinary Daytona network egress, and the same Volume ID mounted at recursive/<workspace-id>/<run-id>/<call-index>. That private sibling scope cannot reach the Root workspaces/<workspace-id> mount. The child receives no Fleet Tools or credentials; strict cleanup purges its scope and deletes its Sandbox before Root success can commit.

Autonomous memory

rlm.autonomous_memory_categories is a TOML-only list of canonical Workspace Memory category names and defaults to [], which omits propose_memory from the Root Tool inventory entirely. A non-empty allowlist enables a Root-only, Run-scoped candidate collector and permits best-effort promotion only after a successful durable Turn commit; it does not change explicit-user memory behavior.

Environment inputs

Only variables named by the selected profile are read.

Terminal-only setting

FLEET_API_URL changes the standalone pi-tui API base URL from http://127.0.0.1:8000. It is not a backend Settings field and is unnecessary when the supervised fleet cli command supplies the local API URL.

Local terminal editing

The pi-tui /settings command reads and edits the non-secret policy in config/fleet.toml. It is available only to a loopback API client, including when an operator has explicitly exposed the normal API on another interface. The selector supports [defaults] and every existing named profile, and offers choice, text/number, and boolean child panels. Edits are revision-checked, validated against every profile, and saved as one atomic batch; either every change lands or none do. You can also reset a profile override so the field inherits its [defaults] value again. The /settings panels never read or display .env values or provider credentials; database and provider values are represented only by their environment-variable names. A saved policy applies only after Fleet is restarted; existing runtime composition and active Turns are never changed in place. The companion pi-tui /profiles command writes the chosen name to config.default_profile through the same loopback policy. It labels the active profile as running and a different default_profile as selected for restart.

Example .env

Copy the shipped template and fill only variables named by the selected profile:
.env
Never commit .env, credentials, raw provider failures, or evidence containing secrets.

See also

Last modified on September 6, 2026