Skip to main content
Fleet runs on native dspy.RLM (DSPy 3.3.1). Fleet keeps one resident native dspy.RLM and one caller-owned interpreter per healthy Session and reuses them across sequential clean Turns; a tainted or incompatible runtime rotates before the next Turn. There is no dspy.ReAct chat wrapper and no Fleet-owned RLM variable-mode monkeypatch. The maintained integration surface is the shipped Signatures, the delegation ladder, and the instruction fragments that compose the Root prompt.

Certified DSPy version

Fleet certifies exactly one published DSPy release: dspy==3.3.1. The runtime fails closed on anything else, including neighboring patches (3.3.0, 3.3.2), prereleases, post releases, and local builds. The guard runs before any other startup work. Backend startup and every serving or diagnostic CLI command exit with a bounded error before any provider, database, or Daytona resource is constructed and before any listener binds:
Argument parsing and --help stay reachable on any runtime. To restore the certified version, resync from the lock:

Where DSPy lives

Configure models through profiles

Model, provider, token, and recursion policy live in config/fleet.toml under the selected profile. There are no DSPY_* environment variables. Select the profile at the top of config/fleet.toml:
Provider environment variables are set by profile:
  • The shipped policy calls an OpenAI-compatible Chat Completions endpoint. The committed defaults use the Databricks Unity AI Gateway: DATABRICKS_TOKEN, FLEET_LLM_BASE_URL.
  • To route through OpenAI or another compatible gateway, update the selected profile’s model, api_key_env, and base_url_env in config/fleet.toml.
  • The committed policy runs databricks-deepseek-v4-flash-0731 for both Root and Sub.
See the configuration reference for the full matrix.

The Root Signature

The default Fleet Root Signature carries a bounded, strict input surface. Every Signature receives:
  • request text.
  • history: dspy.History with the complete committed Session conversation.
  • Bounded session_context.
  • Bounded skill_cards.
  • Bounded Attachment metadata.
The history field is the canonical conversation input: one ordered {"request": ..., "answer": ...} record per committed Turn, excluding hidden reasoning, generated Python, raw Tool results, and uncommitted candidates. The bounded previews in session_context and the read_session_history Tool remain as compatible navigation and retrieval surfaces. The Signature uses strict local Pydantic DTOs; conversion and JSON serialization happen once immediately before native dspy.RLM.acall(). Custom Skill Signatures retain JSON-compatible common input annotations. Only one selected Skill may provide a validated custom Signature per Turn; data-analysis is the only bundled Skill that does so.

Instructions are composed, not monolithic

src/fleet_rlm/rlm/program.py owns the default Fleet Root instruction fragments: base, REPL, tool, optional recursion, verification, and bounded-context guidance. Fragments are composed directly; disabling recursion under a non-recursive profile omits recursion guidance rather than deleting text from one large monolithic docstring.

Delegation ladder

The Root selects the cheapest sufficient primitive:
  1. Python in the interpreter for deterministic work.
  2. Native llm_query / llm_query_batched for semantic work.
  3. rlm_query for one iterative isolated subproblem.
  4. Root-only rlm_query_batched for ordered independent child RLMs.
Recursive children remain one native level deep. RLM_NATIVE_CHILD_DEPTH = 1 is a fixed product invariant, not a policy value. Fleet reserves the shared recursive budget atomically before starting a child and bounds sibling concurrency through recursion_max_parallel_children. See the Recursive RLM concept page for the isolation contract and the [rlm] bounds.

Large inputs

Long documents, workspace bundles, and durable Attachments cannot be inlined into the Root prompt. DSPy 3.3.1 ships the SandboxSerializable contract; Fleet builds host-constructed capsules that DSPy injects into the interpreter as REPL variables. For example, authorized Attachment context is packaged into AttachmentContextCapsule before the Turn runs:
Inside the sandbox, DSPy reconstructs the value through sandbox_assignment(...) while the LM sees only a short rlm_preview() summary. Skill authors should not import capsule classes directly; the host builds them.

Interpreter reuse across Turns

Within one Run, interpreter calls reuse one context, so Python state persists across RLM iterations. A healthy Session also reuses that caller-owned interpreter across sequential clean Turns: ordinary Python globals, imports, and helper functions may carry forward while the resident runtime stays healthy and compatible. DSPy’s REPLHistory stays fresh per invocation even while the interpreter is reused sequentially; that is the upstream DSPy 3.3.1 contract Fleet relies on. Failure, cancellation, timeout, or uncertain settlement taints the runtime; Fleet then rotates to a fresh interpreter and Sandbox and rehydrates only durable state. Replacing a Daytona Sandbox remounts the Workspace Volume Scope without preserving Python globals.

Bounded provider re-asks

DSPy 3.3.1 raises AdapterParseError when a provider returns an empty or unparsable action response, which previously failed the Turn immediately. Fleet now re-asks the same LM with corrective feedback appended to the prompt, up to 2 additional attempts, before the original AdapterParseError propagates and fails the Turn. LM timeout and transport errors receive the same bounded re-ask treatment. This behavior is automatic and has no configuration surface. Re-ask attempts appear as extra LM calls in MLflow traces.

Tracing

config/fleet.toml [mlflow] policy controls DSPy tracing. Fleet enables MLflow DSPy inference autologging for the selected experiment; compile and evaluator traces stay disabled for live Turn observability. mlflow.trace_content_max_chars bounds each readable field, and mlflow.async_logging = true keeps trace export off the Turn critical path. See the observability page for the full policy.

What is not here

  • No dspy.GEPA optimization API surface. Committed profiles are deterministic; there is no optimize subcommand and no POST /api/v1/optimization/* endpoints.
  • No dspy.ReAct FleetAgent chat wrapper.
  • No BYOK per-caller LM selection.
  • No Fleet-owned RLM variable-mode wrappers. Large inputs use DSPy’s native SandboxSerializable.

See also

Recursive RLM

One-level recursive child boundary and [rlm] bounds.

Agent model

The Turn model, resident Session runtime, and Skill disclosure.

HTTP API

Turn SSE contract that every Signature runs behind.

Configuration

Profile matrix and provider environment variables.
Last modified on September 4, 2026