dspy.RLM (DSPy 3.3.1). Fleet keeps one resident native dspy.RLM and one caller-owned interpreter per healthy Session and reuses them across sequential clean Turns; a tainted or incompatible runtime rotates before the next Turn. There is no dspy.ReAct chat wrapper and no Fleet-owned RLM variable-mode monkeypatch. The maintained integration surface is the shipped Signatures, the delegation ladder, and the instruction fragments that compose the Root prompt.
Certified DSPy version
Fleet certifies exactly one published DSPy release:dspy==3.3.1. The runtime fails closed on anything else, including neighboring patches (3.3.0, 3.3.2), prereleases, post releases, and local builds.
The guard runs before any other startup work. Backend startup and every serving or diagnostic CLI command exit with a bounded error before any provider, database, or Daytona resource is constructed and before any listener binds:
--help stay reachable on any runtime. To restore the certified version, resync from the lock:
Where DSPy lives
Configure models through profiles
Model, provider, token, and recursion policy live inconfig/fleet.toml under the selected profile. There are no DSPY_* environment variables. Select the profile at the top of config/fleet.toml:
- The shipped policy calls an OpenAI-compatible Chat Completions endpoint. The committed defaults use the Databricks Unity AI Gateway:
DATABRICKS_TOKEN,FLEET_LLM_BASE_URL. - To route through OpenAI or another compatible gateway, update the selected profile’s
model,api_key_env, andbase_url_envinconfig/fleet.toml. - The committed policy runs
databricks-deepseek-v4-flash-0731for both Root and Sub.
The Root Signature
The default Fleet Root Signature carries a bounded, strict input surface. Every Signature receives:requesttext.history: dspy.Historywith the complete committed Session conversation.- Bounded
session_context. - Bounded
skill_cards. - Bounded Attachment metadata.
history field is the canonical conversation input: one ordered {"request": ..., "answer": ...} record per committed Turn, excluding hidden reasoning, generated Python, raw Tool results, and uncommitted candidates. The bounded previews in session_context and the read_session_history Tool remain as compatible navigation and retrieval surfaces. The Signature uses strict local Pydantic DTOs; conversion and JSON serialization happen once immediately before native dspy.RLM.acall().
Custom Skill Signatures retain JSON-compatible common input annotations. Only one selected Skill may provide a validated custom Signature per Turn; data-analysis is the only bundled Skill that does so.
Instructions are composed, not monolithic
src/fleet_rlm/rlm/program.py owns the default Fleet Root instruction fragments: base, REPL, tool, optional recursion, verification, and bounded-context guidance. Fragments are composed directly; disabling recursion under a non-recursive profile omits recursion guidance rather than deleting text from one large monolithic docstring.
Delegation ladder
The Root selects the cheapest sufficient primitive:- Python in the interpreter for deterministic work.
- Native
llm_query/llm_query_batchedfor semantic work. rlm_queryfor one iterative isolated subproblem.- Root-only
rlm_query_batchedfor ordered independent child RLMs.
RLM_NATIVE_CHILD_DEPTH = 1 is a fixed product invariant, not a policy value. Fleet reserves the shared recursive budget atomically before starting a child and bounds sibling concurrency through recursion_max_parallel_children.
See the Recursive RLM concept page for the isolation contract and the [rlm] bounds.
Large inputs
Long documents, workspace bundles, and durable Attachments cannot be inlined into the Root prompt. DSPy 3.3.1 ships theSandboxSerializable contract; Fleet builds host-constructed capsules that DSPy injects into the interpreter as REPL variables.
For example, authorized Attachment context is packaged into AttachmentContextCapsule before the Turn runs:
sandbox_assignment(...) while the LM sees only a short rlm_preview() summary. Skill authors should not import capsule classes directly; the host builds them.
Interpreter reuse across Turns
Within one Run, interpreter calls reuse one context, so Python state persists across RLM iterations. A healthy Session also reuses that caller-owned interpreter across sequential clean Turns: ordinary Python globals, imports, and helper functions may carry forward while the resident runtime stays healthy and compatible. DSPy’sREPLHistory stays fresh per invocation even while the interpreter is reused sequentially; that is the upstream DSPy 3.3.1 contract Fleet relies on. Failure, cancellation, timeout, or uncertain settlement taints the runtime; Fleet then rotates to a fresh interpreter and Sandbox and rehydrates only durable state. Replacing a Daytona Sandbox remounts the Workspace Volume Scope without preserving Python globals.
Bounded provider re-asks
DSPy 3.3.1 raisesAdapterParseError when a provider returns an empty or unparsable action response, which previously failed the Turn immediately. Fleet now re-asks the same LM with corrective feedback appended to the prompt, up to 2 additional attempts, before the original AdapterParseError propagates and fails the Turn. LM timeout and transport errors receive the same bounded re-ask treatment. This behavior is automatic and has no configuration surface. Re-ask attempts appear as extra LM calls in MLflow traces.
Tracing
config/fleet.toml [mlflow] policy controls DSPy tracing. Fleet enables MLflow DSPy inference autologging for the selected experiment; compile and evaluator traces stay disabled for live Turn observability. mlflow.trace_content_max_chars bounds each readable field, and mlflow.async_logging = true keeps trace export off the Turn critical path. See the observability page for the full policy.
What is not here
- No
dspy.GEPAoptimization API surface. Committed profiles are deterministic; there is nooptimizesubcommand and noPOST /api/v1/optimization/*endpoints. - No
dspy.ReActFleetAgentchat wrapper. - No BYOK per-caller LM selection.
- No Fleet-owned RLM variable-mode wrappers. Large inputs use DSPy’s native
SandboxSerializable.
See also
Recursive RLM
One-level recursive child boundary and
[rlm] bounds.Agent model
The Turn model, resident Session runtime, and Skill disclosure.
HTTP API
Turn SSE contract that every Signature runs behind.
Configuration
Profile matrix and provider environment variables.