> ## Documentation Index
> Fetch the complete documentation index at: https://docs.qredence.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Reasoning and recursion

> How Fleet RLM uses native dspy.RLM, Root and Sub models, semantic calls, recursive child RLMs, and Turn budgets.

Fleet RLM runs the native DSPy `dspy.RLM` loop. DSPy owns the loop, the REPL history, trajectory semantics, and provider retries. Fleet supplies Turn-scoped tools, budgets, and finalization. There is no second planner or model router.

## Root and Sub models

Fleet configures two model roles in `config/fleet.toml`.

| Role | Table | Used for |
| - | - | - |
| Root | `[llm.root]` | The main RLM loop that plans, writes Python, and produces the answer. |
| Sub | `[llm.sub]` | Semantic calls from Sandbox code through `llm_query` and `llm_query_batched`. |

Both roles call an OpenAI-compatible Chat Completions endpoint through a stock `dspy.LM`. See [Configure models](/fleet-rlm/guides/models-and-providers).

## Three ways to do sub-work

Pick the lightest option that solves the problem.

| Option | Runs | Use it for |
| - | - | - |
| Sandbox Python | In the root Sandbox | Deterministic inspection, parsing, math, and reduction. |
| `llm_query`, `llm_query_batched` | Native DSPy semantic calls on the Sub model | Bounded semantic questions inside the current invocation, such as classifying or summarizing a chunk. |
| `rlm_query`, `rlm_query_batched` | A separate child RLM in its own Sandbox | An independent investigation that needs its own iterative loop. |

## Recursive child RLMs

`rlm_query` investigates one bounded task. `rlm_query_batched` runs several independent tasks in parallel and returns results in input order. Both are available only to the root RLM.

Each call takes a `task`, a list of staged relative `inputs`, and an optional `context`. Fleet then:

1. Resolves the selected inputs under the Session's authority.
2. Stages bounded copies in private scratch.
3. Runs the child in a Volume-less Daytona Sandbox, using `FLEET_DAYTONA_CHILD_SNAPSHOT` when set.
4. Validates and harvests declared result files.
5. Cleans up the child Sandbox.

Children stop at depth one and share the parent Turn budget. They receive no parent Workspace tools, memory, task checkpoint, publication capability, credentials, or writable Volume. The root RLM verifies child findings and owns every durable update. A Turn can't commit successfully while a child or its cleanup is unresolved.

Recursion is on in the shipped configuration. Set `rlm.recursion_enabled = false` and restart for native-only operation.

## Turn budgets

Every Turn has a shared budget. When a dimension runs out, Fleet stops new work and finalizes.

| Setting | Shipped value | Limits |
| - | - | - |
| `rlm.max_iters` | 12 | RLM loop iterations. |
| `rlm.max_llm_calls` | 32 | Prompts sent through `llm_query` and `llm_query_batched`. |
| `rlm.max_tool_calls` | 256 | Host tool calls. |
| `rlm.recursion_max_calls` | 4 | Child RLM calls. |
| `rlm.recursion_max_parallel_children` | 4 | Concurrent children. The maximum is 8. |
| `rlm.recursion_child_max_iters` | 8 | Iterations per child. |
| `rlm.recursion_child_max_llm_calls` | 12 | Semantic calls per child. |
| `runtime.turn_timeout_seconds` | 1800 | Wall-clock time for one Turn. |
| `rlm.wrap_up_seconds` | 300 | Part of the Turn deadline held back for finalization. |

Fleet enforces no separate per-Turn deadline on the models themselves. See the [configuration reference](/fleet-rlm/reference/configuration) for every limit.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.