Every agent in a pool obeys a budget: how many turns it may take, how much of the shared context it may use, how long it may run, and what happens to it if it is stopped before it reports. You give those numbers as one row of data, and the framework derives the policy that enforces them.
When to use what
| You want to | Use |
|---|---|
| Set limits — turns, context, time — and what an agent is told when it hits one | A budget row on agentPool or useAgent |
| Add a rule at one moment of a tool call — refuse a call, retry a failure, refuse an early report | hooks — see Tool hooks and guards |
| Change a decision the numbers cannot express — "stop only for room, never for time" | A policy of your own, extending DefaultAgentPolicy |
Start with a budget. Most harnesses never need more.
Quickstart
import { agentPool } from "@lloyal-labs/lloyal-agents";
import type { Budget, NudgeInput } from "@lloyal-labs/lloyal-agents";
/** What an agent is told when it must wind up. The framework supplies `reason` and `words`; the sentence is yours. */
const NUDGE = ({ reason, words }: NudgeInput): string =>
reason === "result" ? `Tool result too large for the remaining context. Report your findings now within ${words} words.`
: `Report your findings now within ${words} words.`;
const budget: Budget = {
maxTurns: 8,
context: { softLimit: 2048, hardLimit: 1024 },
time: { softLimit: 90_000, hardLimit: 150_000 },
nudge: NUDGE,
};
const pool = yield* agentPool({ tools, parent: spine, terminal: citedReport.tool, budget, orchestrate });useAgent({ budget, … }) takes the same row for a single agent.
What a budget row holds
Every field is optional; the framework's default stands where the row is silent.
| Field | What it decides | Default |
|---|---|---|
maxTurns |
Tool-use turns before the hard cut | 100 |
context |
softLimit: tokens remaining at which agents are nudged to report and new work stops. hardLimit: the floor at which an agent must stop — at least 512, the decode batch, or the pool refuses to start |
1024 / 512 |
time |
The same two limits in milliseconds since the agent started | none |
nudge |
The words an agent is told when it must wind up, given { reason, terminal, words }. Without it nothing is said: an agent over budget goes idle, and a result that will not fit is dropped |
none |
recovery |
For an agent stopped before it reported: prompt (its words, given the report's word budget), and the floors minTokens (100) and minToolCalls (2) below which it is not worth asking |
pruned |
recoveryShape |
"staggered" recovers one agent at a time, each with all the room freed so far; "parallel" recovers them together inside the loop, each capped |
"staggered" |
recoveryBudget |
The token cap for a recovery report, and for a voluntary report in flight | adaptive |
shouldExplore |
When retrieval tightens from explore to exploit: context (fraction of context still free) and time (fraction of the time limit used) |
0.4 / 0.5 |
maxToolRetries |
Retries of a transient tool failure before the call fails | 1 |
The framework speaks no words of its own to the model: the nudge and the recovery prompt are always yours. A row may also carry numbers of your own beside these — the research template keeps maxTasks in the same row.
Keep a table of rows
A harness with several kinds of agent keeps its rows in one table and hands the right one to each pool. The research template keeps one per effort level, plus one per stage:
export const BUDGETS = {
effort: {
high: {
maxTasks: 6, maxTurns: 10,
context: { softLimit: 2048, hardLimit: 1024 },
time: { softLimit: 240_000, hardLimit: 360_000 },
shouldExplore: { context: 0.4 },
recoveryShape: "staggered", // serial, full-headroom, lossless
},
low: {
maxTasks: 2, maxTurns: 10,
context: { softLimit: 10240, hardLimit: 8192 },
time: { softLimit: 90_000, hardLimit: 150_000 },
shouldExplore: { context: 1.0 }, // always exploit: strict on-topic retrieval from the first turn
recoveryShape: "parallel",
},
},
/** The settling pass: a turn cap and nothing else, so it writes for as long as the answer needs. */
settle: { maxTurns: 10 },
} as const;Accept prose as a result
An agent with tools normally finishes by calling its terminal tool; prose before any tool call is an idle stop, which keeps an agent from answering without evidence. An agent whose prose is its result — a settling pass, a passthrough answer — says so:
const agent = yield* useAgent({ parent: spine, ...prompt, budget: BUDGETS.settle, acceptFreeText: true });Write a policy of your own
When a decision cannot be written as numbers, extend DefaultAgentPolicy and override that one decision. Hand the policy to the pool in place of the row: with policy, the pool takes no budget or guards, and hooks and acceptFreeText go to the policy's constructor instead of the pool.
import { DefaultAgentPolicy } from "@lloyal-labs/lloyal-agents";
import type { Agent, ContextPressure } from "@lloyal-labs/lloyal-agents";
class Patient extends DefaultAgentPolicy {
shouldExit(agent: Agent, pressure: ContextPressure): boolean {
// Only ever for room, never for time.
return pressure.critical && super.shouldExit(agent, pressure);
}
}
const pool = yield* agentPool({ /* … */ policy: new Patient() });The decisions a policy makes, and when the pool asks:
| Method | Asked | It answers |
|---|---|---|
onProduced(agent, parsed, pressure, config) |
The model stopped, after the gates admitted any call | Dispatch the call, return the result, nudge, or go idle |
shouldExplore(agent, pressure) |
Before a tool that admits content runs | Score results against the agent's own query (explore), or against the original question too (exploit) |
shouldExit(agent, pressure) |
Before the agent produces another token | Stop it now; its branch stays for recovery |
onRecovery(agent, pressure, budget?) |
An agent was stopped without a result | extract with a prompt, or skip |
recoveryShape, recoveryBudget, pressureThresholds |
Once, when the pool starts | The shape and cap of recovery; the soft and hard limits |
hooks, guardOverrides |
At every moment of a tool call | See Tool hooks and guards |
DefaultAgentPolicy takes the same options a budget row derives — budget.context, budget.time, nudge, recovery, recoveryShape, recoveryBudget, shouldExplore, maxToolRetries, acceptFreeText, hooks, guardOverrides — so super keeps every default you did not override.
End work early
A reader's "Wrap up" is not the same event as a closed window. The platform already gives every pool the signals for each; your harness only needs to accept the commands and pass them to run, the execution from useExecution():
| To | Command | Handler | What happens |
|---|---|---|---|
| Wrap up and keep what the agents have | wrap_up |
run.wrapUp() |
No new agents; tool calls in flight finish and settle; every agent is recovered into a report. A tool parked on a rate-limit retry is abandoned with an honest failure. |
| Drop one line of work | cancel_agent |
run.cancel(agentId) |
That agent's tool call is aborted, it ends with user_cancel and no recovery, and its branch is pruned so its siblings get the room |
| Pause, then resume | pause, resume |
run.pause(), run.resume() |
Decoding holds at the next tick; branches stay resident; paused time does not count against time budgets |
| Stop everything | stop |
run.stop() |
Everything the run started unwinds, with nothing recovered |
The basic template accepts only stop. To add the rest, widen its Command to the whole of rig's RunCommand and add the handlers, as the research template does:
import type { RunCommand } from "@lloyal-labs/rig";
export type Command =
| { type: "submit_query"; query: string }
| { type: "open_doc"; docId: DocId | null }
| RunCommand
| { type: "quit" };handlers: {
// …
*wrap_up() { run.wrapUp(); },
*pause() { run.pause(); },
*resume() { run.resume(); },
*cancel_agent({ agentId }) { run.cancel(agentId); },
},The dev pane (LLOYAL_DEV=1) already has a Wrap up button and a cancel on every agent's lane, and sends these commands.
Common mistakes
- Treating the soft limit as a kill line. It is the new-work and nudge boundary; recovery may use the reserve down to the hard limit.
- Raising the soft limit to get longer recovery reports. It nudges earlier; recovery stays bounded by the hard-limit reserve.
- Leaving out
nudgeand expecting agents to be told to report. Without it the framework says nothing: an over-budget agent goes idle. - Conflating exploit with exit.
shouldExplore() === falsenarrows retrieval; it does not end the agent. - Using a halt for "wrap up". A halt removes the owner. Use
wrap_up(run.wrapUp()) when the reader wants the best available result.
Tuning by workload
Interactive assistant
Prefer:
- fewer Agents;
- earlier exploit;
- tighter time limits;
- parallel recovery;
- modest report budgets;
- low retry count.
Broad research
Prefer:
- exploration while context is healthy;
- enough soft reserve for synthesis;
pruneOnReturn;- terminal Tool contracts;
- parallel recovery for lower effort levels.
Deep research
Prefer:
- dynamic orchestration;
- larger context;
- later exploit thresholds;
- staggered recovery;
- generous hard time;
- explicit downstream synthesis reserve.
Casework
Prefer:
- policy guards for institutional rules;
- domain-specific exit conditions;
- recoverable intermediate findings;
- protected action Tools;
- grants and audit events;
- procedural orchestration.
Policy should reflect the product’s meaning of done.
Rules for coding agents
## Lloyal AgentPolicy invariants
- AgentPolicy is a synchronous decision strategy, not a scheduler.
- BranchStore accounts, ContextPressure observes, policy decides, and
AgentPool enforces.
- Do not mutate branches, prune KV, or call native decode from a policy hook.
- Treat ContextPressure as an immutable decision-boundary snapshot.
- `softLimit` is the new-work and nudge boundary.
- `hardLimit` is the mechanical decode floor and recovery boundary.
- Recovery budgets from `remaining - hardLimit`, not only from `headroom`.
- `shouldExplore` changes retrieval strategy; it does not terminate an Agent.
- `shouldExit` stops an Agent but leaves its branch available for recovery.
- `onRecovery` chooses extract or skip; AgentPool enforces grammar and budget.
- `beforeAdmit` is asked only at the stall-break, after ordinary deferral cannot resolve.
- Wind-down, per-Agent cancellation, and Effection scope halt are different.
- Prefer configuring or extending `DefaultAgentPolicy` over replacing it.Related
- Adaptive compute — how pressure, limits, recovery and pruning work together.
- Tool hooks and guards — the five moments of a tool call.