Documentation / Concepts
One runtime, one loop
Why Seshat keeps sessions, prompts, tools, permissions and storage in a single runtime, and what happens in each turn.
Many agent projects are a loop glued to separate pieces: a prompt builder here, a tool wrapper there, a permission check bolted on, a database somewhere else. Seshat takes the opposite approach, sometimes called mono-run: one runtime owns the whole life of a run, so every part sees the same state.
- SessionsHistory, status, resume
- PromptsBuilt every turn from layers
- ToolsResolved, validated, executed
- PermissionsChecked before any tool runs
- ProvidersRetries and fallbacks
- StorageSessions and artifacts
- CompactionKeeps the context in budget
- EventsStreaming and hooks
Why it matters
- Permissions cannot be skipped. A tool call always goes through the same pipeline, whether it comes from the terminal, the SDK or the gRPC server.
- The prompt knows the state. The prompt is rebuilt each turn with the current session, tools and memory, so it never describes tools that are not there.
- Failures are handled in one place. Retries, fallbacks and recovery live next to the loop that needs them.
- One thing to embed. You adopt a runtime, not a set of parts to wire together.
The price is that the runtime is opinionated. If you want full control of every step, a framework of small parts may suit you better.
What happens in a turn
Each turn repeats these steps until the model has nothing left to do:
- Compaction. If the context window is filling up, older messages are summarised first.
- Model call. The conversation is sent to the provider, streaming when possible. Recoverable errors are retried, and a failing model or provider can be replaced by a fallback.
- Tool calls. If the model asked for tools, they run. Calls that are safe to run together run in parallel; the others run one after another.
- Results. Tool results are added to the conversation and the loop goes round again.
- Stop. When the model ends its turn, stop hooks get a chance to ask for one more cycle. Otherwise the run ends and returns its messages, tool calls and token usage.
When something goes wrong
| Situation | What the runtime does |
|---|---|
| Rate limit or timeout | Waits with exponential backoff and retries |
| The answer is cut off at the token limit | Adds a continuation message and goes on |
| The model announces a next step but stops | Nudges it to continue, up to a limit you can set |
| A model keeps failing | Tries the fallback models, then the fallback providers |
| Everything failed | Returns an error with the recovery context |
Where to read more
The layers are described in Architecture, the tool pipeline in Tools and providers, and the full design notes in the repository: architecture.md.
Updated on 2026-10-07