1.9 Durable orchestration engines
Business process management is not new. Software for tracking a case, recording where it sits, and managing how it executes has existed for decades, and durable workflow state is not new either. What is new is the prominence of the word in front of it — durable — in agent tooling, and the word that needs defining is not execution but orchestration, because what an orchestrator runs is not one program. It is a system of components, each holding state of its own.
A durable orchestration engine is software that can stand an orchestrated system back up in a particular state, at any time in the future — to recreate it after a failure, or to continue a long-lived process. The durability question for a single workflow is what the existing products answer: the state of the process itself. The durability question for an orchestrated system is broader, because every component ready to help — agents, memory, retrieval, tools — carries state too. Durable orchestration means all of it persists: component state saved, memory tracked, context rebuildable, nothing tied to a particular machine. Can you stand up this system, in this state, at any time in the future? That is the question this layer exists to answer.
What kind of work is this? Not the calls — the processes. The work shows up as business processes with a shape in time: a customer onboarding that takes a week; an insurance claim that takes a month; a loan application waiting on an appraisal; a government permit waiting on a third-party response. This is the machinery of banking, government, healthcare and any other domain where a customer or citizen interacts with an organization and the interaction cannot finish in one sitting. What makes it hard is not any single step. It is that complicated context has to be maintained across all of the steps: what this customer filed, which documents have arrived, who reviewed what, what is still outstanding, what happens next. An agent that can wake up and respond to an event is not enough here — the context belongs to one particular customer or situation, and it has to be tracked reliably for weeks.
The naive way to build this is a scheduled job plus a state table: run the process a step at a time, write where it is into a database, and write code to pick it back up after each wait. Every step needs resume logic. Every wait needs a poller. Every crash needs recovery code. The glue is where it breaks, and the glue is most of the program. A durable execution engine does that work once, for every process it runs.
Temporal records the workflow as a durable history of events: started, called this, got that result, timer set, signal received. The workflow code runs as ordinary code, but it does not hold its own state — the history does. If the process dies mid-wait, the engine starts the code again on another machine and replays the history into it, handing back each recorded result as the code asks for it, until the code arrives exactly where it stopped. A three-day wait for a human review is not a process that stays alive for three days; it is a timer recorded in the history, and when the timer fires, the workflow resumes on whatever machine is available. From the developer’s point of view, the function pauses at that line and resumes weeks later with every variable intact, as if the failure never happened.38
Run the shape against something concrete: a loan application, from arrival to decision. Day one, the workflow starts, verifies the applicant’s data, and requests documents. Day three, the documents arrive, recorded as events delivered into the running workflow. Day four, the workflow requests a human underwriter review and sets a timer. Day nine, the review is recorded, the workflow orders a third-party appraisal, and sets another timer. Three weeks in, the appraisal returns, the decision step runs, the customer is notified, and the workflow closes. In between every one of those steps, machines rebooted, deployments shipped, instances died and were replaced. The workflow did not notice, because its state was never in any of those machines.
The durable orchestration engine described above does not exist yet, and it is worth recording what does. What exists today is durable execution — Temporal and Restate, whose mechanics are above — which persists the state of a workflow. Restate extends the idea to persistent agent sessions, and the agent frameworks persist what they can. Among the products reviewed here, I found no integrated system that automatically snapshots every external dependency, credential, memory layer, and runtime component as one portable state object — the components, the memory, the context, the whole arrangement, restorable on demand.39 Gas Town can look like it, and it persists more than it first appears to: task state, hooks, agent identity, history, and orchestration state live in its Beads and Git-backed storage, so worker restarts do not erase everything. What it does not provide is the portable whole — the state that survives is the state it was designed to persist, not a restorable snapshot of the entire arrangement. What most people run today is in the same position, from an agent on a laptop to agent clients spread across a bunch of machines: some state persists, and none of it is a restorable whole. A durable orchestration engine is an entirely separate approach: everything persisted, nothing tied to a particular VM, systems that go dormant and restore, work that continues reliably across restarts and regions. The foundation is arriving — the execution engines are real — but the layer itself is still to be built, and an orchestrator designing for the long term in 2026 assembles it by hand, if at all.
This is the last infrastructure layer in the survey — the sections after it describe assembled agents, not components — and it is the layer that keeps state distinct from memory. Memory (Section 5.5) is what the system knows — facts, decisions, corrections. Durable orchestration is where the system is — which step, waiting on what, what happens next. A system can have one without the other; an orchestrated business process needs both. And durability is the floor of this layer, not the ceiling: once a system’s state can be stood up at will, replaying a decision for an auditor, evaluating a system against its own history, and improving it over months rather than sessions all become buildable. Everything this book describes above this layer assumes it will exist.
1.9.1 Where durable orchestration shows up for an orchestrator
The encounters are decisions about what runs on this layer and what does not. Choosing the boundary: which parts of an orchestrated flow need durable guarantees (the approval wait, the payment step, the permit application) and which can stay ephemeral. Wiring the waits: the multi-day pause for human review is a durable timer, not a script that sleeps. And replaying: after a failure or an audit, the history is a first-class record of what step the workflow was on, what it had already done, and what it was waiting for, queryable step by step. The one discipline to know before adopting these engines: when a workflow is replayed to rebuild its state, the code has to make exactly the same decisions it made the first time. Anything that could answer differently on a second run — asking what time it is, generating a random number, calling an outside service directly — cannot live inside the workflow. It goes through the engine instead, which records the answer once and returns the same recorded answer on every replay; real external work goes in activities, which the engine retries.
This layer also carries the audit burden asymmetry that Chapter 7 develops: operational work — the runs that touch customers, money, and production — is exactly the work that needs durability most, because it is the work whose interruption is expensive and whose record is demanded later. The question to ask of any long-running design: if this workflow dies at step seven, what is lost, what resumes, and who finds out?
Temporal’s durability model per the company’s documentation and engineering writeups: workflow state is durably recorded as an event history and reconstructed by replay on worker failure, with durable timers, signals delivered as recorded events, and at-least-once activity execution with idempotency mechanisms. Temporal, “What is Durable Execution,” https://temporal.io/blog/what-is-durable-execution, and https://docs.temporal.io (including the workflow-determinism requirements: workflow code must be deterministic, and side effects belong in activities). The loan-application walkthrough is this book’s illustration of the documented mechanics, not a reported case study.↩︎
Temporal, “Durable AI,” https://docs.temporal.io/ai (verified August 29, 2026) — workflows resume automatically after crashes, network timeouts, or multi-day waits for human approval. Restate, “AI / Agents,” https://docs.restate.dev/ai (verified August 29, 2026) — durable execution, persistent sessions, approvals with pause and resume, cancel/kill/rollback for stuck agents.↩︎
- 1.1 The technology under discussion
- 1.2 The survey: connecting the taxonomy to the market
- 1.3 Inference engines and models
- 1.4 Retrieval systems
- 1.5 Memory systems
- 1.6 Tool interfaces and actuation
- 1.7 Identity and access control
- 1.8 Policy and permission systems
- 1.9 Durable orchestration engines
- 1.10 Task agents
- 1.11 Everything is called an agent
- 1.12 Multi-agent coding systems
- 1.13 The architecture of participation
- 1.14 The chapter in one picture