Sagas: durable coordination

An aggregate can make a correct decision only from state it owns. Alice's account can decide whether she has the money to send, but it cannot decide whether Bob's account can take it. That rule belongs to Bob's account.

Aggregates cannot reference each other, and no transaction spans two of them. Each one processes its own commands against its own state, which is what keeps it safe to run, recover, and relocate independently. A rule that needs answers from two aggregates therefore needs a coordinator standing between them.

A saga is that coordinator: the event-sourced counterpart of a distributed transaction coordinator. It stores where the conversation has reached, listens for outcomes, and sends the next command. It cannot lock the participants or roll them back together; it drives the workflow forward one durable step at a time. It does not move another aggregate's rules into one large object.

Start with a workflow that can stop halfway

Sending 30 from Alice to Bob takes several steps:

  1. Alice's account records TransferSent, which debits her;
  2. the saga asks Bob's account to take the money with ReceiveTransfer;
  3. Bob's account replies with TransferReceived, or rejects the transfer;
  4. after a rejection, the saga asks Alice's account for a RefundTransfer.

The process can stop after any step. A node may restart after Alice's account is debited but before Bob's account is credited.

Motivation: A chain of in-memory callbacks forgets where it was when the process stops. A saga stores that progress so the conversation can continue from its last durable step.

A saga is a state machine

The transfer workflow can be written as a table before any FCQRS code:

Current state Incoming event Persisted next state Command after persistence
not started TransferSent Delivering ReceiveTransfer to the target
Delivering TransferReceived Completed none; stop
Delivering Rejected from the target Refunding RefundTransfer to the source
Refunding TransferRefunded Completed none; stop

The state names say what the workflow is waiting for. The state also carries the identifiers needed to repeat the next command after recovery.

A saga receives an event, persists its next workflow state, sends commands for that state, and can recover and safely repeat those commands

The saga has two functions

An aggregate separates deciding an event from folding that event into state. A saga separates accepting an incoming event from performing the work associated with its persisted state.

handleEvent      : incoming event + current saga state -> persisted next state
applySideEffects : persisted saga state + recovering   -> transition + commands

handleEvent is the saga's event-driven transition function. For example, a rejection from the target is accepted only while the saga is Delivering. It returns StateChangedEvent (Refunding transfer). FCQRS stores that state change before running commands for the new state.

applySideEffects runs after a state change is durable. In Refunding, it returns RefundTransfer to the originating account. FCQRS also calls it after recovery, with recovering = true, so the workflow can safely resume from the state it last stored.

This ordering is the core guarantee:

incoming event
  -> handleEvent chooses next state
  -> next state is stored
  -> applySideEffects issues commands

The state is stored first, so a restart has a durable answer to “what should happen next?”

Motivation: Keeping state selection separate from command emission makes the next action recoverable. The journal records the saga's intent before delivery introduces uncertainty.

StateChangedEvent and SagaTransition are different

The two functions return different control values because they answer different questions.

handleEvent normally returns:

  • StateChangedEvent next: accept the event and persist next as saga progress;
  • UnhandledEvent: this event is not valid for the current saga state.

applySideEffects returns commands together with one of these transitions:

Transition Meaning
Stay Keep the current persisted state after issuing the commands
NextState next Persist another state immediately, without waiting for an incoming event
StopSaga Issue any returned commands, then complete and passivate the saga

Delayed commands returned with StopSaga are still delivered as the saga's final act. The exception is Self-targeted delayed commands, which FCQRS cancels with a warning, so a completed saga cannot be resurrected by its own final message. Saga regions use Akka.NET's remembered entities: a passivated saga restarts when a message arrives, and recovery would re-drive the same state. Final delayed commands should target other entities.

Most workflows use StateChangedEvent for business events and Stay while waiting for a reply. NextState is useful for an internal step that should advance immediately. Use it carefully: every automatically entered state runs applySideEffects again, so a cycle of NextState transitions can loop without waiting for new information.

Aggregates and sagas have different authority

Aggregate Saga
Receives commands Reacts to events
Protects rules inside one identity Coordinates independent identities
Emits outcome events Emits follow-up commands and stores workflow states
Owns domain decision state Owns only process progress

The transfer saga does not inspect Bob's account. It sends ReceiveTransfer, and Bob's account decides whether it can take the money. The saga reacts to that owner's answer.

Why a saga needs an explicit start

Ordinary actors already exist logically before a caller sends a command. A saga instance is different: it represents one particular workflow and should exist only when its starting business event occurs.

The StartOn predicate declares that boundary. For the transfer saga it matches TransferSent and ignores every other account event. The aggregate that produced this event is the originator. Commands such as toOriginator (RefundTransfer ...) route back to that exact account.

Only a persisted event can start a saga. A deferred reply is not journaled, so a repeated request answered with DeferEvent or persistIf false does not start a second workflow.

Application code does not construct SagaStartingEvent directly. FCQRS wraps the matched originator event in that internal envelope so the saga can retain:

  • the event that created this workflow;
  • the originator identity and version used by the startup handshake;
  • the correlation and metadata context needed during recovery.

The distinction matters: TransferSent is the domain fact; SagaStartingEvent is FCQRS runtime evidence about how this saga instance began.

Motivation: StartOn is more than an event filter. It marks the creation boundary FCQRS needs to install the new saga before releasing the event that gives it work.

Starting is itself a race

The first event must create the saga and also reach it. Publishing before the new saga subscribes can lose the one event that moves it out of its initial state.

FCQRS closes that race with a handshake between the originator and the saga:

  1. StartOn matches an event the originator is about to store;
  2. the originator sends the saga its starting message and waits before storing the event;
  3. the saga stores its starting envelope and subscribes to the event's correlation topic;
  4. the saga tells the originator it is ready;
  5. the originator stores and publishes the domain event;
  6. handleEvent (Start in C#) receives it with no user-defined state yet and stores the first state.
The originator matches TransferSent with StartOn, sends the transfer saga its starting message, and stores and publishes the event only after the saga reports it is ready

Fcqrs.wireSagaStarters in F#, or AddSaga in C# with each saga's StartsOn, installs these start rules. A saga is named after its originator's aggregate ID and the event's correlation ID, so it starts once per correlation ID and aggregate. A second start is refused, and the command that would cause it stores nothing (Correlation IDs). An empty F# application still calls wireSagaStarters api [] so runtime startup follows one explicit path.

Waiting costs no thread

In step 2 the originator waits as an actor does: it handles the saga's readiness as a message and stashes the commands that arrive meanwhile, then processes them in order once the event is stored. It holds no thread while it waits, so many aggregates can start sagas at the same moment. An event that starts no saga is stored at once.

A saga that restarts before it answers loses the reference it would answer to. While the originator waits, it sends the starting message again from time to time, and a saga that already stored its start only answers. The originator waits at most config:akka:fcqrs:saga-start-timeout, 30 seconds by default. Past that, FCQRS stops the process, because the saga may already have stored its start while the event it waits for was never stored.

Concurrent saga starts are bounded by how fast the journal stores the sagas' first records, not by threads. Commands whose events start sagas wait for those writes, so their callers' command timeout applies to the whole handshake too.

Resumption means re-drive, not rewind

FCQRS stores each accepted saga state transition. On restart it loads a snapshot if available, replays later state changes, restores the starting-event context, and subscribes the saga again. It then calls applySideEffects for the recovered state with recovering = true.

Suppose the last durable state is Delivering. FCQRS knows the saga must send ReceiveTransfer to Bob's account; it cannot know whether the previous process sent that command just before it stopped. The correct recovery action is therefore to re-drive the state with a retry-safe command.

stored state: Delivering t1
unknown:      was ReceiveTransfer t1 delivered before the crash?
resume:       send ReceiveTransfer t1 again safely

Bob's account treats a repeated transfer ID as the same business outcome: it replies with the first answer again and credits the money once. At an external boundary, use a stable idempotency key or query the operation's status. The recovering flag can select that status-check path when repeating the normal command is unsafe.

Returning no commands for every recovered state is usually wrong. It strands a saga when the original command was never delivered. Resumability comes from durable state plus safe re-delivery, not from exactly-once messaging.

Motivation: At a crash boundary, the sender cannot prove whether its last command arrived. Persisting the intended step and making that command safe to repeat turns this uncertainty into a workflow that can resume.

The starting event also protects resumption

A saga starts before its originator stores the starting event: the saga-start handshake subscribes the saga first, so it cannot miss the originator's publication. If the originator's write then fails, the saga holds a starting event that was never stored. A saga recovered before it leaves Started therefore asks its originator whether the starting event is stored.

The originator compares the event stored at the starting event's version with the starting event, by version and event ID. The ID matters after a failed save, when another event can hold the same version. If that version is the originator's latest, it compares with the event it holds in memory. If it has stored later events, it reads that one event from its journal, through the same journal plugin it recovers from, and retries while the journal cannot be read; the saga waits in Started meanwhile. An originator recovered from a snapshot with no later events does not know which event holds its version, so it compares only the version.

If the starting event is stored, the originator sends it to the saga, and the saga continues as it would have without the restart. If another event holds that version, or none does, the originator answers with an AbortedEvent, visible as an Abort: span and in the message-flow logs, and the saga passivates. Both answers go only to the saga that asked. An AbortedEvent that arrives after the saga has stored a newer state answers an earlier check, and the saga ignores it.

This check protects the FCQRS actor conversation. It does not make an independent database, payment provider, or HTTP service part of the saga journal transaction. External operations still require their own idempotency and reconciliation rules.

Delivery and persistence cannot share one transaction

For any outgoing saga command, failure may occur:

  • before delivery;
  • after delivery but before the receiver acts;
  • after the receiver acts but before its event reaches the saga;
  • after the saga receives the event but before its next state is stored.

No local transaction spans the saga journal and every participant. Design each step so repetition is safe, give every wait a timeout, and make failed or compensating states visible to operators. A waiting state can declare its timeout with StayExpecting: the framework re-sends the declared commands on a schedule and, past a deadline anchored to the persisted state-entry time, delivers an ExpectationExhausted message the saga must answer with a transition. Write a saga shows the declared and the hand-rolled form.

Compensation is a new action

A distributed workflow cannot generally erase work that already succeeded. A refund is not time travel; it is another command with its own outcome and possible failure. A sent e-mail may have no useful compensation at all.

Model compensation explicitly: what triggers it, which state stores its progress, whether it is idempotent, and what happens if it fails. Some workflows should stop for human resolution rather than pretend every action is reversible.

Decide whether you need a saga

Use a saga when work crosses independent consistency boundaries and its progress must survive restart. Use an aggregate command when one owner can decide the rule. Use an async effect for best-effort work that may safely be lost and does not need durable progress.

The tutorial's transfer runs this saga, and Write a saga is the compact F# and C# recipe. Consistency and recovery places saga resumption beside the other durable boundaries.