Sagas: durable coordination
An aggregate can make a correct decision only from state it owns. Alice's account can decide whether she has the money to send, but it cannot decide whether Bob's account can take it. That rule belongs to Bob's account.
Aggregates cannot reference each other, and no transaction spans two of them. Each one processes its own commands against its own state, which is what keeps it safe to run, recover, and relocate independently. A rule that needs answers from two aggregates therefore needs a coordinator standing between them.
A saga is that coordinator: the event-sourced counterpart of a distributed transaction coordinator. It stores where the conversation has reached, listens for outcomes, and sends the next command. It cannot lock the participants or roll them back together; it drives the workflow forward one durable step at a time. It does not move another aggregate's rules into one large object.
Start with a workflow that can stop halfway
Sending 30 from Alice to Bob takes several steps:
- Alice's account records
TransferSent, which debits her; - the saga asks Bob's account to take the money with
ReceiveTransfer; - Bob's account replies with
TransferReceived, or rejects the transfer; - after a rejection, the saga asks Alice's account for a
RefundTransfer.
The process can stop after any step. A node may restart after Alice's account is debited but before Bob's account is credited.
Motivation: A chain of in-memory callbacks forgets where it was when the process stops. A saga stores that progress so the conversation can continue from its last durable step.
A saga is a state machine
The transfer workflow can be written as a table before any FCQRS code:
| Current state | Incoming event | Persisted next state | Command after persistence |
|---|---|---|---|
| not started | TransferSent |
Delivering |
ReceiveTransfer to the target |
Delivering |
TransferReceived |
Completed |
none; stop |
Delivering |
Rejected from the target |
Refunding |
RefundTransfer to the source |
Refunding |
TransferRefunded |
Completed |
none; stop |
The state names say what the workflow is waiting for. The state also carries the identifiers needed to repeat the next command after recovery.
The saga has two functions
An aggregate separates deciding an event from folding that event into state. A saga separates accepting an incoming event from performing the work associated with its persisted state.
handleEvent : incoming event + current saga state -> persisted next state
applySideEffects : persisted saga state + recovering -> transition + commands
handleEvent is the saga's event-driven transition function. For example, a rejection from the target
is accepted only while the saga is Delivering. It returns StateChangedEvent (Refunding transfer).
FCQRS stores that state change before running commands for the new state.
applySideEffects runs after a state change is durable. In Refunding, it returns RefundTransfer to
the originating account. FCQRS also calls it after recovery, with recovering = true, so the workflow
can safely resume from the state it last stored.
This ordering is the core guarantee:
incoming event
-> handleEvent chooses next state
-> next state is stored
-> applySideEffects issues commands
The state is stored first, so a restart has a durable answer to “what should happen next?”
Motivation: Keeping state selection separate from command emission makes the next action recoverable. The journal records the saga's intent before delivery introduces uncertainty.
StateChangedEvent and SagaTransition are different
The two functions return different control values because they answer different questions.
handleEvent normally returns:
StateChangedEvent next: accept the event and persistnextas saga progress;UnhandledEvent: this event is not valid for the current saga state.
applySideEffects returns commands together with one of these transitions:
| Transition | Meaning |
|---|---|
Stay |
Keep the current persisted state after issuing the commands |
NextState next |
Persist another state immediately, without waiting for an incoming event |
StopSaga |
Issue any returned commands, then complete and passivate the saga |
Delayed commands returned with StopSaga are still delivered as the saga's final act. The exception
is Self-targeted delayed commands, which FCQRS cancels with a warning, so a completed saga cannot be
resurrected by its own final message. Saga regions use Akka.NET's remembered entities: a passivated
saga restarts when a message arrives, and recovery would re-drive the same state. Final delayed
commands should target other entities.
Most workflows use StateChangedEvent for business events and Stay while waiting for a reply.
NextState is useful for an internal step that should advance immediately. Use it carefully: every
automatically entered state runs applySideEffects again, so a cycle of NextState transitions can
loop without waiting for new information.
Aggregates and sagas have different authority
| Aggregate | Saga |
|---|---|
| Receives commands | Reacts to events |
| Protects rules inside one identity | Coordinates independent identities |
| Emits outcome events | Emits follow-up commands and stores workflow states |
| Owns domain decision state | Owns only process progress |
The transfer saga does not inspect Bob's account. It sends ReceiveTransfer, and Bob's account
decides whether it can take the money. The saga reacts to that owner's answer.
Why a saga needs an explicit start
Ordinary actors already exist logically before a caller sends a command. A saga instance is different: it represents one particular workflow and should exist only when its starting business event occurs.
The StartOn predicate declares that boundary. For the transfer saga it matches TransferSent and
ignores every other account event. The aggregate that produced this event is the originator.
Commands such as toOriginator (RefundTransfer ...) route back to that exact account.
Only a persisted event can start a saga. A deferred reply is not journaled, so a repeated request
answered with DeferEvent or persistIf false does not start a second workflow.
Application code does not construct SagaStartingEvent directly. FCQRS wraps the matched originator
event in that internal envelope so the saga can retain:
- the event that created this workflow;
- the originator identity and version used by the startup handshake;
- the correlation and metadata context needed during recovery.
The distinction matters: TransferSent is the domain fact; SagaStartingEvent is FCQRS runtime
evidence about how this saga instance began.
Motivation:
StartOnis more than an event filter. It marks the creation boundary FCQRS needs to install the new saga before releasing the event that gives it work.
Starting is itself a race
The first event must create the saga and also reach it. Publishing before the new saga subscribes can lose the one event that moves it out of its initial state.
FCQRS closes that race with a handshake between the originator and the saga:
StartOnmatches an event the originator is about to store;- the originator sends the saga its starting message and waits before storing the event;
- the saga stores its starting envelope and subscribes to the event's correlation topic;
- the saga tells the originator it is ready;
- the originator stores and publishes the domain event;
handleEvent(Startin C#) receives it with no user-defined state yet and stores the first state.
Fcqrs.wireSagaStarters in F#, or AddSaga in C# with each saga's StartsOn, installs these start
rules. A saga is named after its originator's aggregate ID and the event's correlation ID, so it starts
once per correlation ID and aggregate. A second start is refused, and the command that would cause it
stores nothing (Correlation IDs). An empty F# application still calls wireSagaStarters api [] so runtime startup follows one
explicit path.
Waiting costs no thread
In step 2 the originator waits as an actor does: it handles the saga's readiness as a message and stashes the commands that arrive meanwhile, then processes them in order once the event is stored. It holds no thread while it waits, so many aggregates can start sagas at the same moment. An event that starts no saga is stored at once.
A saga that restarts before it answers loses the reference it would answer to. While the originator
waits, it sends the starting message again from time to time, and a saga that already stored its start
only answers. The originator waits at most config:akka:fcqrs:saga-start-timeout, 30 seconds by
default. Past that, FCQRS stops the process, because the
saga may already have stored its start while the event it waits for was never stored.
Concurrent saga starts are bounded by how fast the journal stores the sagas' first records, not by threads. Commands whose events start sagas wait for those writes, so their callers' command timeout applies to the whole handshake too.
Resumption means re-drive, not rewind
FCQRS stores each accepted saga state transition. On restart it loads a snapshot if available, replays
later state changes, restores the starting-event context, and subscribes the saga again. It then calls
applySideEffects for the recovered state with recovering = true.
Suppose the last durable state is Delivering. FCQRS knows the saga must send ReceiveTransfer to
Bob's account; it cannot know whether the previous process sent that command just before it stopped.
The correct recovery action is therefore to re-drive the state with a retry-safe command.
stored state: Delivering t1
unknown: was ReceiveTransfer t1 delivered before the crash?
resume: send ReceiveTransfer t1 again safely
Bob's account treats a repeated transfer ID as the same business outcome: it replies with the first
answer again and credits the money once. At an external boundary, use a stable idempotency key or query the operation's status.
The recovering flag can select that status-check path when repeating the normal command is unsafe.
Returning no commands for every recovered state is usually wrong. It strands a saga when the original command was never delivered. Resumability comes from durable state plus safe re-delivery, not from exactly-once messaging.
Motivation: At a crash boundary, the sender cannot prove whether its last command arrived. Persisting the intended step and making that command safe to repeat turns this uncertainty into a workflow that can resume.
The starting event also protects resumption
A saga starts before its originator stores the starting event: the saga-start handshake subscribes the
saga first, so it cannot miss the originator's publication. If the originator's write then fails, the
saga holds a starting event that was never stored. A saga recovered before it leaves Started therefore
asks its originator whether the starting event is stored.
The originator compares the event stored at the starting event's version with the starting event, by
version and event ID. The ID matters after a failed save, when another event can hold the same version.
If that version is the originator's latest, it compares with the event it holds in memory. If it has
stored later events, it reads that one event from its journal, through the same journal plugin it
recovers from, and retries while the journal cannot be read; the saga waits in Started meanwhile. An
originator recovered from a snapshot with no later events does not know which event holds its version,
so it compares only the version.
If the starting event is stored, the originator sends it to the saga, and the saga continues as it would
have without the restart. If another event holds that version, or none does, the originator answers
with an AbortedEvent, visible as an Abort: span and in the message-flow logs, and the saga
passivates. Both answers go only to the saga that asked. An AbortedEvent that arrives after the saga
has stored a newer state answers an earlier check, and the saga ignores it.
This check protects the FCQRS actor conversation. It does not make an independent database, payment provider, or HTTP service part of the saga journal transaction. External operations still require their own idempotency and reconciliation rules.
Delivery and persistence cannot share one transaction
For any outgoing saga command, failure may occur:
- before delivery;
- after delivery but before the receiver acts;
- after the receiver acts but before its event reaches the saga;
- after the saga receives the event but before its next state is stored.
No local transaction spans the saga journal and every participant. Design each step so repetition is
safe, give every wait a timeout, and make failed or compensating states visible to operators. A
waiting state can declare its timeout with StayExpecting: the framework re-sends the declared
commands on a schedule and, past a deadline anchored to the persisted state-entry time, delivers an
ExpectationExhausted message the saga must answer with a transition.
Write a saga shows the declared and the hand-rolled form.
Compensation is a new action
A distributed workflow cannot generally erase work that already succeeded. A refund is not time travel; it is another command with its own outcome and possible failure. A sent e-mail may have no useful compensation at all.
Model compensation explicitly: what triggers it, which state stores its progress, whether it is idempotent, and what happens if it fails. Some workflows should stop for human resolution rather than pretend every action is reversible.
Decide whether you need a saga
Use a saga when work crosses independent consistency boundaries and its progress must survive restart. Use an aggregate command when one owner can decide the rule. Use an async effect for best-effort work that may safely be lost and does not need durable progress.
The tutorial's transfer runs this saga, and Write a saga is the compact F# and C# recipe. Consistency and recovery places saga resumption beside the other durable boundaries.