Skip to main content

Sagas & compensations

The Saga pattern trades distributed transactions for compensating actions: each step registers how to undo itself, and on failure the engine rolls completed steps back in reverse order.

Action-level compensation​

The primary style attaches an undo to each step:

public function handle(string $orderId): void
{
$this->action(ChargeCard::class, $orderId)
->compensateWith(RefundCard::class, $orderId)
->run();

$this->action(ReserveStock::class, $orderId)
->compensateWith(ReleaseStock::class, $orderId)
->run();

// If this throws, ReleaseStock then RefundCard run automatically.
$this->action(ShipOrder::class, $orderId)->run();
}

Compensation can also be a closure:

$this->action(MakeReservation::class, $id)
->compensateWith(fn () => Reservation::release($id))
->run();

A child workflow takes compensateWith() as well. Its compensation joins the same stack once the child completes.

Grouped sagas​

saga() expresses a compensation boundary explicitly and exposes group-level policies:

use DiscoveryUkraine\SagaLaraFlow\Enums\CompensationFailurePolicy;

$this->saga()
->onCompensationFailure(CompensationFailurePolicy::Continue)
->compensateInParallel()
->step(ChargeCard::class, $orderId)->compensateWith(RefundCard::class, $orderId)
->step(ReserveStock::class, $orderId)->compensateWith(ReleaseStock::class, $orderId)
->run();
  • compensateInParallel() runs the group's undos concurrently (a single rollback level: together via Bus::batch when queued, sequentially under runSync).
  • compensateStepOnSelfFailure() also compensates a step that itself failed (for non-atomic actions that may leave partial effects) — such compensations must be idempotent.

Failure policies​

CompensationFailurePolicy:

  • Stop (default) — halt the rollback on the first compensation that does not complete.
  • Continue — keep rolling back even if one undo does not complete.

Precedence is action > group > config (sagas.default_compensation_failure_policy), and a child's onCompensationFailure() stands where an action's does. If a compensation itself fails under Stop, a CompensationFailedException surfaces.

"Does not complete" covers more than a throw. A compensation whose worker was killed, or whose job never arrived, is left Pending or Running when its level finishes, and Stop halts the rollback for that too: unwinding further on top of a step that may still stand is what the policy exists to prevent. The run records it under flow_run.exception['compensation'] as a CompensationUnfinishedException rather than a CompensationFailedException — there is no cause to report, since it never got far enough to have one. A rollback therefore never finalizes looking clean while one step was silently left undone.

The exception states what was observed: the compensation had not finished when its rollback level ended. Usually it never does. It can also mean an at-least-once queue closed the batch a moment before the live worker recorded its success — see Reclaim & recovery for that race and how to spot it in your logs.

To have such a compensation retried rather than only reported, enable sagas.reclaim.stale_running — see Reclaim & recovery.

While a rollback runs​

A run rolling back is in Cancelling, which is not terminal — and nothing new begins under it. The plan is settled by then, so anything started afterwards would finish outside it: its compensation in no stack, never run, under a run reporting a complete unwind. That covers a step whose job arrives to claim its row, a child workflow and a side effect a pass still replaying reaches for the first time, and a replacement the doctor would otherwise send.

That is what moving the run first is for. A plan is drawn before the move as well, but only to find out whether one can be drawn at all: it writes nothing, so a run whose replay throws is left exactly where it was found. The plan that is unwound is the one made afterwards, and it holds the step whose owed queue attempt completed while the first was being drawn.

The later plan is not always the longer one. An attempt that claimed a failed step in the same gap leaves it Running, and what compensateStepOnSelfFailure() registers for a failed step is not what a replay reads off a running one — and a parallel block can lose an ordinal that way while finding another in the same pass. So an ordinal the later plan is missing is restored from the earlier one rather than the later plan being discarded, and the difference is journalled as replan_incomplete. A replay that throws leaves nothing to merge and is journalled as replan_failed; the rollback then goes ahead on the plan already in hand.

Settling what already started is the other question, and it carries on: a step past its own deadline is still expired, and the rollback's own compensations still run. See Statuses for the boundary in full.

Manual compensation​

You can trigger a rollback from outside the workflow through the handle:

SagaFlow::loadFlow($runId)->compensate(); // roll back completed steps, then cancel

The rollback is planned before anything is undone: the handle replays handle() only to learn which completed steps and children carry compensations. Every seam is guarded for that pass, so it runs no business logic, starts no work, and settles no step — one it has not seen is a place to stop, not a place to schedule. (Your own tag() calls still rewrite their rows, as they do on every replay.) The pass reads the history from the write connection, so a lagging read replica cannot cut the plan short. It ends on the frontier it stopped at, and on a step failure, an expiry, a signal timeout or an awaited child's failure, expiry or cancellation already in the run's history — each of them raised by the seam that read it. Anything else is a fault, not an ending: an argument expression reading a record that has since been deleted, say, or your own code raising one of those business exceptions itself. The plan is then incomplete, so compensate() surfaces the throw and leaves the run as it found it, rather than unwinding part of it and reporting a finished rollback. Fix the cause and call it again. That is the plan drawn before the run is moved; a throw from the one made after it is journalled instead, because by then the run has been taken and there is a plan to unwind either way.

Not inside a transaction of your own

The compensations execute before your transaction closes, so a rollback afterwards discards the record while the undo work is already done — and leaves the run compensatable a second time. See what a host transaction leaves behind.

Postponing a rollback​

Not every failure deserves a rollback. When a step failed only because the world was not ready — a declined card, a service still provisioning — retryOnSignal() parks that step and waits instead of compensating, and re-runs it alone when the signal arrives. Compensation happens only if the retry policy eventually gives up. See Retry on signal.