← All posts

Pausing Is Easy. Resuming Is the Hard Part.

When an AI process sleeps for a week and wakes up, is its context rechecked — or trusted as still true? That question is the line between a demo and infrastructure.

I got a sharp question in a comment thread recently: when an AI process sleeps for a week and wakes up, is its context rechecked — or is it trusted as still true?

That question is the line between a demo and infrastructure. Here's how we answer it, and the design decision the whole thing rests on.


Two things most systems blur together

When a process runs, two things are in play, and keeping them apart is the entire trick:

The context — one live document that every node reads and every action enriches. It's what is true now.

The snapshot — where the process was: its position in the graph, what it has already done, what it's still waiting on.

Snapshot  →  where I WAS   (position, history, pending waits)
Context   →  what is TRUE  (one live document, always current)

We never put the context inside the snapshot. The snapshot is where I was. The context is what is true now. They stay separate entities.


What happens when a process sleeps

A flow is a graph of nodes. Most nodes do work. Some are continue-after nodes — the points where a process is meant to sleep and wait for an event, a timer, or a human.

When the flow reaches one:

  • The node records exactly what it's waiting for.
  • The runtime gets a stop signal and the process ends — nothing sits in memory for a week.
  • The process is stored as scheduled, with a snapshot of its position — not its context.
  • The context stays where it always lived: its own entity, outside the snapshot.
… → Continue-After node
      • record what we're waiting for (event / timer / human)
      • runtime stops — process ends, nothing parked in memory
      • store as SCHEDULED with a snapshot of POSITION
      • context stays outside the snapshot, as always

What happens when it wakes

The runtime loads the position from the snapshot and continues from exactly the right spot. But the context is read live — it may already be different, because another process may have changed it while this one slept. We don't pretend it's still true. It's read fresh, every time.

Wake → load POSITION from snapshot (exact spot)
     → read CONTEXT live           (may already be different)
     → continue

Two guarantees the runtime does make:

  1. It resumes from the exact right point.
  2. It keeps a signature of the graph's shape, so if the flow itself changed while the process slept, it knows.

What a changed context or a changed flow means for the process is not something the runtime guesses. That's the designer's call.

The runtime promises position and integrity. The designer owns meaning.


Fan-out, and the join that waits

Inside a flow, a node can branch to several outbound edges at once, and they run concurrently — that's the default, not an opt-in. So the model needs a way to say wait for all of these before continuing. Every node carries a depends declaration for exactly that: it's the join, the barrier. If you've used a WaitGroup in Go or Promise.all in JavaScript, it's the same idea — expressed as an edge in the graph instead of code.


A loop is just a continue-after that points backward

Say a flow needs to run something every 24 hours. You don't add an end node. You add a continue-after whose next node points back to an earlier node. The process sleeps for a day, wakes, runs the segment again, and loops. Each turn is a fresh resumed process against the live context — a durable loop, not a thread parked for a day.

   ┌─────────────────────────────┐
   │                             │
   ▼                             │
[ work ] → … → [ continue-after ]┘   sleep 24h, wake, run again

And loops stay traceable across all of it, because the snapshot doesn't just remember where — it remembers which pass. The same generation bookkeeping that powers the join (so a join inside a loop waits on this iteration, not a stale one) is what the snapshot carries forward. Loops and joins compose because they run on the same state.


The designed refresh

This is where the shared context and the "designer owns meaning" line meet. When a context is shared, a value another process wrote can go stale under you. Instead of the runtime guessing whether that's safe, the designer drops in a continue-after that loops back to the node which re-reads and re-derives that part of the context. The process sleeps, wakes, and refreshes its own view against what's now true.

The revalidation the runtime deliberately doesn't do — the designer expresses it as a step in the graph, exactly where the person who understands the process can decide how strict it should be.


It goes all the way down

One of our node types runs sub-flows. So a continue-after — or a loop, or a refresh — can happen many layers deep inside a nested process, and the parent doesn't stay blocked waiting on it. The entire nested state is saved as scheduled the same way, and resumes against the same live context. Depth doesn't change the contract.


Closer to a language than to no-code

A lot of no-code tools buy their simplicity by taking expressiveness away — you get what the vendor decided you'd need, and the moment your problem doesn't fit, you're stuck. We went the other way. The goal was to keep as much of what real programming gives you as a canvas can hold — concurrency, joins, loops, sub-routines, shared state — through a small set of generic nodes composed together, not a big catalog of narrow ones.

Join      = WaitGroup            (wait for all branches)
Loop      = continue-after ↩     (edge that points backward)
Refresh   = designed re-read     (loop back to the reading node)
Sub-flow  = callable routine     (continue-after, layers deep)

That's why a join is a WaitGroup, a loop is an edge that points backward, and a sub-flow is a callable routine. They aren't features bolted on — they're the primitives you'd expect from a language, drawn on a graph.

And I'll be honest about the ceiling: a workflow canvas is still a paradigm, and every paradigm has its own restrictions. A graph is not a text file; some things trivial in code are awkward as edges, and some are deliberately out of reach. The aim isn't to pretend the canvas is a general-purpose language — it's to get as close as the paradigm honestly allows, and to keep the places it stops visible, instead of hiding them behind "you can't do that here."


Where the runtime stops and the supervisor starts

The runtime itself knows nothing about time. It runs to a stop and emits its snapshot — that's all. The scheduling, the week-long wait, the waking-up belong to the layer above it: FloMorphic, the supervisor. The runtime resumes cleanly; the supervisor decides when.

Runtime      → runs to a stop, emits a snapshot   (no clock, no meaning)
Supervisor   → schedules, waits, wakes it back up  (owns the clock)

The engine executes graphs and guarantees position. It doesn't own the clock, and it doesn't own meaning.

Someone reached this through Elixir supervision trees. We reached it through graph snapshots and a separate live context. Same conclusion, opposite directions:

Durable execution should be a property of the runtime — but what a resumed process should believe belongs to whoever designed it.


Repo: github.com/FloMorphic/getting-started

Concepts and docs: inflowenger.com/flomorphic

This is the mechanism underneath a lot of what this blog keeps circling. Resumption as an execution primitive is what lets a flow wait days on a human and resume as if no time had passed, and the same separation of durable state from live context is why context is the unit of execution, not just an input you hand the model.