← All posts

FloMorphic v0.4.3: Windows, and a Decision Node That Reads Its Documents

Two halves, and they have nothing to do with each other. FloMorphic installs on Windows now, and so do the plugins written for it. And the Jev node became the AI Decision node, which takes the evidence a decision rests on. Building that took an afternoon. Believing it took two controlled experiments, one of which proved nothing.

FloMorphic v0.4.3 is out, tagged v0.4.3 and latest. It has two halves that have nothing to do with each other, which is roughly what a release looks like when the project is still finding its shape: one half is an installer, the other is the node I spent the last two weeks on.

The installer half is that FloMorphic runs on Windows now, and so do the plugins written for it. The node half is that the Jev node is no longer called jev. It's called ai-decision, and it now takes evidence — the retrieved chunks a decision rests on, injected per run.

The rename is the boring part. The evidence is where it got interesting, because building it took an afternoon and believing it took two controlled experiments, and I'll spend most of this post on those.

All of that is in v0.4.3, which means it's in the image you'd install today — the rename, the evidence rows, the reference data, the retry budget, and every run and experiment below. One section near the end is the exception, and it's marked.

If you just want the thing:

curl -fsSL https://raw.githubusercontent.com/FloMorphic/getting-started/main/install.sh | bash
irm https://raw.githubusercontent.com/FloMorphic/getting-started/main/install.ps1 | iex

The Windows half

FloMorphic is containers, and Docker on Windows is Docker Desktop on the WSL 2 backend. So the prerequisite was never "Docker" — it was Docker and the Linux environment underneath it, and the one-liner couldn't install either. For months the honest answer to "can I run this on Windows?" was "yes, once you've set up three things yourself."

install.ps1 is that missing step. It isn't a port of install.sh; it's the step before it. It checks the Windows host, asks before it installs anything, puts WSL 2 and Docker Desktop in place, and then runs the very same install.sh inside the distro — so Windows and Linux end up with one stack from one source of truth, and there's no second installer to keep in sync.

What it actually checks, in order:

StepWhat it does
Windows hostbuild 19041+, architecture, and an explicit warning when virtualization is off in firmware
WSL 2wsl --install with your consent, Ubuntu if there's no distro, converts a WSL 1 distro, tells you when a reboot is needed
Docker Desktopwinget, else the installer from docker.com; started, waited on, and a walk-through of Settings → Resources → WSL integration
install.shrun in the distro, with every env var you set in PowerShell forwarded through

That virtualization warning is there because it's the classic silent WSL 2 failure: everything reports success, and then nothing works. Docker Desktop is Docker Inc.'s product under its own licence, so the script installs it only when you say yes. Anything needing administrator rights asks before relaunching itself elevated, -Yes makes the whole run unattended, and -NoInstall checks a host without changing it.

The stack lands in the distro's own filesystem (~/flomorphic), not on C:. That's not a preference — the SQLite database is bind-mounted, and SQLite over the /mnt/c bridge is slow and prone to locking errors. Docker Desktop publishes the ports to Windows anyway, so http://localhost:8088 works in the Windows browser with nothing else to configure.

A plugin is a different problem

A plugin isn't a container we hand you. It's a process you run — go build, npm start, docker — and two of those three are native on Windows. So the Windows path for a plugin isn't a hand-off into WSL at all. The API now renders a real PowerShell installer and lifecycle helper beside the bash pair, with the same four steps in the same order (clone, write the dotenv, drop the helper, build and start) and the same verbs: build, start, stop, restart, status, logs. A plugin that installs on Linux installs on Windows, with no WSL in the picture.

Three things the bash version never had to care about, all found the hard way:

  • The generated script never calls exit on a failure path. Under irm | iex that closes the operator's PowerShell window with the error still unread. A throw caught by a wrapper does the job instead.
  • The dotenv is written UTF-8 with no BOM. PowerShell 5.1's Set-Content -Encoding UTF8 emits one, and a BOM makes the first key unparseable to every dotenv reader there is.
  • stop kills the process tree, because npm start runs node as a child. Stopping only npm leaves the plugin itself connected to Infra, cheerfully serving actions nobody thinks are running.

The Extensions page offers a Linux/macOS ⟷ Windows toggle on the install hand-off, guessed from the browser you're on and then remembered, on the theory that whoever installs one plugin from Windows will install the next one there too. Read the script before running it now also shows the lifecycle helper, which carries no credential and is worth reading.

That's the whole installer half. The rest of this post is the node.


Why the node stopped being named after a vendor

Two weeks ago I added Jev to the palette and wrote that a bounded-decision model slots into FloMorphic without the runtime learning anything. That held. What I didn't expect was how fast the thing I'd named the node after stopped being the only one of its kind.

Laya is an open System One decision model — Apache-2.0, from Convai Innovations, runnable on your own hardware, in ONNX, even in a browser. The part that matters for a node: it serves POST /v1/systemone and returns Jev-shaped choice / score / noul answers. Same request body, same reply.

Then, while I was building this, I went looking for where my own API key had come from and found aggregators — a gateway fronting a dozen deciders behind one key, OpenRouter-style, with a catalogue like this:

typesafe/jev-1.13 · convaiinnovations/laya · cloudflare/clef · cloudflare/clef-flash
perplexity/pplx-decider-v1-27b · liquid/d1 · togethercomputer/tev1-4b-experimental
inception/mercury-decide · upstage/solar-decide · respan/span-01 · jaredpalmer/kev-4b

Two weeks ago I was writing about a bounded-decision model. That list is a market. Cloudflare, Perplexity, Together, Liquid, Upstage, and vendor-prefixed ids — which is what a category looks like once it has enough entrants to need namespacing.

So the node was never really about Jev. It was about a protocol, and the protocol turned out to be the thing worth naming:

BeforeNow
Node kindjevai-decision
Label on canvasJevAI Decision
Hosted JevAPI key, doneAPI key, done
Local Layanot expressiblepoint url at your server, name its model
A dozen others via an aggregatornot expressibleits url + key, vendor-prefixed model

A settings profile now decides which model answers, and nothing else in the flow changes. The canvas, the questions, the ports, the record: identical. Swapping Jev for Laya is one field, and so is swapping either for whatever ships next quarter.

That's the point I'd defend hardest. A node that bakes in a vendor ages at the speed of that vendor. This one ages at the speed of a protocol, and protocols tend to outlive the companies that publish them.

Two small things fell out of making that real. A local endpoint usually has no API key, so the node stopped requiring one and now omits the Authorization header entirely rather than sending an empty bearer, which some servers reject outright. And a model id is only guessable for the hosted default, so naming your own url now requires naming a model — a readable error here beats a 422 from the far side.

Flows saved as jev still compile. The old kind is kept as an alias at every layer it touches, which is the part of a rename nobody blogs about and everybody needs. One visible consequence on an upgraded install: the node registry lists both the old Jev builtin and the new AI Decision one, because the seed is keyed by name and the old row isn't pruned. New nodes come from the new row; the stale one is safe to delete.

One upgrade note worth reading

That node's default endpoint changed. It used to default to thejevai.com, which is an aggregator; it now defaults to TypeSafe's own api.typesafe.ai. The aggregator is still a supported url, and still the only way to reach several of those models under one account, but it's now something a profile opts into rather than inherits — a workflow product shouldn't route a customer's state through an unaffiliated third party because nobody chose otherwise.

The two differ in more than a hostname. They bill differently (per input token vs credits), they need their own keys, and an aggregator wants a vendor-prefixed model id where TypeSafe takes its own aliases. Keys are not interchangeable, so a profile that was relying on the old default has to name that url explicitly beside its key now. The node decodes both reply shapes — flat from TypeSafe, enveloped from an aggregator, which is where credits_used and elapsed_ms come from — so a metered gateway still accounts for itself on the canvas.

The profile also carries a retry budget now, as the LLM and HTTP nodes already did. max_retries is how many further attempts a failed call gets, and only for the two statuses the service asks callers to back off on: 429 and 529. A 401, a 422 or a malformed reply comes back immediately, because a second identical request can't fix any of them. Waits honour Retry-After in both documented forms and otherwise back off from 500 ms, capped at 8 seconds, because a decision node sits on the hot path and one that parks a flow for minutes isn't helping. Absent means 2; an explicit 0 means decide once, for a flow where a late decision is worse than none.


The thing I got wrong about the API

I assumed a decision API with RAG ambitions would have an evidence parameter. It does not.

The entire request body is three fields:

{ "model": "…", "state": "…", "questions": { … } }

No evidence, no context, no documents. I checked the schema twice because I didn't believe it.

What it does have is a state that accepts a string, a JSON object or an array, and a question that can point at any part of that state by backticked path. From the primitives reference: "the state is often a JSON object with several parts: a conversation, a record, a policy. When a question is about one of those parts, name it in the instructions with a dot-and-index path to its key, including the backticks." Their own example is `ticket.messages[0].text`.

So evidence doesn't go beside the state. It goes inside it. The node now assembles this:

"state": {
  "case": { "customer": "ABC Ltd", "problem": "Early contract termination", … },
  "evidence": [
    { "source": "contract-s12.pdf",   "text": "Section 12 allows termination with 30 days…" },
    { "source": "regulation-s44.pdf", "text": "Section 44 requires a written statement…" },
    { "source": "precedent-2025.pdf", "text": "Harlow Systems v. Denby Holdings (2025)…" }
  ]
}

and a question cites whichever part it means: `case.problem`, `evidence[0].text`.

With no evidence rows, the state is sent exactly as it was before any of this existed. That mattered more than it sounds — it's why every already-drawn flow kept its behaviour.

Keep both filtered, by the way. The state and all the questions share roughly 64k tokens, with the state plus the longest single question inside about 32k, and accuracy falls as the state fills with material the questions don't need. Evidence is not a place to dump a document. Retrieve, filter, then inject.

Is that shape actually sanctioned?

Worth addressing, because I went looking for a blessed "evidence" example and there isn't one. No page in either the Jev or the Laya docs shows a field called evidence, and the word RAG appears attached to a different pattern — scoring retrieved passages one at a time to decide which reach a generator.

But the shape itself is documented by example. The State page presents this as one state:

{
  "ticket": { "subject": "…", "messages": [ {"from": "customer", "text": "…"}, … ] },
  "order":  { "id": "A-104", "charges": [ {"amount_usd": 49, "status": "captured"}, … ] },
  "refund_policy": "Duplicate charges are eligible for a refund."
}

An object holding nested arrays of objects, plus a policy. {case, evidence: [{source, text}]} is the same construction with different nouns. The page's own framing is "think of state as the material you would present to a panel of experts before asking them to make a judgment", and it says plainly that the state holds "the content and supporting facts". That is evidence injection; it simply isn't named.

So evidence is FloMorphic's convention on top of documented primitives, not a vendor field — and the live runs below are why I'm comfortable with it rather than taking the docs' word for it.


The full scenario

Here's the flow I built to exercise all of it. A contract-termination review, which is the kind of bounded decision that's genuinely hard to rule-engine, because the answer depends on documents.

The case. ABC Ltd gave 30 days' written notice to terminate a 36-month managed-services contract, 14 months in. Annualised value £82,000.

The evidence, three retrieved chunks, each pulled from the run's Context by path:

  • contract-s12.pdf → {{$.kb.contract_clause}} — termination for convenience on 30 days' notice
  • regulation-s44.pdf → {{$.kb.regulation}} — supplier must issue a statement of cause
  • precedent-2025.pdf → {{$.kb.precedent}} — a customer need not prove breach to rely on a convenience clause

Two questions. The first routes; the second is data only:

{
  "id": "termination_valid",
  "type": "noul",
  "instructions": "Does `case.problem` qualify for termination under `evidence[0].text`, and does the notice given satisfy `evidence[1].text`? Weigh `policy` as binding.",
  "references": [{ "name": "policy", "value": "{{$.kb.policy}}" }],
  "options": [{ "name": "yes", … }, { "name": "no", … }]
}

That references row is the second thing the node learned. The API accepts instructions as an object — the question in one field, the data it needs in the others, cited by backticked name. Writing a JSON object in a drawer is miserable, so the drawer keeps the question as text and collects named reference rows beside it; the compiler assembles the two into the object form. policy here is an internal rule that exists nowhere in the evidence: terminations above £50,000 require legal sign-off.

The second question, risk, is a score with three levels and route: false — it contributes no ports, just a distribution for a later node to read.

On the canvas:

   context ──►  ┌──────────────┐── termination_valid.yes ──►  Permitted         (skipped)
                │ AI Decision  │── termination_valid.no ───►  Not permitted     ✔ fired
                │   (1 call)   │── _exception ─────────────►  Decision failed   (skipped)
                └──────────────┘
                   risk → data only, no port

What the run said

One call, one credit, 2.3 seconds:

"termination_valid": { "answer": "no",  "confidence": 0.71, "probabilities": { "no": 0.71, "yes": 0.29 } },
"risk":              { "answer": "low", "confidence": 0.88, "score": 0.08,
                       "probabilities": { "low": 0.93, "medium": 0.05, "high": 0.02 } },
"routed": ["termination_valid.no"],
"usage":  { "input_tokens": 862, "output_tokens": 34 },
"credits_used": 1

Not permitted, at 0.71, with low legal risk at 0.88. Which is a defensible read of a genuinely split case: the clause allows it, the precedent supports it, and the £82,000 trips an internal policy that says a lawyer has to sign first. I'd deliberately built a scenario where both answers were arguable, because a test case the model can ace tells you nothing.

Two branches that could have fired didn't, and the runtime's log names every one of them. The whole distribution landed in the Context, so a rule node downstream can act on 0.71 without a redeploy.

Satisfying, and completely insufficient as evidence that the evidence did anything.


The uncomfortable question

A model reading only the case — 30 days' notice, £82,000, 14 months in — could produce a plausible answer to every question I asked. So what had I actually proved? That the chunks reached the prompt, or only that the node didn't crash?

This is the part I think gets skipped a lot in RAG write-ups. "I injected context and the answer looked right" isn't a finding, because the answer looks right in the control too.

So I built the control. One flow, two decision nodes in sequence: identical state, identical questions, evidence in one and not the other. One run, so there's no cross-run drift to argue about.

The trick is choosing a question the control cannot answer. I put the decisive facts only in the chunks, and made them arbitrary:

  • contract-s12-notice.pdf — termination requires 90 days notice; anything shorter is void
  • contract-s19-exit-fee.pdf — early exit costs 40% of the remaining balance

Nothing in the case hints at 90 or 40. No industry convention supplies them. A correct answer without reading the chunk is essentially impossible, which is the only thing that makes the comparison mean anything.

Question: "Is the written notice period in case long enough to validly terminate?"

AnswerConfidencep(yes)
With evidenceno ✓ (30 < 90)0.960.04
Without evidenceyes ✗0.610.61

Look at the shape of that, not just the flip. Without the clause it sat at 0.61 — a hedge, barely off a coin toss — because 30 days' written notice reads as unremarkable. With the clause it went to 0.96 against. That's what reading a decisive fact looks like, and a model that hadn't read it could not land there.

Corroborated independently: +178 input tokens with evidence (673 vs 495). The chunks physically entered the prompt, regardless of what the model concluded.


Then the same test, one layer in

Evidence rows were now established. The reference values inside instructions were not — different code path, same question outstanding.

Same design, one variable. Both arms got the identical case, the identical single evidence chunk and the identical question. One carried a policy reference pointing at {{$.kb.hard_cap}}: any termination above £5,000 is void without the sector regulator's countersignature. Absurdly low on purpose, and naming a regulator that appears nowhere else.

AnswerConfidencep(yes)
With the referenceno ✓ (£82k > £5k, no countersignature)0.950.05
Without ityes ✓ (clause satisfied)0.890.89

This one is cleaner than the first, because both arms were confident and correct for what they were given. No hedge anywhere: 0.89 yes on the clause alone, 0.95 no once the policy arrived. A decisive reversal from a single variable.

And the token delta settles which text arrived: +99 tokens. An unresolved {{$.kb.hard_cap}} would have cost about 8. So the reference didn't merely get read, it got resolved to the policy wording before the call.


The test that proved nothing

I ran two questions in that first experiment. Only one of them earned its place, and being honest about the other is probably the most useful thing in this post.

The second was "does this termination make an early exit charge payable?" — and it came back yes in both arms, 0.69 with evidence and 0.65 without. Shift: +0.04. Nothing.

That isn't evidence being ignored. It's that "early termination incurs a fee" is a strong general prior, because most contracts do charge one. The control guessed right for the wrong reason, and its low 0.65 confidence shows it was unsure rather than informed.

I'd picked a fact that wasn't actually unguessable. The 90-day number was; the exit fee wasn't. If I'd only run that question, I'd have concluded the evidence was inert and gone looking for a bug that didn't exist.

The lesson generalises past this node: when you A/B a retrieval pipeline, your control has priors. If the question is answerable from general knowledge, a null result tells you nothing about your plumbing. Pick facts that are arbitrary, specific and unguessable, or you're measuring the model's background knowledge rather than your own system.


Two kinds of braces

One real bug surfaced while building this, and it's worth naming because it hides in plain sight.

There are two template syntaxes in play, and they belong to different people:

SyntaxWhoseResolved byMust arrive as
`evidence[0].text`, `policy`the model'sthe model, server-sideliteral text
{{$.kb.policy}}, {{$.case}}FloMorphic'sthe node, before the callthe resolved value

The backticked paths are the decision model's own pointer notation — it walks the state object we sent and finds that chunk. Resolving those ourselves would break them.

{{$.path}} is ours, and this is where the bug was: it was resolved in the state, but not in a question's instructions or in an option's description. Write Does this breach {{$.policy}}? as your question and the model received the characters {{$.policy}}. Worse, the compiler was flattening a structured instructions object through a string accessor, which silently turned the whole thing into "" and dropped the reference data on the floor.

Both fixed, and both now verified against the live service rather than only against my own tests. Every string a designer authors — instructions at any depth, reference values, option descriptions, evidence rows — is resolved before the call. A value fetched back from the Context deliberately is not re-resolved: templates are authored on a canvas, while a scope value is runtime data, and re-resolving a retrieved chunk would let that chunk read your flow's Context.


Already landed, shipping next: a second protocol

One more thing, and it is the only part of this post you can't run today.

Everything above shipped in v0.4.3. What follows landed on main this week, after v0.4.3 was baked from its component commits, so it is deliberately not in the image above — it's under test for the next release.

The decision-model space has settled into two wire protocols, not one. OpenAI shipped the Decisions API — POST /v1/decisions, public beta since DevDay — serving gpt-6-luna. It asks the same kind of question System One does and answers with the same kind of distribution. Only the encoding differs:

System OneDecisions
the subjectstateinput
questionsobject keyed by idarray, each with a name
booleannoul + criteriapredicate, no criteria at all
choicecriteria {name: desc}choices [{value, description}]
scorecriteria [desc, …]levels [{label, description}]
answersobject keyed by idarray, echoing name
refusal—{"type":"refusal","name":…}

So the node grew a provider field — systemone (the default, forever, because profiles are snapshotted onto nodes and flows built before this field existed carry provider:"") or decisions. It names the protocol, not the vendor, deliberately: each shape already carries several vendors, and a gateway can serve one vendor's model over the other vendor's protocol. Vercel's AI Gateway implements the Decisions shape and routes Jev through it.

Nothing downstream of the protocol differs. The typed questions, the declared options, the <question>.<option> port tags, the confidence floor, the _exception branch — one implementation, shared. Which means a flow can be re-pointed from a local decider in development, to hosted Jev in production, to gpt-6-luna for an A/B, by editing a settings profile, with the canvas wiring untouched.

One genuine asymmetry, and it's the kind you want to know before you switch a node over: a System One model cannot decline. It always places the state in one of your declared options, and when none of them fit, the probabilities look more decisive rather than less — which is why the advice has always been to draw the no-match case as an other option. The Decisions protocol does allow a refusal, per question, while the others answer normally. A refused routed question therefore leaves through _exception with code refused. Draw that branch before you point a node at decisions.

It's under test now and ships in v0.4.4, the next release. Until then, provider is not a field on the node you install — v0.4.3 speaks System One, and only System One.


What this adds up to

A decision model that can't answer outside your options is a good primitive. One that can't answer outside your options and reads the documents you retrieved for it is a different thing: a bounded judgment over live evidence, with the whole distribution and the provenance landing in the Context.

Every chunk travels with its source. Every answer carries its probabilities. The question that decided it is on the canvas, and so is every route it could have taken. When someone asks why the process went left instead of right, the answer isn't a log line, it's the run's own record.

And because the node is named for the protocol rather than the vendor, none of that is tied to one supplier's roadmap. The market for bounded deciders is a few weeks old, already has a dozen entrants and now two wire formats, and a flow drawn today shouldn't have to be redrawn when the thirteenth arrives.

And the next time you wire up retrieval, run the control. Not the happy path twice. The control, with a fact your model could not possibly guess.


The ai-decision node ships with the builtin plugins. Install FloMorphic with the script from the getting-started repo — bash on Linux and macOS, PowerShell on Windows:

curl -fsSL https://raw.githubusercontent.com/FloMorphic/getting-started/main/install.sh | bash
irm https://raw.githubusercontent.com/FloMorphic/getting-started/main/install.ps1 | iex