Trace + Live Observability
Trace + Live Observability
Flow Studio turns your running flow into a live business dashboard. Every published flow ships with two per-flow canvas modes, both free:
- Trace mode — pick any single run, scrub through it like a video, inspect the exact packet that crossed each edge and the latency of each node. One event, one nanosecond-resolution answer.
- Live mode — subscribe to a per-flow metrics frame every second.
The canvas heat-maps rate, latency, error rate, or saturation;
packet dots slide down each edge at a speed proportional to real
traffic; a rules engine writes recommendations like "this
directedge is at 800 packets/s — considerephemeralorstream" directly onto the canvas.
Neither mode asks you to leave the canvas, context-switch to Grafana, or stand up a separate dashboard. The flow graph is the dashboard.
Two cross-flow views sit above the single-flow canvas, in the admin under Flows:
- Lifecycle — the event graph of how all your published flows connect (emit → trigger), end to end.
- Alerts — a durable inbox of every node/run failure across every flow, captured automatically.
What's new at a glance
| Capability | Mode | What you see |
|---|---|---|
| Packet dot sliding along an edge in the direction of travel | both | Which way traffic moved, how fast |
| Per-node heat border + glow ring | both | Which node is the bottleneck right now |
Per-node rate · Σtotal · p95 · error chip in the header | live | Live telemetry + cumulative count without a popup |
Per-edge 500/s · Σ12k · 42ms label | live | Transport-level throughput + total delivered |
| Scrubber (play / pause / 0.25×–4×) | trace | Replay any historical run at human speed |
| Copy-to-clipboard on every payload | trace | One-click export of node inputs/outputs |
| Server-computed recommendation cards | live | "Switch direct → stream" with the reason |
| SLO violation badges | live | When a node breaches p95_latency_ms |
Prerequisites
- A running Redelay app with
flowexecwired in (see Running Flows). - Flow Studio reachable at
/flowdslon the admin-api (FLOWSTUDIO_UI_PATH). - At least one published flow with some runs under it.
No new setup — both modes are enabled by default on any flowexec deployment.
Trace mode — one run, inspected
Trace mode is the "why did this specific run do that?" tool. Pick a run from the Runs tab on the left sidebar. The canvas flips to Trace mode with the finished run pre-loaded at the final frame.
Anatomy of a traced run
┌────────┐ ┌────────┐ ┌────────┐
│ start │ ────▶ │ chat │ ────▶ │ end │
│ ✓ 2ms │ ●●● │ ✓ 4.0s │ ●●● │ ✓ 1ms │
└────────┘ └────────┘ └────────┘
▲ ▲ ▲
└ green border: │ └ same treatment for terminal
node finished │
└ amber border mid-run, green on completion.
Packet dots (●●●) slide source→target at
a speed derived from real event timestamps.
The scrubber dock
Press play — the canvas re-animates the run at real time. Changeable at the bottom-centre:
- 0.25× / 0.5× / 1× / 2× / 4× — slow down for presentations or speed up for long runs.
- Scrub bar — drag to any point; earlier nodes un-highlight, edges that hadn't fired yet go dark.
- Elapsed counter —
0.52 s / 4.03 sreads real wall-clock time, not an abstract event index.
The scrubber only appears when you're looking at a historical run — live runs pace themselves off the SSE stream.
The payload drawer
Click any node or any edge:
- Node click — right-side drawer opens showing
status,durationMs,input,output,error. Inputs and outputs are pretty-printed JSON with a copy button (flashes a green check for a second on success). - Edge click — shows the packet that crossed the edge, its
type ref, and the delivery mode it used. The packet body comes
from whichever endpoint actually captured it (source
node.doneoutput, else targetnode.startedinput) — you never see "not captured" unless the engine genuinely never materialised the packet.
The Events popup
Clicking a run row also opens Events (N) — a flat, sorted, row
event list you can expand inline. Each expanded row shows the raw
payload as JSON with its own copy button. Events are always in
emission order regardless of storage quirks (see
Event ordering below).
Live mode — aggregate health
Live mode is the "what's happening across all runs right now?" tool.
Flip the mode switch in the header from trace to live. The canvas
subscribes to GET /flows/{id}/live and paints a new frame every
second.
What each overlay shows
The bottom dock picks which metric drives the heat map:
- rate — packets per second (log-scaled so a 1/s edge is still visible and a 10 000/s edge doesn't dominate)
- p95 — 95th-percentile node latency; latency heat ramps fast because users feel it
- err — error rate; 10% = full red
- sat — handler saturation;
1.0= 100% busy
Colour palette is unified across edges and nodes: green (healthy) → amber (warm) → red (look here). One colour language; no legend to memorise.
Recommendations
In the top-right, Flow Studio shows a collapsible stack of server-computed recommendation cards. The rule engine runs on every frame before it hits the wire, so the cards update in real time and never require a Studio redeploy to iterate the thresholds. The current rule catalogue:
| Rule | Severity | Trigger | Action |
|---|---|---|---|
| Capacity | warn | saturation ≥ 80% | raise concurrency / shard / queue |
| Reliability | warn / critical | error rate > 1% / > 10% | check recent deploys / add retries |
| SLO | critical | p95_latency_ms exceeded | profile / split / relax SLO |
| Transport: direct hot | warn | direct edge ≥ 500/s | switch to ephemeral / stream |
| Transport: extreme rate | warn | any edge ≥ 5000/s | move to stream + partition |
| Transport: backpressure | critical | send-blocked events > 0 | scale consumer / use queued transport |
| Transport: non-durable failures | warn | failures on direct/ephemeral | consider checkpoint / durable |
Click a card — the canvas jumps focus to the offending node or edge.
Declaring SLOs
Service-level objectives live on the flow node, not in a sidecar config. Add any combination of the three dimensions:
# in your flow document
nodes:
- id: chat
kind: action
action_ref: llm.chat
slo:
p95_latency_ms: 3000 # red alert above this
max_error_rate: 0.01 # 1%
max_saturation: 0.8 # 80% busy headroom
A breach fires an slo-kind Suggestion with severity critical and
a ring pulse on the node. Zero on any field means "no SLO on that
dimension".
Lifecycle — the whole flow graph {#lifecycle}
Trace and Live look inside one flow. The Lifecycle view zooms out to how your published flows connect to each other.
Reaction flows are deliberately decoupled — each subscribes to an event
and does one job. That keeps them composable, but it also means the
end-to-end business process (payment.captured → apply → order.paid → issue invoice → capture → …) is spread across a dozen independent
flows with no single diagram. The Lifecycle view rebuilds that diagram
automatically.
- Where: Admin → Flows → Lifecycle (base admin layer), backed by
GET /api/v1/flows/lifecycleon the admin-api. - How it's derived: each published flow's consumed trigger is its
source node's
eventName; its emitted events come from each node'semitsdeclaration (plus any node that publishes a config-driven topic). A directed edgeA →(event) Bmeans flow A emits the event that triggers flow B — so one emitter fans out to every consumer. - What it shows: each flow is a card with a trigger badge
(
event/http/schedule/manual) and its consume/emit event chips; flows are laid out left-to-right by event-dependency depth (viadagre), so entry points sit on the left and the cascade flows right. Three summary panels bucket the loose ends:- external entry points — events flows wait for that no flow emits (webhooks, direct API, upstream services);
- module-delivered — events a flow emits that a Go module's
Consumers()handles rather than another flow (e.g.email.send,sms.send); these are delivered, not dead ends, so they render green; - dead-end emits — events a flow emits that nothing reacts to yet — a hook for a step you could add.
- It's interactive. The graph is a Vue Flow
canvas: pan / zoom, drag cards to rearrange, a minimap and fit-to-view
controls. Search any flow, event, or node from the box top-left — the
matching flow and the full upstream+downstream path it participates in
highlight (BFS over the edges) while everything else dims. Hover a card
to spotlight its one-hop neighbourhood (direct emitters + consumers). Deep
links work too:
/flows/lifecycle?focus=<flow-id|workflow-id|name>opens with that flow pre-selected and its path lit — the Alerts inbox links each alert straight to its flow this way.
The graph is only as complete as your node emits declarations — a node
that publishes an event but doesn't declare it won't draw an edge. See
FlowDSLNode.Emits.
Alerts — the failure inbox {#alerts}
Every node failure across every flow is captured into a durable,
filterable alerts inbox — with zero per-flow wiring. An AlertSink
decorates the flowexec event sink and records the failure telemetry the
executor already emits:
| Event | Severity | Note |
|---|---|---|
node.failed | error | the actionable, node-level failure |
run.failed | critical | a run-level failure with no attributable node (deduped when a node already failed the run) |
- Where: Admin → Flows → Alerts, backed by
GET /api/v1/flows/alerts(filter byseverity,flow_id,resolved; returns anopenCountfor the nav badge) andPOST /api/v1/flows/alerts/{id}/resolve. Alerts are resolve-only — acknowledging keeps an audit trail; there is no hard delete. Each alert row links its flow into the Lifecycle graph (?focus=<flow>) so you can see where the failing flow sits in the end-to-end process in one click. - Storage: a Mongo
flow_alertscollection when a DB is configured, in-memory otherwise. No setup — capture is on by default on any flowexec deployment.
Recording a custom alert — alerts/record
Beyond automatic failure capture, a flow can raise its own alert from an explicit branch — a big order, a risky payment, a failed external check:
nodes:
- id: flag-big-order
kind: action
action_ref: alerts/record
config:
severity: warning # warning | error | critical
source: risk
message: "Large order flagged for review" # an input `message` overrides this
The node emits on the Recorded port (with output.alert_id) or on
Error if the store write fails — it never fails the run. The message
is taken from the incoming packet first, so an upstream edge can template
it ({{ .payload.number }}). Recorded alerts carry kind: custom.
Historical retention
The flowexec in-proc aggregator keeps the last 5 minutes of frames in a ring buffer. Plenty for "is anything red right now?".
For longer retention (capacity planning, postmortems, compare-to- yesterday), enable the optional ClickHouse sink:
# infra/docker-compose.yml
docker compose --profile metrics up -d
export CLICKHOUSE_URL=clickhouse://redelay:redelay@localhost:9000/redelay
The flow_metrics table is auto-created at boot with a 30-day TTL;
each frame fans out to one row per node + one per edge + one per
channel. The schema is in
go-flowdsl/flowexec/metrics/clickhouse
— point Grafana or your BI tool at it if you want charts beyond
what the canvas renders.
Small deployments can ignore ClickHouse entirely; the in-proc ring powers both the live stream and the first-paint history endpoint.
Event ordering {#event-ordering}
FlowEvents carry a per-run monotonic seq counter assigned by the
Executor in emission order. MongoDB time-series storage can't
tie-break identical millisecond timestamps (node.started and
node.done on the same instant-pass-through node, for example), so
the sink sorts by (ts, seq) and the client never re-sorts.
The Studio's event list, payload drawer, scrubber timeline, and Live aggregator all consume this canonical order.
Wire format
Live frames are JSON-encoded over SSE; one frame per tick.
{
"ts": "2026-04-17T20:11:12.387Z",
"windowMs": 1000,
"meta": {
"flowId": "assistant",
"versionHash": "cc2891d2…"
},
"nodes": {
"chat": {
"rate": 45.2,
"latencyP50": 820,
"latencyP95": 4120,
"errorRate": 0.012,
"saturation": 0.87,
"inflight": 4,
"completed": 42, "failed": 1, // this window only
"completedTotal": 12843, "failedTotal": 27 // cumulative since aggregator boot
}
},
"edges": {
"start__chat": {
"rate": 45.2,
"delivered": 45, // this window
"deliveredTotal": 12870, // cumulative
"mode": "direct"
}
},
"suggestions": [
{
"severity": "warn",
"kind": "transport",
"target": "edge:start__chat",
"reason": "45/s over `direct` — in-proc delivery is saturating",
"action": "Switch to `ephemeral` (Redis) or `stream` (Kafka)"
}
]
}
Endpoints:
GET /flows/{id}/live— SSE stream (1 frame / s)GET /flows/{id}/metrics/history?limit=N&since=RFC3339Nano— the most recent frames from the RAM ring (used by the Studio for first-paint + dock scrubbing)
Identity — tenantId, flowId, versionHash, assistantId,
userId — rides on the frame payload (meta), never as routing
metadata, so multi-tenant deployments partition on the read side
without carving up topic names.
Run + metric identity: document ID vs flow-record ID {#run-identity}
Two IDs name a flow, and observability keys on the first:
- Workflow document ID — the authored
idinside the flow document (e.g.asyncshop-checkout-standard). The Executor stamps this onto every run record, metric frame, event, and tap hit (FlowRun.flowId,frame.meta.flowId, …). - Flow-record ID —
flow.<uuid>, the storage handle the admin URLs use (/flows/{id}/...).
The /flows/{id}/runs, /metrics/history, and /live endpoints resolve the
record ID in the URL to the flow's published (or head) document ID before
querying, so a flow's history actually surfaces. Consequence: two flow records
built from the same template share one authored document ID and therefore
share run/metric history on these endpoints — give a flow its own document ID
to separate them. Taps resolve the same split via Tap.WorkflowID.
Where a run's output lives. The run row (GET /runs/{runID}) carries only
metadata — status, startedAt/endedAt, durationMs, inputDigest. The
actual packet each node produced is in the event stream
(GET /runs/{runID}/events/history, or live via /events) and in any tap hit
— not on the run summary. /metrics/history reads the in-RAM ring (last
few minutes) unless ClickHouse retention is enabled (see below),
so a flow that hasn't run recently returns no frames even though its run
history persists.
What's next
- Wire the matching
runtime/observabilityhelpers when you embedflowexecdirectly (without the HTTP module) — the Executor accepts anyMetricsRecorder, so you can plug a test double or a remote aggregator. - Declare SLOs on your hot-path nodes and watch the violation badges appear during load tests.
- Reach for the Choosing a Transport
guide when Live mode starts recommending
ephemeralorstreamupgrades.
Further reading
- Running Flows — the execution contract that Trace + Live sit on top of.
- go-flowdsl reference → metrics —
MetricsSink,Frame,Suggestiontypes. - Transports — the five delivery modes the recommendation engine nudges you between.
Running Flows
End-to-end guide for executing FlowDSL flows in a Redelay service — create a flow, save and publish a version, launch a run, and subscribe to live events over SSE.
Project AI Assistant
Wire the in-app AI assistant component to a FlowDSL flow executed by your Redelay backend. The frontend stays framework-agnostic; all LLM logic lives in a flow you can edit without a redeploy.