Guides

Trace + Live Observability

See every packet, spot bottlenecks, and get transport recommendations — without leaving the canvas. Flow Studio's Trace and Live modes, the cross-flow Lifecycle graph, and the failure Alerts inbox, powered by the flowexec pipeline.

Trace + Live Observability

Flow Studio turns your running flow into a live business dashboard. Every published flow ships with two per-flow canvas modes, both free:

  • Trace mode — pick any single run, scrub through it like a video, inspect the exact packet that crossed each edge and the latency of each node. One event, one nanosecond-resolution answer.
  • Live mode — subscribe to a per-flow metrics frame every second. The canvas heat-maps rate, latency, error rate, or saturation; packet dots slide down each edge at a speed proportional to real traffic; a rules engine writes recommendations like "this direct edge is at 800 packets/s — consider ephemeral or stream" directly onto the canvas.

Neither mode asks you to leave the canvas, context-switch to Grafana, or stand up a separate dashboard. The flow graph is the dashboard.

Two cross-flow views sit above the single-flow canvas, in the admin under Flows:

  • Lifecycle — the event graph of how all your published flows connect (emit → trigger), end to end.
  • Alerts — a durable inbox of every node/run failure across every flow, captured automatically.

What's new at a glance

CapabilityModeWhat you see
Packet dot sliding along an edge in the direction of travelbothWhich way traffic moved, how fast
Per-node heat border + glow ringbothWhich node is the bottleneck right now
Per-node rate · Σtotal · p95 · error chip in the headerliveLive telemetry + cumulative count without a popup
Per-edge 500/s · Σ12k · 42ms labelliveTransport-level throughput + total delivered
Scrubber (play / pause / 0.25×–4×)traceReplay any historical run at human speed
Copy-to-clipboard on every payloadtraceOne-click export of node inputs/outputs
Server-computed recommendation cardslive"Switch direct → stream" with the reason
SLO violation badgesliveWhen a node breaches p95_latency_ms

Prerequisites

  • A running Redelay app with flowexec wired in (see Running Flows).
  • Flow Studio reachable at /flowdsl on the admin-api (FLOWSTUDIO_UI_PATH).
  • At least one published flow with some runs under it.

No new setup — both modes are enabled by default on any flowexec deployment.

Trace mode — one run, inspected

Trace mode is the "why did this specific run do that?" tool. Pick a run from the Runs tab on the left sidebar. The canvas flips to Trace mode with the finished run pre-loaded at the final frame.

Anatomy of a traced run

text
 ┌────────┐        ┌────────┐       ┌────────┐
 │ start  │ ────▶  │  chat  │ ────▶ │  end   │
 │ ✓ 2ms  │  ●●●   │ ✓ 4.0s │  ●●●  │ ✓ 1ms  │
 └────────┘        └────────┘       └────────┘
      ▲                ▲                ▲
      └ green border:  │                └ same treatment for terminal
        node finished  │
                       └ amber border mid-run, green on completion.
                         Packet dots (●●●) slide source→target at
                         a speed derived from real event timestamps.

The scrubber dock

Press play — the canvas re-animates the run at real time. Changeable at the bottom-centre:

  • 0.25× / 0.5× / 1× / 2× / 4× — slow down for presentations or speed up for long runs.
  • Scrub bar — drag to any point; earlier nodes un-highlight, edges that hadn't fired yet go dark.
  • Elapsed counter — 0.52 s / 4.03 s reads real wall-clock time, not an abstract event index.

The scrubber only appears when you're looking at a historical run — live runs pace themselves off the SSE stream.

The payload drawer

Click any node or any edge:

  • Node click — right-side drawer opens showing status, durationMs, input, output, error. Inputs and outputs are pretty-printed JSON with a copy button (flashes a green check for a second on success).
  • Edge click — shows the packet that crossed the edge, its type ref, and the delivery mode it used. The packet body comes from whichever endpoint actually captured it (source node.done output, else target node.started input) — you never see "not captured" unless the engine genuinely never materialised the packet.

The Events popup

Clicking a run row also opens Events (N) — a flat, sorted, row event list you can expand inline. Each expanded row shows the raw payload as JSON with its own copy button. Events are always in emission order regardless of storage quirks (see Event ordering below).

Live mode — aggregate health

Live mode is the "what's happening across all runs right now?" tool. Flip the mode switch in the header from trace to live. The canvas subscribes to GET /flows/{id}/live and paints a new frame every second.

What each overlay shows

The bottom dock picks which metric drives the heat map:

  • rate — packets per second (log-scaled so a 1/s edge is still visible and a 10 000/s edge doesn't dominate)
  • p95 — 95th-percentile node latency; latency heat ramps fast because users feel it
  • err — error rate; 10% = full red
  • sat — handler saturation; 1.0 = 100% busy

Colour palette is unified across edges and nodes: green (healthy) → amber (warm) → red (look here). One colour language; no legend to memorise.

Recommendations

In the top-right, Flow Studio shows a collapsible stack of server-computed recommendation cards. The rule engine runs on every frame before it hits the wire, so the cards update in real time and never require a Studio redeploy to iterate the thresholds. The current rule catalogue:

RuleSeverityTriggerAction
Capacitywarnsaturation ≥ 80%raise concurrency / shard / queue
Reliabilitywarn / criticalerror rate > 1% / > 10%check recent deploys / add retries
SLOcriticalp95_latency_ms exceededprofile / split / relax SLO
Transport: direct hotwarndirect edge ≥ 500/sswitch to ephemeral / stream
Transport: extreme ratewarnany edge ≥ 5000/smove to stream + partition
Transport: backpressurecriticalsend-blocked events > 0scale consumer / use queued transport
Transport: non-durable failureswarnfailures on direct/ephemeralconsider checkpoint / durable

Click a card — the canvas jumps focus to the offending node or edge.

Declaring SLOs

Service-level objectives live on the flow node, not in a sidecar config. Add any combination of the three dimensions:

yaml
# in your flow document
nodes:
  - id: chat
    kind: action
    action_ref: llm.chat
    slo:
      p95_latency_ms: 3000      # red alert above this
      max_error_rate: 0.01      # 1%
      max_saturation: 0.8       # 80% busy headroom

A breach fires an slo-kind Suggestion with severity critical and a ring pulse on the node. Zero on any field means "no SLO on that dimension".

Lifecycle — the whole flow graph {#lifecycle}

Trace and Live look inside one flow. The Lifecycle view zooms out to how your published flows connect to each other.

Reaction flows are deliberately decoupled — each subscribes to an event and does one job. That keeps them composable, but it also means the end-to-end business process (payment.captured → apply → order.paid → issue invoice → capture → …) is spread across a dozen independent flows with no single diagram. The Lifecycle view rebuilds that diagram automatically.

  • Where: Admin → Flows → Lifecycle (base admin layer), backed by GET /api/v1/flows/lifecycle on the admin-api.
  • How it's derived: each published flow's consumed trigger is its source node's eventName; its emitted events come from each node's emits declaration (plus any node that publishes a config-driven topic). A directed edge A →(event) B means flow A emits the event that triggers flow B — so one emitter fans out to every consumer.
  • What it shows: each flow is a card with a trigger badge (event / http / schedule / manual) and its consume/emit event chips; flows are laid out left-to-right by event-dependency depth (via dagre), so entry points sit on the left and the cascade flows right. Three summary panels bucket the loose ends:
    • external entry points — events flows wait for that no flow emits (webhooks, direct API, upstream services);
    • module-delivered — events a flow emits that a Go module's Consumers() handles rather than another flow (e.g. email.send, sms.send); these are delivered, not dead ends, so they render green;
    • dead-end emits — events a flow emits that nothing reacts to yet — a hook for a step you could add.
  • It's interactive. The graph is a Vue Flow canvas: pan / zoom, drag cards to rearrange, a minimap and fit-to-view controls. Search any flow, event, or node from the box top-left — the matching flow and the full upstream+downstream path it participates in highlight (BFS over the edges) while everything else dims. Hover a card to spotlight its one-hop neighbourhood (direct emitters + consumers). Deep links work too: /flows/lifecycle?focus=<flow-id|workflow-id|name> opens with that flow pre-selected and its path lit — the Alerts inbox links each alert straight to its flow this way.

The graph is only as complete as your node emits declarations — a node that publishes an event but doesn't declare it won't draw an edge. See FlowDSLNode.Emits.

Alerts — the failure inbox {#alerts}

Every node failure across every flow is captured into a durable, filterable alerts inbox — with zero per-flow wiring. An AlertSink decorates the flowexec event sink and records the failure telemetry the executor already emits:

EventSeverityNote
node.failederrorthe actionable, node-level failure
run.failedcriticala run-level failure with no attributable node (deduped when a node already failed the run)
  • Where: Admin → Flows → Alerts, backed by GET /api/v1/flows/alerts (filter by severity, flow_id, resolved; returns an openCount for the nav badge) and POST /api/v1/flows/alerts/{id}/resolve. Alerts are resolve-only — acknowledging keeps an audit trail; there is no hard delete. Each alert row links its flow into the Lifecycle graph (?focus=<flow>) so you can see where the failing flow sits in the end-to-end process in one click.
  • Storage: a Mongo flow_alerts collection when a DB is configured, in-memory otherwise. No setup — capture is on by default on any flowexec deployment.

Recording a custom alert — alerts/record

Beyond automatic failure capture, a flow can raise its own alert from an explicit branch — a big order, a risky payment, a failed external check:

yaml
nodes:
  - id: flag-big-order
    kind: action
    action_ref: alerts/record
    config:
      severity: warning        # warning | error | critical
      source: risk
      message: "Large order flagged for review"   # an input `message` overrides this

The node emits on the Recorded port (with output.alert_id) or on Error if the store write fails — it never fails the run. The message is taken from the incoming packet first, so an upstream edge can template it ({{ .payload.number }}). Recorded alerts carry kind: custom.

Historical retention

The flowexec in-proc aggregator keeps the last 5 minutes of frames in a ring buffer. Plenty for "is anything red right now?".

For longer retention (capacity planning, postmortems, compare-to- yesterday), enable the optional ClickHouse sink:

shell
# infra/docker-compose.yml
docker compose --profile metrics up -d
export CLICKHOUSE_URL=clickhouse://redelay:redelay@localhost:9000/redelay

The flow_metrics table is auto-created at boot with a 30-day TTL; each frame fans out to one row per node + one per edge + one per channel. The schema is in go-flowdsl/flowexec/metrics/clickhouse — point Grafana or your BI tool at it if you want charts beyond what the canvas renders.

Small deployments can ignore ClickHouse entirely; the in-proc ring powers both the live stream and the first-paint history endpoint.

Event ordering {#event-ordering}

FlowEvents carry a per-run monotonic seq counter assigned by the Executor in emission order. MongoDB time-series storage can't tie-break identical millisecond timestamps (node.started and node.done on the same instant-pass-through node, for example), so the sink sorts by (ts, seq) and the client never re-sorts.

The Studio's event list, payload drawer, scrubber timeline, and Live aggregator all consume this canonical order.

Wire format

Live frames are JSON-encoded over SSE; one frame per tick.

jsonc
{
  "ts": "2026-04-17T20:11:12.387Z",
  "windowMs": 1000,
  "meta": {
    "flowId": "assistant",
    "versionHash": "cc2891d2…"
  },
  "nodes": {
    "chat": {
      "rate": 45.2,
      "latencyP50": 820,
      "latencyP95": 4120,
      "errorRate": 0.012,
      "saturation": 0.87,
      "inflight": 4,
      "completed": 42, "failed": 1,           // this window only
      "completedTotal": 12843, "failedTotal": 27 // cumulative since aggregator boot
    }
  },
  "edges": {
    "start__chat": {
      "rate": 45.2,
      "delivered": 45,            // this window
      "deliveredTotal": 12870,    // cumulative
      "mode": "direct"
    }
  },
  "suggestions": [
    {
      "severity": "warn",
      "kind": "transport",
      "target": "edge:start__chat",
      "reason": "45/s over `direct` — in-proc delivery is saturating",
      "action": "Switch to `ephemeral` (Redis) or `stream` (Kafka)"
    }
  ]
}

Endpoints:

  • GET /flows/{id}/live — SSE stream (1 frame / s)
  • GET /flows/{id}/metrics/history?limit=N&since=RFC3339Nano — the most recent frames from the RAM ring (used by the Studio for first-paint + dock scrubbing)

Identity — tenantId, flowId, versionHash, assistantId, userId — rides on the frame payload (meta), never as routing metadata, so multi-tenant deployments partition on the read side without carving up topic names.

Run + metric identity: document ID vs flow-record ID {#run-identity}

Two IDs name a flow, and observability keys on the first:

  • Workflow document ID — the authored id inside the flow document (e.g. asyncshop-checkout-standard). The Executor stamps this onto every run record, metric frame, event, and tap hit (FlowRun.flowId, frame.meta.flowId, …).
  • Flow-record ID — flow.<uuid>, the storage handle the admin URLs use (/flows/{id}/...).

The /flows/{id}/runs, /metrics/history, and /live endpoints resolve the record ID in the URL to the flow's published (or head) document ID before querying, so a flow's history actually surfaces. Consequence: two flow records built from the same template share one authored document ID and therefore share run/metric history on these endpoints — give a flow its own document ID to separate them. Taps resolve the same split via Tap.WorkflowID.

Where a run's output lives. The run row (GET /runs/{runID}) carries only metadata — status, startedAt/endedAt, durationMs, inputDigest. The actual packet each node produced is in the event stream (GET /runs/{runID}/events/history, or live via /events) and in any tap hit — not on the run summary. /metrics/history reads the in-RAM ring (last few minutes) unless ClickHouse retention is enabled (see below), so a flow that hasn't run recently returns no frames even though its run history persists.

What's next

  • Wire the matching runtime/observability helpers when you embed flowexec directly (without the HTTP module) — the Executor accepts any MetricsRecorder, so you can plug a test double or a remote aggregator.
  • Declare SLOs on your hot-path nodes and watch the violation badges appear during load tests.
  • Reach for the Choosing a Transport guide when Live mode starts recommending ephemeral or stream upgrades.

Further reading