Project AI Assistant
Project AI Assistant
The Redelay project template ships a small "talk-to-your-app" assistant with a deliberate split of concerns:
- The frontend (
AiAssistant.vue+useChat.ts) knows nothing about flows, nodes, or LLM providers. It POSTs{messages}to/api/chatand reads back a server-sent-events stream of{content}/{done}/{error}frames. - A Nitro proxy in the Nuxt website forwards those requests to the Go backend. It enforces rate limits and caps conversation depth but carries no flow or credential logic.
- A project-local Go module (
backend/modules/assistant) resolves the current published version of a FlowDSL flow, launches a run for each incoming request, and translates the run's FlowEvents into the chat SSE wire format. - The flow itself — shipped as an embedded default and stored in the
flowexecstore the first time the backend boots — carries the actual provider, model, prompt, and any extra routing you want. Edit it through the standardflowexecREST API; changes take effect on the next message with no restart.
This document walks through the wiring end-to-end so you can customize or replace any layer independently.
Architecture at a glance
Browser
│ POST /api/chat {messages}
▼
Nuxt Nitro proxy (website/server/api/chat.post.ts)
│ POST /api/v1/assistant/messages {messages}
▼
Go backend — modules/assistant
│ resolves published version of flow "assistant"
│ flowexec.Executor().RunAsync(RunInput{...})
│ flowexec.Subscribe(runID) → FlowEvents
▼
FlowDSL flow "assistant"
start → guard-in → chat (redelay/llm-chat) → guard-out → end
↘ ↘
rejected rejected
The frontend only knows about /api/chat and the SSE frame shape. Swapping
the backing flow for a RAG pipeline, a tool-router, or a multi-step agent
changes nothing about the browser bundle.
Guarded by default
Both the default and production assistant flows wrap the LLM call with
redelay/llm-guard router nodes — one before
the chat (classifies user input) and one after (classifies assistant output).
On a Block verdict the flow routes through a core/template-render
transform (refusal-in / refusal-out) that stamps a configurable refusal
message onto the packet before the rejected terminal, so the client
always receives a useful reply instead of silence. The production variant
additionally fans blocked outputs to a trust-and-safety email via
ephemeral delivery so a send failure cannot block the user-facing stream.
Guard behaviour is tunable without a code change:
- Set
GUARD_PROVIDER=qwen3guardto turn on classification; leave unset to disable (the flow still runs — the nodes simply route everything toAllow). GUARD_MODEL(defaultQwen3Guard-Gen-8B— free on OVH) selects the classifier variant.- Admin settings
ai-guard.default_on_errorandai-guard.block_flaggedtune fail-open vs fail-closed behaviour and whether controversial messages are treated as blocks. - The refusal copy is just the
templatesetting on therefusal-in/refusal-outnodes — edit it in Studio, save a new version, publish. Interpolation works:{{.decision.reason}},{{.decision.categories}}.
The SSE pump streams content from whichever terminal the flow reaches
(end on success, rejected on block) — ASSISTANT_OUTPUT_NODE is a
comma-separated list of terminal ids; the default end,rejected covers
both without a code change.
See the go-ai guard reference for the full contract and the redelay/llm-guard node doc for the refusal wiring pattern.
Step 1 — Backend wiring
backend/cmd/api/main.go blank-imports the three packages that matter:
import (
_ "github.com/redelay/go-flowdsl/flowexec/module" // flow storage + runner
_ "github.com/redelay/go-ai/llm/flowdsl" // llm-chat / llm-embed FlowDSL nodes
_ "github.com/redelay/go-ai/llm/providers/ollama" // at least one provider
_ "github.com/redelay/backend/modules/assistant" // the chat endpoint + default flow
)
Order matters: flowexec registers its factory first, so when the
assistant factory runs it can resolve flowexecmod.Current() and hold a
reference.
After server.Default() returns, main.go initialises the LLM provider
registry from environment variables and registers the redelay/llm-chat
handler with the flow engine:
llm.BootstrapFromEnv()
llmflowdsl.Register(fx.Engine())
At this point the server exposes two new endpoints under ROUTE_PREFIX:
| Method | Path | Purpose |
|---|---|---|
| POST | /assistant/messages | Chat SSE stream |
| POST | /assistant/reset | Re-seed the flow from the embedded default (superuser) |
Step 2 — The default flow
backend/modules/assistant/flows/default.flowdsl.json is embedded into the
binary via go:embed:
{
"id": "assistant",
"name": "Project Assistant (default)",
"nodes": [
{ "id": "start", "kind": "start" },
{
"id": "chat", "kind": "action",
"action_ref": "redelay/llm-chat",
"config": {
"providerID": "${ASSISTANT_PROVIDER:-ollama}",
"model": "${ASSISTANT_MODEL:-mistral-nemo:12b}",
"temperature": 0.2,
"stream": false
}
},
{ "id": "end", "kind": "end" }
],
"edges": [
{ "id": "e1", "from": "start", "to": "chat" },
{ "id": "e2", "from": "chat", "to": "end" }
]
}
On the first startup the module calls ensureFlow(ctx, force=false):
- If no flow with ID
"assistant"exists, it is created, the embedded JSON is saved as version 1, and that version is published. - If a flow already exists, the embedded default is ignored — whatever the project has saved and published is what runs.
This gives you "works out of the box" on a fresh install and "my customizations win" from that point on.
Step 3 — Request lifecycle
Each POST to /assistant/messages:
- Decodes
{messages, meta}. - Resolves the current published version via
flowexec.Store().GetPublishedVersion(ctx, "assistant"). - Builds
RunInput{Workflow, VersionID, VersionHash, Input: {messages: ...}}and launches viaExecutor().RunAsync. - Subscribes to live
FlowEvents for the returned run ID. - Translates each event into a chat-frame:
KindNodeDoneon the configuredASSISTANT_OUTPUT_NODE(chatby default) →{content}frame, content pulled out of the node's payload.KindNodeDoneon any other node →{trace}frame (diagnostics, skipped by the default UI).KindNodeFailed→{error}frame.KindRunCompleted→ terminal{done: true}frame.KindRunFailed→{error}frame and close.
The extractContent helper accepts several payload shapes
(content, message.content, output.content, output.message.content,
or output as a string), so you can swap the chat node for any handler
that writes one of those keys.
Step 4 — Frontend wiring
website/app/components/AiAssistant.vue drops into any Nuxt page and
renders a brand-styled slide-over. It delegates everything to
useChat.ts, which:
- POSTs to
/api/chatwith{messages}. - Streams the SSE body, coalescing
{content}frames into the in-progress assistant message. - Treats
{done}as success,{error}as failure, and ignores{trace}. - Stores conversation state in a Nuxt
useStatesingleton so the composable survives route changes.
The Nitro proxy in website/server/api/chat.post.ts is the only place the
website speaks to the backend:
POST /api/chat → ${REDELAY_BACKEND_URL}${REDELAY_ROUTE_PREFIX}/assistant/messages
It pipes the backend's SSE body straight through to the browser.
Step 5 — Customizing the flow
Because the flow lives in the flowexec store, you edit it the same way
you would any other flow — no code change needed:
# Fetch current version
curl http://localhost:8000/api/v1/flowexec/flows/assistant
# Save a new version with a richer graph (prompt-template → retrieval → llm-chat)
curl -X POST http://localhost:8000/api/v1/flowexec/flows/assistant/versions \
-H 'Content-Type: application/json' \
--data @my-rag-flow.flowdsl.json
# Publish the new version
curl -X POST http://localhost:8000/api/v1/flowexec/flows/assistant/publish \
-H 'Content-Type: application/json' \
-d '{"versionID": "<id from previous response>"}'
The next message your users send runs on the new flow. Old runs remain
pinned to their original versionHash, so historical traces stay
reproducible.
To revert to the module's embedded default, hit the superuser-only reset endpoint:
curl -X POST http://localhost:8000/api/v1/assistant/reset \
-H "Authorization: Bearer $SUPERUSER_TOKEN"
Environment variables
Assistant module
| Variable | Default | Purpose |
|---|---|---|
ASSISTANT_FLOW_ID | assistant | Flow key in the flowexec store |
ASSISTANT_INPUT_KEY | messages | Map key under which the chat history is placed in RunInput.Input |
ASSISTANT_OUTPUT_NODE | end,rejected | Comma-separated list of terminal node ids whose node.done payload is streamed as content — defaults cover the success (end) and guard-block (rejected) terminals |
ASSISTANT_PROVIDER | ollama | Referenced by ${...} interpolation in the default flow's providerID |
ASSISTANT_MODEL | mistral-nemo:12b | Referenced by ${...} interpolation in the default flow's model |
LLM provider registry (go-ai/llm)
BootstrapFromEnv() reads LLM_PROVIDER and calls the matching factory.
Each factory then reads its own env vars:
| Variable | Default | Provider |
|---|---|---|
LLM_PROVIDER | (none — skip) | Which factory to build: ollama, openai, anthropic, gemini, ovh |
OLLAMA_URL | http://localhost:11434 | Ollama API base URL |
OLLAMA_MODEL | llama3.1 | Ollama default model (overridden by flow config) |
OPENAI_API_KEY | (required) | OpenAI API key |
ANTHROPIC_API_KEY | (required) | Anthropic API key |
GEMINI_API_KEY | (required) | Google Gemini API key |
OVH_AI_ENDPOINTS_TOKEN | (required) | OVH AI Endpoints API token |
All providers also accept optional baseURL and defaultModel via the
factory config map, but when launched from BootstrapFromEnv() those are
read from env only (see each provider's godoc for the full list).
Website proxy (Nitro)
| Variable | Default | Purpose |
|---|---|---|
REDELAY_BACKEND_URL | http://localhost:8000 | Go backend base URL |
REDELAY_ROUTE_PREFIX | /api/v1 | Backend route prefix |
Testing
The module ships a unit test (backend/modules/assistant/module_test.go)
that exercises the whole path against the flowexec memstore and a stub
redelay/llm-chat handler. It covers:
- Seeding the default flow on a fresh store.
- Preserving project customizations on subsequent startups.
- The end-to-end SSE stream via
httptest.NewServer(realhttp.Flusher). 400on missing messages and409when no version is published.extractContentfallbacks for the common payload shapes.
cd backend && go test ./modules/assistant/... -v
When to reach past this module
If your assistant needs:
- Multi-tenant flows (different prompts per customer): customize
ASSISTANT_FLOW_IDat request time viametaand add a router node that dispatches on the tenant — no module change. - Tool calls / RAG: add nodes to the flow.
redelay/llm-tool-routerand any node you register on the engine are available. - Different transport (WebSocket, gRPC): replace
routes.go— the rest of the module is transport-agnostic.
The module is intentionally thin so you can fork the entire folder into a project and iterate freely without fighting the framework.
Production patterns
The repo ships three flow templates that showcase how to grow past
the three-node default as your assistant takes on more responsibility.
Load any of them via POST /api/v1/flows/assistant/versions?format=spec
and publish.
1. Human hand-off on escalation
File: spec/examples/assistant-human-handoff.flowdsl.json
The LLM's system prompt instructs it to prefix replies with a literal
[[ESCALATE]] token when it believes the request needs a human
(explicit request, safety-critical topic, low confidence). A fan-out
edge with a condition: contains(input.message.content, "[[ESCALATE]]")
triggers redelay/email-send on the same output — the user still gets
their reply, and the on-call team gets notified.
start ─► chat ─► end (primary: user-facing SSE)
└────► notify-team (fan-out, only on [[ESCALATE]] token,
ephemeral delivery so the email
never blocks the chat response)
2. Tool-calling with a router
File: spec/examples/assistant-with-tools.flowdsl.json
Uses redelay/llm-tool-router (kind: router) to branch on whether
the LLM returned toolCalls. The Tool port loops through a project-
specific run-tool handler that executes the call and appends the
result to messages; the Message port flows straight to the sink.
start ─► chat ─► router ─► end (plain reply path)
└────► run-tool ─► chat (tool loop — append result,
re-query the model)
3. Production default with fan-out audit + escalation
File: backend/modules/assistant/flows/production.flowdsl.json
The full template that the assistant module can ship as its default instead of the minimal three-node file. It adds:
- An inline SLO on the
chatnode (p95_latency_ms,max_error_rate) so Live-mode shows violation badges the moment the LLM provider gets slow. See Trace + Live Observability. - The same
[[ESCALATE]]hand-off pattern as example 1. - A single-file drop-in — copy over the minimal default with
cp backend/modules/assistant/flows/production.flowdsl.json backend/modules/assistant/flows/default.flowdsl.jsonand restart, or load it via the flowexec API without touching code.
What to add next
Node refs used above (myapp/tool-dispatch, any persistence hook)
must be registered on the runtime engine at startup via
flowexec.Engine().RegisterHandler("…", handler). The
Running Flows guide walks through the
handler registration shape.
Pair these with Trace mode and Live mode to watch them execute in real time while you tune prompts, thresholds, and escalation triggers.
Persistence + handoff (production-ready add-ons)
The assistant module ships three things beyond the chat SSE:
- Chat persistence with a TTL (default 30 days). Every message is
written to
assistant_chatsunder asession_idthe client keeps inlocalStorage. Reconnecting clients see their own history viaGET /assistant/chats/me/{sessionID}. - Manager handoff —
POST /assistant/handoffpersists aHandoffRequestrow (permanent) AND publishes theassistant.handoff_requestedevent via theEventBus. A downstream flow delivers the notification (email, Slack, whatever). The ships-with templateassistant/handoff-notificationwires event →email.sendfor the common case. - LLM cost tracking — every LLM call anywhere in the app (chat
or background flows) lands in
llm_usagewith per-call tokens + cost. Powered bygo-ai/ledger; admin UIs query via/llm/usage(admin-api only).
See the dedicated reference page — Assistant — for the full HTTP surface, data model, event contract, and test matrix.
Trace + Live Observability
See every packet, spot bottlenecks, and get transport recommendations — without leaving the canvas. Flow Studio's Trace and Live modes, the cross-flow Lifecycle graph, and the failure Alerts inbox, powered by the flowexec pipeline.
Debug taps — trace live data through any node
Drop a passive tap on any node of a published flow and watch the data flowing through it, live — without changing the flow. Plus the roadmap for the live-debugging feature.