Guides

Project AI Assistant

Wire the in-app AI assistant component to a FlowDSL flow executed by your Redelay backend. The frontend stays framework-agnostic; all LLM logic lives in a flow you can edit without a redeploy.

Project AI Assistant

The Redelay project template ships a small "talk-to-your-app" assistant with a deliberate split of concerns:

  • The frontend (AiAssistant.vue + useChat.ts) knows nothing about flows, nodes, or LLM providers. It POSTs {messages} to /api/chat and reads back a server-sent-events stream of {content} / {done} / {error} frames.
  • A Nitro proxy in the Nuxt website forwards those requests to the Go backend. It enforces rate limits and caps conversation depth but carries no flow or credential logic.
  • A project-local Go module (backend/modules/assistant) resolves the current published version of a FlowDSL flow, launches a run for each incoming request, and translates the run's FlowEvents into the chat SSE wire format.
  • The flow itself — shipped as an embedded default and stored in the flowexec store the first time the backend boots — carries the actual provider, model, prompt, and any extra routing you want. Edit it through the standard flowexec REST API; changes take effect on the next message with no restart.

This document walks through the wiring end-to-end so you can customize or replace any layer independently.

Architecture at a glance

text
Browser
  │  POST /api/chat                  {messages}
  ▼
Nuxt Nitro proxy  (website/server/api/chat.post.ts)
  │  POST /api/v1/assistant/messages {messages}
  ▼
Go backend — modules/assistant
  │  resolves published version of flow "assistant"
  │  flowexec.Executor().RunAsync(RunInput{...})
  │  flowexec.Subscribe(runID) → FlowEvents
  ▼
FlowDSL flow "assistant"
  start → guard-in → chat (redelay/llm-chat) → guard-out → end
                  ↘                           ↘
                    rejected                   rejected

The frontend only knows about /api/chat and the SSE frame shape. Swapping the backing flow for a RAG pipeline, a tool-router, or a multi-step agent changes nothing about the browser bundle.

Guarded by default

Both the default and production assistant flows wrap the LLM call with redelay/llm-guard router nodes — one before the chat (classifies user input) and one after (classifies assistant output). On a Block verdict the flow routes through a core/template-render transform (refusal-in / refusal-out) that stamps a configurable refusal message onto the packet before the rejected terminal, so the client always receives a useful reply instead of silence. The production variant additionally fans blocked outputs to a trust-and-safety email via ephemeral delivery so a send failure cannot block the user-facing stream.

Guard behaviour is tunable without a code change:

  • Set GUARD_PROVIDER=qwen3guard to turn on classification; leave unset to disable (the flow still runs — the nodes simply route everything to Allow).
  • GUARD_MODEL (default Qwen3Guard-Gen-8B — free on OVH) selects the classifier variant.
  • Admin settings ai-guard.default_on_error and ai-guard.block_flagged tune fail-open vs fail-closed behaviour and whether controversial messages are treated as blocks.
  • The refusal copy is just the template setting on the refusal-in / refusal-out nodes — edit it in Studio, save a new version, publish. Interpolation works: {{.decision.reason}}, {{.decision.categories}}.

The SSE pump streams content from whichever terminal the flow reaches (end on success, rejected on block) — ASSISTANT_OUTPUT_NODE is a comma-separated list of terminal ids; the default end,rejected covers both without a code change.

See the go-ai guard reference for the full contract and the redelay/llm-guard node doc for the refusal wiring pattern.

Step 1 — Backend wiring

backend/cmd/api/main.go blank-imports the three packages that matter:

go
import (
    _ "github.com/redelay/go-flowdsl/flowexec/module"            // flow storage + runner
    _ "github.com/redelay/go-ai/llm/flowdsl"              // llm-chat / llm-embed FlowDSL nodes
    _ "github.com/redelay/go-ai/llm/providers/ollama"     // at least one provider
    _ "github.com/redelay/backend/modules/assistant"      // the chat endpoint + default flow
)

Order matters: flowexec registers its factory first, so when the assistant factory runs it can resolve flowexecmod.Current() and hold a reference.

After server.Default() returns, main.go initialises the LLM provider registry from environment variables and registers the redelay/llm-chat handler with the flow engine:

go
llm.BootstrapFromEnv()
llmflowdsl.Register(fx.Engine())

At this point the server exposes two new endpoints under ROUTE_PREFIX:

MethodPathPurpose
POST/assistant/messagesChat SSE stream
POST/assistant/resetRe-seed the flow from the embedded default (superuser)

Step 2 — The default flow

backend/modules/assistant/flows/default.flowdsl.json is embedded into the binary via go:embed:

json
{
  "id": "assistant",
  "name": "Project Assistant (default)",
  "nodes": [
    { "id": "start", "kind": "start" },
    {
      "id": "chat", "kind": "action",
      "action_ref": "redelay/llm-chat",
      "config": {
        "providerID": "${ASSISTANT_PROVIDER:-ollama}",
        "model":      "${ASSISTANT_MODEL:-mistral-nemo:12b}",
        "temperature": 0.2,
        "stream": false
      }
    },
    { "id": "end", "kind": "end" }
  ],
  "edges": [
    { "id": "e1", "from": "start", "to": "chat" },
    { "id": "e2", "from": "chat",  "to": "end"  }
  ]
}

On the first startup the module calls ensureFlow(ctx, force=false):

  1. If no flow with ID "assistant" exists, it is created, the embedded JSON is saved as version 1, and that version is published.
  2. If a flow already exists, the embedded default is ignored — whatever the project has saved and published is what runs.

This gives you "works out of the box" on a fresh install and "my customizations win" from that point on.

Step 3 — Request lifecycle

Each POST to /assistant/messages:

  1. Decodes {messages, meta}.
  2. Resolves the current published version via flowexec.Store().GetPublishedVersion(ctx, "assistant").
  3. Builds RunInput{Workflow, VersionID, VersionHash, Input: {messages: ...}} and launches via Executor().RunAsync.
  4. Subscribes to live FlowEvents for the returned run ID.
  5. Translates each event into a chat-frame:
    • KindNodeDone on the configured ASSISTANT_OUTPUT_NODE (chat by default) → {content} frame, content pulled out of the node's payload.
    • KindNodeDone on any other node → {trace} frame (diagnostics, skipped by the default UI).
    • KindNodeFailed → {error} frame.
    • KindRunCompleted → terminal {done: true} frame.
    • KindRunFailed → {error} frame and close.

The extractContent helper accepts several payload shapes (content, message.content, output.content, output.message.content, or output as a string), so you can swap the chat node for any handler that writes one of those keys.

Step 4 — Frontend wiring

website/app/components/AiAssistant.vue drops into any Nuxt page and renders a brand-styled slide-over. It delegates everything to useChat.ts, which:

  • POSTs to /api/chat with {messages}.
  • Streams the SSE body, coalescing {content} frames into the in-progress assistant message.
  • Treats {done} as success, {error} as failure, and ignores {trace}.
  • Stores conversation state in a Nuxt useState singleton so the composable survives route changes.

The Nitro proxy in website/server/api/chat.post.ts is the only place the website speaks to the backend:

text
POST /api/chat → ${REDELAY_BACKEND_URL}${REDELAY_ROUTE_PREFIX}/assistant/messages

It pipes the backend's SSE body straight through to the browser.

Step 5 — Customizing the flow

Because the flow lives in the flowexec store, you edit it the same way you would any other flow — no code change needed:

shell
# Fetch current version
curl http://localhost:8000/api/v1/flowexec/flows/assistant

# Save a new version with a richer graph (prompt-template → retrieval → llm-chat)
curl -X POST http://localhost:8000/api/v1/flowexec/flows/assistant/versions \
  -H 'Content-Type: application/json' \
  --data @my-rag-flow.flowdsl.json

# Publish the new version
curl -X POST http://localhost:8000/api/v1/flowexec/flows/assistant/publish \
  -H 'Content-Type: application/json' \
  -d '{"versionID": "<id from previous response>"}'

The next message your users send runs on the new flow. Old runs remain pinned to their original versionHash, so historical traces stay reproducible.

To revert to the module's embedded default, hit the superuser-only reset endpoint:

shell
curl -X POST http://localhost:8000/api/v1/assistant/reset \
  -H "Authorization: Bearer $SUPERUSER_TOKEN"

Environment variables

Assistant module

VariableDefaultPurpose
ASSISTANT_FLOW_IDassistantFlow key in the flowexec store
ASSISTANT_INPUT_KEYmessagesMap key under which the chat history is placed in RunInput.Input
ASSISTANT_OUTPUT_NODEend,rejectedComma-separated list of terminal node ids whose node.done payload is streamed as content — defaults cover the success (end) and guard-block (rejected) terminals
ASSISTANT_PROVIDERollamaReferenced by ${...} interpolation in the default flow's providerID
ASSISTANT_MODELmistral-nemo:12bReferenced by ${...} interpolation in the default flow's model

LLM provider registry (go-ai/llm)

BootstrapFromEnv() reads LLM_PROVIDER and calls the matching factory. Each factory then reads its own env vars:

VariableDefaultProvider
LLM_PROVIDER(none — skip)Which factory to build: ollama, openai, anthropic, gemini, ovh
OLLAMA_URLhttp://localhost:11434Ollama API base URL
OLLAMA_MODELllama3.1Ollama default model (overridden by flow config)
OPENAI_API_KEY(required)OpenAI API key
ANTHROPIC_API_KEY(required)Anthropic API key
GEMINI_API_KEY(required)Google Gemini API key
OVH_AI_ENDPOINTS_TOKEN(required)OVH AI Endpoints API token

All providers also accept optional baseURL and defaultModel via the factory config map, but when launched from BootstrapFromEnv() those are read from env only (see each provider's godoc for the full list).

Website proxy (Nitro)

VariableDefaultPurpose
REDELAY_BACKEND_URLhttp://localhost:8000Go backend base URL
REDELAY_ROUTE_PREFIX/api/v1Backend route prefix

Testing

The module ships a unit test (backend/modules/assistant/module_test.go) that exercises the whole path against the flowexec memstore and a stub redelay/llm-chat handler. It covers:

  • Seeding the default flow on a fresh store.
  • Preserving project customizations on subsequent startups.
  • The end-to-end SSE stream via httptest.NewServer (real http.Flusher).
  • 400 on missing messages and 409 when no version is published.
  • extractContent fallbacks for the common payload shapes.
shell
cd backend && go test ./modules/assistant/... -v

When to reach past this module

If your assistant needs:

  • Multi-tenant flows (different prompts per customer): customize ASSISTANT_FLOW_ID at request time via meta and add a router node that dispatches on the tenant — no module change.
  • Tool calls / RAG: add nodes to the flow. redelay/llm-tool-router and any node you register on the engine are available.
  • Different transport (WebSocket, gRPC): replace routes.go — the rest of the module is transport-agnostic.

The module is intentionally thin so you can fork the entire folder into a project and iterate freely without fighting the framework.

Production patterns

The repo ships three flow templates that showcase how to grow past the three-node default as your assistant takes on more responsibility. Load any of them via POST /api/v1/flows/assistant/versions?format=spec and publish.

1. Human hand-off on escalation

File: spec/examples/assistant-human-handoff.flowdsl.json

The LLM's system prompt instructs it to prefix replies with a literal [[ESCALATE]] token when it believes the request needs a human (explicit request, safety-critical topic, low confidence). A fan-out edge with a condition: contains(input.message.content, "[[ESCALATE]]") triggers redelay/email-send on the same output — the user still gets their reply, and the on-call team gets notified.

text
 start ─► chat ─► end            (primary: user-facing SSE)
          └────► notify-team     (fan-out, only on [[ESCALATE]] token,
                                  ephemeral delivery so the email
                                  never blocks the chat response)

2. Tool-calling with a router

File: spec/examples/assistant-with-tools.flowdsl.json

Uses redelay/llm-tool-router (kind: router) to branch on whether the LLM returned toolCalls. The Tool port loops through a project- specific run-tool handler that executes the call and appends the result to messages; the Message port flows straight to the sink.

text
 start ─► chat ─► router ─► end            (plain reply path)
                  └────► run-tool ─► chat  (tool loop — append result,
                                            re-query the model)

3. Production default with fan-out audit + escalation

File: backend/modules/assistant/flows/production.flowdsl.json

The full template that the assistant module can ship as its default instead of the minimal three-node file. It adds:

  • An inline SLO on the chat node (p95_latency_ms, max_error_rate) so Live-mode shows violation badges the moment the LLM provider gets slow. See Trace + Live Observability.
  • The same [[ESCALATE]] hand-off pattern as example 1.
  • A single-file drop-in — copy over the minimal default with cp backend/modules/assistant/flows/production.flowdsl.json backend/modules/assistant/flows/default.flowdsl.json and restart, or load it via the flowexec API without touching code.

What to add next

Node refs used above (myapp/tool-dispatch, any persistence hook) must be registered on the runtime engine at startup via flowexec.Engine().RegisterHandler("…", handler). The Running Flows guide walks through the handler registration shape.

Pair these with Trace mode and Live mode to watch them execute in real time while you tune prompts, thresholds, and escalation triggers.

Persistence + handoff (production-ready add-ons)

The assistant module ships three things beyond the chat SSE:

  • Chat persistence with a TTL (default 30 days). Every message is written to assistant_chats under a session_id the client keeps in localStorage. Reconnecting clients see their own history via GET /assistant/chats/me/{sessionID}.
  • Manager handoff — POST /assistant/handoff persists a HandoffRequest row (permanent) AND publishes the assistant.handoff_requested event via the EventBus. A downstream flow delivers the notification (email, Slack, whatever). The ships-with template assistant/handoff-notification wires event → email.send for the common case.
  • LLM cost tracking — every LLM call anywhere in the app (chat or background flows) lands in llm_usage with per-call tokens + cost. Powered by go-ai/ledger; admin UIs query via /llm/usage (admin-api only).

See the dedicated reference page — Assistant — for the full HTTP surface, data model, event contract, and test matrix.