Assistant (chat + handoff)
Assistant
The assistant module (backend/modules/assistant) is the project-facing chat surface. It sits in front of a FlowDSL flow — picking the published variant, launching a run, pumping flow events as SSE frames — and owns three durable collections:
| Collection | Purpose | Retention |
|---|---|---|
assistant_chats | Per-session conversation history | TTL (default 30 days) |
assistant_handoffs | Customer escalation records | Permanent (business audit) |
llm_usage | Per-call LLM cost ledger (owned by go-ai/ledger) | Time-series, default 30 days |
Chat + handoff are new in this release; LLM usage is cross-cutting and tracks every LLM call the app makes (not just assistant ones).
HTTP surface
Public (api + admin-api)
| Method | Path | Purpose |
|---|---|---|
| POST | /assistant/messages | Chat SSE stream. Server echoes X-Assistant-Session-Id. |
| GET | /assistant/chats/me/{sessionID} | Read-own transcript. Session id is the capability. |
| POST | /assistant/handoff | Request a human callback. Persists + emits assistant.handoff_requested. |
| GET | /assistant/config | Public runtime config — flow id, TTL days, anonymous allowed, and a capabilities object derived from the published flow topology (handoff/escalation/streaming). |
| POST | /assistant/reset | Reseed + publish the live flow from an embedded template (superuser). ?template=default|production. |
Admin-only (admin-api — blank-import backend/modules/assistant/admin)
| Method | Path | Purpose |
|---|---|---|
| GET | /assistant/chats | List chats. Filters: userId, flowId, variantLabel, cursor, limit. |
| GET | /assistant/chats/{id} | Fetch one chat. |
| GET | /assistant/handoffs | List handoffs. Default status=pending. |
| GET | /assistant/handoffs/{id} | Fetch one handoff. |
| PATCH | /assistant/handoffs/{id} | Update status + notes. ContactedAt + ContactedBy stamped on transition to contacted. |
All admin routes are gated on admin:access + auth:active-user.
Flow-driven capabilities
GET /assistant/config walks the currently-published flow document and reports which
features the widget should expose. The widget renders buttons only for capabilities the
flow actually implements, so adding a node in Studio flips the matching UI affordance on
without a frontend change.
{
"flowId": "assistant",
"flowName": "Project Assistant",
"handoffEnabled": true, // back-compat mirror of capabilities.handoff
"anonymousAllowed": true,
"chatTtlDays": 30,
"capabilities": {
"handoff": true, // flow has a redelay/assistant-handoff-request node
"escalation": true, // flow has a redelay/email-send fan-out
"streaming": false // llm-chat node has stream=true
}
}
Detection rules (in capabilities.go):
| Capability | Trigger |
|---|---|
handoff | Any node with action_ref=redelay/assistant-handoff-request |
escalation | Any node with action_ref=redelay/email-send |
streaming | Any redelay/llm-chat node with config.stream=true |
Adding a new capability is a two-line change: add a field to AssistantCapabilities and
teach capabilitiesFromWorkflow which node provides it.
Flow shape
The default assistant flow ships with a router edge so the "Talk to a human" button in the widget surfaces as a real branch in Studio, not a hidden side channel:
start ──► (condition: meta.intent == "handoff") ──► handoff ─┐
\ │
└────────────────────────────────────────────► chat ──►├──► end
The widget POSTs meta.intent = "handoff" when the user clicks the button; the edge
condition routes to the redelay/assistant-handoff-request node which persists the record
and publishes assistant.handoff_requested. The handoff-notification template flow
listens for that event and fans out via email.
Two variants are shipped:
| Template ID | File | Shape |
|---|---|---|
assistant/default | backend/modules/assistant/flows/default.flowdsl.json | start → {chat | handoff} → end — the minimal router + two branches. Seeded on first startup. |
assistant/production | backend/modules/assistant/flows/production.flowdsl.json | Adds an [[ESCALATE]]-token detector that fans out to an email-send node on the chat branch. |
assistant/handoff-notification | backend/modules/assistant/flows/handoff-notification.flowdsl.json | Event-driven: listens for assistant.handoff_requested and emits email.send. Install separately from the main flow. |
Pattern: domain-grounded chat (deterministic search → LLM)
RAG templates embed the question and retrieve prose chunks. When the answer should instead point at rows you own — products, programs, tickets — a deterministic node in front of the LLM is usually a better fit than tool-calling: cheaper, predictable, and the model can't invent an item that isn't in the catalog.
start ─► guard-in ─► domain-search ─► llm-chat ─► guard-out ─► end
│Block │Block
└─► refusal-in ─► rejected └─► refusal-out ─► rejected
The custom node does three things and writes them onto the packet:
func (m *Module) searchHandler(ctx context.Context, step *runtime.Step) error {
msg := latestUserMessage(step.Input) // 1. read the turn
rows := m.search(ctx, msg) // 2. query YOUR data (vector and/or facets)
// 3. hand the LLM a grounded prompt, plus UI affordances
step.Output["messages"] = []any{
map[string]any{"role": "system", "content": promptListing(rows)},
map[string]any{"role": "user", "content": msg},
}
step.Output["sources"] = rowLinks(rows) // clickable cards under the reply
step.Output["actions"] = rowActions(rows) // buttons: adopt, add to cart, book…
return nil
}
Why this composes cleanly:
llm-chatforwardssourcesandactionsverbatim, so whatever the search node attaches reaches the terminal node untouched.- The SSE pump emits them alongside the content frame, and the widget's action
registry dispatches clicks by
type— no frontend change per new action. llm-guardis a pass-through router on Allow/Flag, copying the whole packet, so guards can wrap the chain without the search node knowing.- The prompt instructs the model to recommend only from the supplied list and to reply in the user's language; retrieval stays deterministic and auditable.
Register the node on the shared engine at Startup, exactly like the assistant's own nodes:
flowexecmod.Current().Engine().RegisterHandler("myapp/domain-search", m.searchHandler)
Point the deployment at your flow with ASSISTANT_DEPLOYMENT_ID. Reach for the
tool-calling loop (llm-chat with tools → llm-tool-router) only when the model
genuinely needs to decide whether and how many times to search.
FlowDSL node: redelay/assistant-handoff-request
Package: backend/modules/assistant/flowdsl (factory id assistant-flowdsl,
depends on assistant).
One node — an action kind — wraps handoff.Service.Request. It is the
primitive behind the "Talk to a human" branch of the default flow shape. Drop
it into any flow that should turn a packet into a handoff record.
Settings:
| Setting | Type | Default | Purpose |
|---|---|---|---|
defaultPriority | enum low | normal | high | urgent | normal | Priority to apply when the incoming packet does not carry one. |
defaultReason | string | "User requested human assistance" | Fallback reason written to the admin inbox when the packet has no explicit reason. |
Ports:
| Direction | Name | Shape |
|---|---|---|
| input | Request | {sessionId, email} required; optional phone, reason, priority, userId. |
| output | Handoff | {handoffId, status, duplicate}. duplicate=true when the service deduped against an existing pending handoff for the same session (no new row written). |
| output | Error | {error}. Use this branch to notify the user that their handoff could not be filed. |
Settings-level defaultPriority / defaultReason are only applied when the
packet leaves the corresponding field blank — explicit packet values always win.
LLM tuning from Studio
The redelay/llm-chat node on the assistant flow exposes three Studio-editable settings
that replace hand-written prompt prose:
replyStyle—short | normal | detailedexamplesPolicy—avoid | when_asked | alwayssystemPrompt— free-form textarea
See go-ai — system prompt composition for the exact
ordering rules. The short version: systemPrompt is the base, the two enums append plain-
English directives when they're anything other than the neutral default.
Embedded template catalog
The module ships with 11 flow templates covering dev → production, cost-free →
quality-maxed, stateless Q&A → chat with memory. Each is a fully-realised
.flowdsl.json embedded into the binary; all are selectable via
POST /admin/assistant/reset?template=<name>. Operators can also discover them at
runtime via GET /admin/assistant/templates (returned fields: name, title,
description, tier, recommended, requires).
Tiers
Templates are tiered by quality / sophistication. Higher tier ≈ more infrastructure and cost per turn, usually better answers.
| Tier | Template | Guards | RAG | Handoff | Streaming | LLM calls / turn | Extra infra |
|---|---|---|---|---|---|---|---|
| T0 | echo | — | — | — | — | 0 | — |
| T1 | minimal | — | — | — | — | 1 | — |
| T2 | default | ✓ | — | ✓ | — | 1 + 2 guards | — |
| T2s | default-stream | ✓ | — | ✓ | ✓ | 1 + 2 guards | — |
| T3 | production | ✓ | — | ✓ | — | 1 + 2 guards | SMTP |
| T3+ | multi-guard | ✓✓ | — | — | — | 1 + 4 guards | — |
| T4 | rag | ✓ | ✓ | ✓ | — | 1 + 2 guards + 1 embed | Qdrant + embedder |
| T4b | rag-clean | ✓ | ✓ | ✓ | — | 1 + 2 guards + 1 embed | Qdrant + embedder |
| T4s | rag-stream | ✓ | ✓ | ✓ | ✓ | 1 + 2 guards + 1 embed | Qdrant + embedder |
| T4 | qa-only | ✓ | ✓ | — | — | 1 + 2 guards + 1 embed | Qdrant + embedder |
| T5 | rag-rewrite | ✓ | ✓ | ✓ | — | 2 + 2 guards + 1 embed | Qdrant + embedder |
Per-template guide
echo (T0) — start → core/template-render → end. Zero LLM tokens. Echoes
the user's latest message back in a banner. Purpose: CI smoke tests, UI development
without provider costs, wiring sanity checks when external providers are rate-limited
or down, cost-free customer demos.
minimal (T1) — start → redelay/llm-chat → end. No safety, no grounding, no
handoff. Rawest possible LLM integration. Purpose: local development, model-
comparison bench, fast feedback loops when guard latency hurts iteration.
default (T2) — minimal + redelay/llm-guard on both sides + configurable
refusals + human-handoff branch. Baseline for any production deployment. Unsafe
user input → rejected with a configurable refusal; unsafe model draft → same.
default-stream (T2s) — default with stream: true on llm-chat. Tokens
arrive as the model writes; lower time-to-first-token. Caveat: guard-out can't
suppress output mid-stream — it runs once the full draft is complete. Use for
internal tools and authenticated user bases; prefer default for public anonymous
chat where pre-send moderation matters.
production (T3) — default + trust-and-safety email fan-out on guard-out
blocks + [[ESCALATE]] token detection that short-circuits the chat path to
handoff. Public-facing deployments. Ephemeral email delivery so SMTP failures
never affect the user-facing stream.
multi-guard (T3+) — chains two redelay/llm-guard classifiers on each side
of the LLM (guard-in-1 fails open, guard-in-2 fails closed; same on the output
side). Use when single-classifier recall isn't adequate for your risk tolerance:
regulated industries, child-facing products, high-liability verticals. ~2x guard
latency per turn.
rag (T4) — default + redelay/assistant-rag-context between guard-in and
chat. Retrieves top-K chunks from the configured Qdrant index (default
redelay_docs) and prepends them as a system-prompt excerpt. LLM is instructed
to quote excerpts and cite by their [N] index. Eliminates hallucinated facts.
Requires assistant-ingest-docs CLI to have populated the index first.
rag-clean (T4b) — same retrieval as rag, but the LLM system prompts are
tuned to suppress inline [N] markers and write polished conversational prose.
Sources still ride the SSE stream; the UI renders a "Sources" footer below each
assistant turn. Use when chat tone matters more than inline traceability.
rag-stream (T4s) — rag with stream: true. Same latency tradeoffs as
default-stream; citation pills materialise as the model writes [N], and the
Sources footer appears on run.completed.
qa-only (T4) — stateless single-turn RAG. Guards + retrieval + guard-out,
but the LLM is instructed to treat each user turn independently and NOT reference
prior messages. Useful for FAQ bots, help-centre deployments, and privacy-sensitive
verticals where retaining chat context is a liability.
rag-rewrite (T5) — rag + a small preprocessor llm-chat call that rewrites
the user's question into a search-friendly query before retrieval. Terse follow-ups
("is it PHP?") and pronoun-heavy messages ("how do I configure it?") recall much
better after rewrite. The main chat still sees the ORIGINAL messages so the
conversation reads naturally. Uses the queryKey setting on rag-context to read
the rewrite from a dedicated packet field. Cost: ~50–100 extra tokens per turn.
How to pick
- Building / debugging:
echo(wiring) →minimal(model comparison) - MVP:
default(baseline prod safety + handoff) - Ship to users:
production(+ SMTP monitoring) - Docs bot:
rag(show sources) orrag-clean(polished prose) - Help centre:
qa-only(stateless, no memory liability) - Long chat sessions with follow-ups:
rag-rewrite(best recall on short queries) - Regulated / high-liability:
multi-guard+ your own audit log - Real-time feel:
default-stream/rag-stream
Roadmap — future templates not yet shipped
rag-rerank— cross-encoder rerank after vector retrieval (OVH's freebge-reranker-v2-m3). Measurably higher precision. Requires a newredelay/reranknode.agentic— tool-calling LLM withsearch_docs,read_file,list_modulestools. Uses existingredelay/llm-tool-router; requires tool-action wiring.rag-memory— chat-history summarization before retrieval. Reduces token bloat on long sessions.
Picking which flow(s) the widget uses
The assistant module references a FlowDeployment — a framework-level routing entity
that lives in go-flowdsl/flowexec/store and is shared by every module that binds flows
at run time. The deployment has one or more variants; each variant points at a flowId
(optionally pinned to a versionId) with a routing weight.
ASSISTANT_DEPLOYMENT_ID = "assistant" ← module config
│
▼
┌─── FlowDeployment ───┐
│ stable: flow.v1 90% │
│ canary: flow.v2 10% │
└──────────────────────┘
Sticky bucketing by session_id keeps the same user on the same variant across reloads
and container restarts. See the full framework reference at
go-flowdsl → FlowDeployment for the data model, store API,
and admin HTTP surface (POST /deployments/…, PATCH, DELETE, the one-shot
/deployments/{id}/variants/from-template).
Back-compat
Existing installs set ASSISTANT_FLOW_ID; the module auto-synthesises a deployment with a
single stable 100% variant pointing at that flow on first Startup. No manual migration.
Once you start managing variants through the deployment CRUD, ASSISTANT_FLOW_ID is
ignored.
Five paths to switch flows, in increasing order of flexibility:
1. Studio UI
/flowdsl → select assistant → Versions tab → click Publish on the row you want.
New chat runs go through it immediately; existing SSE connections stay on the old version
until reconnect. To swap shape entirely, use Templates → pick assistant/production →
Save as new version → Publish.
2. POST /assistant/reset?template= — one-liner template switcher
# Swap to the production variant (adds [[ESCALATE]] email fan-out)
curl -X POST -H "Authorization: Bearer $TOKEN" \
http://localhost:8001/api/v1/assistant/reset?template=production
# Back to default (minimal router)
curl -X POST -H "Authorization: Bearer $TOKEN" \
http://localhost:8001/api/v1/assistant/reset?template=default
Superuser only. Saves a new version from the embedded JSON and publishes it atomically.
History is preserved — rollback via /flows/{id}/rollback still works.
3. flowexec API — publish any existing version
# List versions
curl -H "Authorization: Bearer $TOKEN" \
http://localhost:8001/api/v1/flows/assistant/versions
# Publish one of them
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
http://localhost:8001/api/v1/flows/assistant/publish \
-d '{"versionID":"ver.XXXX"}'
4. Weighted variants within one flow — A/B between versions
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
http://localhost:8001/api/v1/flows/assistant/publish \
-d '{
"variants": [
{"label": "stable", "versionID": "ver.minimal", "weight": 80},
{"label": "canary", "versionID": "ver.production", "weight": 20}
]
}'
80/20 split across versions of the same flow, sticky per session_id.
5. Weighted variants across different flows — deployment-level A/B
Use this when variants differ in shape, not just parameters — e.g. a plain chat flow vs a RAG-enabled flow. Build or promote the flows separately, then bind them in a deployment:
# One call: create a new flow from the production template, publish v1, and append it
# as a 10% canary on the "assistant" deployment. Sticky per session_id so every user
# stays on one variant across reloads.
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
http://localhost:8001/api/v1/deployments/assistant/variants/from-template \
-d '{
"templateId": "assistant/production",
"label": "canary",
"weight": 10,
"pinVersion": false
}'
# Adjust weights later without provisioning new flows:
curl -X PATCH -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
http://localhost:8001/api/v1/deployments/assistant \
-d '{
"variants": [
{"label": "stable", "flowId": "assistant", "weight": 50},
{"label": "canary", "flowId": "flow.assistant-production-XXXXXX", "weight": 50}
]
}'
The widget picks up the new split on the next /assistant/config call. GET /assistant/config?sessionId=X
now returns activeVariant: "canary" for sessions that bucket into the canary, so telemetry
can join on the label without a second round-trip.
Flow version upgrades
ensureFlow auto-upgrades a module-owned published version when the embedded
meta.schemaVersion advances past the stored one. User-owned versions (anything with a
CreatedBy other than assistant-module) are left alone — never overwritten. Disable the
auto-upgrade entirely with ASSISTANT_FLOW_SKIP_AUTO_UPGRADE=true for fully managed
deployments.
Session management
The client holds a UUID session id in localStorage["redelay.assistant.session"]. On each POST /assistant/messages:
- Client sends
sessionIdin the body (omit on first ever message). - Server generates one if missing.
- Server echoes it via
X-Assistant-Session-Idso the client can persist. - Every subsequent request uses the same id — the
Chatdocument gets appended in place, TTL rolls forward on each turn, and sticky-by-session variant routing (see Flow Templates) routes the caller to the same variant for the life of the chat.
clear() rotates the id so a "new chat" UI button gets a fresh document rather than appending to the old one.
Chat persistence
chats.Service owns assistant_chats. Indexes:
{session_id: 1}unique{user_id: 1, last_at: -1}{flow_id: 1, variant_label: 1}{expires_at: 1}TTL (expireAfterSeconds=0)
Every append bumps expires_at = now + TTL. Dormant chats are reaped automatically by Mongo's TTL monitor. The associated HandoffRequest (if any) is not tied to the TTL — handoff rows survive TTL expiry because they're business audit.
Env config:
| Env | Default | Purpose |
|---|---|---|
ASSISTANT_CHAT_TTL_DAYS | 30 | Chat retention |
ASSISTANT_ANONYMOUS_ALLOWED | true | Gate /messages + /handoff on auth |
ASSISTANT_HANDOFF_ENABLED | true | Master switch for the handoff path |
Handoff
When a user clicks "Talk to a human" in the widget (or a flow node explicitly publishes the event), handoff.Service.Request does three things:
- Persists a
HandoffRequestrow with the transcript snapshot. Permanent — no TTL. - Attaches the handoff id back to the chat so admin UIs can navigate either way.
- Publishes
assistant.handoff_requestedvia the frameworkEventBus.
Notification is entirely decoupled — the service never sends email. A downstream flow subscribes to the event and does whatever the project needs (email, Slack, PagerDuty, CRM push). The assistant/handoff-notification flow template ships pre-wired for email: event-source → redelay/emit-event (email.send with a rendered template) → end. Customise in Studio, or replace entirely.
Dedup: a second pending handoff for the same session returns 409 Conflict. Users hitting the button repeatedly don't flood the admin inbox.
Event contract
name: assistant.handoff_requested
entityType: assistant_handoff
action: requested
topic: assistant.handoff_requested
payload:
handoffId: string (ObjectID hex)
chatId: string (ObjectID hex, "" when no chat)
sessionId: string
userId: string (optional — authed callers only)
email: string (always set)
phone: string (optional)
reason: string (optional, free-form)
priority: "low" | "normal" | "high"
transcript: string (plain-text dump at request time)
requestedAt: string (RFC3339)
Handoff schemas + the full HandoffRequest type are defined in backend/modules/assistant/handoff/model.go.
Vue integration
const {
messages, isOpen, isLoading,
sessionID, handoffRequested,
send, clear, requestHandoff,
} = useChat()
// Chat normally.
await send("How do flows work?")
// User asks for a human.
const res = await requestHandoff({
email: '[email protected]',
reason: 'billing question',
priority: 'normal',
})
if (res.ok) {
// show res.message
}
requestHandoff returns {ok, message} so the UI can render inline feedback without try/catch. The handoff button gates itself on handoffRequested so users can't double-submit.
Data model summary
Chat (assistant_chats, TTL)
├─ SessionID unique index
├─ FlowID + VersionHash + VariantLabel stamped at first create
├─ Messages[] append-only (user + assistant turns)
├─ RunIDs[] set of flow-run ids that served this chat
├─ ExpiresAt bumped on every append
└─ HandoffID set when user escalated
HandoffRequest (assistant_handoffs, permanent)
├─ ChatID + SessionID
├─ Email + Phone + Reason
├─ Transcript snapshot — readable after chat TTL
├─ Priority low | normal | high
├─ Status pending | contacted | resolved | dismissed
└─ ContactedAt + ContactedBy + Notes stamped by admin PATCH
Testing
Three test packages cover the full surface:
backend/modules/assistant/chats/service_test.go— idempotent create, TTL bump, message persistence, run-id set tracking, list pagination, handoff attach.backend/modules/assistant/handoff/service_test.go— persist-and-publish happy path, dedup, validation, no-chat fallback, status filter, state transitions, authed user-context propagation, RFC3339 format.backend/modules/assistant/public_handlers_test.go—/config,/chats/me/{id},/handoffend-to-end including dedup and anonymous gating.backend/modules/assistant/admin/module_test.go— admin listings + status transitions through the real admin handlers.
Plus the existing module_test.go covers the SSE pump, flow seeding, and content extraction.