Reference

Assistant (chat + handoff)

Project-local AI assistant module — FlowDSL-backed chat, persisted conversations with TTL, LLM cost tracking, and a human-handoff path that fans out via the event bus.

Assistant

The assistant module (backend/modules/assistant) is the project-facing chat surface. It sits in front of a FlowDSL flow — picking the published variant, launching a run, pumping flow events as SSE frames — and owns three durable collections:

CollectionPurposeRetention
assistant_chatsPer-session conversation historyTTL (default 30 days)
assistant_handoffsCustomer escalation recordsPermanent (business audit)
llm_usagePer-call LLM cost ledger (owned by go-ai/ledger)Time-series, default 30 days

Chat + handoff are new in this release; LLM usage is cross-cutting and tracks every LLM call the app makes (not just assistant ones).

HTTP surface

Public (api + admin-api)

MethodPathPurpose
POST/assistant/messagesChat SSE stream. Server echoes X-Assistant-Session-Id.
GET/assistant/chats/me/{sessionID}Read-own transcript. Session id is the capability.
POST/assistant/handoffRequest a human callback. Persists + emits assistant.handoff_requested.
GET/assistant/configPublic runtime config — flow id, TTL days, anonymous allowed, and a capabilities object derived from the published flow topology (handoff/escalation/streaming).
POST/assistant/resetReseed + publish the live flow from an embedded template (superuser). ?template=default|production.

Admin-only (admin-api — blank-import backend/modules/assistant/admin)

MethodPathPurpose
GET/assistant/chatsList chats. Filters: userId, flowId, variantLabel, cursor, limit.
GET/assistant/chats/{id}Fetch one chat.
GET/assistant/handoffsList handoffs. Default status=pending.
GET/assistant/handoffs/{id}Fetch one handoff.
PATCH/assistant/handoffs/{id}Update status + notes. ContactedAt + ContactedBy stamped on transition to contacted.

All admin routes are gated on admin:access + auth:active-user.

Flow-driven capabilities

GET /assistant/config walks the currently-published flow document and reports which features the widget should expose. The widget renders buttons only for capabilities the flow actually implements, so adding a node in Studio flips the matching UI affordance on without a frontend change.

jsonc
{
  "flowId": "assistant",
  "flowName": "Project Assistant",
  "handoffEnabled": true,                 // back-compat mirror of capabilities.handoff
  "anonymousAllowed": true,
  "chatTtlDays": 30,
  "capabilities": {
    "handoff":    true,    // flow has a redelay/assistant-handoff-request node
    "escalation": true,    // flow has a redelay/email-send fan-out
    "streaming":  false    // llm-chat node has stream=true
  }
}

Detection rules (in capabilities.go):

CapabilityTrigger
handoffAny node with action_ref=redelay/assistant-handoff-request
escalationAny node with action_ref=redelay/email-send
streamingAny redelay/llm-chat node with config.stream=true

Adding a new capability is a two-line change: add a field to AssistantCapabilities and teach capabilitiesFromWorkflow which node provides it.

Flow shape

The default assistant flow ships with a router edge so the "Talk to a human" button in the widget surfaces as a real branch in Studio, not a hidden side channel:

text
start ──► (condition: meta.intent == "handoff") ──► handoff ─┐
      \                                                       │
       └────────────────────────────────────────────► chat ──►├──► end

The widget POSTs meta.intent = "handoff" when the user clicks the button; the edge condition routes to the redelay/assistant-handoff-request node which persists the record and publishes assistant.handoff_requested. The handoff-notification template flow listens for that event and fans out via email.

Two variants are shipped:

Template IDFileShape
assistant/defaultbackend/modules/assistant/flows/default.flowdsl.jsonstart → {chat | handoff} → end — the minimal router + two branches. Seeded on first startup.
assistant/productionbackend/modules/assistant/flows/production.flowdsl.jsonAdds an [[ESCALATE]]-token detector that fans out to an email-send node on the chat branch.
assistant/handoff-notificationbackend/modules/assistant/flows/handoff-notification.flowdsl.jsonEvent-driven: listens for assistant.handoff_requested and emits email.send. Install separately from the main flow.

Pattern: domain-grounded chat (deterministic search → LLM)

RAG templates embed the question and retrieve prose chunks. When the answer should instead point at rows you own — products, programs, tickets — a deterministic node in front of the LLM is usually a better fit than tool-calling: cheaper, predictable, and the model can't invent an item that isn't in the catalog.

text
start ─► guard-in ─► domain-search ─► llm-chat ─► guard-out ─► end
             │Block                       │Block
             └─► refusal-in ─► rejected   └─► refusal-out ─► rejected

The custom node does three things and writes them onto the packet:

go
func (m *Module) searchHandler(ctx context.Context, step *runtime.Step) error {
    msg := latestUserMessage(step.Input)     // 1. read the turn
    rows := m.search(ctx, msg)               // 2. query YOUR data (vector and/or facets)

    // 3. hand the LLM a grounded prompt, plus UI affordances
    step.Output["messages"] = []any{
        map[string]any{"role": "system", "content": promptListing(rows)},
        map[string]any{"role": "user", "content": msg},
    }
    step.Output["sources"] = rowLinks(rows)   // clickable cards under the reply
    step.Output["actions"] = rowActions(rows) // buttons: adopt, add to cart, book…
    return nil
}

Why this composes cleanly:

  • llm-chat forwards sources and actions verbatim, so whatever the search node attaches reaches the terminal node untouched.
  • The SSE pump emits them alongside the content frame, and the widget's action registry dispatches clicks by type — no frontend change per new action.
  • llm-guard is a pass-through router on Allow/Flag, copying the whole packet, so guards can wrap the chain without the search node knowing.
  • The prompt instructs the model to recommend only from the supplied list and to reply in the user's language; retrieval stays deterministic and auditable.

Register the node on the shared engine at Startup, exactly like the assistant's own nodes:

go
flowexecmod.Current().Engine().RegisterHandler("myapp/domain-search", m.searchHandler)

Point the deployment at your flow with ASSISTANT_DEPLOYMENT_ID. Reach for the tool-calling loop (llm-chat with tools → llm-tool-router) only when the model genuinely needs to decide whether and how many times to search.

FlowDSL node: redelay/assistant-handoff-request

Package: backend/modules/assistant/flowdsl (factory id assistant-flowdsl, depends on assistant).

One node — an action kind — wraps handoff.Service.Request. It is the primitive behind the "Talk to a human" branch of the default flow shape. Drop it into any flow that should turn a packet into a handoff record.

Settings:

SettingTypeDefaultPurpose
defaultPriorityenum low | normal | high | urgentnormalPriority to apply when the incoming packet does not carry one.
defaultReasonstring"User requested human assistance"Fallback reason written to the admin inbox when the packet has no explicit reason.

Ports:

DirectionNameShape
inputRequest{sessionId, email} required; optional phone, reason, priority, userId.
outputHandoff{handoffId, status, duplicate}. duplicate=true when the service deduped against an existing pending handoff for the same session (no new row written).
outputError{error}. Use this branch to notify the user that their handoff could not be filed.

Settings-level defaultPriority / defaultReason are only applied when the packet leaves the corresponding field blank — explicit packet values always win.

LLM tuning from Studio

The redelay/llm-chat node on the assistant flow exposes three Studio-editable settings that replace hand-written prompt prose:

  • replyStyle — short | normal | detailed
  • examplesPolicy — avoid | when_asked | always
  • systemPrompt — free-form textarea

See go-ai — system prompt composition for the exact ordering rules. The short version: systemPrompt is the base, the two enums append plain- English directives when they're anything other than the neutral default.

Embedded template catalog

The module ships with 11 flow templates covering dev → production, cost-free → quality-maxed, stateless Q&A → chat with memory. Each is a fully-realised .flowdsl.json embedded into the binary; all are selectable via POST /admin/assistant/reset?template=<name>. Operators can also discover them at runtime via GET /admin/assistant/templates (returned fields: name, title, description, tier, recommended, requires).

Tiers

Templates are tiered by quality / sophistication. Higher tier ≈ more infrastructure and cost per turn, usually better answers.

TierTemplateGuardsRAGHandoffStreamingLLM calls / turnExtra infra
T0echo————0—
T1minimal————1—
T2default✓—✓—1 + 2 guards—
T2sdefault-stream✓—✓✓1 + 2 guards—
T3production✓—✓—1 + 2 guardsSMTP
T3+multi-guard✓✓———1 + 4 guards—
T4rag✓✓✓—1 + 2 guards + 1 embedQdrant + embedder
T4brag-clean✓✓✓—1 + 2 guards + 1 embedQdrant + embedder
T4srag-stream✓✓✓✓1 + 2 guards + 1 embedQdrant + embedder
T4qa-only✓✓——1 + 2 guards + 1 embedQdrant + embedder
T5rag-rewrite✓✓✓—2 + 2 guards + 1 embedQdrant + embedder

Per-template guide

echo (T0) — start → core/template-render → end. Zero LLM tokens. Echoes the user's latest message back in a banner. Purpose: CI smoke tests, UI development without provider costs, wiring sanity checks when external providers are rate-limited or down, cost-free customer demos.

minimal (T1) — start → redelay/llm-chat → end. No safety, no grounding, no handoff. Rawest possible LLM integration. Purpose: local development, model- comparison bench, fast feedback loops when guard latency hurts iteration.

default (T2) — minimal + redelay/llm-guard on both sides + configurable refusals + human-handoff branch. Baseline for any production deployment. Unsafe user input → rejected with a configurable refusal; unsafe model draft → same.

default-stream (T2s) — default with stream: true on llm-chat. Tokens arrive as the model writes; lower time-to-first-token. Caveat: guard-out can't suppress output mid-stream — it runs once the full draft is complete. Use for internal tools and authenticated user bases; prefer default for public anonymous chat where pre-send moderation matters.

production (T3) — default + trust-and-safety email fan-out on guard-out blocks + [[ESCALATE]] token detection that short-circuits the chat path to handoff. Public-facing deployments. Ephemeral email delivery so SMTP failures never affect the user-facing stream.

multi-guard (T3+) — chains two redelay/llm-guard classifiers on each side of the LLM (guard-in-1 fails open, guard-in-2 fails closed; same on the output side). Use when single-classifier recall isn't adequate for your risk tolerance: regulated industries, child-facing products, high-liability verticals. ~2x guard latency per turn.

rag (T4) — default + redelay/assistant-rag-context between guard-in and chat. Retrieves top-K chunks from the configured Qdrant index (default redelay_docs) and prepends them as a system-prompt excerpt. LLM is instructed to quote excerpts and cite by their [N] index. Eliminates hallucinated facts. Requires assistant-ingest-docs CLI to have populated the index first.

rag-clean (T4b) — same retrieval as rag, but the LLM system prompts are tuned to suppress inline [N] markers and write polished conversational prose. Sources still ride the SSE stream; the UI renders a "Sources" footer below each assistant turn. Use when chat tone matters more than inline traceability.

rag-stream (T4s) — rag with stream: true. Same latency tradeoffs as default-stream; citation pills materialise as the model writes [N], and the Sources footer appears on run.completed.

qa-only (T4) — stateless single-turn RAG. Guards + retrieval + guard-out, but the LLM is instructed to treat each user turn independently and NOT reference prior messages. Useful for FAQ bots, help-centre deployments, and privacy-sensitive verticals where retaining chat context is a liability.

rag-rewrite (T5) — rag + a small preprocessor llm-chat call that rewrites the user's question into a search-friendly query before retrieval. Terse follow-ups ("is it PHP?") and pronoun-heavy messages ("how do I configure it?") recall much better after rewrite. The main chat still sees the ORIGINAL messages so the conversation reads naturally. Uses the queryKey setting on rag-context to read the rewrite from a dedicated packet field. Cost: ~50–100 extra tokens per turn.

How to pick

  • Building / debugging: echo (wiring) → minimal (model comparison)
  • MVP: default (baseline prod safety + handoff)
  • Ship to users: production (+ SMTP monitoring)
  • Docs bot: rag (show sources) or rag-clean (polished prose)
  • Help centre: qa-only (stateless, no memory liability)
  • Long chat sessions with follow-ups: rag-rewrite (best recall on short queries)
  • Regulated / high-liability: multi-guard + your own audit log
  • Real-time feel: default-stream / rag-stream

Roadmap — future templates not yet shipped

  • rag-rerank — cross-encoder rerank after vector retrieval (OVH's free bge-reranker-v2-m3). Measurably higher precision. Requires a new redelay/rerank node.
  • agentic — tool-calling LLM with search_docs, read_file, list_modules tools. Uses existing redelay/llm-tool-router; requires tool-action wiring.
  • rag-memory — chat-history summarization before retrieval. Reduces token bloat on long sessions.

Picking which flow(s) the widget uses

The assistant module references a FlowDeployment — a framework-level routing entity that lives in go-flowdsl/flowexec/store and is shared by every module that binds flows at run time. The deployment has one or more variants; each variant points at a flowId (optionally pinned to a versionId) with a routing weight.

text
ASSISTANT_DEPLOYMENT_ID = "assistant"    ← module config
                          │
                          ▼
              ┌─── FlowDeployment ───┐
              │ stable: flow.v1 90%  │
              │ canary: flow.v2 10%  │
              └──────────────────────┘

Sticky bucketing by session_id keeps the same user on the same variant across reloads and container restarts. See the full framework reference at go-flowdsl → FlowDeployment for the data model, store API, and admin HTTP surface (POST /deployments/…, PATCH, DELETE, the one-shot /deployments/{id}/variants/from-template).

Back-compat

Existing installs set ASSISTANT_FLOW_ID; the module auto-synthesises a deployment with a single stable 100% variant pointing at that flow on first Startup. No manual migration. Once you start managing variants through the deployment CRUD, ASSISTANT_FLOW_ID is ignored.

Five paths to switch flows, in increasing order of flexibility:

1. Studio UI

/flowdsl → select assistant → Versions tab → click Publish on the row you want. New chat runs go through it immediately; existing SSE connections stay on the old version until reconnect. To swap shape entirely, use Templates → pick assistant/production → Save as new version → Publish.

2. POST /assistant/reset?template= — one-liner template switcher

shell
# Swap to the production variant (adds [[ESCALATE]] email fan-out)
curl -X POST -H "Authorization: Bearer $TOKEN" \
  http://localhost:8001/api/v1/assistant/reset?template=production

# Back to default (minimal router)
curl -X POST -H "Authorization: Bearer $TOKEN" \
  http://localhost:8001/api/v1/assistant/reset?template=default

Superuser only. Saves a new version from the embedded JSON and publishes it atomically. History is preserved — rollback via /flows/{id}/rollback still works.

3. flowexec API — publish any existing version

shell
# List versions
curl -H "Authorization: Bearer $TOKEN" \
  http://localhost:8001/api/v1/flows/assistant/versions

# Publish one of them
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  http://localhost:8001/api/v1/flows/assistant/publish \
  -d '{"versionID":"ver.XXXX"}'

4. Weighted variants within one flow — A/B between versions

shell
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  http://localhost:8001/api/v1/flows/assistant/publish \
  -d '{
    "variants": [
      {"label": "stable", "versionID": "ver.minimal",    "weight": 80},
      {"label": "canary", "versionID": "ver.production", "weight": 20}
    ]
  }'

80/20 split across versions of the same flow, sticky per session_id.

5. Weighted variants across different flows — deployment-level A/B

Use this when variants differ in shape, not just parameters — e.g. a plain chat flow vs a RAG-enabled flow. Build or promote the flows separately, then bind them in a deployment:

shell
# One call: create a new flow from the production template, publish v1, and append it
# as a 10% canary on the "assistant" deployment. Sticky per session_id so every user
# stays on one variant across reloads.
curl -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  http://localhost:8001/api/v1/deployments/assistant/variants/from-template \
  -d '{
    "templateId": "assistant/production",
    "label":      "canary",
    "weight":     10,
    "pinVersion": false
  }'

# Adjust weights later without provisioning new flows:
curl -X PATCH -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  http://localhost:8001/api/v1/deployments/assistant \
  -d '{
    "variants": [
      {"label": "stable", "flowId": "assistant",                         "weight": 50},
      {"label": "canary", "flowId": "flow.assistant-production-XXXXXX",  "weight": 50}
    ]
  }'

The widget picks up the new split on the next /assistant/config call. GET /assistant/config?sessionId=X now returns activeVariant: "canary" for sessions that bucket into the canary, so telemetry can join on the label without a second round-trip.

Flow version upgrades

ensureFlow auto-upgrades a module-owned published version when the embedded meta.schemaVersion advances past the stored one. User-owned versions (anything with a CreatedBy other than assistant-module) are left alone — never overwritten. Disable the auto-upgrade entirely with ASSISTANT_FLOW_SKIP_AUTO_UPGRADE=true for fully managed deployments.

Session management

The client holds a UUID session id in localStorage["redelay.assistant.session"]. On each POST /assistant/messages:

  1. Client sends sessionId in the body (omit on first ever message).
  2. Server generates one if missing.
  3. Server echoes it via X-Assistant-Session-Id so the client can persist.
  4. Every subsequent request uses the same id — the Chat document gets appended in place, TTL rolls forward on each turn, and sticky-by-session variant routing (see Flow Templates) routes the caller to the same variant for the life of the chat.

clear() rotates the id so a "new chat" UI button gets a fresh document rather than appending to the old one.

Chat persistence

chats.Service owns assistant_chats. Indexes:

  • {session_id: 1} unique
  • {user_id: 1, last_at: -1}
  • {flow_id: 1, variant_label: 1}
  • {expires_at: 1} TTL (expireAfterSeconds=0)

Every append bumps expires_at = now + TTL. Dormant chats are reaped automatically by Mongo's TTL monitor. The associated HandoffRequest (if any) is not tied to the TTL — handoff rows survive TTL expiry because they're business audit.

Env config:

EnvDefaultPurpose
ASSISTANT_CHAT_TTL_DAYS30Chat retention
ASSISTANT_ANONYMOUS_ALLOWEDtrueGate /messages + /handoff on auth
ASSISTANT_HANDOFF_ENABLEDtrueMaster switch for the handoff path

Handoff

When a user clicks "Talk to a human" in the widget (or a flow node explicitly publishes the event), handoff.Service.Request does three things:

  1. Persists a HandoffRequest row with the transcript snapshot. Permanent — no TTL.
  2. Attaches the handoff id back to the chat so admin UIs can navigate either way.
  3. Publishes assistant.handoff_requested via the framework EventBus.

Notification is entirely decoupled — the service never sends email. A downstream flow subscribes to the event and does whatever the project needs (email, Slack, PagerDuty, CRM push). The assistant/handoff-notification flow template ships pre-wired for email: event-source → redelay/emit-event (email.send with a rendered template) → end. Customise in Studio, or replace entirely.

Dedup: a second pending handoff for the same session returns 409 Conflict. Users hitting the button repeatedly don't flood the admin inbox.

Event contract

yaml
name: assistant.handoff_requested
entityType: assistant_handoff
action: requested
topic: assistant.handoff_requested
payload:
  handoffId:   string (ObjectID hex)
  chatId:      string (ObjectID hex, "" when no chat)
  sessionId:   string
  userId:      string (optional — authed callers only)
  email:       string (always set)
  phone:       string (optional)
  reason:      string (optional, free-form)
  priority:    "low" | "normal" | "high"
  transcript:  string (plain-text dump at request time)
  requestedAt: string (RFC3339)

Handoff schemas + the full HandoffRequest type are defined in backend/modules/assistant/handoff/model.go.

Vue integration

ts
const {
  messages, isOpen, isLoading,
  sessionID, handoffRequested,
  send, clear, requestHandoff,
} = useChat()

// Chat normally.
await send("How do flows work?")

// User asks for a human.
const res = await requestHandoff({
  email: '[email protected]',
  reason: 'billing question',
  priority: 'normal',
})
if (res.ok) {
  // show res.message
}

requestHandoff returns {ok, message} so the UI can render inline feedback without try/catch. The handoff button gates itself on handoffRequested so users can't double-submit.

Data model summary

text
Chat            (assistant_chats, TTL)
├─ SessionID         unique index
├─ FlowID + VersionHash + VariantLabel    stamped at first create
├─ Messages[]        append-only (user + assistant turns)
├─ RunIDs[]          set of flow-run ids that served this chat
├─ ExpiresAt         bumped on every append
└─ HandoffID         set when user escalated

HandoffRequest  (assistant_handoffs, permanent)
├─ ChatID + SessionID
├─ Email + Phone + Reason
├─ Transcript        snapshot — readable after chat TTL
├─ Priority          low | normal | high
├─ Status            pending | contacted | resolved | dismissed
└─ ContactedAt + ContactedBy + Notes   stamped by admin PATCH

Testing

Three test packages cover the full surface:

  • backend/modules/assistant/chats/service_test.go — idempotent create, TTL bump, message persistence, run-id set tracking, list pagination, handoff attach.
  • backend/modules/assistant/handoff/service_test.go — persist-and-publish happy path, dedup, validation, no-chat fallback, status filter, state transitions, authed user-context propagation, RFC3339 format.
  • backend/modules/assistant/public_handlers_test.go — /config, /chats/me/{id}, /handoff end-to-end including dedup and anonymous gating.
  • backend/modules/assistant/admin/module_test.go — admin listings + status transitions through the real admin handlers.

Plus the existing module_test.go covers the SSE pump, flow seeding, and content extraction.