All example architectures
SaaS / B2B Support

Cut first-response time to seconds without sacrificing CSAT

Scenario: a 40-seat SaaS platform replaces the first hour of its ticket queue with a FlowDSL-powered assistant. The flow handles billing lookups, docs questions, and status checks; anything outside that scope publishes a handoff event that routes a scored callback to the ops inbox.

Illustrative scenario — not a customer deployment. Numbers are targets for comparable workloads.

Target outcomes (illustrative)

First response

Before42 min
After0 min

Tier-1 deflection

0%

of incoming tickets resolved without a human reading them

CSAT 30-day trend

Illustrative target trajectory after the flow ships

Cost per resolved ticket

Before3.8 USD
After0.0 USD

The flow

Click any node to inspect it. Traveling dots are live — each colour is one packet path through the pipeline.

POST /messagesin scopestreamout of scopeeventpersistWeb widgetVue / React embed~180/hrIntent classifierllm-chat · qwen2.5Answer LLMgpt-4o-miniReply to userSSE streamHandoff routerpriority + routingOps inboxemail + SlackChat ledgerMongo · TTL 30d
Click any node
The inspector shows what it does and what edges connect it to the rest of the flow.

The scenario

A typical support queue in this shape reaches 40+ minutes to first reply at weekend peaks. Most tickets are questions answerable from the docs or a billing page, yet ops spends the day on them — and the genuine escalations get buried in the flood.

The flow

One FlowDSL flow with two LLM nodes: an intent classifier and an answer generator. Anything the classifier flags as out-of-scope takes a second edge into a handoff router that persists a callback record and publishes `assistant.handoff_requested`. A separate flow picks that event up and fans out to email + Slack.

Why events matter

Notification is a separate flow from detection. If ops later wants urgent handoffs in PagerDuty, that is one new flow subscribing to the same event — zero changes to the assistant path.

Where the targets come from

Cost per resolution can fall by an order of magnitude because the cheap classifier gates the expensive answer model. Deflecting tier-1 questions frees ops to answer the real escalations in minutes instead of hours, which is what the CSAT target assumes.

The stack

  • Chat flow: redelay/event-source → redelay/llm-chat (intent) → redelay/llm-chat (answer) → SSE
  • Escalation flow: redelay/event-source on assistant.handoff_requested → redelay/email-send + Slack webhook
  • Persistence: go-modules/assistant/chats (TTL 30d) + go-modules/assistant/handoffs
  • Cost tracking: go-ai/ledger records every LLM call