All example architectures
Internal Tools

A company assistant that actually knows the company

Scenario: a 200-person engineering org replaces "ask on Slack" with an assistant that pulls context from the wiki, the ticket tracker, and a Postgres schema catalog — all through a single FlowDSL flow. Every answer cites sources. Every LLM call is costed to the asking team.

Illustrative scenario — not a customer deployment. Numbers are targets for comparable workloads.

Target outcomes (illustrative)

Median answer time

Before18 min
After0.0 min

Resolved without human

0%

Questions the assistant answers end-to-end with citations

Weekly question volume

Illustrative target adoption curve

Cost per answer

Before0.31 USD
After0.0 USD

The flow

Click any node to inspect it. Traveling dots are live — each colour is one packet path through the pipeline.

questionveckeywordentitycontexttokensusageSlack / Webuser asksEmbed queryllm-embedWiki searchvector storeTickets searchmongoops/findSchema lookuppostgres/queryCompose contexttop-k mergeAnswer LLMwith citationsStream replySSE + citationsCost attributionper-team
Click any node
The inspector shows what it does and what edges connect it to the rest of the flow.

The scenario

The wiki exists, the ticket history exists, the schema catalog exists — but "where does the billing code live" still turns into a 4-hop Slack thread that lands on whoever happens to be online.

The flow

One retrieval flow that fans out to three knowledge sources in parallel, merges the top-k results, and feeds an answer LLM that's instructed to cite every claim. Each source node is swappable — adding Confluence is one node, not a project.

Cost attribution

The assistant writes `team_id` into `runInput.meta` on every question. The LLM ledger records it on every call. Finance generates the internal chargeback report from the same collection.

Where the targets come from

The target takes median time from ~18 minutes (Slack ask → colleague replies) to under a minute (the assistant streams an answer with citations). Cost per answer stays low because most of the work is retrieval, not generation, and the ledger shows where a cheaper answer model is good enough.

The stack

  • Query flow: redelay/event-source → redelay/llm-embed → parallel wiki/tickets/schema lookups → compose → redelay/llm-chat
  • Retrieval nodes: custom vector-search handler + mongoops/find + postgres/query
  • Session + cost: assistant/chats ties team_id to runInput.meta; go-ai/ledger records usage