go-ai — LLM Providers, Guard, Ledger
go-ai — LLM Providers, Guard, Ledger
github.com/redelay/go-ai provides a single llm.Provider interface implemented by every
supported LLM backend, a guard.Guard interface for safety classification, and a usage
ledger that records token spend per run. Applications pick a provider by ID at runtime;
FlowDSL nodes make the choice configurable per-step.
go get github.com/redelay/go-ai
Package layout
github.com/redelay/go-ai
├── llm Provider interface, registry, types, middleware, usage
│ ├── providers/openai OpenAI Chat + Responses + Embeddings
│ ├── providers/openaicompat Base for OpenAI-compatible endpoints
│ ├── providers/ollama Local Ollama runtime
│ ├── providers/anthropic Claude (Messages API, SSE tool-use)
│ ├── providers/gemini Google Gemini (generateContent, streamGenerateContent)
│ ├── providers/ovh OVH AI Endpoints (openai-compatible preset)
│ └── flowdsl FlowDSL nodes: llm-chat, llm-embed, llm-tool-router, llm-guard
├── guard Safety-classification abstraction (pre/post filters)
│ ├── qwen Qwen3Guard-Gen backend (OVH-hosted, free)
│ ├── noop Always-allow backend (default when GUARD_PROVIDER unset)
│ └── module Framework module — reads config + admin settings, builds guard
└── ledger Usage + cost recorder (admin-api submodule for reports)
└── admin Admin HTTP surface — /admin/llm/usage, /admin/llm/pricing
The Provider interface
type Provider interface {
ID() string
Capabilities() Capabilities
Chat(ctx, ChatRequest) (*ChatResponse, error)
ChatStream(ctx, ChatRequest) (<-chan Chunk, error)
Embed(ctx, EmbedRequest) (*EmbedResponse, error)
}
ChatRequest carries Messages, optional Tools, sampling knobs (Temperature, TopP,
MaxTokens, Seed), a ResponseFormat hint, and free-form Labels that flow into the
usage record. Chunk is the transport-agnostic streaming envelope — the same shape feeds
SSE and WebSocket adapters without per-transport provider code.
Providers are registered via init():
import _ "github.com/redelay/go-ai/llm/providers/openai"
// Build an instance from config and store it in the registry.
p, err := llm.Build("openai", map[string]any{"apiKey": os.Getenv("OPENAI_API_KEY")})
// Later:
prov, _ := llm.Get("openai")
llm.Providers() and llm.Factories() return sorted lists of registered instance IDs and
factory IDs respectively — consumed by the FlowDSL subpackage to populate the Studio dropdown.
FlowDSL nodes (llm/flowdsl)
Blank-import to expose the nodes in /modules and make them available to the FlowDSL
runtime:
import _ "github.com/redelay/go-ai/llm/flowdsl"
| Node ID | Kind | Purpose |
|---|---|---|
redelay/llm-chat | action | Send a chat completion to a registered provider and emit the response |
redelay/llm-embed | action | Compute embeddings for one or more input strings |
redelay/llm-tool-router | router | Inspect a chat response and route to Tool (toolCalls present) or Message (plain text) |
redelay/llm-guard | router | Classify a message with a safety model and route to Allow / Flag / Block |
Live providerID / guardID enums
llm-chat and llm-embed declare a required providerID setting; llm-guard declares a
required guardID setting. At spec-build time the module overrides FlowDSLNodes() and
injects the live lists from llm.Providers() / guard.Guards() (falling back to their
factory lists if no instance has been built yet) into each setting's enum. Studio sees
dropdowns that reflect the actual providers + guards wired up in the running app, not a
hard-coded list.
Registering handlers
Declaring the nodes in YAML is not enough — the runtime still needs handler functions that
invoke the Provider. Call flowdsl.Register once after bootstrapping the flow engine:
import (
llmflowdsl "github.com/redelay/go-ai/llm/flowdsl"
"github.com/redelay/go-flowdsl/runtime"
)
eng := runtime.NewEngine(runtime.EngineConfig{})
llmflowdsl.Register(eng)
Register wires four handlers keyed by node ID (redelay/llm-chat,
redelay/llm-embed, redelay/llm-tool-router, redelay/llm-guard). It resolves the
configured providerID through llm.Get and the configured guardID through guard.Get.
For tests or apps that manage their own provider cache, use RegisterWith and supply a
custom lookup:
llmflowdsl.RegisterWith(eng, func(id string) (llm.Provider, error) {
return myCache.Resolve(id)
})
Handler behaviour
llm-chat reads providerID, model, temperature, maxTokens, systemPrompt,
replyStyle, and examplesPolicy from node config; input fields (model, temperature,
maxTokens, responseFormat) override the config on a per-run basis. messages and
optional tools come from the packet input. The output carries message, finishReason,
and usage.
System prompt composition — the effective system message is built at call time from three node settings:
| Setting | Type | Effect |
|---|---|---|
systemPrompt | string (textarea) | Base prompt. Written verbatim as the first paragraph of the system message. |
replyStyle | enum short | normal | detailed | Appends a plain-English directive about reply length. normal (default) appends nothing. |
examplesPolicy | enum avoid | when_asked | always | Appends a directive about whether to include runnable examples. when_asked (default) appends nothing. |
If the incoming packet's first message already has role=system, node-level composition is
skipped entirely — explicit callers always win. Otherwise the three parts are joined with
blank lines and prepended as a single system message. The composed prompt is empty when
every input is empty, which lets the handler skip injecting a system role entirely.
Example — non-engineer tuning from Studio:
nodes:
- id: chat
action_ref: redelay/llm-chat
config:
providerID: ollama
model: qwen2.5:3b
replyStyle: short # → "Keep replies short: aim for 1-3 sentences..."
examplesPolicy: avoid # → "Do not include code or runnable examples..."
systemPrompt: "You are the support assistant for Acme."
Tuning these three fields in FlowDSL Studio + publishing a new version is enough to change assistant tone — no code changes, no env-var churn.
llm-embed reads providerID and model from node config, input (string array) from
the packet. Output carries vectors (same length as input) and usage.
llm-tool-router reads message from the packet input, inspects toolCalls, and
writes route = "Tool" + toolCall (the first call) when present, otherwise
route = "Message". The runtime engine dispatches outbound edges whose condition matches
the route name.
Provider notes
- openai, openaicompat, ovh — share an OpenAI-compatible transport; OVH is a thin preset
over
openaicompatdefaulting to the Meta-Llama-3.3-70B endpoint. - anthropic — translates
RoleSystemmessages into the top-levelsystemfield and emitstool_useblocks through the Messages API.EmbedreturnsErrNoEmbed. - gemini — rewrites
RoleAssistant→model, foldsRoleSystemintosystemInstruction, and mapsRoleToolresponses into user-rolefunctionResponseblocks. Supports generateContent, streamGenerateContent, and embedContent. - ollama — local-only; no API key required; default model is whatever Ollama has pulled.
ai-llm framework module
The llm/module/ subpackage is a framework module that reads provider
credentials from env + admin settings and builds every provider that has
credentials on Startup. Use it instead of calling llm.BootstrapFromEnv
directly.
Blank-import to register:
import (
_ "github.com/redelay/go-ai/llm/module" // the module itself
_ "github.com/redelay/go-ai/llm/providers/anthropic" // factories
_ "github.com/redelay/go-ai/llm/providers/gemini"
_ "github.com/redelay/go-ai/llm/providers/ollama"
_ "github.com/redelay/go-ai/llm/providers/openai"
_ "github.com/redelay/go-ai/llm/providers/ovh"
)
Config (ENV — startup defaults)
| Key | Default | Purpose |
|---|---|---|
LLM_PROVIDER | ollama | Default provider id |
OPENAI_API_KEY | (empty) | Enables openai provider |
ANTHROPIC_API_KEY | (empty) | Enables anthropic provider |
GEMINI_API_KEY | (empty) | Enables gemini provider |
OVH_AI_ENDPOINTS_TOKEN | (empty) | Enables ovh provider |
OLLAMA_URL | http://localhost:11434 | Ollama HTTP endpoint |
OLLAMA_MODEL | llama3.1 | Default Ollama model |
Admin settings (runtime, override env)
All keys live under the ai-llm module. Groups:
| Group | Keys |
|---|---|
ai-llm.defaults | default_provider |
ai-llm.openai | openai_api_key, openai_base_url, openai_default_model |
ai-llm.anthropic | anthropic_api_key, anthropic_default_model |
ai-llm.gemini | gemini_api_key, gemini_default_model |
ai-llm.ovh | ovh_api_key, ovh_base_url, ovh_default_model |
ai-llm.ollama | ollama_base_url, ollama_default_model |
The module implements SettingsResolver
so the admin UI pre-populates every field with the env-resolved effective
value instead of showing blanks on first boot. Setting a value in the UI
overrides the env fallback; clearing it reverts to env.
Profiles
Rather than hand-configuring providerID + model + temperature on every
FlowDSL llm-chat node, use a profile — a named preset bundling
providerID (live options_ref: llm.providers) + apiKey + baseURL +
model + tuning. The ai-llm module declares the llm-chat profile kind
in its module.yaml profiles: block; operators create and edit profiles of
that kind in the admin Profiles page (/profiles).
# FlowDSL node picks a profile instead of listing knobs
- id: chat
action_ref: redelay/llm-chat
config:
profile: assistant-fast # single dropdown, rest derived
# Any explicit field still overrides the profile:
systemPrompt: "You are the Acme support assistant."
The profile setting is an enum populated live from the registered profiles at
spec-build time, matching the pattern already used for providerID and
guardID. At runtime the redelay/llm-chat handler resolves the profile by id
via modules.ResolveProfile("llm-chat", id) and layers it under the node config
with modules.MergeProfileIntoConfig — so individual node settings (model,
temperature, systemPrompt) still win over profile defaults. Profiles are
presets, not hard locks.
Non-node consumers resolve profiles too. Any module built on go-ai — not
just FlowDSL nodes — can resolve its own llm-chat profile with
modules.ResolveProfile. go-modules/ai (the translate/generate module) picks
its own profile, so different features run different models: translations on a
best-quality model, a cheaper/faster model elsewhere — each configured
independently.
See the Profiles reference for the full kind/consumer model, admin API, and cross-worker cache invalidation.
OVH AI Endpoints — multi-model catalog
The ovh provider is a single registration that calls any model OVH hosts behind its
shared OpenAI-compatible endpoint (https://oai.endpoints.kepler.ai.cloud.ovh.net/v1).
Pick the model per call — node config.model or ChatRequest.Model — the same provider
instance serves all of them.
| Category | Model | Input (USD/M) | Output (USD/M) | Notes |
|---|---|---|---|---|
| Chat — flagship | Meta-Llama-3_3-70B-Instruct | 0.74 | 0.74 | Default when OVH provider is built without override |
| Chat — budget | Mistral-7B-Instruct-v0.3 | 0.11 | 0.11 | |
| Chat — multilingual | Mistral-Nemo-Instruct-2407 | 0.14 | 0.14 | |
| Reasoning | gpt-oss-120b | 0.09 | 0.44 | |
| Reasoning | gpt-oss-20b | 0.04 | 0.17 | |
| Reasoning | Qwen3-32B | 0.09 | 0.25 | |
| Code | Qwen3-Coder-30B-A3B-Instruct | 0.07 | 0.24 | 256K context |
| Vision | Mistral-Small-3.2-24B-Instruct-2506 | 0.10 | 0.31 | |
| Vision | Qwen2.5-VL-72B-Instruct | 1.00 | 1.00 | |
| Embedding | bge-m3 | 0.01 | — | 1024-dim; cheapest embedding |
| Embedding | bge-multilingual-gemma2 | 0.01 | — | |
| Embedding | Qwen3-Embedding-8B | 0.11 | — | Higher-quality variant |
| Guard — moderation | Qwen3Guard-Gen-8B | free | free | Default guard model |
| Guard — moderation | Qwen3Guard-Gen-0.6B | free | free | Faster, slightly less accurate |
| ASR | whisper-large-v3-turbo | free | — | Non-token modality |
Prices are loaded into ledger/pricing.go builtinRates(). Admins can override any row from
/admin/llm/pricing without a code change. Rows with 0, 0 are intentionally free (local
Ollama, OVH free-tier) — removing a row stops cost calculation for that model.
Cross-check against the OVH public catalog before invoicing — prices drift.
guard package — safety classification
import (
"github.com/redelay/go-ai/guard"
_ "github.com/redelay/go-ai/guard/qwen" // registers "qwen3guard" factory
_ "github.com/redelay/go-ai/guard/noop" // registers "noop" factory
_ "github.com/redelay/go-ai/guard/module" // framework module — config + settings driven
)
Guard interface
type Guard interface {
ID() string
Check(ctx context.Context, input Input) (*Decision, error)
}
type Input struct {
Role Role // "user" | "assistant"
Content string
Context []Input // optional previous turns
}
type Decision struct {
Action Action // Allow | Flag | Block
Categories []string // policy tags from the classifier
Reason string // short human-readable explanation
Raw string // raw model output for debugging
}
Allow passes through; Flag records for review but continues; Block short-circuits.
The ai-guard.block_flagged admin setting collapses Flag into Block when strict
compliance is required.
Backends
guard/qwen— Qwen3Guard-Gen generative classifier, called via anyllm.Provider(typicallyovh). Factory config keys:providerID(llm provider id, resolved lazily at Check time),model(defaults toQwen3Guard-Gen-8B). The backend parses Qwen'sSafety: <Safe|Unsafe|Controversial>/Categories: ...response into aDecision.guard/noop— always returnsAllow. Registered so callers never need a nil check.
Framework module ai-guard
The guard/module/ subpackage registers an ai-guard framework module that reads
configuration from env + admin settings and builds the configured backend at Startup. See
Base module reference — AI modules for
admin-side details.
Config (ENV, startup-only):
| Key | Default | Purpose |
|---|---|---|
GUARD_PROVIDER | (empty) | Backend factory id — qwen3guard, noop, or empty to disable |
GUARD_UPSTREAM | value of LLM_PROVIDER | LLM provider id used to call the guard model |
GUARD_MODEL | Qwen3Guard-Gen-8B | Model name passed to the upstream provider |
Admin settings (runtime-overridable):
| Key | Default | Purpose |
|---|---|---|
ai-guard.default_on_error | allow | Fail-open (allow) vs fail-closed (block) when the guard model errors |
ai-guard.block_flagged | false | Treat Flag as Block for strict compliance |
redelay/llm-guard FlowDSL node
Router with three output ports: Allow, Flag, Block. Per-node settings:
| Setting | Type | Default | Purpose |
|---|---|---|---|
guardID | enum (live) | — | Registered guard id; populated from guard.Guards() |
role | enum user | assistant | user | Whose message is being classified |
onError | enum allow | block | allow | Override the module-level default per node |
Input shape: { content?: string, messages?: object[], message?: object } — the handler
picks content first, then falls back to the last message's content (handy when
chaining after llm-chat). Every output port emits { decision, content }.
The assistant flows now guard both directions — start → guard-in → chat → guard-out → end, with unsafe messages routed to a dedicated rejected end node and
fanned out to a trust-and-safety email in the production variant.
ledger package — usage + cost
At bootstrap the ledger auto-discovers a UsageSink (flowexec's) via Configure and
installs a go-ai recorder middleware that writes one row per LLM call into a MongoDB
time-series collection (llm_usage). Rows carry runID, flowID, nodeID, provider,
model, token counts, latency, and computed costUSD from the price table.
LLM_TRACK_USAGE toggles recording on or off.
Admin endpoints live in go-ai/ledger/admin/ and mount under /admin/llm/... — grouped
there so the public cmd/api binary never sees them:
| Method | Path | Purpose |
|---|---|---|
| GET | /admin/llm/usage | Paginated usage rows with filters by flow, run, provider |
| GET | /admin/llm/usage/by?field= | Aggregate usage by a dimension (provider, model, flow, node, user, profile) |
| GET | /admin/llm/breakdowns | Aggregated by provider × model × day |
| GET | /admin/llm/pricing | Current price table snapshot (per-{provider,model}, per-1M-token) |
| PUT | /admin/llm/pricing/{provider}/{model} | Override a single rate row |
| DELETE | /admin/llm/pricing/{provider}/{model} | Drop a row (cost calc stops) |
The admin-facing UI is the @redelay/js-admin base-layer AI module — nav AI →
Usage / Breakdowns / Pricing pages, backed by /admin/llm/usage,
/admin/llm/usage/by?field=, and /admin/llm/pricing (per-{provider,model} rate
overrides, per-1M-token). See Base module reference.