Reference

go-ai — LLM Providers, Guard, Ledger

Provider-agnostic LLM access for Redelay — chat, streaming, embeddings, tool-use, safety classification, and cost ledger. Includes openai, ollama, anthropic, gemini, and ovh providers, a Guard abstraction backed by Qwen3Guard, and FlowDSL nodes for all of it.

go-ai — LLM Providers, Guard, Ledger

github.com/redelay/go-ai provides a single llm.Provider interface implemented by every supported LLM backend, a guard.Guard interface for safety classification, and a usage ledger that records token spend per run. Applications pick a provider by ID at runtime; FlowDSL nodes make the choice configurable per-step.

shell
go get github.com/redelay/go-ai

Package layout

text
github.com/redelay/go-ai
├── llm                     Provider interface, registry, types, middleware, usage
│   ├── providers/openai    OpenAI Chat + Responses + Embeddings
│   ├── providers/openaicompat  Base for OpenAI-compatible endpoints
│   ├── providers/ollama    Local Ollama runtime
│   ├── providers/anthropic Claude (Messages API, SSE tool-use)
│   ├── providers/gemini    Google Gemini (generateContent, streamGenerateContent)
│   ├── providers/ovh       OVH AI Endpoints (openai-compatible preset)
│   └── flowdsl             FlowDSL nodes: llm-chat, llm-embed, llm-tool-router, llm-guard
├── guard                   Safety-classification abstraction (pre/post filters)
│   ├── qwen                Qwen3Guard-Gen backend (OVH-hosted, free)
│   ├── noop                Always-allow backend (default when GUARD_PROVIDER unset)
│   └── module              Framework module — reads config + admin settings, builds guard
└── ledger                  Usage + cost recorder (admin-api submodule for reports)
    └── admin               Admin HTTP surface — /admin/llm/usage, /admin/llm/pricing

The Provider interface

go
type Provider interface {
    ID() string
    Capabilities() Capabilities
    Chat(ctx, ChatRequest) (*ChatResponse, error)
    ChatStream(ctx, ChatRequest) (<-chan Chunk, error)
    Embed(ctx, EmbedRequest) (*EmbedResponse, error)
}

ChatRequest carries Messages, optional Tools, sampling knobs (Temperature, TopP, MaxTokens, Seed), a ResponseFormat hint, and free-form Labels that flow into the usage record. Chunk is the transport-agnostic streaming envelope — the same shape feeds SSE and WebSocket adapters without per-transport provider code.

Providers are registered via init():

go
import _ "github.com/redelay/go-ai/llm/providers/openai"

// Build an instance from config and store it in the registry.
p, err := llm.Build("openai", map[string]any{"apiKey": os.Getenv("OPENAI_API_KEY")})
// Later:
prov, _ := llm.Get("openai")

llm.Providers() and llm.Factories() return sorted lists of registered instance IDs and factory IDs respectively — consumed by the FlowDSL subpackage to populate the Studio dropdown.

FlowDSL nodes (llm/flowdsl)

Blank-import to expose the nodes in /modules and make them available to the FlowDSL runtime:

go
import _ "github.com/redelay/go-ai/llm/flowdsl"
Node IDKindPurpose
redelay/llm-chatactionSend a chat completion to a registered provider and emit the response
redelay/llm-embedactionCompute embeddings for one or more input strings
redelay/llm-tool-routerrouterInspect a chat response and route to Tool (toolCalls present) or Message (plain text)
redelay/llm-guardrouterClassify a message with a safety model and route to Allow / Flag / Block

Live providerID / guardID enums

llm-chat and llm-embed declare a required providerID setting; llm-guard declares a required guardID setting. At spec-build time the module overrides FlowDSLNodes() and injects the live lists from llm.Providers() / guard.Guards() (falling back to their factory lists if no instance has been built yet) into each setting's enum. Studio sees dropdowns that reflect the actual providers + guards wired up in the running app, not a hard-coded list.

Registering handlers

Declaring the nodes in YAML is not enough — the runtime still needs handler functions that invoke the Provider. Call flowdsl.Register once after bootstrapping the flow engine:

go
import (
    llmflowdsl "github.com/redelay/go-ai/llm/flowdsl"
    "github.com/redelay/go-flowdsl/runtime"
)

eng := runtime.NewEngine(runtime.EngineConfig{})
llmflowdsl.Register(eng)

Register wires four handlers keyed by node ID (redelay/llm-chat, redelay/llm-embed, redelay/llm-tool-router, redelay/llm-guard). It resolves the configured providerID through llm.Get and the configured guardID through guard.Get.

For tests or apps that manage their own provider cache, use RegisterWith and supply a custom lookup:

go
llmflowdsl.RegisterWith(eng, func(id string) (llm.Provider, error) {
    return myCache.Resolve(id)
})

Handler behaviour

llm-chat reads providerID, model, temperature, maxTokens, systemPrompt, replyStyle, and examplesPolicy from node config; input fields (model, temperature, maxTokens, responseFormat) override the config on a per-run basis. messages and optional tools come from the packet input. The output carries message, finishReason, and usage.

System prompt composition — the effective system message is built at call time from three node settings:

SettingTypeEffect
systemPromptstring (textarea)Base prompt. Written verbatim as the first paragraph of the system message.
replyStyleenum short | normal | detailedAppends a plain-English directive about reply length. normal (default) appends nothing.
examplesPolicyenum avoid | when_asked | alwaysAppends a directive about whether to include runnable examples. when_asked (default) appends nothing.

If the incoming packet's first message already has role=system, node-level composition is skipped entirely — explicit callers always win. Otherwise the three parts are joined with blank lines and prepended as a single system message. The composed prompt is empty when every input is empty, which lets the handler skip injecting a system role entirely.

Example — non-engineer tuning from Studio:

yaml
nodes:
  - id: chat
    action_ref: redelay/llm-chat
    config:
      providerID: ollama
      model: qwen2.5:3b
      replyStyle: short            # → "Keep replies short: aim for 1-3 sentences..."
      examplesPolicy: avoid        # → "Do not include code or runnable examples..."
      systemPrompt: "You are the support assistant for Acme."

Tuning these three fields in FlowDSL Studio + publishing a new version is enough to change assistant tone — no code changes, no env-var churn.

llm-embed reads providerID and model from node config, input (string array) from the packet. Output carries vectors (same length as input) and usage.

llm-tool-router reads message from the packet input, inspects toolCalls, and writes route = "Tool" + toolCall (the first call) when present, otherwise route = "Message". The runtime engine dispatches outbound edges whose condition matches the route name.

Provider notes

  • openai, openaicompat, ovh — share an OpenAI-compatible transport; OVH is a thin preset over openaicompat defaulting to the Meta-Llama-3.3-70B endpoint.
  • anthropic — translates RoleSystem messages into the top-level system field and emits tool_use blocks through the Messages API. Embed returns ErrNoEmbed.
  • gemini — rewrites RoleAssistant → model, folds RoleSystem into systemInstruction, and maps RoleTool responses into user-role functionResponse blocks. Supports generateContent, streamGenerateContent, and embedContent.
  • ollama — local-only; no API key required; default model is whatever Ollama has pulled.

ai-llm framework module

The llm/module/ subpackage is a framework module that reads provider credentials from env + admin settings and builds every provider that has credentials on Startup. Use it instead of calling llm.BootstrapFromEnv directly.

Blank-import to register:

go
import (
    _ "github.com/redelay/go-ai/llm/module"              // the module itself
    _ "github.com/redelay/go-ai/llm/providers/anthropic" // factories
    _ "github.com/redelay/go-ai/llm/providers/gemini"
    _ "github.com/redelay/go-ai/llm/providers/ollama"
    _ "github.com/redelay/go-ai/llm/providers/openai"
    _ "github.com/redelay/go-ai/llm/providers/ovh"
)

Config (ENV — startup defaults)

KeyDefaultPurpose
LLM_PROVIDERollamaDefault provider id
OPENAI_API_KEY(empty)Enables openai provider
ANTHROPIC_API_KEY(empty)Enables anthropic provider
GEMINI_API_KEY(empty)Enables gemini provider
OVH_AI_ENDPOINTS_TOKEN(empty)Enables ovh provider
OLLAMA_URLhttp://localhost:11434Ollama HTTP endpoint
OLLAMA_MODELllama3.1Default Ollama model

Admin settings (runtime, override env)

All keys live under the ai-llm module. Groups:

GroupKeys
ai-llm.defaultsdefault_provider
ai-llm.openaiopenai_api_key, openai_base_url, openai_default_model
ai-llm.anthropicanthropic_api_key, anthropic_default_model
ai-llm.geminigemini_api_key, gemini_default_model
ai-llm.ovhovh_api_key, ovh_base_url, ovh_default_model
ai-llm.ollamaollama_base_url, ollama_default_model

The module implements SettingsResolver so the admin UI pre-populates every field with the env-resolved effective value instead of showing blanks on first boot. Setting a value in the UI overrides the env fallback; clearing it reverts to env.

Profiles

Rather than hand-configuring providerID + model + temperature on every FlowDSL llm-chat node, use a profile — a named preset bundling providerID (live options_ref: llm.providers) + apiKey + baseURL + model + tuning. The ai-llm module declares the llm-chat profile kind in its module.yaml profiles: block; operators create and edit profiles of that kind in the admin Profiles page (/profiles).

yaml
# FlowDSL node picks a profile instead of listing knobs
- id: chat
  action_ref: redelay/llm-chat
  config:
    profile: assistant-fast       # single dropdown, rest derived
    # Any explicit field still overrides the profile:
    systemPrompt: "You are the Acme support assistant."

The profile setting is an enum populated live from the registered profiles at spec-build time, matching the pattern already used for providerID and guardID. At runtime the redelay/llm-chat handler resolves the profile by id via modules.ResolveProfile("llm-chat", id) and layers it under the node config with modules.MergeProfileIntoConfig — so individual node settings (model, temperature, systemPrompt) still win over profile defaults. Profiles are presets, not hard locks.

Non-node consumers resolve profiles too. Any module built on go-ai — not just FlowDSL nodes — can resolve its own llm-chat profile with modules.ResolveProfile. go-modules/ai (the translate/generate module) picks its own profile, so different features run different models: translations on a best-quality model, a cheaper/faster model elsewhere — each configured independently.

See the Profiles reference for the full kind/consumer model, admin API, and cross-worker cache invalidation.

OVH AI Endpoints — multi-model catalog

The ovh provider is a single registration that calls any model OVH hosts behind its shared OpenAI-compatible endpoint (https://oai.endpoints.kepler.ai.cloud.ovh.net/v1). Pick the model per call — node config.model or ChatRequest.Model — the same provider instance serves all of them.

CategoryModelInput (USD/M)Output (USD/M)Notes
Chat — flagshipMeta-Llama-3_3-70B-Instruct0.740.74Default when OVH provider is built without override
Chat — budgetMistral-7B-Instruct-v0.30.110.11
Chat — multilingualMistral-Nemo-Instruct-24070.140.14
Reasoninggpt-oss-120b0.090.44
Reasoninggpt-oss-20b0.040.17
ReasoningQwen3-32B0.090.25
CodeQwen3-Coder-30B-A3B-Instruct0.070.24256K context
VisionMistral-Small-3.2-24B-Instruct-25060.100.31
VisionQwen2.5-VL-72B-Instruct1.001.00
Embeddingbge-m30.01—1024-dim; cheapest embedding
Embeddingbge-multilingual-gemma20.01—
EmbeddingQwen3-Embedding-8B0.11—Higher-quality variant
Guard — moderationQwen3Guard-Gen-8BfreefreeDefault guard model
Guard — moderationQwen3Guard-Gen-0.6BfreefreeFaster, slightly less accurate
ASRwhisper-large-v3-turbofree—Non-token modality

Prices are loaded into ledger/pricing.go builtinRates(). Admins can override any row from /admin/llm/pricing without a code change. Rows with 0, 0 are intentionally free (local Ollama, OVH free-tier) — removing a row stops cost calculation for that model.

Cross-check against the OVH public catalog before invoicing — prices drift.

guard package — safety classification

go
import (
    "github.com/redelay/go-ai/guard"
    _ "github.com/redelay/go-ai/guard/qwen"   // registers "qwen3guard" factory
    _ "github.com/redelay/go-ai/guard/noop"   // registers "noop" factory
    _ "github.com/redelay/go-ai/guard/module" // framework module — config + settings driven
)

Guard interface

go
type Guard interface {
    ID() string
    Check(ctx context.Context, input Input) (*Decision, error)
}

type Input struct {
    Role    Role   // "user" | "assistant"
    Content string
    Context []Input // optional previous turns
}

type Decision struct {
    Action     Action   // Allow | Flag | Block
    Categories []string // policy tags from the classifier
    Reason     string   // short human-readable explanation
    Raw        string   // raw model output for debugging
}

Allow passes through; Flag records for review but continues; Block short-circuits. The ai-guard.block_flagged admin setting collapses Flag into Block when strict compliance is required.

Backends

  • guard/qwen — Qwen3Guard-Gen generative classifier, called via any llm.Provider (typically ovh). Factory config keys: providerID (llm provider id, resolved lazily at Check time), model (defaults to Qwen3Guard-Gen-8B). The backend parses Qwen's Safety: <Safe|Unsafe|Controversial> / Categories: ... response into a Decision.
  • guard/noop — always returns Allow. Registered so callers never need a nil check.

Framework module ai-guard

The guard/module/ subpackage registers an ai-guard framework module that reads configuration from env + admin settings and builds the configured backend at Startup. See Base module reference — AI modules for admin-side details.

Config (ENV, startup-only):

KeyDefaultPurpose
GUARD_PROVIDER(empty)Backend factory id — qwen3guard, noop, or empty to disable
GUARD_UPSTREAMvalue of LLM_PROVIDERLLM provider id used to call the guard model
GUARD_MODELQwen3Guard-Gen-8BModel name passed to the upstream provider

Admin settings (runtime-overridable):

KeyDefaultPurpose
ai-guard.default_on_errorallowFail-open (allow) vs fail-closed (block) when the guard model errors
ai-guard.block_flaggedfalseTreat Flag as Block for strict compliance

redelay/llm-guard FlowDSL node

Router with three output ports: Allow, Flag, Block. Per-node settings:

SettingTypeDefaultPurpose
guardIDenum (live)—Registered guard id; populated from guard.Guards()
roleenum user | assistantuserWhose message is being classified
onErrorenum allow | blockallowOverride the module-level default per node

Input shape: { content?: string, messages?: object[], message?: object } — the handler picks content first, then falls back to the last message's content (handy when chaining after llm-chat). Every output port emits { decision, content }.

The assistant flows now guard both directions — start → guard-in → chat → guard-out → end, with unsafe messages routed to a dedicated rejected end node and fanned out to a trust-and-safety email in the production variant.

ledger package — usage + cost

At bootstrap the ledger auto-discovers a UsageSink (flowexec's) via Configure and installs a go-ai recorder middleware that writes one row per LLM call into a MongoDB time-series collection (llm_usage). Rows carry runID, flowID, nodeID, provider, model, token counts, latency, and computed costUSD from the price table. LLM_TRACK_USAGE toggles recording on or off.

Admin endpoints live in go-ai/ledger/admin/ and mount under /admin/llm/... — grouped there so the public cmd/api binary never sees them:

MethodPathPurpose
GET/admin/llm/usagePaginated usage rows with filters by flow, run, provider
GET/admin/llm/usage/by?field=Aggregate usage by a dimension (provider, model, flow, node, user, profile)
GET/admin/llm/breakdownsAggregated by provider × model × day
GET/admin/llm/pricingCurrent price table snapshot (per-{provider,model}, per-1M-token)
PUT/admin/llm/pricing/{provider}/{model}Override a single rate row
DELETE/admin/llm/pricing/{provider}/{model}Drop a row (cost calc stops)

The admin-facing UI is the @redelay/js-admin base-layer AI module — nav AI → Usage / Breakdowns / Pricing pages, backed by /admin/llm/usage, /admin/llm/usage/by?field=, and /admin/llm/pricing (per-{provider,model} rate overrides, per-1M-token). See Base module reference.