Localized content
Localized content
Three capabilities that only make sense together, and that span four modules —
go-framework/modules/settings, go-modules/content, go-modules/i18n and
go-modules/legalpages:
- Settings tokens — content holds the prose, settings hold the facts.
- Localized settings — for the facts that are themselves prose.
- A shared translation lifecycle — is this ready, has it gone stale, does someone want it redone.
- Localized fragments — reusable content blocks, per language.
- Two CLIs that fill all of it without an admin session.
The problem they solve as a set: a privacy policy translated into 25 languages has the operator's registered address baked into all 25. Change the address and every translation is wrong, so either you retranslate 25 documents or you ship 25 documents with the old address in them.
Settings tokens
Any string served by the content module or the i18n catalogue may contain
[[ module.key ]], which is replaced at read time with a public setting.
<p>The controller of your personal data is [[ legalpages.entity_name ]],
[[ legalpages.entity_form ]], VAT [[ legalpages.tax_id ]].</p>
Stored once, resolved per request. Changing the registered address in the admin takes effect on the next read, in every language, with no retranslation — because the token is what every translation contains.
Three rules, all load-bearing
1. Default deny. Only fields whose schema says public: true and not
sensitive: true expand — the same test GET /api/settings/{module}/{key}
applies. Without it, any admin-editable string becomes an exfiltration
primitive: type [[ email.smtp_password ]] into a CMS fragment and the server
renders the secret onto a public page. A nil reader expands nothing, so an app
that never wired the settings module fails closed.
2. Unresolved tokens survive verbatim. A typo, a private field or an unset
value leaves [[ legalpages.entity_nme ]] visible in the output. The
alternative — substituting empty — silently deletes a clause from a legal
document, and nobody reviewing the page can see that it happened. An empty
setting value is treated as unresolved for the same reason.
3. No recursion. Values are never rescanned. Otherwise one level of
indirection defeats rule 1: set a public legalpages.site_name to the literal
text [[ email.smtp_password ]] and a second pass would resolve it.
Why [[ ]] and not {{ }}
i18n messages are run through Go text/template for their parameters
({{.name}}). A settings token sharing those delimiters does not merely fail to
expand — it fails to parse, and text/template abandons the whole template,
so the message loses its parameters too:
"{{ legalpages.entity_name }} greets {{.name}}" → unchanged, both broken
"[[ legalpages.entity_name ]] greets {{.name}}" → "IHAR FINCHUK greets Ihar"
Vue owns the same delimiters on the frontend. [[ ]] collides with neither.
Ordering matters: interpolation runs first, expansion second, so a settings
value that happens to contain {{ is inserted as literal text rather than
executed as a template.
Two variants
| Call | Use for | Escaping |
|---|---|---|
modules.ExpandSettings(ctx, r, s) | plain text — titles, SEO meta, i18n messages | none |
modules.ExpandSettingsHTML(ctx, r, s) | rich HTML rendered with v-html | values escaped, surrounding markup untouched |
Escaping both would double-encode: a company name containing & would reach
the reader as & in an <h1> that Vue already escapes.
Wiring
The expander takes a modules.PublicSettingsReader, discovered optionally in
Configure():
if r, ok := registry.Get("settings").(modules.PublicSettingsReader); ok {
svc.SetTokenSource(r)
}
Deliberately a separate interface from SettingsReader. GetSetting answers
"what is this value", which is right for a module reading its own config and
wrong for rendering into a page — it will happily return an SMTP password. A
distinct interface makes the safe path the only path an expander can take.
Localized settings
A settings field may hold one value per language:
- key: entity_form
type: string
public: true
localized: true # "a sole trader registered in Poland" is a SENTENCE
Declared per field, and that is the whole point. A registered address, a VAT number and a contact email are the same characters in every language and must never be translated; only prose that is read inside a sentence should be. A blanket "translate any setting that appears in prose" would eventually put a tax number through a language model.
Sensitive + Localized is refused at boot: a secret has no translations, and
storing one per locale multiplies the places it can leak from.
Resolution is the same chain pick() uses — requested locale, then the
source, then any non-empty value. The last step matters: a deployment that has
only ever filled Polish should render Polish, because an empty answer deletes
the clause the token sits in.
The locale seam. go-framework cannot import go-modules/i18n, so it owns
only a key and a string — modules.WithLocale / modules.LocaleFromContext —
exactly the shape of the channel seam beside it. i18n's own WithLocale writes
both. "" is a normal answer, not an error: a CLI, a queue consumer and a
single-language deployment all return it, and it means "render the source".
// Anything that needs the reader's language, without depending on i18n.
locale := modules.LocaleFromContext(ctx)
Storage is a locale→string map, and a plain string is still accepted — which
is what every existing row holds, so switching a field to localized needs no
migration. Two shape traps, both of which shipped as bugs before being fixed:
- A nested document round-trips out of Mongo as
bson.D, notmap[string]any. A type switch that handles only the map matches nothing and leaves the token unresolved — in the SOURCE language too, so declaring a field localized breaks the value that worked before. bson.Dmarshals to JSON as[{"Key":…,"Value":…}]— an array. An admin read that returns it raw gives the editor a list, and it renders[object Object],[object Object]. Admin reads project it to an object; public reads pick one string.
In the admin, a localized field renders a default plus sparse per-locale overrides: a chip per language, green where an override exists, and clearing one falls back to the default rather than blanking. If the flag is missing from a UI that predates it — a real risk when the admin layer is versioned separately from the API — the string input detects the object and says so instead of stringifying it.
Translation lifecycle
crud.TranslationState is an embeddable struct, so a page, a fragment and a
domain record all answer the same three questions the same way:
type TranslationState struct {
TranslationReady bool // has anyone said the source text is finished?
SourceHash string // what did the translator last work from?
Retranslate bool // one-shot "do it again"
Units []TranslationUnit // per-leaf: hash + which locales are current
ReviewRequired bool // hold machine output back from the live record
Pending []PendingTranslation
}
Staleness is per unit of work, not per record
SourceHash alone answers "has anything moved", which is the wrong question to
bill against. A fragment holds a dozen independent blocks of prose; fixing a
typo in one moved the record hash, so every block was retranslated into
every language. On an eleven-leaf fragment across 24 locales that was 263
translations replacing text that was already correct.
The cost is the smaller half. Re-rolling a good translation can make it worse: asked three times for the same paragraph, a model returned three different faults, one of which changed the meaning ("every logged session" became "every logged set"). Retranslating what did not change is not a neutral waste — it is a way to lose work that was already right.
So each unit — a leaf path for a fragment, a field for a page — records the hash it was translated from and which locales are current at that hash:
st.MarkUnitTranslated("hero_lead", hash, "ru", "de")
st.UnitCurrent("hero_lead", hash, "ru") // false once the English moves
st.UnitNeedsWork("hero_lead", hash, "sk") // not current AND not awaiting review
Partial success then has somewhere to live. Twenty locales landing and four failing leaves twenty recorded and four to retry, where a record-level hash forced a choice between losing the four and redoing the twenty. A changed hash replaces the locale list rather than merging into it, or translations made from text that no longer exists would be marked current.
Records written before this field exists have no units, so existing text is trusted and adopted on the next run that touches the leaf — otherwise the first run after upgrading would retranslate the entire corpus.
Review gate — translated, but not served
Set ReviewRequired on anything whose wording is a commitment rather than copy.
legalpages sets it on every page it seeds: those state a GDPR one-month
response deadline and cite Articles 11 and 12 DSA, and a machine rendering of
either reaches readers in two dozen languages most teams cannot read.
Translations for such a record are produced and held in Pending; the English
(or an older approved translation) keeps rendering. Three admin endpoints work
the queue:
| Method | Path | Purpose |
|---|---|---|
| GET | /admin/content/translations/pending | What is waiting, with the English beside it and what readers currently see |
| POST | /admin/content/translations/approve | Publish one — optionally with edited text, and force to override staleness |
| POST | /admin/content/translations/reject | Discard, so the next run produces a fresh one |
Approve takes a text override on purpose. The most valuable thing a reviewer does is fix a translation rather than reject it; without an edit field the only choices are publish-it-wrong or spin the wheel again. Approving refuses by default when the English has moved since the translation was made — that is text describing a superseded document.
UnitPending is what stops the queue costing money on a loop: a record awaiting
approval is never "current", so without it every sweep retranslates it and
replaces the pending entry with an equivalent one, billing each time for a page
nobody has looked at.
The admin surface is Content → Translation review (@redelay/js-admin
consumers get it from the project's content module).
type Page struct {
crud.BaseModel `bson:",inline"`
crud.TranslationState `bson:",inline"`
// …
}
Ready gate. Off by default. A draft fanned out to 25 languages has to be redone the moment it is finished, so the default that costs nothing waits. Bulk translators skip anything not marked ready — and should say so, because "0 translated" with no explanation is indistinguishable from a broken translator.
Stale detection. SourceHash records the source the last run worked from.
IsStale(currentHash) is true when the English has changed since. Nobody has
to remember to flag it.
Retranslate. For when the English has not changed but a translation is wrong. Cleared once a translator has acted.
Only MarkTranslated may write SourceHash — an admin write that could set it
would let someone mark a page fresh with nothing translated, silently hiding
every stale locale.
crud.HashSource(parts...) sorts its inputs, so a hash computed over a Go
map is stable. A walk-order-dependent hash differs on every call and marks
everything permanently stale, which reads as a broken indicator and gets
ignored.
Hash the source fields only. A page hashes title["en"] and body["en"]
and deliberately excludes SEO: a reworded meta description should not trigger
24 retranslations of a privacy policy.
A date stamp is not a revision
A module that seeds prose and drops translations when its English changes must
not let a rendered timestamp count as a change. legalpages opens every
template with Last updated {{.Updated}}, filled from the clock at seed time,
and compared whole bodies to decide whether the source had moved — so on the
first boot of any new day, every page "changed", every translated body was
dropped, and the document restamped itself as revised today.
Two failures for the price of one: 24 translations lost per page, and a legal document claiming a revision that never happened. Compare the prose with the volatile part excluded, and skip the write entirely when it matches — that keeps the stored date as well as the translations.
Localized fragments
A Page is localized at the type level — Title and Body are Localized.
A Fragment cannot be: its Data is free-form JSON precisely so each consumer
can shape it, so nothing in the type says which leaves are prose and which are a
URL, a CSS class or an icon name.
Localized leaves therefore carry an explicit marker:
{
"heading": {"_i18n": {"en": "What the app does", "pl": "Co robi aplikacja"}},
"cta_href": "/login?mode=signup"
}
GET /fragments/{key} projects that for the active locale:
{"heading": "Co robi aplikacja", "cta_href": "/login?mode=signup"}
A marker, not a heuristic. "Any map whose keys look like locale codes" reads fine until a fragment legitimately keys data by language or country — a per-locale download link, an app-store URL per region — and then silently collapses it to one string. It is also unstable over time: the same stored document changes meaning the day someone adds a locale whose code collides with an existing key. A marker cannot be triggered by accident.
_i18n and not $i18n because Mongo rejects field names beginning with $.
The marker must be the only key in its map. A map carrying _i18n alongside
siblings is malformed; honouring it would drop them with nothing to show for it.
It is treated as an ordinary map instead, so _i18n stays visible in the
output — the same choice rule 2 above makes for an unresolved token.
Locale first, then tokens. A translated legal line contains
[[ legalpages.entity_name ]] in every language — that is the whole point — so
expanding before picking would scan the locale map's siblings and leave the
chosen text raw.
Projection returns a copy. The store hands back the document it holds and a fragment is read on every render, so projecting in place would freeze one reader's language for everyone after them.
Fallback is the same chain a page uses: requested locale → fallback locale → any non-empty value. A half-translated fragment renders English, not a blank.
Filling the translations
Two commands, both opt-in subpackages — their parents sit on the request path and must not carry an AI provider for work most deployments never do:
# CMS pages and content fragments.
redelayctl content-translate --dry-run
redelayctl content-translate --pages=about --locales=pl,de
redelayctl content-translate --fragments=about-page --force
# The i18n message catalogue.
redelayctl i18n-translate --prefix=marketing.
redelayctl i18n-translate --keys=a.b,c.d --overwrite
Both fill missing locales only unless told otherwise, so a run is resumable
and cheap to repeat. Three independent reasons to redo a locale that already has
text: --force/--overwrite, the one-shot Retranslate flag, and a changed
source hash. Only the first is a flag; staleness is derived, so nobody has to
remember it.
Running it on a schedule
The CLI is a person deciding to translate now. For unattended work, create a
scheduler entry whose event type + action join to content.translate.sweep;
the content/translate module consumes it and does exactly the work that is
outstanding.
An event rather than a ticker inside the module, for a reason that only appears in production: with several replicas a ticker runs in every one of them, so N replicas means N times the model spend for the same work. The scheduler already solves that with leader election.
| Variable | Default | Purpose |
|---|---|---|
CONTENT_TRANSLATE_QUIET_MINUTES | 10 | How long a record must sit untouched before the sweep will translate it. 0 disables |
CONTENT_TRANSLATE_BUDGET | 200 | Max (unit, locale) translations per sweep. 0 = no cap |
CONTENT_TRANSLATE_DOMAIN | — | One phrase describing what this deployment is, prepended to every translation hint |
The quiet period is the debounce, and it costs nothing because the sweep is state-driven: an editor saving six times in two minutes leaves one record with one recent timestamp, and the run that picks it up translates the final wording once. Reacting per save would pay for five drafts nobody reads — and on a review-required record, queue five candidates a human must dismiss. A record with no timestamp counts as settled, or anything seeded without that field would be ineligible forever.
The budget is a ceiling on a bad day, not a target. A bulk import otherwise turns one cron tick into an unbounded bill with nobody watching. Work not done is not lost: nothing here is a queue that can drain incorrectly, so the next sweep sees the same state and continues.
A sweep that did nothing says which of the three reasons applies — waiting for edits to settle, stopped at the budget, or genuinely nothing to do. Printing silence for all three is how an unattended job looks healthy while stuck.
Tell the model what the product is
CONTENT_TRANSLATE_DOMAIN exists because every hint otherwise said some
variant of "a block of prose on a web page", which is true of every site ever
written and leaves the model to guess the sense of any ambiguous word. On a
training log it guessed wrong the first time it was asked: "a set that did not
save" came back as Russian «набор» — a set as in a collection of objects —
where a lifter says «сет».
Name the terms, do not imply them. A general "this is a strength-training app" did not move that word; a glossary naming «сет» and «тренировка» fixed it first try. Readers do not see a mistranslation, they see a product that does not know its own vocabulary.
i18n-translate and the admin endpoint run one shared loop
(i18n.TranslateMissing over a Catalogue interface). Two copies would be two
answers to "what counts as missing" and would drift. They differ in exactly one
respect: the endpoint caps the number of keys because an HTTP request has a
deadline; the CLI passes zero, because a command that stopped at 40 could never
finish a larger catalogue.
Raise CLI_TIMEOUT for long runs. The framework caps a CLI at five minutes
by default. Translating six legal pages into 24 languages is ~144 model calls at
roughly ten seconds each; past the cap every remaining call fails instantly, and
before this was reported properly it looked like the model rejecting everything.
CLI_TIMEOUT=0 redelayctl content-translate # no cap
A failed locale is not "nothing to do". Both commands report attempted-and- failed separately from complete, and neither records a source hash for work that did not land — so a rerun retries exactly what is missing.
Consuming it from a frontend
Use useRequestFetch(), not $fetch, for anything locale-dependent.
On the server $fetch opens a new request carrying no cookies, so a lang
cookie never reaches the proxy that turns it into ?lang= — and every CMS page
renders in the default language however the reader set theirs, on the SSR pass
that crawlers also see.
This failure is invisible without looking for it, because the fallback chain works: a request that carried no locale is indistinguishable from one whose locale is untranslated. Both return a healthy 200 in English.
const requestFetch = useRequestFetch()
const lang = useCookie<string>('lang') // read in setup, not in the key fn
const { data } = await useAsyncData(
() => `page:${slug}:${lang.value || 'en'}`, // locale in the cache key
() => requestFetch(`${apiBase}/pages/${slug}`),
)
Two traps in those four lines. The locale belongs in the cache key, or the
first language rendered is served to everyone afterwards. And useCookie must
be called in setup rather than inside the key function — useAsyncData invokes
that lazily, outside the component's setup context, where it throws
NUXT_E1001 and takes the page down.
Splitting copy between t() and content
| Goes in | Why | |
|---|---|---|
| Short labels — headings, buttons, nav | i18n t() keys | A translator gets a phrase and returns a phrase. |
| Prose, especially with inline links | CMS page or fragment | A catalogue full of <p>…<a href>…</a></p> is one nobody can translate safely: a dropped </a> breaks the page rather than the sentence. |
Give short keys a context hint. A page title is two words with no sentence
around it, which is the worst case for a translator, and some are worse than
they look: English "About" is elliptical for "About us", a convention most
languages do not share, so translated as a word it comes back as the bare
preposition it literally is — О, O, Über, Про. None of those is a title.
The frontend owns these hints in i18n-contexts.json; the extractor posts them
with the keys, and the bulk translator feeds each note to the model:
"about.eyebrow": "Navigation link and page heading. English \"About\" is ELLIPTICAL for \"About us\". Render the natural page-title form, NEVER the bare preposition. Sample — PL: \"O NAS\" (not \"O\"); DE: \"ÜBER UNS\"; UA: \"ПРО НАС\"."
Content change events
content publishes content.page.changed and content.fragment.changed for
anything that derives from a page or fragment: a search index that must not
keep serving the old wording, a cache, an audit trail. Optional — with no bus
wired the module behaves exactly as before.
The payload carries source_changed, so a consumer can ignore a position
change or a translation landing, and a reflowed paragraph that renders
identically wakes nobody.
Translation deliberately does not consume these. It works from stored hashes, so a lost event costs latency and never correctness — which is why the sweep is state-driven rather than queue-driven, and why a missed event is not a class of bug this system can have.
Editing 25 languages in the admin
A localized field renders one editor per locale, which is fine at two languages and unusable at twenty-five. Declare a focus instead:
<LocaleFocusBar :locales="locales" />
<LocalizedField v-model="form.title" :locales="locales" />
<LocalizedRichText v-model="form.body" :locales="locales" />
Every localized field below the bar renders only the focused language. The bar carries a language dropdown and a chip per locale — green complete, amber partial, grey untouched — so the gaps are visible without clicking through the list to find them.
Strictly opt-in: with no bar on the page there is no provider and fields render every locale exactly as before. Fields register themselves with the bar, so per-locale coverage is computed across the whole record without the bar being told the form's shape — a title translated while the body is not reads as partial, which is what it is.
What the translator does to survive a real model
The ai module's Translate grew these one failure at a time. Each is here
because its absence produced a silent, expensive wrong answer.
The token cap is sized from the source, not configured once. A translation
is the one completion whose output length is knowable up front — the source
text, once per locale — so a single AI_MAX_TOKENS is the wrong shape. At the
2048 default, every locale of a 9 KB legal page came back truncated mid-JSON:
correctly rejected by format validation, retried, bisected to a batch of one,
truncated identically, and finally dropped. Eighteen minutes and ~240 paid
calls produced nothing. The estimate assumes the expensive case, because
Cyrillic, Greek and CJK tokenize at roughly two characters per token where
English HTML manages four — a cap sized off the English source still truncates
the Russian reply. The configured value is a floor, never a ceiling.
Batch size comes from an output budget, not a count. Six locales of that same page is a single ~28,000-token reply, near enough the ceiling that one slightly longer page loses all six at once. Short strings still batch six-at-a-time; long documents step down to one locale per call.
Retries carry the reason. A format retry used to re-send a byte-identical
prompt, which cannot help — the model already answered that question. A Slovak
page failed three identical attempts on one paragraph whose <p> came back
without its closing bracket. Each locale's rejection reason now goes into its
next attempt, and the first corrected retry succeeded.
Rejections name the tags. problems() reported "structure mismatch" for
the one check that fires on an HTML body. It now diffs the tag multisets as
counts (dropped </p>×1, <p>×1), which is what made the above diagnosable.
Each call is bounded by AI_CALL_TIMEOUT (default 3m; 0 disables). The
bound belongs at the layer that makes calls: a caller can only bound the whole
operation it asked for, and "this page in 24 locales" is many calls — too short
a limit kills the job, too long lets one stalled request hold it up. That was a
fifteen-minute hang with no output. TRANSLATE_CALL_TIMEOUT still works and
feeds the same setting.
A related lesson worth keeping: CLI_TIMEOUT defaults to five minutes, and a
six-page run that exceeds it reported "no locale answered" 144 times — blaming
the model for a deadline. Long runs want CLI_TIMEOUT=0.