Deploying to production
Deploying to production
This is how the Redelay reference backend (api.redelay.com) and the
gymmer.com backend ship. The app repository owns three files, and nothing
outside it is needed to build or deploy:
| File | Job |
|---|---|
Dockerfile | Builds every binary from this checkout alone, fetching private framework modules over SSH |
docker-swarm-stack.yml | The swarm services: placement, healthchecks, rollout and rollback policy, Traefik labels |
.github/workflows/deploy-production.yml | Push to production → build, push to the registry, docker stack deploy, wait for the rollout |
It relies on the app being Mode B: go.mod pins released framework
tags with no replace (see
Creating a Go App → Mode B).
A build that relies on sibling checkouts breaks as soon as the siblings'
branches drift apart. The Redelay backend's own CI failed for seven
weeks for exactly this reason before it moved to Mode B.
Picking up framework changes
A deployed app gets framework changes only when you move a pin:
# in the library: tag the change
git tag v0.44.0 && git push origin v0.44.0
# in the app: move the pin, build, commit
go get github.com/redelay/[email protected]
go mod tidy && go build ./... && go test ./...
Tagging a library never changes another project's build, because every app keeps its own pins. When several apps share the framework, starting a new app on the same pin set as one already in production gives you a known-good baseline.
The Dockerfile
FROM golang:1.25 AS build
WORKDIR /src
ENV GOPRIVATE=github.com/redelay/*,github.com/flowdsl/*,github.com/FlowDSL/*,github.com/dreplyai/* \
GOTOOLCHAIN=local CGO_ENABLED=0 GOOS=linux
COPY go.mod go.sum ./
RUN --mount=type=ssh \
mkdir -p -m 0700 ~/.ssh && \
ssh-keyscan -t rsa,ecdsa,ed25519 github.com >> ~/.ssh/known_hosts 2>/dev/null && \
git config --global url."[email protected]:redelay/".insteadOf "https://github.com/redelay/" && \
git config --global url."[email protected]:flowdsl/".insteadOf "https://github.com/flowdsl/" && \
git config --global url."[email protected]:dreplyai/".insteadOf "https://github.com/dreplyai/" && \
for i in 1 2 3; do go mod download && exit 0; sleep $((i * 15)); done; exit 1
COPY . .
RUN GOWORK=off go build -trimpath -ldflags="-s -w" -o /out/api ./cmd/api
# … one `go build` per binary, then a slim runtime stage that copies /out/*
Key points:
--mount=type=sshforwards the deploy key for that singleRUN, so the key never lands in an image layer. A key added withCOPYstays recoverable from the layer even in a stage you throw away.- Use one
insteadOfrule per private org. A single rule for all ofhttps://github.com/would also send public dependencies over SSH. - Write the module path's case, not the repo's. The module path
github.com/dreplyai/…is lowercase even though the repo isDreplyAI/….insteadOfmatches the URL string exactly, so a rule written with the repo's casing never fires. - Pin
known_hostswithssh-keyscan. Don't useStrictHostKeyChecking=no: with checking off, the build would hand its key to whatever answered on port 22. - Retry
go mod download. It makes hundreds of fetches, and one hiccup shouldn't fail an otherwise good build. .dockerignoremust exclude.env*,go.workandgo.work.sum.
Build it locally with a key file, the same way CI does:
docker buildx build --ssh default=$HOME/.ssh/id_ed25519 -t myapp-api .
The swarm stack
Three settings decide whether the service recovers on its own:
x-deploy-policy: &deploy-policy
restart_policy:
condition: any # not on-failure, and no max_attempts
delay: 5s
update_config:
parallelism: 1
order: start-first
failure_action: rollback
monitor: 120s # how long a new task must stay healthy
rollback_config:
order: start-first
failure_action: pause
on-failure with max_attempts. If a node reboots and
reaches the registry over Tailscale, image pulls fail until Tailscale is up.
A capped policy uses up its attempts in that window and leaves the service
at 0 replicas with no retry. Swarm won't reschedule it. redelay.com and
flowdsl.com were down for about 45 hours this way in September 2026.monitoris what makesfailure_action: rollbackwork. Without it, a task that crashes after its first healthcheck still counts as a successful update.- Give every service a real healthcheck. Distroless images have no
wgetorcurl, so ship a small static/healthcheckbinary that requests$HEALTH_URL. - Keep the stack name stable.
docker stack deployunder a new name starts a second stack, and both then claim the same Traefik hosts.
The workflow
The workflow runs on a self-hosted runner on the swarm manager:
- Guard. Fail if the checkout contains a
go.work, a.env, or anyreplaceingo.mod. - Find the deploy key. Probe the candidate keys with
git ls-remoteagainst one repo per private org, and use the first key that can read all of them. A repository or org variable (GO_MODULES_SSH_KEY) can override the choice. - Build and push with
docker/setup-buildx-actionusingdriver-opts: network=host. - Deploy. Run
docker stack deploywith secrets sourced from an env file on the manager (never from the repo). - Wait for the rollout. Poll each service's
UpdateStatus.State. Fail onrollback_*orpaused, and print the task errors and recent logs. - Run a health check against the public URL, then clean up the registry (keep the last 5 tags).
network=host is required when the registry is on a Tailscale address.
With the default docker-container driver, BuildKit runs on the Docker
bridge, and the bridge can't route to 100.x addresses. The build succeeds
and then the push fails after minutes of retries with
dial tcp 100.x.x.x:5000: i/o timeout, which looks like a registry outage.Two GitHub quirks:
- Pass
--ssh default=<key path>, not plaindefault. A runner installed as a service has no$SSH_AUTH_SOCK. - The workflow file must exist on the default branch. Otherwise GitHub
never registers it: pushes to
productionstart nothing andworkflow_dispatchreturns 404. Commit it to the default branch as well, even though it only triggers onproduction.
Deploying is then:
git push origin main:production # or your integration branch
Monitoring
Probe the public URL with the Prometheus blackbox exporter. That covers every failure a user would see: 0 replicas, a missing Traefik router, TLS, DNS, or app 5xx errors.
- job_name: public-sites
metrics_path: /probe
params: {module: [http_2xx]}
static_configs:
- targets: ['https://api.example.com/api/v1/health']
labels: {project: myapp, service: api}
relabel_configs:
- {source_labels: [__address__], target_label: __param_target}
- {source_labels: [__param_target], target_label: instance}
- {target_label: __address__, replacement: blackbox-exporter:9115}
- alert: public_site_down
expr: probe_success{job="public-sites"} == 0
for: 3m
labels: {severity: critical}
E2E tests & demo videos
One Playwright harness drives the real Redelay / AsyncShop stack for two jobs — browser E2E tests and branded demo videos for the docs and website. How it works, how to run it in Docker, and how to add a test or a video for FlowDSL, Redelay, or AsyncShop/AsyncCart.
Creating a Go Module
Step-by-step guide to creating a new module for the redelay go-framework.