Guides

Deploying to production

Ship a Redelay Go backend to Docker Swarm — a Mode B image built with SSH-forwarded private modules, a swarm stack that survives node reboots, and a push-to-deploy GitHub Actions workflow that fails on an automatic rollback.

Deploying to production

This is how the Redelay reference backend (api.redelay.com) and the gymmer.com backend ship. The app repository owns three files, and nothing outside it is needed to build or deploy:

FileJob
DockerfileBuilds every binary from this checkout alone, fetching private framework modules over SSH
docker-swarm-stack.ymlThe swarm services: placement, healthchecks, rollout and rollback policy, Traefik labels
.github/workflows/deploy-production.ymlPush to production → build, push to the registry, docker stack deploy, wait for the rollout

It relies on the app being Mode B: go.mod pins released framework tags with no replace (see Creating a Go App → Mode B). A build that relies on sibling checkouts breaks as soon as the siblings' branches drift apart. The Redelay backend's own CI failed for seven weeks for exactly this reason before it moved to Mode B.

Picking up framework changes

A deployed app gets framework changes only when you move a pin:

shell
# in the library: tag the change
git tag v0.44.0 && git push origin v0.44.0
# in the app: move the pin, build, commit
go get github.com/redelay/[email protected]
go mod tidy && go build ./... && go test ./...

Tagging a library never changes another project's build, because every app keeps its own pins. When several apps share the framework, starting a new app on the same pin set as one already in production gives you a known-good baseline.

The Dockerfile

dockerfile
FROM golang:1.25 AS build
WORKDIR /src

ENV GOPRIVATE=github.com/redelay/*,github.com/flowdsl/*,github.com/FlowDSL/*,github.com/dreplyai/* \
    GOTOOLCHAIN=local CGO_ENABLED=0 GOOS=linux

COPY go.mod go.sum ./
RUN --mount=type=ssh \
    mkdir -p -m 0700 ~/.ssh && \
    ssh-keyscan -t rsa,ecdsa,ed25519 github.com >> ~/.ssh/known_hosts 2>/dev/null && \
    git config --global url."[email protected]:redelay/".insteadOf "https://github.com/redelay/" && \
    git config --global url."[email protected]:flowdsl/".insteadOf "https://github.com/flowdsl/" && \
    git config --global url."[email protected]:dreplyai/".insteadOf "https://github.com/dreplyai/" && \
    for i in 1 2 3; do go mod download && exit 0; sleep $((i * 15)); done; exit 1

COPY . .
RUN GOWORK=off go build -trimpath -ldflags="-s -w" -o /out/api ./cmd/api
# … one `go build` per binary, then a slim runtime stage that copies /out/*

Key points:

  • --mount=type=ssh forwards the deploy key for that single RUN, so the key never lands in an image layer. A key added with COPY stays recoverable from the layer even in a stage you throw away.
  • Use one insteadOf rule per private org. A single rule for all of https://github.com/ would also send public dependencies over SSH.
  • Write the module path's case, not the repo's. The module path github.com/dreplyai/… is lowercase even though the repo is DreplyAI/…. insteadOf matches the URL string exactly, so a rule written with the repo's casing never fires.
  • Pin known_hosts with ssh-keyscan. Don't use StrictHostKeyChecking=no: with checking off, the build would hand its key to whatever answered on port 22.
  • Retry go mod download. It makes hundreds of fetches, and one hiccup shouldn't fail an otherwise good build.
  • .dockerignore must exclude .env*, go.work and go.work.sum.

Build it locally with a key file, the same way CI does:

shell
docker buildx build --ssh default=$HOME/.ssh/id_ed25519 -t myapp-api .

The swarm stack

Three settings decide whether the service recovers on its own:

yaml
x-deploy-policy: &deploy-policy
  restart_policy:
    condition: any        # not on-failure, and no max_attempts
    delay: 5s
  update_config:
    parallelism: 1
    order: start-first
    failure_action: rollback
    monitor: 120s         # how long a new task must stay healthy
  rollback_config:
    order: start-first
    failure_action: pause
Don't use on-failure with max_attempts. If a node reboots and reaches the registry over Tailscale, image pulls fail until Tailscale is up. A capped policy uses up its attempts in that window and leaves the service at 0 replicas with no retry. Swarm won't reschedule it. redelay.com and flowdsl.com were down for about 45 hours this way in September 2026.
  • monitor is what makes failure_action: rollback work. Without it, a task that crashes after its first healthcheck still counts as a successful update.
  • Give every service a real healthcheck. Distroless images have no wget or curl, so ship a small static /healthcheck binary that requests $HEALTH_URL.
  • Keep the stack name stable. docker stack deploy under a new name starts a second stack, and both then claim the same Traefik hosts.

The workflow

The workflow runs on a self-hosted runner on the swarm manager:

  1. Guard. Fail if the checkout contains a go.work, a .env, or any replace in go.mod.
  2. Find the deploy key. Probe the candidate keys with git ls-remote against one repo per private org, and use the first key that can read all of them. A repository or org variable (GO_MODULES_SSH_KEY) can override the choice.
  3. Build and push with docker/setup-buildx-action using driver-opts: network=host.
  4. Deploy. Run docker stack deploy with secrets sourced from an env file on the manager (never from the repo).
  5. Wait for the rollout. Poll each service's UpdateStatus.State. Fail on rollback_* or paused, and print the task errors and recent logs.
  6. Run a health check against the public URL, then clean up the registry (keep the last 5 tags).
network=host is required when the registry is on a Tailscale address. With the default docker-container driver, BuildKit runs on the Docker bridge, and the bridge can't route to 100.x addresses. The build succeeds and then the push fails after minutes of retries with dial tcp 100.x.x.x:5000: i/o timeout, which looks like a registry outage.

Two GitHub quirks:

  • Pass --ssh default=<key path>, not plain default. A runner installed as a service has no $SSH_AUTH_SOCK.
  • The workflow file must exist on the default branch. Otherwise GitHub never registers it: pushes to production start nothing and workflow_dispatch returns 404. Commit it to the default branch as well, even though it only triggers on production.

Deploying is then:

shell
git push origin main:production        # or your integration branch

Monitoring

Probe the public URL with the Prometheus blackbox exporter. That covers every failure a user would see: 0 replicas, a missing Traefik router, TLS, DNS, or app 5xx errors.

yaml
- job_name: public-sites
  metrics_path: /probe
  params: {module: [http_2xx]}
  static_configs:
    - targets: ['https://api.example.com/api/v1/health']
      labels: {project: myapp, service: api}
  relabel_configs:
    - {source_labels: [__address__], target_label: __param_target}
    - {source_labels: [__param_target], target_label: instance}
    - {target_label: __address__, replacement: blackbox-exporter:9115}
yaml
- alert: public_site_down
  expr: probe_success{job="public-sites"} == 0
  for: 3m
  labels: {severity: critical}