Deploying Stalwart Mail v0.16 declaratively on Kubernetes — GitOps with no config.toml
Stalwart v0.16 removed TOML config files. Here's why pinning to v0.15.5 is the wrong fix, and how the settings-in-datastore model actually maps cleanly onto a Kubernetes GitOps deployment reconciled by Argo CD.
The crash
We bumped the image tag, Argo CD synced, and the pod went into CrashLoopBackOff. The log was short:
Failed to parse data store settings: TOML parse error at line 1 column 1
|
1 | {
| ^
expected valueThe parser was reading our config.toml and finding a { — because the file it actually wanted was JSON. Stalwart v0.16 removed the TOML configuration file format. The server no longer reads a monolithic config.toml at all.
The easy fix is to pin the image to v0.15.5, the last release that still speaks TOML. It works — and it strands you on an old release of an actively-developed mail server, with a migration you'll have to do anyway, later, under worse conditions. It's the wrong fix. Understand what replaced TOML instead.
What v0.16 actually does
In v0.16, the on-disk config shrinks to a single small file that describes only where the datastore lives. Everything else — listeners, TLS, directories, queues, DKIM, spam rules — is stored as structured objects inside the datastore, managed through stalwart-cli. Our config.json is four lines:
{"@type":"RocksDb","path":"/opt/stalwart/data","blobSize":16834,"bufferSize":134217728,"poolWorkers":null}jsonThat's it. On boot with a valid config.json pointing at an initialized datastore, the server comes straight up in normal mode and binds its listeners — SMTP on 25, submissions on 465, IMAPS on 993, POP3S on 995, ManageSieve on 4190, and the HTTP/JMAP API on 443.
The settings themselves are managed like Vault policies or Terraform state: as a declarative plan you apply. stalwart-cli apply consumes an NDJSON plan (one JSON object per line) and reconciles the running server toward it, idempotently — re-applying the same plan is a no-op. That's the same infrastructure-as-code pattern we already use everywhere else, which is why the migration ended up simplifying our setup rather than complicating it.
Mapping it onto Kubernetes GitOps
The pieces line up cleanly once you stop thinking "config file" and start thinking "datastore + a plan":
config.jsonas a ConfigMap, mounted withsubPath. Mounting the single file (not the directory) means the pod boots directly into normal mode and skips the interactive setup wizard. No init container, no first-run web form to click through.- The datastore on a PVC. RocksDb at
/opt/stalwart/datais the source of truth for every setting, so it has to survive pod restarts. This is the one piece of durable state. - Settings seeded and reconciled from a git-committed plan. An Argo CD PostSync Job runs
stalwart-cli applyof the NDJSON plan against the live service on every sync. We've verified the idempotency claim in production: upserts converge, and replaying the full plan against an already-configured server reports zero failures.
The authoring loop that actually works (we verified this end to end):
- Create objects with
stalwart-cli create Object/Variant— e.g.stalwart-cli create Directory/Oidc --field issuerUrl=.... Do NOT hand-write@typevariant discriminators inside an apply-planvalueobject for creates — that path fails; the CLI's builder constructs the correct wire format for you (the WebUI works too). stalwart-cli snapshotthe configured server. The snapshot output is a valid apply plan. One wrinkle: snapshot refuses to serialize unresolved references — pass--allow-unresolved <Ref>for each reference type you haven't populated (Tenant,DnsServer,Certificate, ...).- Filter any secret material out of the snapshot (see the next section for why there usually is none), commit it as
plan.ndjson, and let the PostSync Job replay it forever after.
Create → snapshot → filter → commit → apply beats writing the plan by hand, every time.
TLS, DKIM, and secrets — the File variant is the Kubernetes-native answer
The part that worried us most — "does this mean private keys end up in the datastore or, worse, in git?" — turned out to have a clean answer. Stalwart's secret-bearing fields (SecretText, PublicText, SecretKey) are @type-tagged enums with three variants: Text (inline), EnvironmentVariable, and File:
{"@type":"File","filePath":"/certs/tls.key"}json(The serde rename file_path → filePath lives in crates/registry/src/schema/structs.rs if you want to check the source.)
The File variant is the one you want on Kubernetes. Point the Certificate object's key/cert fields at a cert-manager-managed Secret mounted into the pod, and: (a) no key material ever enters the datastore, the plan, or git — the plan only carries a file path; (b) cert-manager renewals are picked up on the next pod restart. This replaced our initial instinct of pasting PEM into the plan, which would have been strictly worse on every axis.
Two related findings, both verified live:
- TLS binding: a fresh v0.16 server serves an rcgen self-signed certificate until a
Certificateobject exists. Create one withFilerefs as above — SANs are parsed server-side from the cert — and a pod restart makes all the TLS listeners serve it. We confirmed Let's Encrypt certs on:465and:993this way. - DKIM: with
generateDkimKeys=truein theBootstrapobject, creating aDomainauto-mints RSA and Ed25519 DKIM signatures. The TXT records you need to publish are handed to you in the Domain's server-setdnsZoneFilefield. The DKIM private keys stay datastore-resident — the datastore backup covers them, and they never belong in git.
How the platform made this clean
None of this needed bespoke tooling — the mail server is just another workload in our deploy-mesh GitOps repo:
- Argo CD reconciles it alongside everything else — ConfigMap, PVC, Deployment, and apply-plan are all declarative YAML/NDJSON in git.
- The recovery admin secret comes from HashiCorp Vault via External Secrets Operator — never committed to git, injected as a Kubernetes Secret at sync time.
- Users authenticate via Keycloak OIDC (
sso.kraftware.dev), with Stalwart pointed at it as an external directory. Our universal rule: every account-bearing service uses one SSO, so there are no per-service passwords to manage. - It runs on our grow-as-needed sovereign k3s edge nodes — one env file plus one command stands up a new node and Argo CD populates it.
The nice surprise: v0.16's "settings live in the datastore, reconciled by a plan" model aligns with a config-only-via-git doctrine instead of fighting it. TOML-in-a-file was always an awkward fit — you baked it into the image or mounted a big ConfigMap and hoped nothing drifted. A git-committed apply-plan that idempotently reconciles a running server is exactly the shape everything else in our repo already has.
Gotchas that cost us time
- The tracer variant is
Stdout, notConsole. If you're configuring logging to land on stdout forkubectl logs, the enum value isStdout.Consoleis rejected. --dry-runis client-side only. It validates that your plan parses; it does not catch server-side enum validation (like theStdout/Consolecase above). Apply-and-verify beats dry-run alone — run the apply against a real server and check the result, don't trust--dry-runto be a full gate.- The
Bootstrapobject is only applyable in bootstrap mode. Trying to apply a plan containing aBootstrapobject against a server already running in normal mode is rejected. Keep bootstrap-only objects out of your steady-state reconcile plan. stalwart-cliships from a separate repo. It is not in the server container image. It's released independently at [stalwartlabs/cli](https://github.com/stalwartlabs/cli) — pull it from there for your init/reconcile job rather than expecting it alongside the server binary.primaryKeyViolationon create usually means you already succeeded. If an earlier create looked like it failed and the retry reportsprimaryKeyViolation, the first attempt very likely went through. Check with a read before assuming the object is missing.- Hand-written variant creates fail. As above: embedding
@typevariant discriminators in an apply-planvaluefor a create doesn't work — go throughstalwart-cli create Object/Variant(or the WebUI) and snapshot the result.
Where the platform stands
For context on why a clean mail migration matters to us: we're building an edge-first, code-as-infrastructure platform on sovereign single-node k3s clusters, provisioned grow-as-needed — one env file plus one command spins up a new node. It's provider-agnostic; we've proven the same pattern on Oracle Cloud ARM and Hetzner x86 in different regions. There's no stretched etcd and no cross-WAN consensus: each site is sovereign, and workloads deploy by git commit through an Argo CD app-of-apps.
The current beta runs two clusters — an OCI control hub and a Helsinki "eu1" edge node — with Keycloak SSO across every UI, drift-alerting and clock-discipline monitoring, a Cloudflare-R2-backed container registry, nightly backups to R2, and now declarative Stalwart mail, all reconciled from one GitOps repo with secrets from Vault via External Secrets. This is an honest beta — nearly ready for test clients. The grow-as-needed pattern, SSO, GitOps, and backups are proven; we're finishing per-tenant onboarding. Helsinki also doubles as a dogfood target for remote/cellular edge deployments: the ~120ms link is a decent stand-in for a real tenant site.