Migrating a live mail server from hand-edited configuration files to Ansible is not a rewrite. It is a controlled sequence of small, reversible changes. The goal is not to eliminate risk; it is to make each step observable and recoverable before the next one begins.
This article describes a migration pattern for a single mail server or a small cluster. It assumes Postfix and Dovecot, a VPS or homelab host, and an operator who can tolerate a few minutes of elevated attention but not a bounced-message incident. The pattern is conservative by design. It uses Ansible’s documented execution controls, Postfix’s own validation commands, and file-level staging rather than in-place edits.
Why not just template everything at once
The temptation is to write a role that renders main.cf, master.cf, and Dovecot’s configuration tree, then run it against production. That approach conflates two problems: learning what the current configuration actually is, and changing it. On a mail server, the first problem is harder than it looks. Hand-edited files accumulate comments, local overrides, and parameters that were added during an incident and never removed.
A safer sequence is:
- Capture the current state as data.
- Reproduce that state with Ansible in a non-destructive way.
- Introduce changes one at a time, with validation and rollback at each step.
Each step should be independently useful. If the migration stalls after step two, the server is no worse off than before.
Step 1: Capture the current configuration
Postfix provides postconf for reading effective parameters. Running postconf -n prints only parameters that differ from compiled defaults, which is the useful subset for a migration. The output is stable enough to diff across runs.
Dovecot’s equivalent is doveconf -n, which prints non-default settings. The Dovecot documentation site has moved and reorganized over time; the configuration manual is the authoritative reference for the current version, and the doveconf man page documents the -n flag. If the documentation URL you have bookmarked returns a 404, check the version-specific path for your installed release rather than assuming the command has changed.
Store these outputs in version control before writing any Ansible. They are the baseline. A diff against this baseline is the only reliable way to know whether a playbook run changed something you did not intend.
Step 2: Stage files, do not edit in place
The ansible.posix.synchronize module wraps rsync and is documented as originating on the local host where Ansible runs, with the destination being the host it connects to. It supports delegate_to, which allows copying between two remote hosts or entirely on one remote machine. It also enables --delay-updates by default, which the documentation describes as avoiding leaving a destination in a broken in-between state if the underlying rsync process encounters an error.
That default matters for mail configuration. A partially written main.cf is not a valid configuration. Using synchronize to push a rendered tree into a staging directory such as /etc/postfix/staged/ keeps the live files untouched until a separate task moves them into place.
A minimal pattern:
- name: Stage rendered Postfix configuration
ansible.posix.synchronize:
src: "{{ role_path }}/files/postfix/"
dest: /etc/postfix/staged/
delete: true
rsync_opts:
- "--chown=root:root"
- "--chmod=D755,F644"
The delete: true option removes files in the destination that no longer exist in the source. The documentation notes it requires recursive=true and behaves like --delete-after. For a staging directory that is fully managed by Ansible, this is the correct behavior. For a directory containing operator-local files, it is not.
Step 3: Validate before switching
Postfix documents postfix check as a command that causes Postfix to report file permission and ownership discrepancies. The same documentation recommends running it nightly before log rotation, alongside a grep for reject, warning, error, fatal, and panic lines. That recommendation is about routine hygiene, but the command is equally useful as a pre-switch gate.
A validation task can run postfix check against the staged configuration by temporarily pointing Postfix at it, or by using postconf -c to read from an alternate configuration directory. The exact invocation depends on your Postfix version; check postconf(1) for the -c flag on your system. The principle is to fail the play before the live files change, not after.
For Dovecot, doveconf -n reads the active configuration. To validate a staged tree, run doveconf -c /etc/dovecot/staged/dovecot.conf -n and compare the output to the baseline. If the command exits non-zero, stop.
Step 4: Reload, do not restart
Postfix’s basic configuration documentation states plainly: whenever you make a change to main.cf or master.cf, execute postfix reload as root to refresh a running mail system. That is the documented mechanism. A full restart is not required for configuration changes and is more disruptive to active SMTP sessions.
Dovecot similarly supports reloading configuration without dropping IMAP connections, though the exact signal or command depends on your init system and Dovecot version. The doveadm reload command is the documented interface in current releases. Verify against your installed version’s man pages rather than relying on older forum posts.
In Ansible, this maps to a handler. Handlers are designed to run only once per play; verify against the current Ansible handlers documentation for the exact semantics in your Ansible version.
The pattern is:
- name: Deploy Postfix main.cf
ansible.builtin.copy:
src: /etc/postfix/staged/main.cf
dest: /etc/postfix/main.cf
owner: root
group: root
mode: '0644'
backup: true
notify: reload postfix
- name: Deploy Dovecot configuration
ansible.builtin.copy:
src: /etc/dovecot/staged/dovecot.conf
dest: /etc/dovecot/dovecot.conf
owner: root
group: root
mode: '0644'
backup: true
notify: reload dovecot
The backup: true option preserves the previous file with a timestamp suffix. That is your rollback artifact. It is not a substitute for version control, but it is available on the host without network access.
Handlers are defined separately:
handlers:
- name: reload postfix
ansible.builtin.command: postfix reload
- name: reload dovecot
ansible.builtin.command: doveadm reload
Using command rather than service is deliberate. service with state: reloaded may fall back to a restart on some init systems if the reload operation is not defined. postfix reload is unambiguous.
Step 5: Control the blast radius
On a single mail server, serial has no effect. On a small cluster, it is the difference between a controlled rollout and a simultaneous outage. Ansible’s documentation describes serial as completing the play on a specified number or percentage of hosts before starting the next batch. It also notes that setting the batch size changes the scope of failures to the batch size, not the entire host list, and that max_fail_percentage can modify this behavior.
For a three-node mail cluster, a conservative play might use:
- hosts: mailservers
serial: 1
max_fail_percentage: 0
tasks:
# ...
serial: 1 means one host at a time. max_fail_percentage: 0 means any failure stops the play before the next host is touched. This is slower than a parallel run and appropriate when a failed reload on one node could cascade.
The documentation also describes throttle, which limits the number of workers for a particular task. It can be set at the block and task level and is useful for tasks that are CPU-intensive or interact with a rate-limiting API. A DNS zone transfer or an API call to a monitoring service are candidates. A configuration file copy is not.
Step 6: Delegate the checks that should not run on the mail server
Ansible’s delegation documentation describes delegate_to as a way to perform a task on one host with reference to other hosts. The canonical example is removing a web server from a load balancer pool before updating it. The same pattern applies to mail: a task that checks whether a node is still accepting connections should run from the control node or a monitoring host, not from the node being updated.
Delegation also has a documented concurrency caveat. Tasks are executed in parallel by default, and delegating a task does not change this. Multiple forks writing to the same file on a delegated host will overwrite each other. The documentation suggests run_once: true with a loop, or an intermediate play with serial: 1, or throttle: 1 at the task level. For a migration playbook that writes a single summary file or updates a single DNS record, this matters.
Step 7: Rollback is a file copy
The rollback path should be shorter than the forward path. If the staged configuration fails validation, nothing has changed. If the live configuration fails after reload, the previous file is available from the backup option or from version control.
A rollback task is not a separate playbook. It is a conditional branch:
- name: Restore previous Postfix configuration
ansible.builtin.copy:
src: "{{ postfix_backup_path }}"
dest: /etc/postfix/main.cf
remote_src: true
when: postfix_reload_failed | default(false)
notify: reload postfix
The postfix_reload_failed variable would be set by a register on the reload task combined with failed_when or a subsequent check. The exact mechanics depend on how you detect failure. A reload that exits zero but leaves the service unable to accept connections is a different problem; that is what monitoring is for.
What this pattern does not solve
It does not solve configuration drift that predates the migration. If the hand-edited files contain parameters that are not in your baseline capture, the first Ansible run will either preserve them or remove them depending on how the template is written. Capture first, diff second, template third.
It does not solve the problem of a mail server that is already unhealthy. Migrating a broken configuration to Ansible produces a broken configuration managed by Ansible. Fix the underlying issue before or during the migration, not after.
It does not eliminate the need for out-of-band access. If a reload leaves the server unreachable over SSH, you need console access or a rescue mode. Ansible cannot help with that.
FAQ
Can I use service with state: reloaded instead of command?
You can, but the behavior depends on the init system and the service unit. On systemd, systemctl reload postfix maps to the ExecReload directive if one is defined. If it is not, systemd may return an error or fall back to restart depending on the unit. postfix reload is documented by Postfix and does not depend on the init system’s interpretation.
How do I know whether a reload actually took effect?
Check the logs. Postfix logs to syslog, and the basic configuration documentation describes the logging classes and levels. A successful reload produces a log line indicating that the master daemon has re-read its configuration. Absence of that line is a signal to investigate. For Dovecot, the equivalent depends on your logging configuration; doveadm log errors or the configured log file is the place to look.
Should I run postfix check before or after the reload?
Before. The Postfix documentation recommends it as a routine check for file permission and ownership discrepancies. Running it against the staged configuration before the live files change is the point. Running it after the reload tells you that something is wrong, but by then the live configuration is already active.
What about Dovecot’s configuration validation?
doveconf -n prints non-default settings from the active configuration. To validate a staged tree, use the -c flag to point at the staged configuration file. The exact syntax is documented in the doveconf(1) man page for your installed version. If the command exits non-zero, do not proceed.
Is serial: 1 necessary for a single mail server?
No. serial controls how many hosts Ansible manages at a time. With one host, it has no effect. It becomes relevant when you have two or more mail servers and want to avoid a simultaneous reload.
How do I handle secrets in the Ansible repository?
Ansible Vault is the documented mechanism for encrypting sensitive data at rest. The alternative is to keep secrets out of the repository entirely and inject them at runtime from a separate source. Either approach works; the important thing is that the repository does not contain plaintext credentials. This article does not cover vault setup in detail because the Ansible documentation already does.
Summary
The migration from hand-edited configs to Ansible on a live mail server is a sequence of small, validated steps. Capture the current state with postconf -n and doveconf -n. Stage rendered files with synchronize rather than editing in place. Validate with postfix check and doveconf -c before switching. Reload with postfix reload and doveadm reload rather than restarting. Use handlers so reloads happen once per play. Use serial and max_fail_percentage on clusters. Keep the rollback path shorter than the forward path.
None of this is novel. It is the documented behavior of the tools, applied in an order that keeps the mail flowing while you work.