Incident Overview
I took over this emergency ticket after an overnight batch server migration halted incoming email for dozens of corporate clients. Company management initially accused the infrastructure manager of an operational design mistake during account transfer. I ran a forensic proof-of-concept in an isolated lab and proved that the outage was caused by a native bug in cPanel/WHM, exonerating the colleague and restoring the mail service.
Technical Lead: Percio Andrade Castelo Branco (Linux Infrastructure & Forensics).
Environment and Operational Scope
- Affected Cluster: cPanel/WHM servers on CentOS/CloudLinux running Exim MTA.
- Mail Topology: Corporate accounts routed through a cloud AntiSpam gateway.
- Routing Rule: Mandatory Remote Mail Exchanger policy saved in
/etc/remotedomains. - Migration Window: Batch migration via WHM Transfer Tool (overnight, 00:27 to 04:13).
image_02 - Remote Mail Exchanger routing parameter in cPanel for external AntiSpam delivery.
The Failure Chain and the Accusation
- During the overnight maintenance window, dozens of accounts were transferred to the new server.
- To avoid overwriting custom DNS zones and third-party pointers, the infrastructure manager unchecked "Update DNS Zone" in the Transfer Tool.
- In the morning, clients began filing urgent tickets reporting that no external emails were arriving.
- A quick inspection revealed that migrated domains had switched on their own to Local Mail Exchanger.
- Management blamed the infrastructure manager, claiming he had misconfigured the migration or forgotten to check the accounts.
Production Log Audit
To move away from assumptions and inspect actual system data, I checked the raw service logs:
# 1. Checking the Exim local routing table
grep client-domain.com /etc/localdomains
# 2. Testing mail host network resolution
ping mail.client-domain.com
# 3. Checking for manual edits in cPanel's Zone Editor
grep cpanel_user /usr/local/cpanel/logs/access_log | grep zone_editor/index.html
- Access Logs: The cPanel
access_logconfirmed that nobody had modified the Zone Editor or touched mail routing settings after the migration. - Transfer Logs: The migration log contained silent warnings:
The system could not restore the zone [...] because it does not match any domain on this account.
image_03 - Terminal evidence collection: checking /etc/localdomains and reading raw cPanel logs.
Laboratory Proof of Concept (PoC)
To recreate the behavior without touching production accounts, I configured an identical test environment:
- Created a test domain (
dominio123.com.br) configured as Remote Mail Exchanger, verifying its record in/etc/remotedomains. - Deleted the MX and CNAME records in the Zone Editor to mimic the missing DNS zone left by the migration tool.
- Manually executed cPanel's periodic MX verification script:
/scripts/checkalldomainsmxs --yes
image_04 - The evidence: /scripts/checkalldomainsmxs automatically flipped delivery to LOCAL MAIL EXCHANGER when it could not resolve local MX records.
image_05 - The disk state: domain removed from /etc/remotedomains and written into /etc/localdomains.
Confirmed Root Cause (RCA)
- Software Bug (cPanel Ticket #95782358): The WHM Transfer Tool dropped the mail routing policy on the destination host when "Update DNS Zone" was disabled.
- Aggressive Fallback Script: The cron script
/scripts/checkalldomainsmxsfound the missing MX records and automatically forced the route to Local Mail Exchanger. - Forensic Conclusion: The infrastructure manager made the correct decision to protect custom customer DNS records. The failure was caused by cPanel software.
Standard Workaround Procedure
To normalize production and safeguard the remaining migration schedule, I standardized this CLI workaround:
# 1. Rebuild the DNS zone in BIND
/scripts/rebuilddnszone client-domain.com
# 2. Force Remote routing via WHM API with automatic detection disabled
whmapi1 set_manual_mx_redirects domain=client-domain.com mx=remote
# 3. Verify domain presence in Exim's remote routing table
grep -E '^client-domain.com$' /etc/remotedomains
# 4. Run the MX check script to confirm stability
/scripts/checkalldomainsmxs --yes
Operational Results
- 100% of domains restored to proper external AntiSpam routing.
- Infrastructure Manager fully vindicated: The technical report and video evidence were presented to management, closing the dispute.
- Workaround Adopted: The procedure was implemented across subsequent maintenance windows without any further routing issues.
Discuss this service
Do you want to apply this incident response model in your environment with evidence-driven technical execution?