· 11 min read

The upgrade that re-opened our front door.

Every compromise story sounds sophisticated until you trace it back to the start. Ours starts with a package upgrade doing exactly what it is documented to do, a broken network route, and a throwaway password that was meant to live for five minutes. There was no zero-day and no exploit chain. A few ordinary things went wrong on the same afternoon and lined up into a hole.

This one is from work, not from the box this site runs on. I am writing it up the way I would want to read someone else’s, with the identifying details filed off. The host, the network and the people are anonymised. The attacker addresses and malware hashes are real, in case they help someone spot the same thing.

The host

The machine that got hit was one of our more important ones. It was the internal certificate authority, the configuration-management server and the SSH jump host we use to reach other environments. It also ran a self-hosted Git service and some monitoring. If you are keeping score on blast radius, that list is the whole story. More on that later.

It was hardened the way you would expect: no root login over SSH, no password authentication, keys only. That mattered until the afternoon it stopped being true.

How it started

Someone ran an OS release upgrade on it, walking it up two major versions in an afternoon. Normal maintenance. Partway through, the OpenSSH server package did something it is fully entitled to do: its post-install script noticed that our sshd_config differed from the package default and put the default back. That default re-enabled root login and turned password authentication on again.

Our hardened config was not gone. It was sitting next to the live one as a .ucf backup file, which is how I later proved this was a regression from the upgrade and not a choice anyone made. But the running config now said that root may log in with a password.

On its own that is a latent risk, not a breach. The second thing the upgrade broke turned it into one: networking. The box came back with no default route and was unreachable over SSH, so it had to be recovered through the provider’s serial console. Pasting into that console was not working, which meant a long random password could not be typed in. To get logged in and fix the routing, my colleague set root’s password to a throwaway, something like test1234, fully intending to change it back within minutes.

The change back did not take. A laggy serial console is exactly where a mistyped confirmation fails silently, and the logs suggest that is what happened: root’s password was changed once during recovery and not again until we reset everything the next morning. Nobody noticed. The password that was supposed to live for five minutes was live for hours, on a host where a bot hammers root login around the clock.

The third thing: the SSH config was corrected a few hours later, but the SSH service was never restarted. The process kept running with the bad config and kept accepting root password logins all day and into the night.

The break-in

A bot had been trying to brute-force root on this box every day for weeks. That is background noise on any host with SSH open to the internet. It had always failed because password authentication was off. It works through common-password wordlists, the kind that have test1234 near the top.

With a wordlist-trivial password and root password login both live for hours, the only question was which brute-force run would get there first. About two hours after the password went on, one did: three failures, then Accepted password for root in the same session. Within about a minute a second address logged in as root and installed the malware. Two more addresses showed up over the next several hours to confirm the credentials and stage tools. Late that night a fourth uploaded the main payload over SFTP.

The timeline, in local time:

  • About 10:52 to 12:35: recovering the box from the serial console after the upgrade. Root’s password is set to a throwaway to get in. The default route is re-added at 12:22.
  • 12:35: the brute-force bot guesses the throwaway password while root password login is on.
  • 12:36: a second host logs in as root and installs the first-stage malware, which starts beaconing from cron every three minutes.
  • 15:40: two more hosts log in as root, check the credentials and stage tools.
  • 23:05 to 23:12: a fourth host uploads a payload of about 30 MB over SFTP.
  • 23:12: the SSH service finally gets restarted and root password login closes, eleven hours too late.

Closing the door did not evict anyone. The malware kept running as root until we cleaned up the next day, so the attackers had root for roughly 29 hours.

One nice detail from the investigation: the payload came in over SFTP, and OpenSSH logs SFTP file operations, so the upload is in the auth log byte for byte. You can watch them mkdir a hidden directory, open the file for write, and close it after writing exactly 30,304,472 bytes. There is no guessing about what landed.

What they installed

Two families, both commodity.

The first was XorDDoS, an old but stubborn Linux DDoS bot. Its footprint was textbook: a 32-bit payload dropped as a fake system library, a cron job relaunching it every three minutes, an init script for persistence, and a running process masquerading as syslogd. It also disabled a set of cloud-provider agent services by symlinking their unit files to /dev/null, a known XorDDoS housekeeping move that keeps cloud security agents out of its way.

The second, and the reason the box got noticed at all, was a Go-based SSH password scanner. It ran from a hidden directory with a long list of target IPs as arguments, opened thousands of concurrent outbound connections to port 22 across the internet, and pegged a core at 95% CPU. That outbound scanning is what an upstream noticed and flagged. The binary was also installed as a systemd service with 51 target addresses baked into its unit file, so it would come back after a reboot.

They also left renamed copies of wget and curl under innocuous names, presumably as download helpers, and a second init-script persistence with a random name.

If you want to check your own fleet, the indicators are at the bottom.

Finding it and cleaning up

The tell was the outbound port-22 storm. A single process with thousands of established connections to random hosts on port 22 is not something a jump host does on its own. From there it was process triage, reading the quarantined binaries, and then the slow part: reconstructing the whole thing from logs.

That reconstruction was harder than it should have been, which is a lesson in itself. The system log had been empty for years because of a misconfiguration, the journal retained about a day, and the auth log was so flooded with monitoring-exporter noise that every query needed heavy filtering. We got the full story out of it, but with any less retention we would have been guessing.

Our on-call engineer killed the malicious processes and stripped the persistence while I ran a read-only forensic pass in parallel. One of us restored service; the other preserved evidence and worked out scope. After a reboot the host came back clean: no scanning, no leftover persistence, core system binaries intact, and the mail relay checked and confirmed not abused. We pulled a full evidence bundle off the box before wiping it.

The part that kept me up

Cleaning malware off a box is easy. The hard question on a host like this is what root could have reached from there.

This is why a jump host is such a good target. Ours had, readable by root: the private key of an internal certificate authority that signs host and user certificates across our estate; the configuration-management server’s own keys, which effectively control what runs on every managed machine; VPN keys; the secrets and database of the Git service; cloud and registry credentials; and a scattering of TLS private keys and personal SSH keys in home directories.

The uncomfortable part, and the thing worth taking away even if you never meet XorDDoS: hardware security keys on your engineers’ laptops do not protect any of that. A YubiKey protects the secrecy of one private key. It does nothing for a certificate authority key sitting on the server, and nothing for a config-management server that can push a root-level command to every node it manages. An attacker with root on a box like this does not need to steal anyone’s protected credential. They can mint their own certificate from the CA, push a config change, or ride an already-authenticated session while an engineer is connected. The reverse tunnels and agent forwarding that make a jump host convenient are what make its compromise expensive.

We found no evidence that any of that happened. The intrusion looked like an automated botnet grabbing another host to scan from, not a human operator pivoting deeper. But “no evidence” from a box with a one-day log horizon is not the same as “it did not happen”. We treated every one of those secrets as exposed, rotated them, and checked the far ends for logins from the compromised host during the window.

The boring lessons

None of this is clever. That is the point.

Your hardening can be silently reverted by a package upgrade. Config that lives in the main sshd_config is fair game for a package’s post-install script to overwrite. Put your SSH hardening in a drop-in under /etc/ssh/sshd_config.d/, where the package manager will not touch it, and it survives major upgrades.

A config change is not applied until the service restarts. Fixing the file at 13:16 did nothing, because the daemon kept running with the old config until 23:12. If you edit sshd_config, restart sshd in the same breath and check with sshd -T that the effective policy is what you think it is. The running daemon’s policy is what matters, not the file on disk. Editing and moving on is how you get an eleven-hour gap you do not know about.

There is no such thing as a temporary weak password on an internet-facing host. This is the one that got us. test1234 was supposed to live for five minutes. The change back did not take, nobody verified it, and a bot that brute-forces you around the clock only needs one of the hours that followed. There are two lessons in that. First, do not set a stopgap credential on a host that is still exposed: cut it off (pull it behind the firewall, or drop inbound 22), set whatever you need, and lock root down again before you re-expose it. Second, never assume a change applied because you typed it, especially over a laggy console. Confirm it. Better still, keep root with no password at all and recover through a real console account, so a stopgap like this never has to exist.

Rate-limit and scope SSH. The box had a decent firewall, but port 22 was open to the whole internet. That is how a bot got thousands of attempts a day and was ready the moment a weak password appeared. Restrict SSH to known networks or a VPN, and run something like fail2ban so a brute-force campaign does not get unlimited tries.

Keep enough logs to investigate. Use persistent journald with real retention, and do not let noisy exporters bury your auth log. The difference between a confident report and a shrug is whether the logs were there.

Know your blast radius before you need to. The most valuable artifact from this whole thing was the honest inventory of what a single host could reach. Write that down for your important hosts now, while it is calm. If one is ever compromised you will already know what to rotate, and you will not be discovering it at 2am.

Indicators of compromise

Attacker addresses seen in this incident:

  • 45.148.10.151: the brute-force source that guessed the password
  • 172.82.91.35 and 163.128.235.62: payload delivery
  • 193.47.62.69: credential validation
  • 109.160.32.0/24: the bulk of the brute-force volume

Malware artifacts (SHA-256):

  • Go SSH scanner (ran from a hidden directory, also installed as a fake systemd-worker service): 94f2e4d8d4436874785cd14e6e6d403507b8750852f7f2040352069a75da4c00
  • XorDDoS payload (dropped as a fake libudev.so, 32-bit ELF): 99ff974e391f7a3c879097d47ccde7578cc7a38dbeb257ccd983f7a376e5ecb2
  • Cron relauncher shell script: 74d31cac40d98ee64df2a0c29ceb229d12ac5fa699c2ee512fc69360f0cf68c5
  • Init-script persistence: eee3a8cea44d28591133025523d5fbb4c52f877a5af9e72f8f6974f1648b6805

Host-side signs worth grepping for: a process claiming to be syslogd that is pegging a CPU; a cron entry firing a script out of /etc/cron.hourly every few minutes; a fake libudev.so outside the normal library path; system unit files symlinked to /dev/null; and a burst of outbound connections to port 22 across many unrelated hosts.