The server compromise
A framework RCE, a botnet, three failed rebuilds and the outbound connection that found it
- Role
- Solo, incident response
- Status
- Shipped
- Source
- Private repository
- Stack
- Ubuntu 24.04Next.jsOpenSSHUFWfail2bansysctlssTailscale
In numbers
3
clean-ups and rebuilds re-infected before the root cause
4h
from first deploy to root cause
42s
to install the patched Next.js 16.0.11
18
crontab entries the malware wrote for persistence
10
hardening steps, then 10 verification checks
The problem
I deployed the ValoxVSL video tool to a fresh cloud server. About an hour in, the server was found running x86_64.kok, a Mirai variant disguised as systemd-logind and persisted through crontab (18 entries), rc.local, init.d and bashrc. A clean-up in rescue mode did not hold, and neither did two rebuilds, one with new SSH keys and one behind a cloud firewall that limited SSH to a single IP. Each time it was back within five minutes.
The approach
I stopped cleaning and looked at what the server was talking to. ss -tnp state established showed the Next.js process itself holding a connection to a command-and-control server, which moved the search from SSH to the application. The npm dependency tree checked out clean. The cause was React2Shell, an unauthenticated remote code execution in the React Server Components protocol, which the incident report records as CVE-2025-66478 (CVSS 10.0). The app ran Next.js 16.0.1, and the RondoDox botnet was exploiting it over ports 80 and 443; each rebuild had redeployed the same version. Upgrading to 16.0.11 took 42 seconds. The server was rebuilt once more, its secrets were rotated, and a 15-minute automated watch came back clean.
How it works
Outbound connections as the signal
Inbound firewall rules cannot see a connection the server opens itself. ss -tnp state established lists each established connection with the process that owns it, and an application process connected to an unknown host is where to start. It is now a line of the verification checklist, and the expected output is my own SSH session and nothing else.
Two attacks at once
auth.log showed SSH brute force, so three rounds of work went into SSH. That attack was real and concurrent; the persistent one came through the application's own ports, which have to stay open. Ubuntu 24.04 added noise: ssh.socket's default TriggerLimitBurst of 20 kills the SSH listener after 20 connection attempts, which is why SSH kept dying during the brute force.
The playbook
Written from the incident: ten steps, from disabling ssh.socket, through key-only SSH with Mozilla's modern ciphers, a non-root user and a locked root, UFW deny-incoming, fail2ban, unattended upgrades with automatic reboot off, sysctl hardening and a login banner, to egress blocks for known C2 addresses. Ten checks verify it afterwards. Its rule for a compromised server: do not clean it; dump the database, rebuild fresh and rotate every secret, since malware inside the app process can read its environment.
A second lesson, later
Later I locked myself out of the game server: UFW and the cloud firewall both pinned SSH to one IP, and the match failed. The playbook now says key-only SSH plus fail2ban, with no IP pinning in UFW. The game server was rebuilt to it almost line for line, down to the banner text. Later hosts start from it and adapt: admin surfaces bind to the tailnet or loopback, and the Biogard stack adds container isolation.
What I chose, and what lost
Chose
Patch the framework, rebuild once more and rotate every secret
Over
A fourth round of SSH hardening and clean-up
Each rebuild had redeployed the vulnerable Next.js, so the server was re-infected over HTTP whatever the SSH setup was.
Chose
Key-only SSH plus fail2ban, with no IP pinning in UFW
Over
Restricting SSH to one IP in UFW and the cloud firewall
That setup later locked me out of the game server.
Outcome
Resolved once the framework was patched. The report records the patched deployment clean on a 15-minute watch and after two hours of uptime, with the botnet's later attempts showing up in the logs as rejected Server Action calls. The post-mortem runs to 647 lines in 21 sections, with a narrative version beside it. The playbook is the baseline for new servers and is applied unevenly: the game server follows it almost exactly, and the others take parts of it and substitute container isolation and tailnet-only admin.
What comes next
Bring the oldest box in the fleet up to the playbook it inspired.