·   ·  6 posts
  • 1 members
  • 10 friends

Lock down your self-hosted AI in an afternoon: lessons from what's hitting other AI operators

I'm a big believer in running your own AI. It's the whole reason IRL exists. But owning the stack means owning the security too, and over the last few weeks a lot of other people and businesses running their own AI have found that out the hard way. None of this happened to IRL, but every one of these incidents is worth learning from before it happens to you.

None of this needs a security team, just an afternoon on a box you already run.

What others have run into

PoeLLM, a crypto-mining botnet aimed at AI servers. This week Lumen's Black Lotus Labs reported a botnet it calls PoeLLM that has compromised more than 3,400 servers since April, mostly in the US and Western Europe. Many of the victims were running exposed LiteLLM and Ollama instances, alongside Gotenberg and Gitea. Once in, it mines cryptocurrency and uses the server to hunt for the next victim. On LiteLLM it goes after an endpoint covered by CVE-2026-42271, which was fixed in LiteLLM 1.83.7.

A critical flaw in GitLab's self-hosted AI Gateway. On 2 October GitLab disclosed CVE-2026-90970, rated 9.9 out of 10. A logged-in user with Duo Agent Platform access could craft a custom flow that escaped the prompt template sandbox and ran commands on the gateway. It's fixed in AI Gateway 19.2.4, 19.3.2 and 19.4.1. Only self-hosted gateways need action, but if that's you, it's urgent.

A coding agent that reportedly deleted 48,000 files. Last month a Reddit user reported that a Claude Code sub-agent, asked to rebuild a test copy of a project, wrote a cleanup script that followed Windows directory junctions into the live project. By the user's account and the agent's own log, it deleted about 48,000 files in 103 seconds, including the Git object store, so Git couldn't restore anything. It hasn't been independently verified, but the failure mode is real.

None of these is exotic. The people caught out had an exposed port, unpatched software, or an agent with far more reach than its task needed.

The checklist

  1. Keep model endpoints off the public internet. Ollama's API has no built-in authentication, so anyone who can reach port 11434 can use your models and your hardware. Treat LiteLLM, Open WebUI and llama.cpp's server the same way.
  2. Bind to localhost. Set OLLAMA_HOST=127.0.0.1:11434 and do the same for other services. In Docker, publish ports as 127.0.0.1:11434:11434, not 11434:11434. Docker's published ports can bypass ufw rules.
  3. Put a reverse proxy with authentication in front if you need to share. Caddy or Nginx with TLS and proper auth is enough for most small setups. In Open WebUI, turn off open sign-ups once your accounts exist.
  4. Use a VPN for remote access. WireGuard or Tailscale lets you reach your models from anywhere without opening a port. With Tailscale, share services with Serve (inside your tailnet), not Funnel (the public internet), and use ACLs so only the devices that need the model can reach it.
  5. Default-deny your firewall. Block all inbound traffic, then allow only what you need. Test from outside your network, for example from your phone on mobile data, rather than trusting the config.
  6. Patch on a schedule. Upgrade LiteLLM to 1.83.7 or later, update any self-hosted GitLab AI Gateway, and keep Ollama, Open WebUI and llama.cpp current.
  7. Rotate keys if anything was exposed. A gateway like LiteLLM holds your cloud provider API keys. If it was ever reachable from the internet, assume those keys are burnt.
  8. Give agents the least privilege that works. Run coding agents under their own user account, with no sudo, no production credentials and API tokens scoped to the task.
  9. Sandbox anything that can run commands. Use a container or VM, mount only the project folder, and restrict outbound network access. On Windows, check for junctions and symlinks that point outside the sandbox. That's exactly what reportedly went wrong in the file-wipe story.
  10. Back up before an agent touches files. Push to a remote Git repository and take a snapshot (ZFS, Btrfs or a VM snapshot) before the run. Keep at least one backup the agent can't reach at all, because in the reported wipe the local .git folder went too.
  11. Require approval for destructive actions. Most agent tools can pause for sign-off before deleting, overwriting or running shell commands. Turn it on.
  12. Watch for miners. Warning signs are CPU or GPU pinned at 100% when nobody's using it, unfamiliar processes (XMRig is a common one), new cron jobs or systemd services, and outbound connections to mining pools. Lumen has published indicators of compromise for PoeLLM that are worth checking against.

A 15-minute self-audit

  • Run ss -tlnp and look for anything listening on 0.0.0.0 or :: that shouldn't be.
  • Run docker ps and check which ports are published to all interfaces.
  • Check nvidia-smi or top for load you can't explain.
  • Look through crontab -l and your systemd services for anything you didn't add.
  • Compare each AI tool's version with its latest release.

If any of that turns something up, take the box off the network first and investigate second.

Own it properly

Self-hosting is still the right call for privacy, cost and control. It just comes with the responsibilities a cloud provider used to quietly carry for you.

Want to build it properly from day one? Join the 30-Day Sovereign Builder Challenge. Over 30 days you go from little or no technical background to running your own secure Linux environment, hardening production-grade servers and deploying your own AI agents, alongside other builders in the IRL community. See the full day-by-day blueprint here: https://imreal.life/page/30-day-challenge

Got a question about locking down your own setup? Post it in the comments below.

Sources

  • More
Comments (0)
Login or Join to comment.

IMREAL.LIFE

Close