n8n Workflow Engine
The central nervous system. A self-hosted n8n instance routes scheduled health checks, inbound webhook events and SSH executions between every other component, so there is no bespoke API glue to maintain.
Human-in-the-loop enforced
AutOps watches your containers, reads the logs when something breaks, works out why, and writes the fix. Then it waits for you to say yes. Detection in under a minute, resolution in a five-second reply from your phone — and not one state-changing command without explicit approval.
docker ps againTraditional server management runs on a reactive loop that costs a human being their evening.
A monitor flags an issue. An alert fires. Somebody connects to a VPN, authenticates over SSH, hunts for the failing container, pulls the error logs, reads them, works out what happened, and finally types a fix. Every one of those steps is manual, and every one of them adds minutes to an outage.
The result is a high mean time to detect, a high mean time to resolve, and the slow attrition of alert fatigue. Existing tools do not close the gap: Nagios, Zabbix and Prometheus are excellent at telling you that something is wrong, but they leave the questions of why and what to do about it entirely to you. Full-blown AIOps platforms answer those questions, but they are priced and sized for large enterprises.
AutOps closes that gap for small and medium infrastructure — and it does so without handing an autonomous agent the keys to your production servers.
Each layer is a container, each container has one job, and none of them trust the others by default.
The central nervous system. A self-hosted n8n instance routes scheduled health checks, inbound webhook events and SSH executions between every other component, so there is no bespoke API glue to maintain.
A persistent Python daemon built on discord.py holds the stateful gateway
connection Discord requires and translates chat into stateless webhooks n8n can consume.
Your phone becomes the control terminal.
A LangChain agent driven by Nvidia Nemotron reads stdout and
stderr together, distinguishes a database timeout from a misconfiguration
from an OOM kill, and writes the exact command that fixes it.
Past incidents, remediation outcomes and your own runbooks are embedded as vectors. When something breaks, the agent retrieves what happened last time — so its advice fits your architecture, not a generic one.
Detection cannot execute. Reasoning cannot execute. Only the execution module can, and only after the decision module has recorded your approval.
Starts the pipeline. A cron schedule opens an SSH session, runs docker ps -a,
parses the variable-width output, filters known-ephemeral containers, and raises a formal
event only when a real service is down.
Ingests logs and metrics, builds a situational snapshot, queries the vector store for historical context, and produces a root-cause assessment with a proposed remediation.
Evaluates the proposal against automation policy, writes it to the
pending_actions store, and blocks. Nothing proceeds without a recorded
administrator decision.
Delivers the incident, the diagnosis and the exact proposed command to your Discord channel as a formatted report, then listens for the reply.
Validates the approved command against a whitelist, runs it over SSH with key-based auth, captures both output streams, and confirms the fix actually worked before reporting success.
An application crashes, runs out of memory, or loses its database connection.
The scheduled workflow opens an SSH session, enumerates container state, and finds one
that is no longer Up. Known short-lived containers — a certbot renewal
job, for instance — are whitelisted, so this does not fire on noise.
The system pulls the tail of the failing container's logs, capturing
stdout and stderr together — Docker routes application
logs through the latter, and reading only the former loses the actual error.
The LLM parses the trace, retrieves similar past incidents from the vector store, and settles on a root cause and a specific remediation command.
The proposal is written to the pending-actions store and delivered to Discord under a Proposed Action Required heading. The pipeline stops here. It will wait as long as it has to.
!approve
The command is validated against the whitelist, executed over SSH, and verified from its
own output. You get a confirmation, not an assumption. Replying !deny
discards the proposal instead.
AutOps is not a replacement for a metrics platform. It is the reasoning and remediation layer those platforms leave out.
| Capability | Nagios | Zabbix | Prometheus | AutOps |
|---|---|---|---|---|
| Availability monitoring | Yes | Yes | Yes | Yes |
| AI root-cause analysis | No | No | No | Yes |
| Remediation execution | No | No | No | Approval gated |
| Mobile approval workflow | No | Limited | No | Yes |
| Grounded in your own runbooks | No | No | No | Yes, via RAG |
| Workflow customisation | Config files | Templates | Rules | Visual, in n8n |
Silent when healthy. Custom parsing filters expected short-lived processes, and an upsert strategy means a container that stays down does not generate a fresh alert every cycle. When AutOps speaks, it matters.
Diagnostics capture stdout and stderr simultaneously,
eliminating the blind spot that makes most log automation useless against Docker.
The agent is prompt-constrained against state-changing commands, and the execution path is physically separate from the reasoning path. Safety is not a setting you can accidentally turn off.
No VPN, no terminal, no laptop. A formatted incident report arrives in Discord and a single command resolves it — from a train, a restaurant, or bed.
Because the logic lives in n8n rather than compiled code, adding a check, a channel or an integration is a visual edit — not a fork and a rebuild.
The reference deployment is a single 1 vCPU / 2 GB VPS. Lightweight enough for a homelab, structured enough for a small production estate.
AutOps exists to prove a specific point: that the useful part of AIOps does not require an enterprise budget or a cluster to run on. A lightweight, Docker-based, workflow-driven system can take over the most time-expensive part of server administration — the constant watching, the log archaeology, the 3 a.m. context switch — while leaving every consequential decision with a human.
The roadmap runs from single-host Docker environments toward multi-node estates, with
deduplicated alerting, automatic incident lifecycle cleanup, and on-demand
!status reporting already specified.