Human-in-the-loop enforced

AutOps
Autonomous AI
System Administrator

AutOps watches your containers, reads the logs when something breaks, works out why, and writes the fix. Then it waits for you to say yes. Detection in under a minute, resolution in a five-second reply from your phone — and not one state-changing command without explicit approval.

The AutOps control loop

  1. 01 Detect A container stops responding
  2. 02 Diagnose The agent reads the logs and names the cause
  3. 03 Approve Nothing runs until you reply yes
  4. 04 Remediate The fix runs and the service comes back

n8n orchestration Nemotron LLM Pinecone RAG Docker native

Measured results

< 1 min Mean time to detect Detection logic runs in seconds; the polling interval sets the ceiling.
> 50% MTTR reduction Measured against the manual VPN, SSH and log-hunting cycle.
> 80% Diagnostic accuracy Correct root cause and valid remediation syntax on recurring faults.
100% Approval gated Every state-changing action requires an explicit administrator reply.
The problem

Three in the morning, and you are typing docker ps again

Traditional server management runs on a reactive loop that costs a human being their evening.

A monitor flags an issue. An alert fires. Somebody connects to a VPN, authenticates over SSH, hunts for the failing container, pulls the error logs, reads them, works out what happened, and finally types a fix. Every one of those steps is manual, and every one of them adds minutes to an outage.

The result is a high mean time to detect, a high mean time to resolve, and the slow attrition of alert fatigue. Existing tools do not close the gap: Nagios, Zabbix and Prometheus are excellent at telling you that something is wrong, but they leave the questions of why and what to do about it entirely to you. Full-blown AIOps platforms answer those questions, but they are priced and sized for large enterprises.

AutOps closes that gap for small and medium infrastructure — and it does so without handing an autonomous agent the keys to your production servers.

Core architecture

Four systems, one event-driven pipeline

Each layer is a container, each container has one job, and none of them trust the others by default.

Orchestration

n8n Workflow Engine

The central nervous system. A self-hosted n8n instance routes scheduled health checks, inbound webhook events and SSH executions between every other component, so there is no bespoke API glue to maintain.

Interface

Discord Command Bridge

A persistent Python daemon built on discord.py holds the stateful gateway connection Discord requires and translates chat into stateless webhooks n8n can consume. Your phone becomes the control terminal.

Reasoning

Embedded LLM Agent

A LangChain agent driven by Nvidia Nemotron reads stdout and stderr together, distinguishes a database timeout from a misconfiguration from an OOM kill, and writes the exact command that fixes it.

Memory

Pinecone RAG Store

Past incidents, remediation outcomes and your own runbooks are embedded as vectors. When something breaks, the agent retrieves what happened last time — so its advice fits your architecture, not a generic one.

Module topology

Five modules, strictly separated

Detection cannot execute. Reasoning cannot execute. Only the execution module can, and only after the decision module has recorded your approval.

  1. Trigger Module

    Starts the pipeline. A cron schedule opens an SSH session, runs docker ps -a, parses the variable-width output, filters known-ephemeral containers, and raises a formal event only when a real service is down.

  2. AI Analysis Module

    Ingests logs and metrics, builds a situational snapshot, queries the vector store for historical context, and produces a root-cause assessment with a proposed remediation.

  3. Decision-Making Module

    Evaluates the proposal against automation policy, writes it to the pending_actions store, and blocks. Nothing proceeds without a recorded administrator decision.

  4. Alert & Notification Module

    Delivers the incident, the diagnosis and the exact proposed command to your Discord channel as a formatted report, then listens for the reply.

  5. Execution Module

    Validates the approved command against a whitelist, runs it over SSH with key-based auth, captures both output streams, and confirms the fix actually worked before reporting success.

Read the full architecture →

Incident lifecycle

From crash to confirmed fix

  1. T + 0s

    A container exits

    An application crashes, runs out of memory, or loses its database connection.

  2. Next poll

    The health check notices

    The scheduled workflow opens an SSH session, enumerates container state, and finds one that is no longer Up. Known short-lived containers — a certbot renewal job, for instance — are whitelisted, so this does not fire on noise.

  3. + seconds

    Diagnostics are collected

    The system pulls the tail of the failing container's logs, capturing stdout and stderr together — Docker routes application logs through the latter, and reading only the former loses the actual error.

  4. + seconds

    The agent reasons

    The LLM parses the trace, retrieves similar past incidents from the vector store, and settles on a root cause and a specific remediation command.

  5. Gate

    You are asked

    The proposal is written to the pending-actions store and delivered to Discord under a Proposed Action Required heading. The pipeline stops here. It will wait as long as it has to.

  6. + 5s

    You reply !approve

    The command is validated against the whitelist, executed over SSH, and verified from its own output. You get a confirmation, not an assumption. Replying !deny discards the proposal instead.

Where it fits

Against the tools you already run

AutOps is not a replacement for a metrics platform. It is the reasoning and remediation layer those platforms leave out.

Capability comparison with established monitoring tools
Capability Nagios Zabbix Prometheus AutOps
Availability monitoring YesYesYesYes
AI root-cause analysis NoNoNoYes
Remediation execution NoNoNoApproval gated
Mobile approval workflow NoLimitedNoYes
Grounded in your own runbooks NoNoNoYes, via RAG
Workflow customisation Config filesTemplatesRulesVisual, in n8n
Key capabilities

What you actually get

Zero-fatigue monitoring

Silent when healthy. Custom parsing filters expected short-lived processes, and an upsert strategy means a container that stays down does not generate a fresh alert every cycle. When AutOps speaks, it matters.

Log analysis that reads both streams

Diagnostics capture stdout and stderr simultaneously, eliminating the blind spot that makes most log automation useless against Docker.

Approval as an architectural constraint

The agent is prompt-constrained against state-changing commands, and the execution path is physically separate from the reasoning path. Safety is not a setting you can accidentally turn off.

Operations from your pocket

No VPN, no terminal, no laptop. A formatted incident report arrives in Discord and a single command resolves it — from a train, a restaurant, or bed.

Workflows you can actually change

Because the logic lives in n8n rather than compiled code, adding a check, a channel or an integration is a visual edit — not a fork and a rebuild.

Runs on almost nothing

The reference deployment is a single 1 vCPU / 2 GB VPS. Lightweight enough for a homelab, structured enough for a small production estate.

The vision

Stop fighting fires. Start building.

AutOps exists to prove a specific point: that the useful part of AIOps does not require an enterprise budget or a cluster to run on. A lightweight, Docker-based, workflow-driven system can take over the most time-expensive part of server administration — the constant watching, the log archaeology, the 3 a.m. context switch — while leaving every consequential decision with a human.

The roadmap runs from single-host Docker environments toward multi-node estates, with deduplicated alerting, automatic incident lifecycle cleanup, and on-demand !status reporting already specified.