Product perspective

AutOps is designed to be dropped into an existing server environment, not to replace it. It interacts with six things: the monitored servers themselves, whatever monitoring tools are already collecting metrics, the n8n workflow engine at the centre, the AI analysis module, the messaging platform used for approvals, and the administrators who make the decisions.

Four constraints shaped every decision that follows. The system must express its logic as n8n workflows, which bounds what integrations are possible. It depends on external monitoring tools, AI services and messaging platforms, each of which can impose API limits or go down. It must remain portable and well-documented, because it is open source. And every automated action must pass through a human, which rules out the simpler fully autonomous designs.

System context

Before any of the internals matter, it is worth being precise about what AutOps is and is not responsible for. Five components sit inside the system boundary. Everything else — the servers being watched, the messaging platform, the inference API and the vector store — is a dependency the project consumes but does not own, which is exactly where its failure modes and its API limits come from.

  • Automated flow
  • Human approval gate
  • Verified outcome
  • Rejected or failed
  • Outside the system boundary
AutOps system context and scope boundary A dashed boundary encloses the five components AutOps owns: the n8n orchestration engine, the AI analysis module, the decision and policy engine, the execution module and the n8n data tables holding pending actions. Outside the boundary on the left sit the monitoring tools already in place and the monitored Docker servers. Outside on the right sit the administrators, the Discord platform, the hosted Nemotron inference API and the hosted Pinecone vector store. Arrows cross the boundary for metrics in, SSH both ways to the servers, HTTP both ways to Discord, inference calls out to Nemotron and Pinecone, and an amber approve-or-deny decision from the administrators. AUTOPS SYSTEM BOUNDARY n8n Orchestration Engine WORKFLOWS UNDER VERSION CONTROL AI Analysis Module LANGCHAIN AGENT / READ-ONLY TOOLS Decision & Policy Engine POLICY CHECK / WRITES PENDING_ACTIONS Execution Module WHITELIST / SSH / VERIFY n8n Data Tables PENDING_ACTIONS / DURABLE STATE HUMAN APPROVAL REQUIRED WRITE / READ BACK Monitoring tools EXISTING - METRICS IN Monitored servers DOCKER HOSTS - SSH BOTH WAYS Administrators DECISION MAKERS Discord platform GATEWAY + REST - HTTP Nemotron LLM API HOSTED INFERENCE Pinecone HOSTED VECTOR STORE APPROVE / DENY
Figure 1 — System context. Everything inside the dashed boundary is AutOps; everything outside it is a dependency the project does not own. Amber marks the only path a human decision takes.

The five-module pipeline

Responsibilities are separated so that no single module can both decide and act. Detection cannot execute. Reasoning cannot execute. Only the execution module touches the server, and only when the decision module has recorded an approval.

The AutOps incident pipeline A vertical pipeline of six stages. Monitored servers are polled by the trigger module, which raises an event to the AI analysis module. The AI module also draws on the Pinecone vector store for historical context. Its proposal passes to the decision-making module, which writes it to the pending-actions store and stops at the approval gate, where an administrator replies in Discord. Only on approval does the execution module run the command over SSH, and its verified result feeds back to the monitored servers. Monitored Servers DOCKER SERVICES / IP NODES / ENDPOINTS Trigger Module CRON · SSH POLL · PARSE · NOISE FILTER Pinecone RAG CONTEXT AI Analysis Module STDOUT + STDERR → ROOT CAUSE → PROPOSAL Decision-Making Module POLICY CHECK · WRITE PENDING_ACTIONS Approval Gate — Human DISCORD · !APPROVE / !DENY · PIPELINE BLOCKS Execution Module WHITELIST · SSH · CAPTURE · VERIFY CLOSED-LOOP VERIFICATION
Figure 2 — The incident pipeline. Amber marks the approval gate: nothing crosses it without a human.

1. Trigger module

Initiates the workflow, either from a scheduled cron node or a manual invocation. It opens a secure SSH session, runs docker ps -a, and reconstructs the complete state of every container including ones that have crashed or exited.

Because Docker's output uses variable-width whitespace, a JavaScript node tokenises it with a /\s{2,}/ split to extract container names and states reliably. State-aware filtering then whitelists known-ephemeral containers — the certbot renewal container, for example, whose exited state is normal — so they never raise false positives. If every remaining container is Up, the workflow terminates silently.

2. AI analysis module

A LangChain agent node gives the model genuine tool access rather than static chat behaviour. On receiving a trigger, a data-ingestion component gathers relevant data from the state store, the monitoring layer and the host; a context builder assembles it into a situational snapshot; and the reasoning engine determines what happened.

The agent's toolset lets it:

  • query the Pinecone vector store for historical documentation and past remediations;
  • issue HTTP checks to confirm whether a service is actually reachable;
  • run a restricted set of read-only diagnostic commands, such as tailing logs or checking disk usage.

It cannot run anything that changes state. That restriction is enforced through a structured system prompt and reinforced by the module boundary — the analysis module has no path to the execution node.

3. Decision-making module

Receives the proposal and routes it through policy evaluation. A policy manager decides whether an action may proceed automatically or requires human intervention; for state-changing commands the answer is always the latter. The proposal is persisted and the workflow halts pending an administrator's explicit decision.

4. Alert and notification module

Formats the incident, the diagnosis and the exact proposed command into a rich report and delivers it to the configured administrative channel. The agent is programmed to emit a strict template, placing the remediation command under a Proposed Action Required heading followed by the explicit !approve / !deny prompt, so the message is machine-parseable as well as human-readable.

5. Execution module

Receives the approved action, validates its safety, and runs it. The command is loaded into an SSH execution node that authenticates to the host with cryptographic keys.

Why the split is enforced, not just documented

The five modules fall into three privilege zones, and the boundaries between them are the whole safety argument. A module in the observe zone has no state-changing tool available to it. A module in the decide zone can write a record but has no connector to the execution node. Only the act zone can reach the host, and it will only run a record that already carries an approval. An attacker who fully controls the model’s output still lands in zone two.

Separation of duties across three privilege zones Three dashed zones side by side. Zone one, observe, contains the trigger module and the AI analysis module and is annotated as having no state-changing tool. Zone two, decide, contains the decision and policy engine and the notification module, and only writes a pending-actions record. Zone three, act, contains the execution module and its verification step and is the sole writer to the host. A cyan arrow carries a proposal from zone one to zone two; an amber arrow carries an approved decision from zone two to zone three. A closing note records that no connector runs from zone one or two directly to the host. 1 - OBSERVE (READ ONLY) Trigger module CRON / SSH POLL AI analysis module LANGCHAIN AGENT no state-changing tool READS LOGS / RAG / HTTP CHECKS 2 - DECIDE (NO EXECUTION PATH) Decision & policy POLICY EVALUATION Notification module FORMATS THE REPORT writes a record only PENDING_ACTIONS ROW 3 - ACT (SOLE WRITER TO THE HOST) Execution module WHITELIST / SSH KEYS Verification STDOUT + STDERR CHECKED runs approved records only VALIDATED AGAINST STATE PROPOSAL APPROVED No connector runs from zone 1 or zone 2 to the host DETECTION CANNOT EXECUTE - REASONING CANNOT EXECUTE
Figure 3 — Separation of duties. The privilege to act exists in exactly one zone, and the only arrow into it is amber.

Closed-loop verification

Execution is not fire-and-forget. The module captures the complete output of the command — both stdout and stderr — and evaluates it to confirm the command ran without syntax errors or runtime failures before reporting back. What you receive is a verified result, not an assumption that the command probably worked.

Why both streams matter

Docker routes containerised application logs through stderr, not stdout. An early version of the !logs command captured only stdout and delivered empty messages while appearing to succeed. Reading both streams is what makes log diagnostics work at all — and it materially improved the AI's accuracy, because the model was finally seeing the actual error text.

Incident lifecycle end to end

The module view says who does what; it does not show the one property that shapes the rest of the design. Detection is continuous and machine-paced, while authorisation is human-paced and arrives whenever somebody reads their phone. The sequence below makes that gap explicit.

Incident lifecycle as a sequence, including the asynchronous approval gap Five lifelines: the monitored server, the n8n workflow, the AI agent, the pending_actions table and the administrator on Discord. The workflow polls the server over SSH, receives container states including exited ones, tokenises and filters them, and upserts an incident record without creating a duplicate row. It asks the AI agent to diagnose, the agent retrieves RAG context and runs read-only probes, and returns a root cause with a proposed command. The workflow stores the proposal as pending approval and sends the report to the administrator. A highlighted band across all five lifelines marks the point where the workflow halts entirely, holding nothing in memory, for a wait that may last minutes or hours. When the administrator replies approve, the workflow revalidates that an unresolved record still exists, runs the whitelisted command over SSH, reads both output streams and returns a verified result. Monitored server DOCKER HOST n8n workflow ORCHESTRATOR AI agent LANGCHAIN + RAG pending_actions DURABLE STATE Administrator VIA DISCORD cron - ssh docker ps -a states, incl. exited tokenise - filter ephemeral upsert - no duplicate row diagnose(logs) RAG retrieve - read-only probes cause + proposed command state = pending_approval report + !approve / !deny WORKFLOW HALTS - NOTHING HELD IN MEMORY THE WAIT MAY BE MINUTES OR HOURS; STATE LIVES ONLY IN PENDING_ACTIONS !approve website validate - unresolved record? whitelisted command - SSH stdout + stderr verified result
Figure 4 — The incident lifecycle over time. The band in the middle is the design constraint everything else follows from: because the wait is unbounded, the workflow cannot hold state in memory, which is why pending_actions exists at all.

Everything awkward about the implementation follows from the band in the middle of that diagram. Because the workflow cannot stay resident while it waits, the incident has to be written down somewhere durable, the approval has to be matched back to it later, and the match has to be revalidated at the moment of execution rather than trusted from when the alert was sent.

State management

Continuous background telemetry and asynchronous human authorisation are decoupled, so the system needs durable state to connect a detection to the approval that arrives minutes or hours later. This is held in n8n's native data tables rather than an external database, avoiding an extra dependency.

Schema

A pending_actions table tracks the lifecycle of each detected anomaly with explicit fields: the target container's identifier, the AI-formulated remediation command, the current alert state, and a detection timestamp. When an approval arrives over Discord, the execution module queries this table to retrieve the exact validated parameters.

Idempotent ingestion

A five-minute polling cycle against a container that stays down for two hours would naively produce twenty-four duplicate incident records. Ingestion therefore uses an upsert rather than an insert: if an unresolved entry already exists for that container, only the detection timestamp is refreshed; a new row is appended only when no prior record exists. This keeps the table clean, keeps queries fast, and prevents duplicated alerts at the source.

Lifecycle of a pending_actions record A state machine with eight states. A detected row is stored and becomes pending approval, where the workflow halts. Re-detection while pending only refreshes the timestamp and never creates a new row. A human approval moves it to approved, then to executing once the whitelist validates it. Executing resolves either to verified on a zero exit code or to failed on a non-zero exit or stderr output. A denial moves the record straight to denied. Verified, failed and denied all drain into a terminal closed state. A note records that an approval naming a target with no unresolved row is rejected, which is what stops a stale or repeated approval from firing. detected ROW CREATED pending_approval WORKFLOW HALTED approved !APPROVE PARSED executing SSH RUN verified EXIT 0 failed NON-ZERO / STDERR denied !DENY PARSED closed TERMINAL STORED HUMAN VALID RE-DETECTED: TIMESTAMP REFRESHED, NO NEW ROW EXIT 0 ERROR !DENY An !approve for a target with no unresolved row is rejected here THIS IS WHAT STOPS A STALE OR REPEATED APPROVAL FROM FIRING
Figure 5 — The pending_actions lifecycle. The loop at the top is the idempotent upsert: a container down for two hours produces one row, not twenty-four.

Command routing

A Command Router node classifies inbound messages before any inference happens, which is both a cost and a latency optimisation.

Semantic routing
Natural-language input — something like “analyse the high memory utilisation on the web container” — is forwarded to the AI agent for full processing.
Prefix bypass
A payload beginning with a known operational prefix such as !approve, !deny or !logs skips the inference engine entirely. Direct administrative commands do not need an LLM, and routing them around it cuts both API cost and response latency.
State validation
Before anything reaches the execution module it passes a validation gateway that confirms an outstanding, unresolved incident actually exists for the targeted container. This is what prevents a stale or duplicated approval from executing a command that is no longer relevant.
Command routing and the validation gateway An inbound Discord message reaches the Command Router node, which classifies it before any inference happens. Three paths leave the router. A message with no prefix goes to the AI agent for semantic routing and full inference with RAG context, the highest cost and latency path. An approve or deny prefix goes to state validation. A logs or status prefix goes to a read-only diagnostic that bypasses the language model entirely and spends no tokens. On the validation path a decision asks whether an unresolved record exists for the target: if it does the execution module runs the whitelisted command over SSH, and if it does not the request is rejected, logged and the operator notified. Inbound Discord message Command Router node CLASSIFIES BEFORE ANY INFERENCE AI agent NO PREFIX / SEMANTIC State validation !APPROVE / !DENY Read-only diagnostic !LOGS / !STATUS Full inference + RAG HIGHEST COST / LATENCY Bypasses the LLM NO TOKEN SPEND unresolved record for this target? YES NO Execution module WHITELIST / SSH Reject & log OPERATOR NOTIFIED
Figure 6 — Command routing. Two of the three paths never reach the language model, which is where both the cost and the latency saving come from.

The Discord bridge, and why it is custom

The first prototype used a community-built n8n Discord node. It failed for an architectural reason rather than a bug: the node maintains a stateful, long-lived WebSocket connection to Discord's gateway, while n8n executes on a stateless, trigger-based model. The connections experienced silent timeouts and dropped payloads without raising errors — the worst possible failure mode for a monitoring system, because it fails quietly.

The fix was a middleware daemon. A containerised Python service built on discord.py runs continuously alongside the workflow engine. It natively holds the stateful gateway connection Discord expects, subscribes to the administrative channel, and converts inbound chat events into structured JSON delivered to n8n as ordinary HTTP webhooks. Outbound replies go the other way, through an HTTP request node against Discord's REST API.

The result is a protocol translation layer: stateful where Discord requires it, stateless where n8n requires it, with neither side compromised.

The Discord bridge as a protocol translation layer On the left, inside the external Discord platform, sit the persistent WebSocket gateway and the request-reply REST API. On the right, inside the AutOps n8n engine, sit a stateless webhook trigger and an outbound HTTP request node. Straddling the two sits the containerised Python bridge daemon built on discord.py, annotated as taking a stateful socket in and emitting a stateless webhook out. Gateway events travel both ways over the persistent socket to the daemon, which posts them to the webhook trigger over HTTP. Outbound replies go directly from the HTTP request node to the Discord REST API, bypassing the daemon. A note records the rejected first attempt, in which a community node held the WebSocket connection inside n8n itself and timed out silently. DISCORD PLATFORM (EXTERNAL) AUTOPS - n8n ENGINE Gateway PERSISTENT WSS REST API REQUEST / REPLY Python bridge daemon DISCORD.PY / CONTAINERISED STATEFUL SOCKET IN, STATELESS WEBHOOK OUT Webhook trigger STATELESS HTTP Request node OUTBOUND WSS EVENTS HTTP POST HTTPS - OUTBOUND REPLY BYPASSES THE DAEMON Rejected first attempt: a community node held the WSS socket inside n8n IT TIMED OUT SILENTLY - THE WORST FAILURE MODE FOR A MONITOR
Figure 7 — Protocol translation. The daemon exists because Discord requires a stateful connection and n8n executes statelessly; it is the only component that spans both models.

Deployment topology

The entire ecosystem is described by a single docker-compose.yml acting as a declarative blueprint — infrastructure as code rather than a sequence of manual installation steps. It defines the network topology, persistent volumes, environment variables and inter-service dependencies needed to instantiate the system.

Bringing the stack up provisions, in the correct order:

  • the n8n orchestration engine;
  • an Nginx reverse proxy for TLS termination and secure HTTP routing;
  • the custom Python Discord daemon;
  • the vector store for retrieval-augmented context;
  • a controlled suite of sample web applications and REST APIs, which act as monitoring targets so the system's own detection and remediation can be validated empirically.

Every service is configured with restart: always. After a host failure, a service crash or a scheduled VPS reboot, the Docker daemon restores the whole stack without manual intervention — which matters more for a monitoring tool than for most software, since a monitor that dies silently is worse than no monitor at all.

Deployment topology on a single VPS A dashed boundary marks a single Ubuntu 24.04 VPS running one docker-compose stack with restart always on every service. Four containers sit in a row: nginx terminating TLS, n8n running the engine and the agent, the Python discord-bridge daemon, and a set of sample applications acting as monitored targets. All four attach to an internal bridge network named autops_net, which also carries SSH and HTTP health checks. Three named volumes hang below the network: n8n_data holding workflows and the pending_actions table, letsencrypt holding TLS certificates, and bridge_logs holding daemon output. Crossing into the boundary from outside are inbound TLS traffic on port 443 to nginx, calls out to the hosted Nemotron and Pinecone services from n8n, and two-way traffic between the Discord platform and the bridge daemon. Internet TLS 443 Hosted AI NEMOTRON + PINECONE Discord platform GATEWAY + REST SINGLE VPS - UBUNTU 24.04 LTS - docker-compose.yml - restart: always nginx TLS TERMINATION n8n ENGINE + AGENT discord-bridge PYTHON DAEMON sample-apps MONITORED TARGETS autops_net - BRIDGE NETWORK SSH + HTTP HEALTH CHECKS n8n_data WORKFLOWS / PENDING_ACTIONS letsencrypt TLS CERTIFICATES bridge_logs DAEMON OUTPUT NAMED VOLUMES SURVIVE A REBUILD
Figure 8 — Deployment topology. One Compose file describes all of it, which is what makes the stack reproducible and what brings it back after a reboot.

Resource requirements

Reference deployment versus recommended production sizing
Resource Reference prototype Recommended
CPU 1 vCPU 4 cores
Memory 2 GB RAM 8 GB RAM
Storage VPS default 50 GB
Network Public IPv4 100 Mbps minimum, 1 Gbps preferred
Operating system Ubuntu 24.04 LTS Ubuntu Server LTS
GPU None — inference is hosted Optional, only for self-hosted heavy models

The prototype's 1 vCPU / 2 GB allocation was determined empirically to be sufficient for the orchestration engine, the messaging bridge and the diagnostic test applications without hitting bottlenecks. Inference runs on a hosted API, which is why the compute footprint stays this small.

Design decisions and trade-offs

Every choice below had a cheaper or simpler alternative. They are recorded with what was given up, because a decision with no cost attached to it usually means the alternative was never seriously examined.

Architectural decisions, the alternative rejected, and what each choice costs
Decision Alternative considered Why, and what it costs
Logic lives in n8n workflows A standalone application in Python or Go Auditable, forkable and modifiable without recompiling anything. The cost is that the set of possible integrations is bounded by what n8n nodes exist.
A human approves every state change Fully autonomous remediation A wrong diagnosis can never become a wrong action. The cost is real latency and the requirement that somebody is reachable.
A custom Python Discord daemon The community n8n Discord node The node held a stateful WebSocket inside a stateless engine and timed out silently. The cost is one more container to run and maintain.
n8n native data tables for state An external Postgres or Redis instance Avoids a dependency and a second backup story for what is one small table. It would not hold up under high write volume.
Hosted inference A self-hosted open-weight model Keeps the footprint at 1 vCPU and 2 GB with no GPU. The cost is an external API dependency, a per-token bill, and sending log text off the host.
Upsert on ingestion An insert for every detection One row per unresolved incident instead of one per poll. It requires a well-defined uniqueness rule per target to be correct.
A whitelist in the execution module Trusting the model’s proposed command Bounds the blast radius even if the model is fully compromised. Every legitimate new command needs an explicit addition.
Prefix bypass for direct commands Routing every message through the LLM Removes both cost and latency from the most frequent operations. It adds a second code path that has to stay in step with the first.

Engineering problems worth documenting

Three issues consumed disproportionate effort and are recorded here because each has a non-obvious cause.

Protocol mismatch

Stateful gateway versus stateless engine

Community Discord nodes held WebSocket connections that silently timed out under n8n's stateless execution model. Resolved by introducing the persistent Python daemon described above, which converts gateway events into webhooks.

Data corruption

A phantom equals sign

Approvals stopped matching pending records. The cause was a parsing quirk in the data table module, which prepended = to dynamic string variables — storing =website instead of website and making every strict-equality lookup return null. A sanitisation node now strips leading equals signs and whitespace before any query.

I/O streams

Logs that arrived empty

The !logs command reported success and delivered nothing, because the execution node captured only stdout while Docker routes application logs through stderr. Rewritten to capture, concatenate and return both streams.

API limits

The 2,000-character ceiling

Long diagnostic payloads exceeded Discord's message limit and returned HTTP 400. Outbound messages are now intercepted, truncated to roughly 1,900 characters and marked with an ellipsis so delivery always succeeds.

How it was validated

Testing followed a bottom-up strategy: verify each data source before verifying the consumers that depend on it. All of it ran on production-grade VPS hardware, with a dedicated Docker Compose test profile provisioning isolated ephemeral containers that shadow live service configurations without port or network conflicts.

  • Bridge integration. Payload serialisation and field mapping between the Python daemon and the n8n webhook endpoint, plus authorisation header correctness against Discord's REST API.
  • Concurrency. Ten simultaneous workflow executions against the local SQLite store, checking that all ten records commit without lock errors, corruption or duplicates.
  • Retrieval accuracy. Confirming the agent generates embeddings from incoming queries and successfully fetches relevant context from the vector index before producing a diagnosis.
  • Session continuity. Verifying conversational memory binds to the correct session identifier so the agent does not lose track of which incident is under discussion.
  • Negative security testing. Input fuzzing with malicious payloads to confirm the execution whitelist blocks and logs them. This is covered in detail on the security page.
  • Failure resilience. Unreachable-service scenarios, confirming SSH timeouts and connection refusals are caught rather than crashing the pipeline, and that the fallback notification still reaches the operator.

Traceability

Each requirement maps to the component that realises it and to the test that demonstrates it. Where a test found a defect, the defect is named rather than smoothed over.

Requirement, the component that satisfies it, and the evidence it works
Requirement Realised by Evidence
Detect a container failure in under a minute Trigger module — cron plus docker ps -a over SSH Failure-resilience runs against isolated ephemeral containers; detection logic completes in seconds, with the polling interval setting the ceiling
Do not alert on a normal exit State-aware filtering with an ephemeral-container whitelist The certbot renewal container, whose exited state is expected, raises no incident
Ground the diagnosis in this deployment, not a generic one Pinecone retrieval inside the AI agent’s toolset Retrieval-accuracy test: embeddings generated from the inbound query and matching context returned before any diagnosis
Diagnose from the actual error text Execution and log capture across both output streams !logs regression: the empty-message defect closed once stderr was read, and accuracy improved measurably
Never change state without a human Decision module, approval gate, and the three privilege zones in Figure 3 Negative security testing with malicious payloads; the whitelist blocks and logs every one
Match an approval to the right open incident pending_actions plus the validation gateway in Figure 6 Concurrency test: ten simultaneous executions commit without lock errors, corruption or duplicates
Keep discussing the same incident across messages Session-bound conversational memory Session-continuity test: memory binds to the correct session identifier
Survive a reboot or a crash unattended restart: always on every Compose service Host reboot and service-crash scenarios; the stack returns without manual intervention
Always reach the operator, even on a bad payload Outbound truncation to roughly 1,900 characters, plus a fallback notification Long-payload test: the HTTP 400 from Discord’s 2,000-character ceiling no longer occurs; SSH timeouts and refusals still deliver an alert
Run on modest hardware Hosted inference, so no GPU and no local model weights Reference deployment on 1 vCPU and 2 GB serves the engine, the bridge and the test applications without bottlenecking

The workflows, system prompts and Compose configuration are all in the public repository. If you want to read the actual implementation rather than a description of it, start there.