The problem we solve

Traditional server management is plagued by operational bottlenecks. When a service goes down, an administrator faces a tedious, multi-step process: connect to a VPN, authenticate over SSH, manually hunt for the failing container, retrieve the error logs, interpret them, and finally deploy a fix.

This manual cycle produces a high mean time to detect (MTTD) and a high mean time to resolve (MTTR), which together mean prolonged downtime. It also produces alert fatigue, which is arguably worse, because a team that has learned to ignore alerts will eventually ignore the one that mattered.

AutOps removes that friction. It brings self-healing automation and intelligent analysis to network engineering, strips away the drudgery of routine maintenance, and lets engineering teams spend their attention on scaling and building instead.

The distinction that matters

AutOps is augmented intelligence, not autonomous intelligence. It is highly autonomous in monitoring and diagnosis, and deliberately powerless in execution. Every state-changing command waits for a human. That constraint is the product.

Core architecture and technology stack

AutOps is built on a microservices-oriented architecture that runs natively inside Docker containerisation. Rather than being a standalone application, it is a set of reusable workflows that integrate with an existing server environment to add operational intelligence to it.

Reference deployment stack and versions
Layer Technology Version Role in the system
Orchestration n8n (self-hosted) 1.121.0 Event-driven routing between triggers, AI, storage and messaging
Containerisation Docker / Docker Compose 29.1.5 / 2.32.0 Service isolation, fault containment, reproducible deployment
Reasoning Nvidia Nemotron via OpenRouter nemotron-3-nano-30b Log parsing, root-cause analysis, remediation synthesis
Contextual memory Pinecone vector database Index: autops-knowledge Semantic retrieval of past incidents and internal documentation
Interface Python + discord.py Python 3.11 Persistent gateway bridge and human-in-the-loop approval channel
State n8n Data Tables / SQLite SQLite 3.51.2 pending_actions lifecycle and incident records
Edge Nginx reverse proxy Alpine TLS termination and secure HTTP routing
Host Ubuntu Server LTS on a VPS 24.04 LTS Reference environment: 1 vCPU, 2 GB RAM

n8n orchestration engine

The central nervous system. n8n handles the complex routing of background health checks, webhook events and SSH command executions, tying disparate microservices into one coherent automated workflow. Because the logic is expressed as workflows rather than compiled code, administrators can extend the system without modifying its core.

Custom Discord integration

A resilient Python daemon connects the server environment directly to an administrator's mobile device. This is not a webhook notifier — it is a bidirectional bridge that turns Discord into a secure, real-time command-and-control terminal. The architecture page explains why a custom daemon was necessary rather than an off-the-shelf integration.

Embedded large language model

The system does not just read logs — it interprets them. The agent analyses standard output and standard error streams together to diagnose the root cause of a container failure, distinguishing a database timeout from a misconfiguration from an out-of-memory kill, and then writes the specific command that addresses it.

Retrieval-augmented generation

Backed by a Pinecone vector database, the agent is grounded in your internal documentation. Past server logs, system documentation and prior remediation outcomes are embedded as high-dimensional vectors; when an anomaly appears, semantic similarity search retrieves the most relevant history and injects it into the prompt. The result is diagnostic advice tailored to your architecture rather than to a generic Linux host.

Key capabilities and features

1. Proactive, zero-fatigue monitoring

AutOps runs background health checks at scheduled intervals, executing and parsing docker ps -a state. Custom parsing logic filters out expected short-lived processes to prevent false positives — a certbot container that starts, renews a certificate and exits is normal, and the system knows it. AutOps stays completely silent while the environment is healthy, which is what makes its alerts worth reading.

2. AI-powered diagnostics

When a failure occurs, AutOps securely fetches the tail of the failing service's diagnostic logs. The agent parses them in seconds and identifies the likely cause, drawing on retrieved context from similar past incidents.

3. Human-in-the-loop execution security

The AI is rigorously restricted from executing state-modifying or destructive actions on its own. Instead it generates a precise remediation proposal: a rich alert containing the failure details and the exact command required. The administrator replies !approve <service> or !deny <service>. Only after explicit human authorisation does the system execute anything.

4. Sub-minute detection and rapid resolution

By reducing the troubleshooting pipeline to one notification and a five-second reply, AutOps cuts response times dramatically. The detection logic itself completes in seconds; the practical MTTD is bounded by the polling interval, which ships at five minutes to conserve resources and can be reduced to one minute where an SLA demands it.

Measurable objectives

The project set numeric targets rather than aspirations, so that success or failure would be assessable. These are the targets and where the prototype landed.

Project objectives and prototype outcomes
Objective Target Outcome
Mean time to detect Under 1 minute Detection logic completes in seconds; interval-bound
Mean time to resolve 50% reduction Exceeded — multi-minute cycle reduced to a 5-second interaction
Root-cause accuracy 80% of recurring faults Sustained above threshold across tested fault scenarios
Diagnostic data collection 100% of incidents Automatic log and metric capture on every detected incident
Administrative control 100% approval gated Enforced architecturally, not by configuration
Server downtime 40% reduction Driven by earlier detection and faster authorised response
Manual effort 60% reduction Log retrieval, analysis and command authoring all automated

Why existing tools were not enough

Nagios, Zabbix and Prometheus are mature, capable and widely deployed. The gap is not in what they measure — it is in what happens after the alert fires.

No AI-assisted recommendation
Established platforms rely on rule-based alerts and predefined thresholds, leaving log analysis and root-cause identification to a human. AutOps adds an LLM-assisted module that produces intelligent analysis and an actionable recommendation.
Limited automation in incident handling
Existing tools focus on detection and alerting, leaving investigation and remediation entirely manual. AutOps automates collection, analysis and the remediation workflow while keeping the administrator in control through approval-based execution.
Poor mobile accessibility
Most monitoring solutions assume you are at a desk. AutOps assumes you are not, and makes full incident response possible from a phone.
Cost and complexity
Advanced AIOps platforms are expensive and designed for large enterprises. AutOps is a lightweight, Docker-based, modular system sized for small and medium deployments.
Limited customisation
Fixed workflows make adaptation hard without redevelopment. Because AutOps is built on n8n, administrators can customise workflows, add integrations and extend functionality without touching the core.

Limitations and constraints

A monitoring tool that oversells itself is a liability. These are the real boundaries of the system as it stands.

Scope of automation

AutOps is a reactive incident-response system, not a full infrastructure lifecycle platform. It is optimised for recovering container services — restarting a stalled service, for instance — not for provisioning servers, configuring networks or orchestrating multi-node clusters. It is not a replacement for Kubernetes or Terraform.

AI accuracy boundaries

The agent reasons over what it can retrieve. It performs well on known and documented failure modes and remains vulnerable to out-of-distribution problems: novel multi-layered failures or genuinely new classes of fault may exceed what it can correctly diagnose. Administrator verification is not optional.

Connectivity dependence

The approval path runs over Discord's API. A local network failure or a Discord outage leaves the administrator without real-time notifications, and high packet loss makes response latency non-deterministic. The five-second benchmark assumes a healthy network.

Administrator availability

Approval-gated execution is a safety property with a cost: if nobody is available to approve, resolution waits. This is the correct trade-off for destructive operations, but it is a trade-off.

The intuition gap

The agent is an effective diagnostic filter. It does not replicate the holistic system intuition of an experienced site reliability engineer facing a novel cascading failure, and it is not designed to.

Credential surface

The system holds SSH keys and API credentials by necessity. Misconfiguration or a compromised credential could permit unauthorised server access. See the security model for the controls that mitigate this.

The team

AutOps was developed as a final year design project for the Bachelor of Science in Information Technology, supervised by Prof. Khizer Butt.

  • Ali Asfar Malik Project team
  • Muhammad Shahzaib Arif Project team
  • Abdullah Project team

The core architecture is open source and available for review, audit and deployment at github.com/MrAsfar79/AutOps.

The vision for the future

AutOps represents the next step for AIOps in environments that were previously priced out of it. Traditional AIOps frameworks are resource-heavy and enterprise-oriented; AutOps demonstrates that the same class of automated, AI-assisted operations can run in a lean context without excessive computational power or multi-level orchestration.

The roadmap runs from single-host Docker environments toward orchestrating multi-node cloud estates. Nearer-term work is already specified: deduplicated alerting so a long-running incident does not re-notify every cycle, automatic lifecycle cleanup of resolved records, interception of AI-suggested commands mid-conversation into the formal approval gate, and an on-demand !status command for real-time health summaries between scheduled checks.

By taking over the repetitive, time-consuming parts of administration, AutOps lets complex architectures run efficiently and securely on autopilot — and lets human engineers stop fighting fires and start building the future.