The problem we solve
Traditional server management is plagued by operational bottlenecks. When a service goes
down, an administrator faces a tedious, multi-step process: connect to a VPN, authenticate
over SSH, manually hunt for the failing container, retrieve the error logs, interpret them,
and finally deploy a fix.
This manual cycle produces a high mean time to detect (MTTD) and a high mean time to resolve
(MTTR), which together mean prolonged downtime. It also produces alert fatigue, which is
arguably worse, because a team that has learned to ignore alerts will eventually ignore the
one that mattered.
AutOps removes that friction. It brings self-healing automation and intelligent analysis to
network engineering, strips away the drudgery of routine maintenance, and lets engineering
teams spend their attention on scaling and building instead.
The distinction that matters
AutOps is augmented intelligence, not autonomous intelligence. It is
highly autonomous in monitoring and diagnosis, and deliberately powerless in execution.
Every state-changing command waits for a human. That constraint is the product.
Core architecture and technology stack
AutOps is built on a microservices-oriented architecture that runs natively inside Docker
containerisation. Rather than being a standalone application, it is a set of reusable
workflows that integrate with an existing server environment to add operational
intelligence to it.
Reference deployment stack and versions
| Layer |
Technology |
Version |
Role in the system |
| Orchestration |
n8n (self-hosted) |
1.121.0 |
Event-driven routing between triggers, AI, storage and messaging |
| Containerisation |
Docker / Docker Compose |
29.1.5 / 2.32.0 |
Service isolation, fault containment, reproducible deployment |
| Reasoning |
Nvidia Nemotron via OpenRouter |
nemotron-3-nano-30b |
Log parsing, root-cause analysis, remediation synthesis |
| Contextual memory |
Pinecone vector database |
Index: autops-knowledge |
Semantic retrieval of past incidents and internal documentation |
| Interface |
Python + discord.py |
Python 3.11 |
Persistent gateway bridge and human-in-the-loop approval channel |
| State |
n8n Data Tables / SQLite |
SQLite 3.51.2 |
pending_actions lifecycle and incident records |
| Edge |
Nginx reverse proxy |
Alpine |
TLS termination and secure HTTP routing |
| Host |
Ubuntu Server LTS on a VPS |
24.04 LTS |
Reference environment: 1 vCPU, 2 GB RAM |
n8n orchestration engine
The central nervous system. n8n handles the complex routing of background health checks,
webhook events and SSH command executions, tying disparate microservices into one coherent
automated workflow. Because the logic is expressed as workflows rather than compiled code,
administrators can extend the system without modifying its core.
Custom Discord integration
A resilient Python daemon connects the server environment directly to an administrator's
mobile device. This is not a webhook notifier — it is a bidirectional bridge that
turns Discord into a secure, real-time command-and-control terminal. The
architecture page explains why a custom daemon was
necessary rather than an off-the-shelf integration.
Embedded large language model
The system does not just read logs — it interprets them. The agent analyses standard
output and standard error streams together to diagnose the root cause of a container
failure, distinguishing a database timeout from a misconfiguration from an out-of-memory
kill, and then writes the specific command that addresses it.
Retrieval-augmented generation
Backed by a Pinecone vector database, the agent is grounded in your internal documentation.
Past server logs, system documentation and prior remediation outcomes are embedded as
high-dimensional vectors; when an anomaly appears, semantic similarity search retrieves the
most relevant history and injects it into the prompt. The result is diagnostic advice
tailored to your architecture rather than to a generic Linux host.
Key capabilities and features
1. Proactive, zero-fatigue monitoring
AutOps runs background health checks at scheduled intervals, executing and parsing
docker ps -a state. Custom parsing logic filters out expected short-lived
processes to prevent false positives — a certbot container that starts, renews a
certificate and exits is normal, and the system knows it. AutOps stays completely silent
while the environment is healthy, which is what makes its alerts worth reading.
2. AI-powered diagnostics
When a failure occurs, AutOps securely fetches the tail of the failing service's diagnostic
logs. The agent parses them in seconds and identifies the likely cause, drawing on retrieved
context from similar past incidents.
3. Human-in-the-loop execution security
The AI is rigorously restricted from executing state-modifying or destructive actions on its
own. Instead it generates a precise remediation proposal: a rich alert
containing the failure details and the exact command required. The administrator replies
!approve <service> or !deny <service>. Only after
explicit human authorisation does the system execute anything.
4. Sub-minute detection and rapid resolution
By reducing the troubleshooting pipeline to one notification and a five-second reply, AutOps
cuts response times dramatically. The detection logic itself completes in seconds; the
practical MTTD is bounded by the polling interval, which ships at five minutes to conserve
resources and can be reduced to one minute where an SLA demands it.
Measurable objectives
The project set numeric targets rather than aspirations, so that success or failure would be
assessable. These are the targets and where the prototype landed.
Project objectives and prototype outcomes
| Objective |
Target |
Outcome |
| Mean time to detect |
Under 1 minute |
Detection logic completes in seconds; interval-bound |
| Mean time to resolve |
50% reduction |
Exceeded — multi-minute cycle reduced to a 5-second interaction |
| Root-cause accuracy |
80% of recurring faults |
Sustained above threshold across tested fault scenarios |
| Diagnostic data collection |
100% of incidents |
Automatic log and metric capture on every detected incident |
| Administrative control |
100% approval gated |
Enforced architecturally, not by configuration |
| Server downtime |
40% reduction |
Driven by earlier detection and faster authorised response |
| Manual effort |
60% reduction |
Log retrieval, analysis and command authoring all automated |
Why existing tools were not enough
Nagios, Zabbix and Prometheus are mature, capable and widely deployed. The gap is not in
what they measure — it is in what happens after the alert fires.
- No AI-assisted recommendation
-
Established platforms rely on rule-based alerts and predefined thresholds, leaving log
analysis and root-cause identification to a human. AutOps adds an LLM-assisted module
that produces intelligent analysis and an actionable recommendation.
- Limited automation in incident handling
-
Existing tools focus on detection and alerting, leaving investigation and remediation
entirely manual. AutOps automates collection, analysis and the remediation workflow while
keeping the administrator in control through approval-based execution.
- Poor mobile accessibility
-
Most monitoring solutions assume you are at a desk. AutOps assumes you are not, and makes
full incident response possible from a phone.
- Cost and complexity
-
Advanced AIOps platforms are expensive and designed for large enterprises. AutOps is a
lightweight, Docker-based, modular system sized for small and medium deployments.
- Limited customisation
-
Fixed workflows make adaptation hard without redevelopment. Because AutOps is built on
n8n, administrators can customise workflows, add integrations and extend functionality
without touching the core.
Limitations and constraints
A monitoring tool that oversells itself is a liability. These are the real boundaries of the
system as it stands.
Scope of automation
AutOps is a reactive incident-response system, not a full infrastructure lifecycle
platform. It is optimised for recovering container services — restarting a stalled
service, for instance — not for provisioning servers, configuring networks or
orchestrating multi-node clusters. It is not a replacement for Kubernetes or Terraform.
AI accuracy boundaries
The agent reasons over what it can retrieve. It performs well on known and documented
failure modes and remains vulnerable to out-of-distribution problems: novel
multi-layered failures or genuinely new classes of fault may exceed what it can
correctly diagnose. Administrator verification is not optional.
Connectivity dependence
The approval path runs over Discord's API. A local network failure or a Discord outage
leaves the administrator without real-time notifications, and high packet loss makes
response latency non-deterministic. The five-second benchmark assumes a healthy network.
Administrator availability
Approval-gated execution is a safety property with a cost: if nobody is available to
approve, resolution waits. This is the correct trade-off for destructive operations, but
it is a trade-off.
The intuition gap
The agent is an effective diagnostic filter. It does not replicate the holistic system
intuition of an experienced site reliability engineer facing a novel cascading failure,
and it is not designed to.
Credential surface
The system holds SSH keys and API credentials by necessity. Misconfiguration or a
compromised credential could permit unauthorised server access. See the
security model for the controls that mitigate this.
The team
AutOps was developed as a final year design project for the Bachelor of Science in
Information Technology, supervised by Prof. Khizer Butt.
-
AA
Ali Asfar Malik
Project team
-
MS
Muhammad Shahzaib Arif
Project team
-
A
Abdullah
Project team
The core architecture is open source and available for review, audit and deployment at
github.com/MrAsfar79/AutOps.
The vision for the future
AutOps represents the next step for AIOps in environments that were previously priced out
of it. Traditional AIOps frameworks are resource-heavy and enterprise-oriented; AutOps
demonstrates that the same class of automated, AI-assisted operations can run in a lean
context without excessive computational power or multi-level orchestration.
The roadmap runs from single-host Docker environments toward orchestrating multi-node
cloud estates. Nearer-term work is already specified: deduplicated alerting so a
long-running incident does not re-notify every cycle, automatic lifecycle cleanup of
resolved records, interception of AI-suggested commands mid-conversation into the formal
approval gate, and an on-demand !status command for real-time health
summaries between scheduled checks.
By taking over the repetitive, time-consuming parts of administration, AutOps lets complex
architectures run efficiently and securely on autopilot — and lets human engineers
stop fighting fires and start building the future.