Deployment & operations guide
No two infrastructure environments are identical, so AutOps supports two deployment paths: granular self-hosted control for engineers who want it, and a managed instance for teams who would rather not run the orchestration layer themselves. This guide covers both, then explains day-to-day operation.
Before you start
| Requirement | Minimum | Notes |
|---|---|---|
| Host | Linux VPS or VM | Ubuntu Server 24.04 LTS is the tested reference |
| Compute | 1 vCPU / 2 GB RAM | Sufficient for the reference stack; 4 cores / 8 GB recommended for wider estates |
| Runtime | Docker + Docker Compose | Everything runs in containers; nothing is installed on the host directly |
| Network | Public IPv4, 100 Mbps | Needed for messaging integration and hosted inference |
| SSH access | Key-based, to each monitored host | Generate a dedicated key for AutOps rather than reusing a personal one |
| LLM access | An OpenRouter API key | Or any compatible endpoint if you prefer to self-host inference |
| Vector store | A Pinecone index | Holds your runbooks and incident history for retrieval |
| Messaging | A Discord bot token and channel | This becomes your privileged approval interface |
The approval channel is a privileged interface. Anyone who can post in it can authorise a command against your servers. Create a dedicated, restricted channel before you start — not a general team channel you will tidy up later. The security model sets out the rest of your responsibilities as an operator.
Step 1 — Get the source code
Clone the repository
AutOps is developed in the open. The core architecture is free to review, audit and deploy, and the repository contains everything needed to build the system from the ground up: the master deployment configuration, the n8n orchestration templates, the AI system prompts, and the custom messaging bridge.
git clone https://github.com/MrAsfar79/AutOps.git
cd AutOps
Read the workflow templates and the system prompt before you deploy. They define exactly what the agent is permitted to do, and they are the files you are most likely to want to adjust for your environment.
Step 2 — Choose your deployment strategy
Two models are supported. Pick based on how much of the operational surface you want to own.
Option A — Self-hosted
For server administrators, lab engineers and privacy-conscious teams who prefer to keep everything on their own hardware or VPS. Four steps:
-
Prepare the environment. Confirm Docker and Docker Compose are
installed, then populate your environment variables from the provided
.envtemplates — API keys, database credentials and secure tokens. Keep this file out of version control and restrict its permissions. -
Deploy the stack. Run the build. Compose pulls the required
microservices, builds the custom integrations, and establishes isolated internal
networks for the backend components.
cp .env.example .env # edit .env with your credentials, then: docker compose up -d - Access the engine. Once the stack is live, open your local n8n interface and import the AutOps workflow templates.
- Connect your interface. Bind the Discord credentials so the bridge can reach your administrative channel, and send a test message to confirm the round trip.
Option B — Managed cloud and enterprise integration
If you would rather bypass infrastructure setup, server maintenance and backend security entirely, a managed instance provides a turn-key experience.
- Zero-maintenance hosting. Server provisioning, uptime monitoring, TLS certificate renewal and security patching are handled for you.
- Custom development and integrations. Proprietary hardware, legacy internal tools or specific security matrices can be accommodated with custom n8n nodes, tailored diagnostic prompts and bespoke workflows.
- Dedicated support. Priority troubleshooting and architecture guidance directly from the development team.
Ready to discuss it? Get in touch through the contact form with your architectural requirements and we will scope an instance.
Step 3 — System initialisation
First boot
On a self-hosted deployment the first boot is the critical phase. Because every core service — the orchestration engine, the vector store for retrieval memory, and the messaging bridge — runs in its own isolated container, initialisation has real work to do.
During this phase the system will:
- verify its dependencies are present and correctly versioned;
- establish the encrypted internal networks between services;
- initialise the state tables that back the human-in-the-loop approval checks.
Once the terminal reports all services healthy, log in to the n8n dashboard and finalise the connection between the AI agent and your messaging interface. Verify the bridge before you rely on it: post a message in the administrative channel and confirm it reaches the workflow engine.
docker compose ps # every service should read healthy
docker compose logs -f # watch initialisation in real time
Step 4 — Connect your infrastructure
Register what you want watched
AutOps acts as the central nervous system for your network, but it needs to know what it is managing. Map your environment into the system by adding devices and endpoints to the active inventory.
- Standard servers. Input the addresses, SSH keys and target services for your web servers, databases and application backends.
- Specialised hardware. Register dedicated IP nodes, IoT devices or network-attached hardware you need kept under observation.
Once registered, AutOps begins its passive telemetry cycles immediately. It polls the endpoints on the configured schedule, stays silent while everything is healthy, and stands ready to diagnose and propose a fix the moment an anomaly appears.
Tuning the polling interval
The health check ships with a five-minute cron interval. That is a deliberate default: it keeps resource consumption and bandwidth low, at the cost of a detection ceiling of five minutes. The detection logic itself completes in seconds.
If you have an SLA that requires sub-minute detection, reduce the interval to one minute. Testing confirmed the execution pipeline handles that frequency without instability. Choose based on your actual requirement rather than on instinct — a one-minute interval on a large estate multiplies SSH sessions accordingly.
Operating AutOps day to day
Once deployed, almost all interaction happens in your messaging channel. When a service fails you receive a formatted report containing the failure details, the diagnosis and the exact command proposed to fix it.
| Command | Effect | Availability |
|---|---|---|
!approve <service> |
Authorises the pending remediation for that service. The command is whitelist-checked, executed over SSH, and verified from its own output before you get a confirmation. | Current |
!deny <service> |
Rejects the proposal and clears the pending action without touching the server. | Current |
!logs <service> |
Returns recent diagnostic output for a container, capturing both stdout and stderr. |
Current |
| Natural language | Any message without a command prefix is routed to the AI agent, which can investigate using read-only diagnostics and answer questions about the environment. | Current |
!status |
On-demand health summary of all monitored containers, without waiting for the next scheduled check. | Planned |
Read the proposed command, not just the diagnosis. The diagnosis is the agent's reasoning; the command is what will actually run. They are usually consistent — and the one time they are not is exactly the case the approval gate exists for.
If something is not working
- No alerts arriving, but containers are down
- Check the bridge first. Confirm the Python daemon container is running, then verify the webhook endpoint is reachable from it. A bridge that has lost its gateway connection is the most common cause of silence.
- Approvals are not matching pending actions
-
Confirm the sanitisation step is present in your workflow. Dynamic string values can
arrive with a leading
=prepended, which breaks strict-equality lookups against the pending-actions table. The architecture page documents this. - Log commands return empty messages
-
The execution node must capture both output streams. Docker routes application logs
through
stderr, so capturing onlystdoutproduces an empty but apparently successful response. - Long diagnostics are truncated
-
Expected. Discord rejects messages over 2,000 characters, so outbound payloads are
truncated to roughly 1,900 with an ellipsis. Use
!logsfor the fuller picture. - Repeated alerts for the same incident
- A container that stays down is re-detected on each cycle. Upsert ingestion prevents duplicate database rows, and deduplicated notification is on the roadmap; until then, deny or resolve the pending action to stop the cycle.
Need help?
For deployment questions, enterprise implementations or custom integration requirements, get in touch. If you have found a bug, the repository's issue tracker is the fastest route. If you have found a security problem, please read the disclosure guidance first and report it privately.