The SRE agent you can embed anywhere

Relionaut brings 2026-generation SRE agent capabilities — autonomous incident investigation, tool governance, bring-your-own LLM, MCP tools, persistent memory — in a portable kernel you drop into your own production systems. No platform lock-in.

Get started View on GitHub
Apache-2.0 CI green v0.1.0 BYO LLM — Anthropic / OpenAI / Ollama MCP tools
Why Relionaut

Agent capabilities, without the platform

Cloud SRE agents ship impressive capabilities bound to their platform. Open-source investigators bind you to their stack. Relionaut splits the difference: a self-contained agent kernel with the capabilities, and every integration point behind a port you own.

🧩

Embeddable kernel

A single core/ directory — zero npm dependencies, zero upward imports, enforced by a purity gate. Vendor it, implement ports, assemble with createAgent().

🛡️

Governance in the loop

Every tool declares risk; a pure govern() choke point blocks write-risk tools in read-only mode — structurally, not by prompt. Unknown MCP tools default to write (fail-closed).

🧠

Memory that compounds

Investigations distill into lessons; the next incident starts with the relevant history injected into context.

🔌

BYO everything

Any Anthropic- or OpenAI-compatible endpoint, Ollama for air-gapped. Your telemetry via one interface. Your toolchain via MCP.

📜

Auditable by construction

Every run ends in a persisted terminal outcome — answered, step_budget_exhausted, or llm_error. Nothing is dropped from the audit trail.

📦

All-in-one demo

Prefer batteries included? docker compose up starts the full self-hosted loop: collect → alert → investigate → auto-recover, with a web console.

Architecture

The kernel and its five ports

The ReAct investigation loop, governance, memory, and audit live in the kernel. Everything environment-specific sits behind five interfaces — that's the whole integration surface.

core/

The kernel — investigation loop, tool governance, prompt assembly, trajectory replay, memory injection.

  • LlmProviderstreaming chat + tool calls; provider protocols never leak past this port
  • TelemetrySourcelogs & metrics queries; sources describe themselves into the prompt
  • ToolSourcestandard tools, whitelisted shell, or any MCP server
  • RunStorebatched, block-paired trajectory persistence with terminal outcomes
  • MemoryStorelesson record & retrieval

platform/

The first host — the self-hosted all-in-one is just one consumer of the kernel.

  • anthropic providerGLM / Claude-style endpoints
  • PG telemetrybuilt-in collector: docker logs & stats, host metrics
  • docker toolsread-only whitelisted host forensics
  • PG run storeagent_messages with structured blocks
  • alert enginerule evaluation → incidents → auto-investigate → auto-recover
Get started

Two ways to run it

Self-host in one command

Full stack: PostgreSQL, Redis, collector, alert engine, investigation agent, web console.

git clone https://github.com/Joshwong1908/relionaut.git
cd relionaut
echo 'GLM_API_KEY=your-key' >> .env
docker compose up -d --build
# open http://localhost:8082 — telemetry flows in 30s

Embed in your Node 24 service

Vendor the kernel directory, implement the ports you need, run an investigation.

import { createAgent, openaiProvider,
         staticTools, telemetryTools,
         inMemoryRunStore } from './agent-core/mod.ts';

const agent = createAgent({
  llm: openaiProvider({ baseUrl: 'http://llm.internal:11434/v1',
                        model: 'qwen3:32b' }),
  tools: [staticTools(...telemetryTools(myTelemetry))],
  store: inMemoryRunStore(),
  system: { /* role, environment, output format */ },
});
for await (const ev of agent.run({ runId, input })) { /* SSE events */ }
Position

Where it sits today

Open-source SRE agents cluster on an autonomy spectrum. Relionaut ships as a read-only investigator by default — the red line is structural — with the interface reserved for approved remediation as trust grows.

L1 · ExplainerDeterministic analyzers, LLM explains findings.
L2 · Relionaut v0.1Read-only investigator: ReAct loop, dynamic tool choice, read-only by design.
L3 · Suggests fixesInterface reserved (advise mode).
L4 · Approved remediationInterface reserved (approve-write + Approval).