Documentation

Run it, or embed it

Everything you need for the all-in-one stack or the kernel. For design rationale and acceptance criteria, see the portable-agent-core spec in the repository.

Quickstart — self-hosted

Three steps to a live stack

01

Clone and configure

The stack runs without a key (collection, queries, alerting); the agent needs one.

git clone https://github.com/Joshwong1908/relionaut.git && cd relionaut
cat >> .env <<'EOF'
GLM_API_KEY=your-key            # any Anthropic-compatible endpoint
GLM_BASE_URL=https://open.bigmodel.cn/api/anthropic
GLM_MODEL=glm-4.6
EOF
02

Start the stack

docker compose up -d --build

Open http://localhost:8082. Within 30 seconds the dashboard shows live telemetry; the alert engine evaluates rules every 60s, opens incidents, and — with a key configured — runs an automatic investigation on each new incident.

03

Talk to the agent

Open any incident and chat: the agent investigates with read-only tools (query_logs, query_metrics, whitelisted host commands, HTTP probes) and streams its reasoning. After a run, distill it into a lesson with POST /api/incidents/:id/distill.

Embedding

Embedding the kernel

The kernel is a single directory — api/src/core/ — with zero npm dependencies and zero imports outside itself (enforced by a purity gate in CI). Vendor it into any Node 24 project (native TypeScript, no build step).

// 1. copy api/src/core → your project (e.g. ./agent-core)
// 2. implement the ports you need
const myTelemetry: TelemetrySource = {
  describe: () => ({ logsAvailable: false, metricsAvailable: true,
                     metricNames: ['checkout_error_rate'] }),
  queryMetrics: async (q) => /* call your time-series store */,
};

// 3. assemble and run
const agent = createAgent({
  llm: openaiProvider({ baseUrl, apiKey, model }),
  tools: [staticTools(...telemetryTools(myTelemetry))],
  store: inMemoryRunStore(),          // or your own RunStore
  system: { role, environment, dataSources, outputFormat },
  maxSteps: 12,
});
for await (const ev of agent.run({ runId, input, mode: 'read-only' })) { … }

Ports you don't implement simply don't exist: a telemetry source without queryLogs produces no query_logs tool — there is no "configured but unavailable" state, and the system prompt reflects reality automatically.

Governance

Governance & run modes

Every tool carries a risk tag (single source of truth) and a provenance string. A pure function decides each call inside the loop — the red line cannot be overridden by policy.

Run moderead-risk toolswrite-risk tools
read-only (default)executealways blocked — no policy exception
adviseexecuteblocked; the agent is told to suggest instead
approve-writeexecuteneeds a matching Approval for the call id (L4 direction, host UI not shipped)

MCP tools without an explicit risk mapping are treated as write-risk — fail-closed. Map trusted tools via riskFor to make them available in read-only runs.

Audit

Audit trail

Every run ends in exactly one persisted terminal outcome. Exhausted the step budget? The store records step_budget_exhausted. The LLM endpoint failed mid-investigation? It records llm_error. Trajectories are stored as structured, block-paired turns — tool calls and their results replay in the exact shape the model saw.

// agent_messages row kinds (Postgres host)
text | tool_use | tool_result          // legacy-compatible audit columns
run_outcome                          // terminal: { outcome, summary, findings }
blocks JSONB                         // structured ChatMessage per turn
Reference

Environment variables

VariableDefaultPurpose
GLM_API_KEY—LLM key (required for diagnosis only)
GLM_BASE_URLbigmodel Anthropic-compatibleAny Anthropic-compatible endpoint
GLM_MODELglm-4.6Model id
MAX_STEPS10Max ReAct steps per investigation
RATE_LIMIT60/60API rate limit (count/window seconds)
ALERT_INTERVAL60Rule evaluation period (seconds)
AUTO_DIAGNOSE1Auto-investigate on alert
Security

Security notes

Read-only by design: host commands pass a whitelist and a dangerous-character gate; the agent never mutates state and only suggests changes. Before exposing the platform to the public internet, read the auth-tunnel spec — token auth is mandatory, unauthenticated endpoints return 404.

Resources

Project resources

Development

TypeScript strict throughout, Node 24 native TS (zero build), Apache-2.0. CI runs typecheck, the kernel purity gate, unit tests, web build and docker build on every push. Implementation plan & shards.