AIOps 03 — Workflows & Automated Remediation

SME Track: AIOps & Alerting

Presenter: Specialist (DSR) — AIOps

Date: 2026-11-25

SLB × Elastic Workshop Program

Overview

Build and test automated remediation workflows triggered from observability alerts.

Where this applies

These labs run on Observability Serverless — a fully managed project so you can practice without cluster operations.

The same capabilities you explore here — ES|QL, Streams, AI Assistant, Agent Builder, Workflows, and SLOs — are available on Elastic Cloud Hosted (ECH) and self-managed deployments.

Serverless mainly saves operational toil (sizing, ILM, Fleet, upgrades). Your observability skills transfer directly.

Session topics

  • Elastic Workflows for alert-driven automation
  • Connector patterns — Slack, PagerDuty, webhooks
  • Safe remediation guardrails and approval steps

Why these features?

WorkflowsAutomate alert response safely
🔔Observability AlertsSignal without the noise
SLOsUser-facing reliability, not just green dashboards

Use → to see why each feature matters for SLB.

Why Workflows?

Automate alert response safely

Without it

Manual Slack pings, ticket copy-paste, and runbook hunts — alerts fire but nothing moves until a human acts.

With Workflows

Workflows chain connectors (Slack, PagerDuty, webhooks) with approval steps when alerts or SLOs breach.

  • Notify the right channel with context automatically
  • Add human-in-the-loop before remediation scripts run
  • Reduce toil without bypassing change control
Alert
Workflow
Notify / act

Why Observability Alerts?

Signal without the noise

Without it

Alert storms, duplicate pages, and rules that never get tuned — on-call learns to ignore the channel.

With Observability Alerts

Threshold, anomaly, and SLO-based rules with grouping, suppression, and AI-assisted triage in one alerts UI.

  • Tune rules to SLB services instead of one-size-fits-all thresholds
  • Deduplicate and acknowledge with context for handoffs
  • Feed Workflows for automated first response
Rule
Alert
Triage

Why SLOs?

User-facing reliability, not just green dashboards

Without it

CPU graphs look fine while customers see errors — no shared error budget or burn-rate language with product teams.

With SLOs

SLOs define availability/latency targets from real traces and metrics, with burn alerts before users flood support.

  • Align SRE and product on measurable reliability
  • Prioritize fixes when error budget is draining
  • Native in Observability on every deployment — no custom PromQL recording rules required
SLI signal
SLO target
Burn alert

Hands-on lab

Your lab uses Elastic Observability Serverless for a zero-ops learning environment.

The steps and features are the same on ECH and on-prem — follow the assignment panel when Kibana opens.

Instruqt track: slb-sme-aiops-alerting

Resources

  • Registration: events.elastic.co/slbworkshops
  • Repo: github.com/poulsbopete/slb-workshops
  • Use ← → arrow keys to navigate slides