FluidifyAI
AI SRE Live

Regen

Agentic on-call management for alert routing, escalation paths, and adaptive learning that drafts the post-mortem

Get started free
Triage
Escalate
Slack
Timeline
Post-mortem

The problem

Alert storms bury the signal, and the post-incident paperwork lands on the same engineer who just lost a night to it.

How Regen solves it

  • Routes alerts on regex, literal or source, so only what deserves a page reaches the on-call engineer.
  • Escalation paths with primary/secondary schedules, overrides, leave plans and multi-timezone support.
  • Communication tools like Slack or MS Teams stay in live sync with the incident timeline.
  • Drafts editable post-mortems, summaries and handoff digests automatically.
Learn more
regen.fluidify.ai / incidents / INC-7451
Live
CRITICALINC-7451

Checkout Service Cascade

checkout-api · payments-api

Auto-triaged

Root cause found

Missing index on payment_intents.webhook_id — pool exhausted, 8,400 users blocked.

Evidence

DatadogKubernetesGitHubSlack

Response

1

Circuit breaker enabled, pool scaled 200 → 800

2

Rollback verified — error rate 68% → 2.1%

Triggered11:21 AM
Acknowledged11:23 AM
Resolved11:33 AM
Before and after

What changes once Regen is in the loop.

Before Regen

    20+

    Pages per incident

    $4.2k

    Budget burned monthly

    1 week

    Post-mortem turnaround

    Spreadsheet

    Escalation policies

    After paging

    Root cause found

    After Regen

      80%

      Faster ack (MTTA)

      $50k

      Saved annually

      Same day

      Post-mortem turnaround

      Automated

      Escalation policies

      Before paging

      Root cause found

      How it works

      Inside Regen's reasoning

      1

      Alert fires

      Webhook arrives from Prometheus, Datadog, Grafana, PagerDuty, or any monitoring tool. Regen ingests it and opens an incident record instantly.

      2

      Neuri - Adaptive RCA Engine

      The Neuri queries your observability stack through MCP — pod logs, recent deploys, error traces, dashboards — and surfaces a probable root cause before anyone is woken up.

      3

      Smart escalation

      The right on-call engineer is paged via Slack, SMS, or voice. Escalation policies define who gets notified and when if there's no acknowledgement.

      4

      Reflex - Autoheal Runbooks

      Shadow mode: the agent proposes runbook steps, you approve each one. Every action is logged to the incident timeline with full audit trail.

      5

      Auto post-mortem

      Once resolved, Regen drafts a post-mortem from the incident timeline — what happened, when, who acted, and what the AI found. Publish with one click.

      What's inside Regen

      Features and Capabilities

      On-call schedules & rotations

      Layer-based schedule model, shift types, override support. Import from Grafana OnCall in one step.

      Escalation policies

      Multi-step escalation paths with configurable timeouts. Pages the right person, every time.

      Slack & Teams native

      Incidents live in your existing channels. Dedicated incident channels, bot commands, status updates.

      Neuri - Adaptive Root Cause Engine

      Calls your observability stack through MCP: Datadog, Kubernetes, GitHub, before your phone rings.

      Incident timeline

      Immutable, agent-and-human timeline for every incident. Full audit trail from alert to resolution.

      AI BYOK

      Bring your own OpenAI or Anthropic key. AI features run under your credentials, zero incident data leaves your infrastructure.

      Post mortem & handoff digest

      Timeline auto-drafts the post-mortem the moment an incident closes. Handoff digest catches the incoming on-call up in seconds.

      Slack pre-investigation

      Before paging anyone, the agent scans channel history and past incident timelines for correlated failures, surfacing patterns humans miss.

      Accuracy benchmarks

      How accurate is Regen

      80%

      Accurate

      Integrated with the production tools.

      Human in the loop

      Every RCA and runbook reasoning undergoes Quality check with manual approval before it enters the library.

      Steering capability

      Edit and redirect the reasoning and tool-usage steps, then re-execute to sharpen the result.

      Guardrails

      Hard limits on reasoning loops, tool-use depth and iterations, token usage and LLM response bounds.

      Redaction

      PII and Sensitive information masking. Strict data governance preventing data leakage to LLM.

      Evaluation benchmarks

      Accuracy gates with confidence scores and validation benchmarks.

      Shadow mode

      Simulate against the live system with no write provision. Runbooks and playbooks move to live only once approved.

      Integrations

      Connects to your entire stack

      Slack

      Slack

      Comms

      Microsoft Teams

      Microsoft Teams

      Comms

      Prometheus

      Prometheus

      Alerting

      Grafana

      Grafana

      Alerting

      Alertmanager

      Alertmanager

      Alerting

      Datadog

      Datadog

      Observability

      PagerDuty

      PagerDuty

      Migration

      Opsgenie

      Opsgenie

      Migration

      GitHub

      GitHub

      Engineering

      GitLab

      GitLab

      Engineering

      Kubernetes

      Kubernetes

      Infrastructure

      Webhook

      Custom

      Security & Reliability

      Built to be trusted with production

      Bring Your Own Key (BYOK)

      Bring your own LLM key, with multi-vendor multi-model support.

      Data Leakage Prevention

      PII and Sensitive information masking with strict data governance layers.

      Multiple Deployment Options

      On-prem, Air-gapped and Private Cloud, Managed Cloud SaaS Deployments, configured and validated by a dedicated Forward Deployment Engineer.

      Access Control and Auditing

      RBAC, MFA, Custom SSO/SCIM provisioning and Audit trails

      Compliant on

      SOC 2GDPRDPDPAEU AI ActHIPAA

      Deploys to

      Self-hosted VPCOn-premPrivate cloudAir-gappedCloud SaaS
      AI SRE Suite

      Where does it sit in the Suite?

      Regen owns detecting and communicating the incident in the AI SRE Suite — alert routing, on-call escalation, and the structured incident record every other module builds on. It hands the timeline to Neuri for root cause, Reflex for the fix, and keeps Gills' answers grounded in what actually happened.

      Self-hostable
      Managed Cloud

      Regen is available both as open-source and as a managed cloud service.For open-source deployment, run it on your Kubernetes cluster, a Docker Compose stack, or spin it up locally in seconds. Your data stays on your infrastructure.

      • Docker Compose quick start up in under 5 minutes
      • Helm chart for Kubernetes with TLS + ingress
      • Published image on ghcr.io/fluidifyai/regen
      • No vendor lock-in your infra, your data

      Don't want to self-host?

      We can manage Regen for you - fully hosted, maintained, and updated. See managed plans →

      Terminal
      git clone https://github.com/FluidifyAI/Regen
      cd Regen
      cp .env.example .env
      # Edit .env with your values, then:
      docker compose up -d
      # Open https://github.com/FluidifyAI/Regen
      FAQ

      Frequently asked questions

      Yes. Regen is AGPLv3 licensed and available at github.com/FluidifyAI/Regen for self-hosting, with a Docker Compose quick start up in under 5 minutes. A fully managed cloud option is also available if you don't want to self-host.

      Newsletter

      Product updates & reliability notes, straight to your inbox.