Regen
Agentic on-call management for alert routing, escalation paths, and adaptive learning that drafts the post-mortem
The problem
Alert storms bury the signal, and the post-incident paperwork lands on the same engineer who just lost a night to it.
How Regen solves it
- Routes alerts on regex, literal or source, so only what deserves a page reaches the on-call engineer.
- Escalation paths with primary/secondary schedules, overrides, leave plans and multi-timezone support.
- Communication tools like Slack or MS Teams stay in live sync with the incident timeline.
- Drafts editable post-mortems, summaries and handoff digests automatically.
Checkout Service Cascade
checkout-api · payments-api
Root cause found
Missing index on payment_intents.webhook_id — pool exhausted, 8,400 users blocked.
Evidence
Response
Circuit breaker enabled, pool scaled 200 → 800
Rollback verified — error rate 68% → 2.1%
What changes once Regen is in the loop.
Before Regen
20+
Pages per incident
$4.2k
Budget burned monthly
1 week
Post-mortem turnaround
Spreadsheet
Escalation policies
After paging
Root cause found
After Regen
80%
Faster ack (MTTA)
$50k
Saved annually
Same day
Post-mortem turnaround
Automated
Escalation policies
Before paging
Root cause found
Inside Regen's reasoning
Alert fires
Webhook arrives from Prometheus, Datadog, Grafana, PagerDuty, or any monitoring tool. Regen ingests it and opens an incident record instantly.
Neuri - Adaptive RCA Engine
The Neuri queries your observability stack through MCP — pod logs, recent deploys, error traces, dashboards — and surfaces a probable root cause before anyone is woken up.
Smart escalation
The right on-call engineer is paged via Slack, SMS, or voice. Escalation policies define who gets notified and when if there's no acknowledgement.
Reflex - Autoheal Runbooks
Shadow mode: the agent proposes runbook steps, you approve each one. Every action is logged to the incident timeline with full audit trail.
Auto post-mortem
Once resolved, Regen drafts a post-mortem from the incident timeline — what happened, when, who acted, and what the AI found. Publish with one click.
Features and Capabilities
On-call schedules & rotations
Layer-based schedule model, shift types, override support. Import from Grafana OnCall in one step.
Escalation policies
Multi-step escalation paths with configurable timeouts. Pages the right person, every time.
Slack & Teams native
Incidents live in your existing channels. Dedicated incident channels, bot commands, status updates.
Neuri - Adaptive Root Cause Engine
Calls your observability stack through MCP: Datadog, Kubernetes, GitHub, before your phone rings.
Incident timeline
Immutable, agent-and-human timeline for every incident. Full audit trail from alert to resolution.
AI BYOK
Bring your own OpenAI or Anthropic key. AI features run under your credentials, zero incident data leaves your infrastructure.
Post mortem & handoff digest
Timeline auto-drafts the post-mortem the moment an incident closes. Handoff digest catches the incoming on-call up in seconds.
Slack pre-investigation
Before paging anyone, the agent scans channel history and past incident timelines for correlated failures, surfacing patterns humans miss.
How accurate is Regen
80%
Accurate
Integrated with the production tools.
Human in the loop
Every RCA and runbook reasoning undergoes Quality check with manual approval before it enters the library.
Steering capability
Edit and redirect the reasoning and tool-usage steps, then re-execute to sharpen the result.
Guardrails
Hard limits on reasoning loops, tool-use depth and iterations, token usage and LLM response bounds.
Redaction
PII and Sensitive information masking. Strict data governance preventing data leakage to LLM.
Evaluation benchmarks
Accuracy gates with confidence scores and validation benchmarks.
Shadow mode
Simulate against the live system with no write provision. Runbooks and playbooks move to live only once approved.
Connects to your entire stack
Slack
Comms
Microsoft Teams
Comms
Prometheus
Alerting
Grafana
Alerting
Alertmanager
Alerting
Datadog
Observability
PagerDuty
Migration
Opsgenie
Migration
GitHub
Engineering
GitLab
Engineering
Kubernetes
Infrastructure
Webhook
Custom
Built to be trusted with production
Bring Your Own Key (BYOK)
Bring your own LLM key, with multi-vendor multi-model support.
Data Leakage Prevention
PII and Sensitive information masking with strict data governance layers.
Multiple Deployment Options
On-prem, Air-gapped and Private Cloud, Managed Cloud SaaS Deployments, configured and validated by a dedicated Forward Deployment Engineer.
Access Control and Auditing
RBAC, MFA, Custom SSO/SCIM provisioning and Audit trails
Compliant on
Deploys to
Where does it sit in the Suite?
Regen owns detecting and communicating the incident in the AI SRE Suite — alert routing, on-call escalation, and the structured incident record every other module builds on. It hands the timeline to Neuri for root cause, Reflex for the fix, and keeps Gills' answers grounded in what actually happened.
Incident coordination
Regen is available both as open-source and as a managed cloud service.
For open-source deployment, run it on your Kubernetes cluster, a Docker Compose stack, or spin it up locally in seconds. Your data stays on your infrastructure.
- Docker Compose quick start up in under 5 minutes
- Helm chart for Kubernetes with TLS + ingress
- Published image on ghcr.io/fluidifyai/regen
- No vendor lock-in your infra, your data
Don't want to self-host?
We can manage Regen for you - fully hosted, maintained, and updated. See managed plans →
git clone https://github.com/FluidifyAI/Regen cd Regen cp .env.example .env # Edit .env with your values, then: docker compose up -d # Open https://github.com/FluidifyAI/Regen
Frequently asked questions
Yes. Regen is AGPLv3 licensed and available at github.com/FluidifyAI/Regen for self-hosting, with a Docker Compose quick start up in under 5 minutes. A fully managed cloud option is also available if you don't want to self-host.
Product updates & reliability notes, straight to your inbox.