Skip to content

Incident Management & On-call Arena

Incident Management & On-call arenaBuyer checklist

Every requirement we judge incident management & on-call products against, as a ready-to-send RFP checklist — with each item's priority, why it matters, and how the top-ranked products score on it today.

54 requirements · 14 themes · verdicts for 5 products · updated 2026-09-07 · priorities mirror the story weights our scoring uses (methodology)

Procurement report →
Show the markdown export
# Incident Management & On-call — buyer checklist (RFP)

Derived from ProductArena's evidence-graded user-story taxonomy for Incident Management & On-call: 54 judged requirements. Priorities mirror story weights (3 = must-have, 2 = should-have, 1 = nice-to-have).

## Agenticness

- [ ] **[must-have]** Plug MCP servers into this product so it can use their tools
- [ ] **[must-have]** Connect an agent via an official MCP server
- [ ] **[must-have]** Drive the product through a documented public API
- [ ] **[must-have]** Delegate tasks to a built-in AI assistant inside the product
- [ ] **[should-have]** Point an agent at llms.txt or agent-oriented docs
- [ ] **[should-have]** Run the product headlessly / in CI for automation
- [ ] **[should-have]** Use an official CLI
- [ ] **[should-have]** Issue scoped/least-privilege API credentials for an agent
- [ ] **[should-have]** Build against official SDKs
- [ ] **[should-have]** Subscribe to events via webhooks
- [ ] **[should-have]** Get AI-generated insights and suggestions from my data inside the product
- [ ] **[should-have]** Set up automations that run autonomously in the background
- [ ] **[should-have]** Operate the product with natural-language commands
- [ ] **[should-have]** Explore an interactive API reference with runnable examples
- [ ] **[should-have]** Download a machine-readable API spec (OpenAPI or equivalent)
- [ ] **[should-have]** Rely on versioned APIs with a documented deprecation policy
- [ ] **[nice-to-have]** Test against a sandbox environment without touching production data

## Ai incident

- [ ] **[must-have]** My agent can acknowledge, escalate, and resolve incidents end to end through a documented API or MCP connection — no dashboard in the loop
- [ ] **[must-have]** AI writes the incident as it happens — live summaries, drafted updates, and scribed call notes — so responders respond instead of typing
- [ ] **[should-have]** An AI investigator digs into the probable cause — correlating changes, telemetry, and similar past incidents — before a human even asks
- [ ] **[should-have]** AI drafts the postmortem from the incident record — timeline, contributing factors, follow-ups — ready for human review

## Alerting escalation

- [ ] **[must-have]** Escalation policies walk unacknowledged pages through multiple steps — delays, fallback responders, and repeat rounds — until someone acknowledges
- [ ] **[must-have]** Alerts from my monitoring tools are ingested through documented sources and routed to the right team by conditions I define
- [ ] **[should-have]** Duplicate and related alerts are deduplicated and grouped so one incident pages one human, not fifty
- [ ] **[should-have]** Pages reach me over the channels I choose — push, SMS, phone call, and email — with per-channel notification rules

## Analytics reliability

- [ ] **[should-have]** I get reliability analytics — MTTA/MTTR trends, incident load, on-call health — to see whether we are actually improving

## Automation depth

- [ ] **[must-have]** Define rules that trigger actions automatically on events
- [ ] **[should-have]** Perform bulk operations across many items at once
- [ ] **[should-have]** Schedule recurring jobs or workflows
- [ ] **[nice-to-have]** Version, review, and roll back my automations

## Automation runbooks

- [ ] **[should-have]** Runbooks attach to incidents and their steps can trigger automatically — creating channels, assigning tasks, running diagnostics
- [ ] **[should-have]** A condition-based workflow engine automates the toil — updates, reminders, field changes — triggered by incident events

## Incident response

- [ ] **[must-have]** Declare and run an incident from chat — Slack or Teams — with channels, roles, and updates created for me
- [ ] **[should-have]** The incident timeline is captured automatically — alerts, actions, and chat decisions — and I can edit or annotate it afterwards
- [ ] **[nice-to-have]** Internal stakeholders get structured incident updates they can subscribe to, without joining the war room
- [ ] **[nice-to-have]** Incidents carry defined roles (commander, comms lead) and task checklists so response stays coordinated under pressure

## Integrations observability

- [ ] **[must-have]** First-party integrations cover my observability stack — Datadog, Grafana, Prometheus, CloudWatch, Sentry — with documented setup
- [ ] **[should-have]** The platform integrates with the tools around the incident — Jira, Slack, Teams, Zoom, GitHub — so state flows both ways

## Mobile experience

- [ ] **[should-have]** A full mobile app lets me acknowledge, escalate, and resolve from my phone at 3am

## On call scheduling

- [ ] **[must-have]** Build on-call schedules with rotations, layers, time zones, and round-robin coverage that match how my teams actually work
- [ ] **[should-have]** Take an override, swap a shift, or request coverage without an admin rebuilding the schedule
- [ ] **[nice-to-have]** My shifts sync to my personal calendar via a feed so I always know when I'm on the hook

## Openness

- [ ] **[must-have]** Export all of my data in open formats and leave
- [ ] **[must-have]** Self-host the core product
- [ ] **[should-have]** Do everything through the API that I can do in the UI
- [ ] **[should-have]** Read the product's source under an open license

## Postmortems learning

- [ ] **[must-have]** Postmortems follow a real workflow — templates, drafting from the timeline, review, and publication
- [ ] **[should-have]** Follow-up actions from incidents are tracked to completion and sync to our issue tracker

## Privacy posture

- [ ] **[must-have]** Prevent my data from being used to train AI models
- [ ] **[should-have]** Choose where my data is stored (region/residency)
- [ ] **[should-have]** Control data retention and deletion
- [ ] **[should-have]** Opt out of telemetry and usage tracking

## Status communication

- [ ] **[should-have]** Publish a hosted public status page — custom domain, subscriber notifications — driven from incident state
- [ ] **[nice-to-have]** Run private or internal status pages with access control for customer-specific or employee-only audiences

---

Source: https://ultrametric.ai/productarena/arena/incident-management (evidence-graded verdicts for 5 products) · methodology: https://ultrametric.ai/productarena/methodology

Chips show the top 5 ranked products' current verdict on each requirement — ✓ full · ~ partial · ! disputed · — none · n/a not applicable.

Agenticness — how well agents can access and operate the productAgenticness· 17 items

How well agents can access and operate the product

Ai incident — stories about ai incident in this arenaAi incident· 4 items

Stories about ai incident in this arena

  • ai-native userMy agent can acknowledge, escalate, and resolve incidents end to end through a documented API or MCP connection — no dashboard in the loop

    Core requirement — weighs 3× in arena scoring · 1 of 5 products fully deliver this today

    must-have
  • ai-native userAI writes the incident as it happens — live summaries, drafted updates, and scribed call notes — so responders respond instead of typing

    Core requirement — weighs 3× in arena scoring · 2 of 5 products fully deliver this today

    must-have
  • ai-native userAn AI investigator digs into the probable cause — correlating changes, telemetry, and similar past incidents — before a human even asks

    Important, not disqualifying — weighs 2× in arena scoring · 2 of 5 products fully deliver this today

    should-have
  • ai-native userAI drafts the postmortem from the incident record — timeline, contributing factors, follow-ups — ready for human review

    Important, not disqualifying — weighs 2× in arena scoring · 3 of 5 products fully deliver this today

    should-have

Alerting escalation — stories about alerting escalation in this arenaAlerting escalation· 4 items

Stories about alerting escalation in this arena

Analytics reliability — stories about analytics reliability in this arenaAnalytics reliability· 1 item

Stories about analytics reliability in this arena

  • engineering leaderI get reliability analytics — MTTA/MTTR trends, incident load, on-call health — to see whether we are actually improving

    Important, not disqualifying — weighs 2× in arena scoring · 1 of 5 products fully deliver this today

    should-have

Automation depth — how much of the product can run unattendedAutomation depth· 4 items

How much of the product can run unattended

Automation runbooks — stories about automation runbooks in this arenaAutomation runbooks· 2 items

Stories about automation runbooks in this arena

  • sreRunbooks attach to incidents and their steps can trigger automatically — creating channels, assigning tasks, running diagnostics

    Important, not disqualifying — weighs 2× in arena scoring · 1 of 5 products fully deliver this today

    should-have
  • sreA condition-based workflow engine automates the toil — updates, reminders, field changes — triggered by incident events

    Important, not disqualifying — weighs 2× in arena scoring · 3 of 5 products fully deliver this today

    should-have

Incident response — stories about incident response in this arenaIncident response· 4 items

Stories about incident response in this arena

  • on-call engineerDeclare and run an incident from chat — Slack or Teams — with channels, roles, and updates created for me

    Core requirement — weighs 3× in arena scoring · 2 of 5 products fully deliver this today

    must-have
  • sreThe incident timeline is captured automatically — alerts, actions, and chat decisions — and I can edit or annotate it afterwards

    Important, not disqualifying — weighs 2× in arena scoring · 2 of 5 products fully deliver this today

    should-have
  • engineering leaderInternal stakeholders get structured incident updates they can subscribe to, without joining the war room

    Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 5 products fully deliver this today

    nice-to-have
  • engineering leaderIncidents carry defined roles (commander, comms lead) and task checklists so response stays coordinated under pressure

    Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 5 products fully deliver this today

    nice-to-have

Integrations observability — stories about integrations observability in this arenaIntegrations observability· 2 items

Stories about integrations observability in this arena

  • sreFirst-party integrations cover my observability stack — Datadog, Grafana, Prometheus, CloudWatch, Sentry — with documented setup

    Core requirement — weighs 3× in arena scoring · no product fully delivers this yet

    must-have
  • sreThe platform integrates with the tools around the incident — Jira, Slack, Teams, Zoom, GitHub — so state flows both ways

    Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet

    should-have

Mobile experience — stories about mobile experience in this arenaMobile experience· 1 item

Stories about mobile experience in this arena

On call scheduling — stories about on call scheduling in this arenaOn call scheduling· 3 items

Stories about on call scheduling in this arena

Openness — open source, data portability, and self-hosting storiesOpenness· 4 items

Open source, data portability, and self-hosting stories

Postmortems learning — stories about postmortems learning in this arenaPostmortems learning· 2 items

Stories about postmortems learning in this arena

Privacy posture — data-handling and privacy storiesPrivacy posture· 4 items

Data-handling and privacy stories

Status communication — stories about status communication in this arenaStatus communication· 2 items

Stories about status communication in this arena

  • engineering leaderPublish a hosted public status page — custom domain, subscriber notifications — driven from incident state

    Important, not disqualifying — weighs 2× in arena scoring · 1 of 5 products fully deliver this today

    should-have
  • engineering leaderRun private or internal status pages with access control for customer-specific or employee-only audiences

    Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 5 products fully deliver this today

    nice-to-have

Full evidence behind every verdict lives on the arena page and each product page — chips above deep-link straight to the judged story.