Insights & updates from our experts
PopClub UPI
Download the Case Study1: Critical alerts arrived as Slack messages that had to be spotted rather than answered
2: Pod-wide on-call sent alerts to engineers who did not own the failing service
3: One escalation policy applied to every monitor, regardless of business impact
4: A P0 response time the team could see but could not tighten
1: Severity-based paging, so only critical alerts become an incident on a five-minute P0 window
2: Service-level on-call rotations that page the engineer who owns the service
3: Escalation policies weighted to what an incident can cost the business
4: Slack visibility for the engineering lead, with no pager to carry
How PopClub turned Datadog alerts into incidents that reach the right engineer within seconds
How PopClub turned Datadog alerts into incidents that reach the right engineer within seconds
Good monitoring answers one question: is something broken. It does not answer the second one, which is who should pick this up right now.
PopClub is a rewards UPI payments app and a Razorpay company. Mohan Kumar Pandian leads the engineering team, with backend leads Ankur and Shivam running the pods under him. They have run Datadog for observability for a while, and the monitors are solid. Severity tiers are clean. Alerts fire when they should.
The gap in the stack was never detection. It was what happened to an alert after it fired.
For a year, Datadog sent every alert into Slack as a message across three states: warning, critical, and recovery. The team had full visibility. What a Slack message cannot do is interrupt you. When a payments-down P0 lands in the same channel as a routine warning, it reads like any other line until someone scrolls to it, and on a payments app, the minutes spent scrolling are minutes customers feel.
This is the coordination problem every fast-moving SRE team meets sooner or later. Monitoring gives you a clean signal. Deciding which signals deserve a human, and getting each one to the person who can actually act, is a separate discipline that a monitoring tool was never built to run.
Xurrent IMR became the layer that sorts each alert by impact and routes it to the engineer who owns the fix
PopClub added Xurrent IMR on top of Datadog and used it to make three decisions that Datadog alone was not going to make for them.
Monitors are sorted P0, P1, and P2 by business impact. Everything still flows to Slack for visibility, so nothing is lost. Only a critical alert crosses the threshold into an incident, on a five-minute P0 window. The noise stays in Slack where it belongs, and the alert means something again.
PopClub's pods each own several services, and every service has an owner who holds the context. The team split on-call along those lines. Each owner now has their own rotation and their own escalation policy, so an alert for a payments service reaches the person who built it. The old pattern, where whoever held the pod pager could get an alert for a service they had never touched, acknowledge it, and then hand it to someone with context, is gone.
That tradeoff is worth stating. It works because service ownership at PopClub is real and current. A team without that would need to fix ownership first, or fine-grained rotations turn into upkeep.
A single escalation policy for every monitor is a decision made by not making one. PopClub weighs escalation by impact. Low-priority monitors escalate gently. High-impact ones escalate hard, so if the first responder misses it, the next engineer is minutes away and a serious incident cannot sit unacknowledged while it spreads. Mohan stays out of the rotation and never carries a pager, but he sees every incident through Slack, so he gets oversight without the 3am noise.
The handoff from a Datadog alert to an incident now takes under a minute
Detection is the number PopClub watches most closely, and it is where Xurrent IMR earns its place. Datadog already burns a confirmation window before it fires, three to five minutes on some monitors. Anything the handoff adds lands on top of a clock that is already running, so the team guards it.
Detection and response live in Xurrent IMR
PopClub keeps a clean line between the two tools. Xurrent IMR handles detection and routing. The incident carries just enough signal to make a first call: the affected API groups or service, the latency or error-rate impact, and the service name. From there the engineer opens Datadog and traces it. For anything with a wide blast radius, the team spins up a Slack channel with a Meet link and pulls responders in by hand.
| Metric | Result |
|---|---|
| Datadog alert to incident | Under 1 minute, often seconds |
| P0 acknowledgment (MTTA) | 10 to 15 minutes |
| P0 recovery (MTTR), typical | 30 to 60 minutes |
| Recovery on a known revert | 10 to 15 minutes |
| Average recovery, harder cases | Under 2 hours |
| Detect target the team is driving to | 10 minutes |
| On-call model | Service-level ownership, escalation weighted by impact |
| Stack | Datadog for observability, Xurrent IMR for detection, alerting, and routing |
Before and after, on the path an alert takes
What the team chose to keep after a year of real on-call
These are the features that survived real on-call at PopClub, not the ones in a pitch deck.
This is the one Mohan names first. Who gets paged, which SLA applies, and how escalation runs are all set per service, and that granularity is what made the whole routing model possible.
Slack and Datadog plug in cleanly, and as a lead, Mohan does not need to live on the platform. The work happens on it and reaches him off it.
The team set up deliberately funny page tones, so the sound itself tells the floor an incident is live before anyone reads a word. It used to be one flat tone for everything.
Where PopClub is taking this next
The backend team is heavily on the platform now, at 25 engineers. Front-end is next. PopClub is wiring its RUM signal so that a crash-rate spike or an unresponsive app maps to the right pod and raises a page, the same way backend does today. When that lands, 40 to 45 developers will be on the platform, with another 10 to 15 covered through escalation, for roughly a 60-person presence.
The targets tighten with scale. PopClub is driving MTTD toward 10 minutes for P0 issues and recovery toward sub-13 minutes on known reverts. For a consumer-facing payments app, every minute shaved off detection is a minute fewer customers spend unable to pay.
PopClub did not add monitors or headcount. They changed where alerts go. An incident instead of a message for anything critical. The service owner instead of the pod for anything routed. Escalation sized to what an incident can actually cost the business.
If your monitoring is already solid and your service ownership is clean, that is the place the fast wins are. Bring the observability you already trust. We will handle the part where the right person picks up.


















.webp)












