Insights & updates from our experts
Over the years there have been a bunch of great talks on site reliability and incident response. Below are a few we thought stood out(in no specific order) and is defintely worth a peek.
SREcon19 Europe/Middle East/Africa - All of Our ML Ideas Are Bad (and We Should Feel Bad) by Todd Underwood(Google)
SREcon18 Americas - The History of Fire Escapes by Tanya Reilly(Squarespace)
Who Destroyed Three Mile Island? - Nickolas Means | The Lead Developer Austin 2018 by Nick Means(Muve Health)
Incidents as we Imagine Them Versus How They Actually Are with John Allspaw
LISA19 - What Connections Can Teach Us about Postmortems by Chastity Blackwell(Truss)
LISA19 - Earthquakes, Forest Fires, and Your Next Production Incident by Alex Hidalgo(Squarespace)
SREcon18 Europe - SRE for Good: Engineering Intersections between Operations and Social Activism by Liz Fong-Jones(Honeycomb)
SREcon18 Europe - Ethics in Computing by Theo Schlossnagle(Circonus)
The SRE I aspire to be SRECon19 EU by Yaniv Aknin(Google)
02 Jul 2020

An AI SRE that knows your incidents
Most AI SREs are pattern matchers trained on public data. They know what a memory leak looks like in the abstract. They don't know that your payments-api has a flaky liveness probe everyone ignores, that the checkout team owns the retry policy, or that the last three "database incidents" were actually cache misconfigurations. That knowledge lives in your postmortems, your Slack channels, and the heads of two senior engineers.

How Long Should ITSM Implementation Really Take in 2026?
Most vendors will tell you ITSM implementation takes six months to a year — but modern, configuration-first platforms have rewritten the math entirely. See what real implementations look like in 2026, and why a long rollout is now a choice, not a given.















.webp)
.webp)














