As financial systems accelerate and modernize, the real breakthrough isn't whether AI can help, it's how fast you can trust autonomous agents with real operational decisions. Recently, Agentic AI has become a powerful force inside Mastercard's observability and reliability strategy, driving faster adaptation and transforming incident response. Federico deep dives on how to make agents work in the real world across fragmented data, strict regulation, and ultra low latency environments.
Payments don't pause while 90 people pile onto a bridge to work out who owns a failure. Traditional L1 triage is an expensive bottleneck which consumes valuable engineering capacity where operations and engineers piece together what happened, who is affected, and who needs to respond. Meanwhile, customers are feeling the impact.
We're using AI to change that: turning observability data, operational knowledge, and business context into actionable incident intelligence before the bridge even starts.
This isn't about replacing the people who know how to fix incidents. It's about making sure we only need to call them when they're actually needed.
Standardisation accelerates teams when it removes friction - but over-centralisation creates bottlenecks. The goal is consistent telemetry and governance without preventing teams from shipping.
Running observability in house isn't the easy choice — it's the bold one. Explore how Marta's team built an entire observability platform using open-source projects, battling scalability challenges, data complexity, team structure, costs, and tough security rules head on. With OTEL and the LGTM stack as their backbone, they're proving that owning your stack unlocks flexibility, accountability, and real engineering power.
Dunelm operates 200+ stores, a fast growing e commerce site, and high traffic mobile apps, so diagnosing issues across systems was anything but straightforward. With inconsistent instrumentation, patchy visibility, and rising observability costs, something had to change. Dan Herd shares how Dunelm tackled the chaos by embracing OpenTelemetry end to end and building a custom service that transforms CDN logs into spans, giving the business a clean, connected view from front end to back end and how adopting a true product mindset helped win over engineering teams.
Observability has become mission critical at Nokia, where cloud native teams, strict regional controls and decades old legacy systems all need to operate with the same level of reliability and security. Drawing on more than 20 years of hands on experience, Frank Adu has been instrumental in building a unified observability approach that weaves together reliability, security monitoring and AI assisted analysis. Nokia employing several strategies such as automation, cutting alert noise with correlation intelligence, and using in house LLMs to help teams identify real issues faster across a regulated global environment.