When Nobody’s Problem Becomes Everybody’s Problem

How to Run a Performance Leadership Forum That Actually Changes Behavior

Most application performance forums have a meeting problem.

Not a shortage of meetings. The opposite: a surplus of them, filled with the wrong people, reviewing the wrong information, producing the wrong outputs. Status updates get mistaken for governance. Attendance gets mistaken for accountability. Dashboards get presented, heads nod, and everyone returns to their desks having made no decision that will change what ships next quarter.

I have built performance governance structures in environments where the consequences of getting this wrong were not abstract. In financial services, a missed regression or an unacknowledged escalation does not affect a single product team. It ripples across every institution sitting downstream. In large-scale healthcare platforms, performance failures during peak periods do not just degrade a user experience metric. They disrupt access for real patients at real moments of need. Those environments teach you very quickly which parts of a governance structure are load-bearing and which parts are decoration.

What follows is what I have learned about the difference.


Why Most Performance Forums Fail Before They Start

The failure mode is almost always the same, and it has nothing to do with the technology stack or the observability tooling. It begins with a category error: treating a performance forum as a reporting venue rather than a decision venue.

A reporting venue exists to surface information. A decision venue exists to convert information into action. Both involve presenting data. Only one produces changed behavior. When organizations set up performance forums without being explicit about which one they are building, they almost always build the former while intending the latter.

The Southwest Airlines holiday meltdown in December 2022 is a clean illustration of what that gap costs. Southwest canceled nearly 17,000 flights over ten days, stranding more than two million passengers. The technology failure, a crew scheduling system that could not handle the volume of disruptions triggered by a winter storm, was real. But the organizational failure had a much longer history. The Southwest Airlines Pilots Association president testified before Congress that pilots had been “sounding the alarm bells for over a decade.” Internal audits had flagged catastrophic risk in the legacy scheduling infrastructure as early as 2018. The CEO acknowledged publicly that executives had “talked a lot” about modernizing but had not acted. What Southwest had was not a shortage of information about the risk. What it lacked was a governance structure that could convert that information into a funded, time-bound decision before the system failed under load. The signals existed for years. The decision venue did not.

That is an extreme case. But the underlying dynamic, information flowing toward an organization that has no mechanism to act on it, is not extreme at all. It is the default state of most cross-team performance programs.


Who Belongs in the Room

The first structural decision is attendance, and it is more consequential than it looks.

Performance forums fail in two directions on this dimension. The first is too broad: everyone with a stake in performance is invited, the meeting grows to fifteen or twenty people, and the social dynamics of a large group suppress the candid exchange that governance requires. No one escalates a concern when escalating it means doing so in front of peers from eight different teams. The second failure is too narrow: only individual contributors attend, the meeting produces good technical conversation, and none of it reaches the people with authority to reprioritize work or resolve cross-team conflicts.

The right attendance model for a performance leadership forum is small, senior, and cross-functional. Small means six to ten people. Senior means engineering leaders with genuine decision authority, not representatives who need to check with someone before committing. Cross-functional means the full set of teams that materially influence performance outcomes: client platforms, infrastructure, data science or analytics, and the program management function that owns cross-team coordination. The goal is a room where every person present can either make a decision or directly inform someone who can.

Working groups and technical forums can run parallel to this at the practitioner level. They serve a different function: they develop the findings that the leadership forum acts on. Conflating those two layers is where many programs lose their effectiveness. When engineering practitioners are asked to be both the analysts and the decision-makers in the same meeting, one role will crowd out the other. Usually analysis wins and decisions wait.


What Gets Reviewed, and How

The agenda design of a performance forum is where governance either builds credibility or loses it.

The credibility-destroying pattern is the dashboard walk-through. Someone shares a screen full of metrics, the group scans for red indicators, and the meeting ends when the red items have been noted. Nothing about this structure forces a decision. Nothing distinguishes a metric that is improving because the underlying system improved from a metric that is improving because the measurement changed. Nothing connects a technical finding to a business consequence that would make an executive understand why it warrants reprioritization.

The Slack incident of February 22, 2022 illustrates what an agenda built around current-state metrics misses entirely. A new cache manager, introduced to improve performance, worked exactly as designed. What it also did was expose a pre-existing scatter query that hit every database shard on cache misses, a pattern rarely exercised under normal conditions because cache hit rates had always been nearly universal. When the new component altered those conditions, the query pattern triggered a metastable failure: a cascading degraded state the system could not exit without external intervention. Slack’s own post-mortem noted that the incident commander was herself locked out of Slack during the outage. The component had been introduced without cross-team visibility into how it would interact with existing data access patterns under stress. The agenda that would have caught this is not a review of current health indicators. It is a forward risk review that asks: what are we shipping in the next cycle, what does it assume about the systems around it, and has anyone in this room stress-tested those assumptions together?

An effective performance forum agenda has three components and only three.

The first is a performance narrative, not a metric dump. Before any numbers are presented, someone with end-to-end context should be able to state, in two or three sentences, what the performance story of the past period was for a real user. Which user journeys improved, which degraded, and why. This anchors the subsequent discussion in user reality rather than technical abstraction.

The second is a decision register. Every open item from prior meetings should be reviewed not for status but for resolution. Either a decision was made and the outcome is visible, or the decision is still open and needs to be made in this meeting. Items that remain unresolved across multiple meetings are a signal that either the forum lacks authority or the forum is not being used correctly. Both problems are solvable, but only if someone names them.

The third is a forward risk review. What performance risks are visible in the next release cycle that the group should make a call on now, before they become incidents? This is the component most forums omit entirely, and it is the most valuable. Reactive performance governance, reviewing what already happened, is necessary but insufficient. The forum that changes behavior is the one that catches the regression before it ships.


How Decisions Get Made

A forum that surfaces information but defers all decisions is not governance. It is a briefing.

The decision framework for a performance leadership forum needs to be explicit before the first meeting, not improvised in the room. There are three categories of decision the forum needs to handle.

The first category is decisions that can be made by the forum itself: reprioritization within an existing roadmap, escalation of a risk to a named owner, adjustment of a performance threshold or guardrail metric. These decisions should be made in the meeting and documented with a single owner and a date.

The second category is decisions that require executive visibility: conflicts between performance investment and product delivery timelines that exceed what engineering leaders can resolve independently, resource allocation questions that affect multiple organizations, changes to a release standard that affect how teams ship. These decisions should be framed, not deferred. The forum’s job is to produce a clear recommendation with the supporting evidence, identify who needs to make the call, and set a deadline for that decision. The forum should not dissolve into open-ended discussion about whether a decision is needed.

The third category is decisions that reveal a structural gap: situations where no one in the room has authority to resolve a conflict, and no escalation path exists to someone who does. When this happens repeatedly, it is a signal that the governance model has a missing layer. Either the forum is not senior enough, or the organization has not assigned end-to-end performance accountability at all.

The Ticketmaster failure during the Taylor Swift Eras Tour presale in November 2022 illustrates what the third category looks like from the outside. Taylor Swift’s team had been explicitly assured the platform could handle demand. When the presale opened, the queue system collapsed, users were ejected mid-purchase, and 15% of all site interactions failed. Live Nation’s president later testified before the Senate Judiciary Committee, deflecting accountability across every party involved: Ticketmaster does not set ticket prices, does not determine how many tickets go on sale, and venues set the fees. Every accountability touchpoint belonged to someone else. What that testimony described, whether intentionally or not, was an organization whose performance governance had no named owner for the end-to-end user experience, and therefore no forum with both the authority and the obligation to answer the question: are we actually ready for this?

A decision framework that assigns named owners, documents open risks against those owners, and requires resolution rather than acknowledgment closes the gap that made that failure possible. The named owner is not a formality. It is what gives the forum the authority to demand an answer rather than accept a report.


How to Stop Meeting Drift

Every performance forum is at risk of drifting over time. Initial meetings have energy because the program is new and the problems are acute. Six months in, the problems that were easy to fix have been fixed, the meeting has found its rhythm, and the rhythm has started to substitute for the outcomes.

The signal that a forum has lost its purpose is consistent: the quality of the conversation is high and the rate of decisions is low. People are engaged. Nothing is changing. The meeting has become a place where performance is discussed rather than a place where the organization acts on it.

The Google Cloud outage of June 12, 2025 offers a useful lens on this dynamic. An automated quota update to Google’s API management system corrupted IAM functionality globally, taking down Spotify, Snapchat, Cloudflare, and GitLab simultaneously for nearly three hours. One post-incident analysis observed plainly: “When things break, it’s not always clear who’s accountable, or even what’s broken.” Google was also publicly criticized for a 30-minute lag before their status page reflected the scope of the incident, leaving downstream engineering teams without a signal to act on. The compounding problem was not the failure itself. It was that the dependencies between Google’s internal services and the platforms built on top of them were not visible in any shared governance layer until they became the blast radius of a production incident. A forum that was reviewing shared dependency changes proactively, before deployment, with the teams sitting downstream of those changes in the room, would have had a different conversation than the one that happened after.

Two practices prevent a forum from drifting into pure retrospective, and both are structural rather than motivational.

The first is publishing outcomes, not minutes. The distribution after each meeting should be a brief summary of decisions made, owners assigned, and open risks carried forward, not a transcript of what was discussed. When the output is a decisions document, the forum creates a durable record of whether it is functioning. When that record shows three consecutive meetings with no new decisions, the forum has diagnosed itself.

The second is applying a sunset criterion. Every six months, the forum’s named owner should ask explicitly whether the governance structure it provides is still matched to the problems the organization faces. Performance programs evolve. The operating model appropriate for a mobilization phase is not the right one for a steady-state phase, and the steady-state model itself needs to evolve as the system’s architecture changes. A forum that never revises its own structure becomes a ritual: present, attended, and inert.


The Forum Is Not the Program

There is a final distinction worth naming, because conflating it with the governance forum is one of the more common ways performance programs stall.

The forum is the decision-making layer. It is not the measurement infrastructure, the instrumentation work, the regression prevention tooling, or the engineering practices that produce the signal the forum acts on. A well-run forum sitting on top of poor measurement will produce well-structured decisions about the wrong things. A well-run forum connected to sound instrumentation will produce decisions grounded in what users actually experience.

The OpenAI ChatGPT degradation in November 2023 illustrates what that distinction means in practice. For roughly five hours, ChatGPT and the API were degraded or unavailable globally. OpenAI attributed the disruption to a DDoS attack, but independent analysis from ThousandEyes found the degradation was concentrated specifically in the initial page load and component population process, where multiple backend services and APIs need to coordinate simultaneously to render the interface. The bottleneck was in the service mesh, not at the network perimeter. The governance implication is direct: a functioning performance forum can still produce the wrong decisions if the measurement layer beneath it cannot distinguish a network-layer attack from an internal coordination failure. The forum is only as good as the signal it acts on.

This is why the instrumentation work, the critical user journey measurement, the SLI and SLO design, the shared observability across component boundaries, is not preparatory scaffolding that gets completed before governance begins. It is the substrate that determines whether the governance structure can function at all. A forum reviewing averages is not governance for a system where the users who matter most live in the tail. A forum reviewing synthetic tests is not governance for a platform whose failure modes only appear under real user conditions.

Performance forums need to be calibrated to the systems they govern. The cadence, the attendance, the agenda design, and the decision framework all need to be matched to the failure modes and the measurement fidelity of the platform underneath. A forum that meets monthly to review quarterly trend data is not governance for a system that can regress in a single deployment cycle.

The organizations that build forums that last are not the ones with the most sophisticated agendas. They are the ones with a named owner who treats the forum as a standing obligation, instrumentation that connects technical findings to user reality, and the discipline to restructure the governance model when the evidence shows it has stopped producing decisions.

Leave a Comment