In Program Governance, Use The Right Tool for the Right Job

Program governance has a tendency to accumulate. A dependency failure produces a new meeting. A missed commitment produces another reporting layer. A planning problem produces a new planning process. Each mechanism may solve a legitimate problem, yet over time an organization can end up carrying a coordination system sized for problems, scale, or complexity it no longer has.

Good program governance should be proportional. The planning horizon should reflect how quickly assumptions change. The coordination load should reflect the complexity of the dependency network. The measurement system should reflect the decisions leadership needs to make. The resilience model should reflect the consequences of dependency failure. Every mechanism should earn its cost through a delivery problem it demonstrably helps solve.

I think of this as proportional governance: coordination cost should scale with delivery risk, dependency complexity, and rate of change.

I learned that principle most clearly while working with SAFe at enterprise scale, where its reason for existing was difficult to miss.

I was leading technology programs inside a Fortune 10 healthcare enterprise during one of the largest corporate integrations in U.S. history. The programs crossed digital products, infrastructure, pharmacy systems, performance engineering, release governance, and production operations. Consumer platforms depended on more than 100 backend systems. I was directing more than ten concurrent programs.

At that scale, coordination becomes an engineering problem of its own. A team can execute its sprint cleanly and still contribute to a failed release. An API arrives late. An infrastructure dependency was assumed, never committed. Security enters the critical path after architecture decisions have hardened. Performance testing surfaces a constraint that would have been trivially addressable three months earlier. Individually successful teams produce a collectively unsuccessful program.

That experience taught me why enterprise delivery frameworks exist. The more valuable lesson came later: understanding which coordination problems they solve, which mechanisms transfer well to other environments, and how much governance each problem actually requires.

As engineering organizations grow, coordination failures eventually exceed what team-level practices can solve. The right response is to identify the specific failures affecting delivery, then implement the minimum sufficient governance mechanisms capable of solving them at the organization’s scale, dependency complexity, and rate of change. Enterprise frameworks provide a valuable body of accumulated experience about which mechanisms matter and why. The implementation should remain proportional to the problem.

The practical rule I carried forward is simple: scale the coordination system to the dependency network, then measure whether delivery gets better.

The First Scaling Problem Appears Between Teams

The Scrum Guide describes Scrum Teams as typically ten or fewer people. That design is intentional. Small teams reduce internal coordination complexity, maintain tight feedback loops, and give bounded groups clear ownership of an objective.

The scaling problem surfaces when multiple teams must coordinate through shared architecture, APIs, infrastructure, security controls, regulatory requirements, vendors, release windows, and production environments. At that point, the organization’s hardest delivery problems exist between teams, not inside them.

At the Fortune 10 healthcare enterprise, individual Scrum teams were executing well. Sprint goals were being met. Velocity was stable. Programs were still failing to deliver on time. The underlying problem was usually a dependency that became visible only after execution exposed it. A team building a consumer-facing feature depended on an API team that had no awareness of the timeline. A compliance requirement entered the critical path six weeks before launch because the compliance team had not been included in planning. Release governance reviews surfaced readiness problems that were solvable in week two and severe in week ten.

Large-scale integrated planning solved a specific problem in that environment: bringing large numbers of teams and their dependencies into a shared planning horizon simultaneously. SAFe’s PI Planning and ART structure provided formal mechanisms because the dependency network was genuinely too complex for ad hoc coordination. The IP Iteration also created protected capacity for integration and readiness work that was difficult to absorb inside already committed team iterations.

The mechanism worth preserving from that experience is broader than any framework: make cross-team dependencies and commitments visible before execution exposes them at full cost. The governance mechanism should scale with the complexity of the dependency network.

A team can have a successful sprint while the delivery system has an unsuccessful quarter. Optimizing individual teams is a different problem from optimizing the system those teams operate inside.

Keep Integrated Planning. Size the Planning Horizon.

Periodic integrated planning across teams is one of the most durable mechanisms in scaled delivery: revisit shared objectives, surface dependencies, reconcile demand with capacity, expose changed assumptions, negotiate tradeoffs, and identify delivery risk before execution hardens around stale information.

The cadence should remain a design variable.

Many enterprise planning models operate on multi-month horizons. At very large scale, that can make economic sense. When hundreds of teams need coordinated alignment, the planning event itself carries substantial transaction cost. Replanning continuously would create organizational burden that exceeds the value of more frequent updates. A smaller organization with lower dependency complexity and a faster rate of change faces a different optimization problem. Priorities and constraints can shift materially within weeks. A months-long planning horizon provides extended stability, but it also provides an extended window to remain wrong.

When I was building an observability program management function at a core banking technology provider serving over 160 regulated financial institutions, a 12-week commitment horizon carried real risk. Release governance was tight. Regulatory exposure was constant. The cost of discovering a wrong assumption in week 10 of a 12-week plan was severe. The operating model increasingly moved toward a shorter integrated planning cadence, closer to monthly, allowing the delivery system to reorient around changed assumptions before they became launch blockers.

Team execution cadence and cross-team integrated planning cadence serve different purposes and should be designed separately. A team can run two-week sprints with full iteration discipline while the program-level planning cycle operates on a monthly horizon.

The question worth asking: how long can the organization afford to operate on an assumption before testing whether that assumption still holds? In a smaller organization with rapidly changing priorities and constraints, the answer may be much shorter than three months. A monthly integrated planning cadence can surface changed constraints earlier, reduce estimate drift, and keep the dependency map current when the environment changes quickly enough to justify the additional planning cost.

The mechanism is integrated replanning. Its cadence should be shaped by organizational scale, coordination cost, and the rate at which important assumptions expire.

Synchronous Time Is a Coordination Cost

Scaled delivery creates a legitimate need for recurring coordination. Planning sessions, cross-team synchronization, system demonstrations, retrospectives, architecture discussions, and readiness reviews can all produce valuable outcomes.

Every one of them also consumes engineering capacity.

At enterprise scale, coordination events can be the price of managing a large dependency network. In a smaller organization with shorter decision paths and lower dependency complexity, carrying the same coordination load can cost more engineering capacity than the problem requires.

Every synchronous event should answer four questions before surviving a calendar review: What information does this produce? What decision does it enable? What dependency or ambiguity does it resolve? Does producing that result require people to be present at the same moment?

Some interactions genuinely require simultaneity. Architecture negotiation, dependency resolution across competing priorities, consequential tradeoffs, and launch decisions often need real-time dialogue because the back-and-forth itself produces the answer. Status transmission, metric reporting, risk register updates, and decision records after a decision has been made are strong async candidates. GitLab’s extensively documented remote operating model makes a similar distinction: written communication is the default for durable information exchange, while synchronous interaction is used when real-time discussion resolves ambiguity or accelerates a decision.

In deadline-driven regulated environments, synchronous time stayed where it resolved dependencies or decisions faster than written coordination could. A recurring coordination event earned its place by producing a specific output. Everything else moved to a written channel.

Synchronous time should be spent where simultaneity creates value: deciding, negotiating, resolving, or designing. Status needs a durable written home. Decisions sometimes need a room.

Measure the Delivery System, Not Just the Teams

As delivery scales beyond individual teams, the measurement system has to move up a level with it.

Velocity, sprint-goal attainment, aging work, and other team-level signals can be useful within a team. They reveal far less about the health of the delivery system.

The problem with using team-level velocity as the primary management signal at organizational scale is compounded when teams are compared. Story points are locally calibrated. A five on one team has no stable mathematical relationship to a five on another. Cross-team velocity comparisons produce noise. Leadership operating on aggregate velocity is operating on a metric designed for a different purpose.

The delivery system requires system-level signals. Cycle time and flow time reveal how long work takes to move from committed to done across the system, including wait states that individual teams do not see. Throughput tracks how much work the system completes over time. Aging work surfaces items that have been in progress long enough to indicate an unresolved problem. Dependency health tracks whether inter-team commitments are being met. One useful predictability signal is the relationship between committed and achieved objectives over time, which shows whether the program is becoming more reliable in its commitments.

At the core banking technology provider, making release and operational performance visible across multiple engineering disciplines contributed to a 40% reduction in unplanned release failures and more than a 50% improvement in mean time to detect. Those outcomes came from instrumenting the delivery system, not from improving individual team velocity.

DORA research reinforces the same principle from a software-delivery perspective. Software delivery performance is better understood through measures such as deployment frequency, lead time, change failure rate, and recovery performance. These measures describe the behavior of the delivery system, not the speed of any individual team.

Leadership needs fewer measures of Agile activity and better measures of delivery system behavior. The abstraction level of management information should increase as organizational scale increases.

Your Backlog May Be Hiding a Capacity Tax

Once the delivery system is measured at the right level, another governance problem becomes visible: the capacity leadership believes it has and the capacity engineering actually has are often different numbers.

A backlog tells you what engineering has queued. Toil tells you what engineering keeps having to do again.

At program scale, capacity itself becomes a system-level dependency. Invisible toil corrupts that planning because the organization is allocating capacity it does not actually have.

A mixed backlog containing features, bugs, security work, technical debt, and operational toil creates a visibility problem before it becomes a prioritization problem. When those categories are undifferentiated, leadership sees one queue. In practice, it is managing several demand classes with different urgency, capacity economics, risk, and strategic consequences.

The failure mode I have seen repeatedly: a team carries a significant recurring operational burden. Engineers spend meaningful capacity on manual workarounds for known defects or infrastructure deficiencies. That consumption can appear in aggregate as a velocity or headcount problem even when recurring toil is consuming the missing capacity. Quantifying that consumption changes the discussion: leadership can see the recurring cost, decide which sources to eliminate, and understand the feature capacity displaced while they remain.

Google’s SRE model treats toil as something to measure and constrain because recurring operational work can scale with service growth and steadily displace engineering capacity. Unmeasured toil gets absorbed into capacity that was never truly available for planned work.

At the Fortune 10 healthcare enterprise, features represented one demand category among several competing for the same engineering capacity. The programs that performed most reliably made those competing demands visible enough for leadership to see where engineering capacity was going and what each category displaced.

The practical sequence: quantify where capacity is actually going. Identify recurring toil sources and measure their frequency. Determine an acceptable threshold for that environment. Allocate capacity deliberately to retire the highest-cost sources. Make visible what feature delivery that investment displaces. The goal is visibility and explicit tradeoffs.

Recurring toil is technical debt collecting interest directly from engineering capacity. Leaving it invisible does not make it less expensive.

Not All Constraints Are Created Equal

Prioritization frameworks are useful once the work belongs in a comparable decision space. The harder governance problem comes first: determining which work is actually comparable.

WSJF, Weighted Shortest Job First, is one useful method for comparing discretionary work by relative economic value. Regulatory deadlines, active security exposure, production impact, and contractual commitments can constrain the decision space before relative economic ranking begins.

Some constraints are hard. A regulatory filing deadline carries a fixed external date. A known security vulnerability with active exposure carries risk that compounds over time. A critical production defect affecting customers or essential operations carries operational cost that accumulates daily. A contractual commitment carries commercial consequence. These constraints change the decision space before discretionary features are economically ranked.

At a federally administered health benefits exchange with immovable enrollment windows, no prioritization framework could treat a regulatory readiness requirement and a discretionary UI enhancement as interchangeable items in a mathematical ranking system. The enrollment window was fixed. Performance and infrastructure readiness were evidence-based launch criteria. Scope flexibility lived entirely in the feature tier, where work could be deferred, phased, or limited to a cohort without affecting the launch commitment.

Scope becomes an important control variable when dates, capacity, or quality thresholds are externally constrained. When a delivery date is hard, legitimate control variables include phased rollout, limited cohort exposure, feature flags, progressive release, and deferred secondary functionality. Defining the minimum viable launch scope is often more useful than optimizing the full backlog.

Priority is partly economic value. It is also partly constraint class. Understanding which category a work item occupies before applying a ranking formula is what makes the formula useful.

Critical Dependencies Need Recovery Paths

Constraint classification determines what must move first. Dependency resilience determines how the program responds when a critical dependency slips, degrades, or fails.

Dependency identification is the first layer of mature program management. Resilience is the second: for each critical dependency, define the owner, downstream impact, commitment date, leading indicators of slippage, escalation threshold, fallback trigger, and recovery path.

During a large commerce-platform implementation I led across international markets, a critical technology supplier became unable to continue delivery mid-program. Because I had maintained the incumbent platform as a viable recovery path throughout the implementation, we reverted without downtime, data loss, or customer impact. The recovery path existed before it was needed.

A dependency map tells you where delivery can break. A resilient program defines what happens when it does.

Make Governance Proportional

Scale the coordination system to the dependency network, introduce the minimum sufficient mechanism for the observed failure, and keep it only while delivery outcomes justify its cost.

Three steps make that principle operational.

First, diagnose the delivery-system failure. What is breaking delivery? Undiscovered dependencies? Stale planning assumptions? Slow decisions? Unclear ownership? Invisible toil consuming engineering capacity? Coordination overhead that exceeds the problem it is solving? Fragile external dependencies? Poor constraint classification? The answer determines which mechanism is warranted. A dependency problem and a toil problem call for different interventions.

Second, introduce the minimum sufficient mechanism that addresses it. For dependency failures: cross-team dependency mapping and commitment negotiation sized to the complexity of the network. For stale planning: integrated replanning on a cadence matched to how fast assumptions expire, not to how large the organization looks on paper. For coordination overload: a payoff test for every synchronous event, with information transmission moved to persistent written channels. For measurement gaps: flow-level signals at the program layer and decision-relevant signals at the leadership layer. For invisible toil: category-level capacity visibility and deliberate retirement of the highest-cost recurring sources. For constraint confusion: classification before ranking, and recovery paths for critical dependencies before they fail. For decision latency: defined decision rights, escalation paths, and time bounds appropriate to the consequence of the decision.

Third, measure the result. Are dependencies discovered earlier in the planning cycle? Are cross-team commitments becoming more predictable? Is blocked work trending down? Is security and compliance work entering planning before it becomes a launch blocker? Is decision latency decreasing? Are fewer risks first discovered during release readiness? Is recurring toil consuming a smaller share of capacity? Are release failures declining?

If those outcomes are improving, the governance mechanisms are earning their cost. If they are not, revisit the mechanism before adding more process.

Working with SAFe at enterprise scale taught me how expensive coordination failures can become. The broader lesson has stayed with me across every environment since: program governance should grow in response to observable delivery complexity, and every mechanism should continue earning the engineering capacity it consumes.

Use the right tool for the right job. Then make sure the tool is the right size.

Leave a Comment