Every team in every PI has been asked to absorb more than it can. Cloud peaked at 240% in PI 8, Mobile at 273%, Data at 281% in PI 9. Web is the lowest and still averages 147%. The overload is falling on three teams (Cloud 240 to 166, Data 251 to 208, Fusion 252 to 187) which is real progress, but the destination is still roughly double what the teams can do.
This single fact explains the pattern documented in the PI 8 to PI 10 trend analysis. Velocity vs Capacity has stayed in a narrow band of 94% to 100% org-wide for three PIs. That org-level narrowness is partly an aggregation effect, since individual team-PIs range from 73% to 133%, but the central tendency tracks available people rather than the plan. Commitment has swung by 49 percentage points because it is a planning decision. Completion Ratio, being the second divided into the first, moves with whichever the team chose, and has therefore stayed flat at 45% to 47% while the plan moved underneath it.
The corollary matters for the PI 11 planning change. Reserving capacity for ad hoc work is the right move and this analysis supports it. But reserving capacity does not reduce demand. It makes the overload visible at planning time instead of discovering it at sprint close. That is a large improvement in honesty and a small one in throughput, and it should be introduced with that expectation set.
AddedWork_PI<n> label convention that identifies work added after PI planning. 212 labelled items across PI 8 to PI 10 carry 326 story points between them, of which 315 fall inside the three PI windows. Categories below are derived from issue type and summary text.AddedWork labels account for 81, 115 and 119 story points in PI 8, 9 and 10, against warehouse creep of 292, 329 and 445. So the label captures roughly a quarter to a third of all creep. Treat the category mix below as a representative sample of what added work looks like, not as a complete census. Extending the labelling discipline to all added work would make this measurable properly, and that is one of the proposals.
The largest single category is client and programme work arriving mid-PI at 18% of added points: xCures phase 1, Kardigan adherence visibility, DANI and ATA reprocessing, EDC schema changes. These are not surprises in kind. Prolaio runs client programmes and they generate work between planning events. The surprise is only ever which client.
Discovery and investigation is 17%. Spikes such as "Investigate lag increase", "Investigate Metabase for Dashboard Layer", "Series Frame Generator Exploration". This is the cost of not knowing enough at planning time, and some of it is irreducible: the alternative to a spike is a wrong estimate.
Design iteration after review is 13%, concentrated almost entirely in Mobile during PI 8: twenty UI Design Tasks with summaries like "Updates after final changes in Calibration", "Update illustrations after Unboxing Sessions", "per 01/29/26 discussion". This is design rework generated by review sessions and it disappeared by PI 10, which suggests it was a phase rather than a standing tax.
Platform operations (12%) and upgrades and tech debt (8%) together make a fifth: Kafka consumer config, Flink 2.x, Kotlin 2.4, processor splitting, renovatebot authentication. Data reprocessing requests are 6% and are pure client-driven interrupt. Governance and compliance enablement is 4%, largely BigQuery policy tags. Security and CVE remediation is small in points but notable in character: three items in PI 10 forced by Orca container scans failing, which is externally timed work the team cannot schedule.
Across PI 8 to PI 10, 28% of resolved story points were compliance and verification artefacts: Software Design Specifications, Quality Verification Tasks, Software Item Specs, Requirements and Test Cases. In item count it is 24%, and the item count went 239, then 208, then 425 across the three PIs. Features and change work is 64% of points.
This is not creep and it is not waste. It is the cost of operating under 21 CFR Part 11 and SOC 2. These points are counted in velocity, so the teams are not understated by the throughput measures. What is invisible is the composition: nothing in the current metric set shows that roughly a quarter of output is regulated documentation, so an observer reading feature progress alone will attribute the gap to the teams rather than to an allocation the organisation has effectively chosen by default.
The PE project carries 424 Support tickets resolved between October 2025 and August 2026, running at 29 to 74 raised per full month and rising. Monthly median time to close ranges from 0 to 4 days, so the queue is being served responsively. None of these tickets carry story points and Platform Engineering does not appear in the five-team agile metric set at all. A meaningful volume of interrupt-driven work is being absorbed by a team that is invisible in this reporting. Whatever is decided about the other metrics, this should be brought into view.
Completion Ratio is committed work done divided by commitment. For a team committing at 200% of capacity, as Mobile did in PI 10, hitting 80% completion requires completing committed work worth 160% of capacity. That is not a stretch goal, it is outside the possible. Three of five teams committed above capacity in PI 10, and the threshold that matters is higher than 100%: 80% completion becomes arithmetically impossible once commitment passes 125% of capacity. Ten of the fifteen team-PIs across PI 8 to PI 10 were above that line. In PI 10 it was Fusion at 148% and Mobile at 199%. The best completion ratio anywhere in the dataset is Web's 64.2% in PI 8, which coincided with the lowest commitment ratio in PI 8 at 114% of capacity. Web then cut commitment much further in PI 10, to 69% of capacity, and its completion ratio fell to 45%, so a low commitment ratio is not on its own sufficient.
If Completion Ratio measured delivery reliability, its stability should track the stability of actual output. It does the opposite. Ranked by how stable each team's Velocity vs Capacity has been across the three PIs, the order is Fusion, Web, Data, Mobile, Cloud. Ranked by Completion Ratio stability the order is Data, Cloud, Web, Mobile, Fusion. The Spearman rank correlation between the two is -0.50. Note what this does and does not prove: on its own a negative correlation only says the two metrics disagree, not which one is right. The reason to trust the Velocity vs Capacity ordering is that it is built from two quantities neither of which is a plan, whereas Completion Ratio has a planning decision in its denominator. The Fusion case below makes the disagreement concrete enough to judge.
Fusion is the clearest case. Its Velocity vs Capacity across three PIs was 104.4%, 104.8% and 109.6%, a coefficient of variation of 2.7%, which is exceptional consistency. Its Completion Ratio was 39%, 52% and 64%, a CV of 23.5%, the worst in the organisation. Both describe the same team in the same period. One says it is the most reliable team we have. The other says it is the least predictable.
The Scrum Patterns group, whose Yesterday's Weather and velocity patterns are the origin of this style of measurement, states the expectation directly. From Notes on Velocity: "the team should expect to finish all items on that backlog only 50 percent of the time. Note that the Scrum Team commits to the Sprint Goal and not their forecast delivery." The same page also states: "The team can use velocity as a forecast of the work it will complete, but it is not a target or a guarantee for stakeholders." The companion Yesterday's Weather pattern makes the distributional point: "the team should expect about half of the Sprints to fall short of achieving Yesterday's Weather, and about half to exceed it."
The Scrum Guide removed the word commitment for exactly this reason in July 2011. The official revision note reads: "Development Teams do not commit to completing the work planned during a Sprint Planning Meeting. The Development Team creates a forecast of work it believes will be done, but that forecast will change as more becomes known throughout the Sprint." The 2020 Guide reintroduced the word, but attached it to the Sprint Goal, the Product Goal and the Definition of Done, never to scope.
The most useful line for our situation is also from Notes on Velocity, and it is about what to target instead: "It is more important to reduce the variance in velocity than to increase its magnitude."
Total resolved work, committed plus unplanned, divided by measured capacity. This is already in the warehouse and already on the team briefs. Crucially it removes the commitment decision from the measure: velocity is what got done, capacity is derived from who was available, and neither is a plan. It is not un-gameable in general. There are two known attack surfaces, and both should be watched: story point inflation raises the numerator, and under-declaring capacity raises the ratio while also shrinking the ad hoc reserve that depends on it.
Target a band of 85% to 115% rather than a floor. Below 85% suggests capacity is being lost to something unmeasured. Above 115% suggests capacity is being under-declared, which matters because capacity is the denominator for the ad hoc reserve. A band makes both failure modes visible; a floor only catches one and rewards under-declaring capacity. Be clear what this would have scored historically: 8 of the 15 team-PIs in PI 8 to PI 10 sit inside 85-115, and 7 sit outside (values range from 73% to 133%). The band width is a judgement, not a derivation, and the honest way to set it is to run one PI reporting against it before attaching any consequence.
Reserve an explicit share of capacity for unplanned work at PI planning, sized from each team's own history. Then measure fill rate: unplanned work actually done divided by the reserve. This is the named practice Illegitimus Non Interruptus from the Scrum Patterns group, and it is what the PI 11 planning change already sets up.
Target fill rate 80% to 100%. Under 80% means the reserve is too large and is hiding slack, so shrink it next PI. Over 100% means the reserve was breached, which is the signal to escalate rather than silently absorb. The pattern's own rule is stronger than what is proposed here: it says the sprint must automatically abort and be replanned. Adapted to a PI cadence, the workable version is that a breach triggers replanning and a conversation with management rather than silent absorption.
The practice is intended to be self-liquidating, which is its most attractive property: "if the team uses Yesterday's Weather to size the buffer and the buffer almost never fills up, the buffer size continuously gets smaller, making the interrupt problem go away." Fusion shows that interrupt load can fall on its own, from 33% of capacity in PI 8 to 16% in PI 10. That is illustrative rather than confirmatory, since Fusion has never run a reserve.
Publish the split of each team's output across Features, Defects, Risks and Debt, per Mik Kersten's Flow Framework. Risks in that taxonomy explicitly means compliance and security work, which gives our Part 11 and SOC 2 artefacts a named home instead of an unlabelled deduction from feature capacity.
This is not a target. It is a disclosure. The point is that when 28% of output is regulated documentation and rising, leadership should be choosing that allocation deliberately rather than discovering it as a shortfall in feature delivery.
Keep this one. It is the measure closest to what Scrum actually says teams should commit to, and it is the only current metric where the shared 80% target is both reachable and being reached: Cloud and Fusion both hit 100% in PI 10.
But fix the mechanics first. Data wrote parseable goals in 2 of 5 sprints and Web in 1 of 5, because both used [] with no space between the brackets, which the warehouse regex does not match. On the goals that did parse, four of five teams scored 100% and Mobile scored 78%, so the metric currently rewards teams whose goals the parser can read and says very little about the rest. Until that is fixed it measures Jira formatting as much as delivery.
- [ ] and - [x], or widen the regex to accept []. The second is a one-line change and needs no behaviour change from teams.| Team | PI 8 | PI 9 | PI 10 | Mean | Max | Trend | Proposed reserve |
|---|---|---|---|---|---|---|---|
| Cloud | 23% | 23% | 37% | 28% | 37% | Rising | 35% |
| Data | 22% | 23% | 50% | 32% | 50% | Rising sharply | 40% |
| Fusion | 33% | 18% | 16% | 22% | 33% | Falling | 20% |
| Mobile | 19% | 45% | 28% | 31% | 45% | Volatile | 30% |
| Web | 15% | 11% | 45% | 24% | 45% | Rising sharply | 35% |
| Organisation | 21% | 25% | 36% | 27% | n/a | Rising | 30% |
Teams with a falling trend take their three-PI mean. Teams with a rising trend take the mean blended with the most recent PI, on the grounds that the recent number is the better predictor of the next one. Fusion gets the smallest reserve because it has earned it, which is the incentive the practice is supposed to create.
These numbers are close to the external precedents, which is reassuring. Google SRE measures actual operational load at 33% against a 50% cap. The Scrum Patterns worked example uses one third, 20 points of a 60 point velocity. Our organisation-wide mean is 27% and our 85th percentile across fifteen team-PIs is 44%. Nobody credible publishes a universal percentage, and we should not either, but landing in the same range as two independent sources suggests the method is sound.
The Kanban Guide's position is that throughput should be "the exact count of work items" rather than a sum of estimates, which would remove estimate inflation from the measure entirely. It is a genuinely attractive idea and it fails on our data. Comparing sprint-level coefficient of variation, item counts are less stable than story points for four of five teams: Cloud 114% against 41%, Data 119% against 52%, Fusion 85% against 41%, Web 57% against 40%. Only Mobile improves, at 30% against 41%.
The cause is item size heterogeneity. A Platform Engineering support ticket and a Software Design Specification both count as one item. Until work items are sliced to comparable size, counting them is noisier than estimating them. Revisit if the organisation adopts a slicing discipline.
This is the strongest candidate in the literature and the one to aim for. A Service Level Expectation states delivery as a distribution with a confidence level, for example "85% of work items finish within eight days", computed from history. It cannot be improved by committing less, which is precisely the property Completion Ratio lacks.
We cannot compute it yet. The issues table records created_date and resolution_date but no timestamp for when work actually started, so the only available interval is creation to resolution, which is lead time including backlog wait. The distortion is severe enough to invert the ranking: Fusion shows the worst 85th-percentile lead time at 239 days despite being the most consistent delivery team, because its tickets sit in the backlog a long time before being picked up.
Prerequisite: capture the transition into In Progress. That is a Jira change plus a warehouse column. Once it exists, an SLE per team becomes the natural primary predictability measure and Velocity vs Capacity can step back to being a capacity-utilisation check.
issues table has no sprint field, and the board-based team mapping in dim_jira_project is unreliable: it assigns FSN (Fusion), PW (Prolaio Web) and DATA all to Cloud. Reconciling ticket story points against warehouse velocity gives 184% to 397% for Cloud and 18% to 60% for Web. Team-level creep composition therefore cannot be reported from tickets today. Adding sprint identity to the issues feed would fix this and would make the creep analysis in this document reproducible per team rather than org-wide.
[] alongside [ ]. One line, no behaviour change asked of any team, and it immediately makes the one target that works actually measurable for Data and Web.None of the above reduces the 212% demand overload. Better measurement makes the overload visible and stops teams being scored for absorbing it, which is worth doing on its own. But the decision about what not to build sits with Product and Engineering leadership together, and no metric change substitutes for it. The honest framing for the discussion is that this proposal fixes the instrument, not the load.
Verified. All agile figures come from sprint_metrics and the capacity table, active sprints only, PI 8 to PI 10. Demand, reserve and stability calculations were computed directly from those figures and rechecked. The Completion Ratio arithmetic and the rank correlation are computed from the same source. Flow distribution and the added-work categories come from the issues table.
High confidence. The commitment-creep relationship documented in the trend analysis, where between PI 9 and PI 10 the rank ordering by commitment cut matches the rank ordering by creep rise across all five teams. The same test over the full PI 8 to PI 10 span gives a positive but imperfect association, so the finding is specific to the most recent transition. The external precedents for reserve sizing are primary sources quoted verbatim.
Assumption, needs a PI to validate. That an explicit reserve changes behaviour rather than just relabelling. That the proposed reserve sizes are right; they are derived from three PIs, which is the minimum the Scrum Patterns group considers meaningful for a running average. That the 85-115% band on Velocity vs Capacity is the right width; it is a judgement, not a finding.
Categorisation method. Added-work categories were assigned by rules over issue type and summary text, not by human review of each ticket. The residual 15% is charted as Other, and four categories below 1% are omitted from the chart. The categories are indicative of composition, not an audited taxonomy. The Flow Framework mapping is also a local judgement: Software Design Specifications, Quality Verification Tasks, Software Item Specs, Requirements and Test Cases were all mapped to Risks, which is defensible for a regulated context but is not dictated by the taxonomy and materially affects the 28% figure.
Where no authority exists. We found no published guidance on measuring engineering throughput where regulated documentation is a large fixed overhead. The Flow Framework's Risks flow item and DORA's approach of measuring the compliance path's own wait time are the closest available substitutes, and both are adaptations rather than established practice for this situation.