Engineering Metrics · Diagnosis and Proposal

Why Completion Ratio cannot work · and what to use instead

A diagnosis of the commitment problem across PI 8 to PI 10, an analysis of what unplanned work actually consists of, and a proposed replacement for the shared 80% Completion Ratio target. Written for the Product and Engineering leadership discussion. Every recommendation is labelled with its confidence level and its source.
The diagnosis in one paragraph
Completion Ratio is not measuring delivery, and no amount of commitment calibration will fix it. Across PI 8 to PI 10 the organisation aimed between 144% and 281% of each team's capacity at them, averaging 212%. When a team commits above capacity the ratio is arithmetically unreachable. When it commits below capacity the surplus demand arrives as creep and the ratio still misses. The metric therefore reports on where a team drew its commitment line, not on how much it delivered. Its ranking of the five teams is close to the inverse of reality: Fusion, the most stable team in the organisation, scores worst on completion stability, and Data, the most volatile, scores best.
Demand vs capacity
212%
mean across 15 team-PIs, range 144-281
Spent on unplanned work
27%
of capacity, mean; range 11-50
Output is compliance work
28%
of story points, and rising
Completion vs reality
-0.50
rank correlation with delivery stability

Root cause: the organisation aims twice its capacity at engineering

Total demand is commitment plus creep, expressed as a share of measured capacity. This is the number that explains every other symptom in the metric set.
Total demand as a percentage of capacity
Commitment plus creep divided by capacity, active sprints only. The 100% line is what the team can actually absorb. Nothing sits near it.
Cloud Data Fusion Mobile Web 100% capacity

Every team in every PI has been asked to absorb more than it can. Cloud peaked at 240% in PI 8, Mobile at 273%, Data at 281% in PI 9. Web is the lowest and still averages 147%. The overload is falling on three teams (Cloud 240 to 166, Data 251 to 208, Fusion 252 to 187) which is real progress, but the destination is still roughly double what the teams can do.

This single fact explains the pattern documented in the PI 8 to PI 10 trend analysis. Velocity vs Capacity has stayed in a narrow band of 94% to 100% org-wide for three PIs. That org-level narrowness is partly an aggregation effect, since individual team-PIs range from 73% to 133%, but the central tendency tracks available people rather than the plan. Commitment has swung by 49 percentage points because it is a planning decision. Completion Ratio, being the second divided into the first, moves with whichever the team chose, and has therefore stayed flat at 45% to 47% while the plan moved underneath it.

The corollary matters for the PI 11 planning change. Reserving capacity for ad hoc work is the right move and this analysis supports it. But reserving capacity does not reduce demand. It makes the overload visible at planning time instead of discovering it at sprint close. That is a large improvement in honesty and a small one in throughput, and it should be introduced with that expectation set.

What the unplanned work actually is

Jira carries an AddedWork_PI<n> label convention that identifies work added after PI planning. 212 labelled items across PI 8 to PI 10 carry 326 story points between them, of which 315 fall inside the three PI windows. Categories below are derived from issue type and summary text.
Coverage limit on this section. The AddedWork labels account for 81, 115 and 119 story points in PI 8, 9 and 10, against warehouse creep of 292, 329 and 445. So the label captures roughly a quarter to a third of all creep. Treat the category mix below as a representative sample of what added work looks like, not as a complete census. Extending the labelling discipline to all added work would make this measurable properly, and that is one of the proposals.
Added work by category
Share of labelled added-work story points, PI 8 to PI 10 combined
Flow distribution of all resolved work
Share of story points by work type, PI 8 to PI 10. Categories follow the Flow Framework's four flow items.

It is not churn. It is mostly foreseeable in category, unforeseeable in specific.

The largest single category is client and programme work arriving mid-PI at 18% of added points: xCures phase 1, Kardigan adherence visibility, DANI and ATA reprocessing, EDC schema changes. These are not surprises in kind. Prolaio runs client programmes and they generate work between planning events. The surprise is only ever which client.

Discovery and investigation is 17%. Spikes such as "Investigate lag increase", "Investigate Metabase for Dashboard Layer", "Series Frame Generator Exploration". This is the cost of not knowing enough at planning time, and some of it is irreducible: the alternative to a spike is a wrong estimate.

Design iteration after review is 13%, concentrated almost entirely in Mobile during PI 8: twenty UI Design Tasks with summaries like "Updates after final changes in Calibration", "Update illustrations after Unboxing Sessions", "per 01/29/26 discussion". This is design rework generated by review sessions and it disappeared by PI 10, which suggests it was a phase rather than a standing tax.

Platform operations (12%) and upgrades and tech debt (8%) together make a fifth: Kafka consumer config, Flink 2.x, Kotlin 2.4, processor splitting, renovatebot authentication. Data reprocessing requests are 6% and are pure client-driven interrupt. Governance and compliance enablement is 4%, largely BigQuery policy tags. Security and CVE remediation is small in points but notable in character: three items in PI 10 forced by Orca container scans failing, which is externally timed work the team cannot schedule.

Separately: a quarter of all output is regulated documentation

Across PI 8 to PI 10, 28% of resolved story points were compliance and verification artefacts: Software Design Specifications, Quality Verification Tasks, Software Item Specs, Requirements and Test Cases. In item count it is 24%, and the item count went 239, then 208, then 425 across the three PIs. Features and change work is 64% of points.

This is not creep and it is not waste. It is the cost of operating under 21 CFR Part 11 and SOC 2. These points are counted in velocity, so the teams are not understated by the throughput measures. What is invisible is the composition: nothing in the current metric set shows that roughly a quarter of output is regulated documentation, so an observer reading feature progress alone will attribute the gap to the teams rather than to an allocation the organisation has effectively chosen by default.

Platform Engineering is absorbing a support queue nobody counts

The PE project carries 424 Support tickets resolved between October 2025 and August 2026, running at 29 to 74 raised per full month and rising. Monthly median time to close ranges from 0 to 4 days, so the queue is being served responsively. None of these tickets carry story points and Platform Engineering does not appear in the five-team agile metric set at all. A meaningful volume of interrupt-driven work is being absorbed by a team that is invisible in this reporting. Whatever is decided about the other metrics, this should be brought into view.

Why Completion Ratio cannot be fixed

Three independent arguments: the arithmetic, the evidence from our own data, and the published position of the frameworks the target was drawn from.

1. The arithmetic makes the target unreachable for four of five teams

Completion Ratio is committed work done divided by commitment. For a team committing at 200% of capacity, as Mobile did in PI 10, hitting 80% completion requires completing committed work worth 160% of capacity. That is not a stretch goal, it is outside the possible. Three of five teams committed above capacity in PI 10, and the threshold that matters is higher than 100%: 80% completion becomes arithmetically impossible once commitment passes 125% of capacity. Ten of the fifteen team-PIs across PI 8 to PI 10 were above that line. In PI 10 it was Fusion at 148% and Mobile at 199%. The best completion ratio anywhere in the dataset is Web's 64.2% in PI 8, which coincided with the lowest commitment ratio in PI 8 at 114% of capacity. Web then cut commitment much further in PI 10, to 69% of capacity, and its completion ratio fell to 45%, so a low commitment ratio is not on its own sufficient.

2. It ranks our teams close to inversely

If Completion Ratio measured delivery reliability, its stability should track the stability of actual output. It does the opposite. Ranked by how stable each team's Velocity vs Capacity has been across the three PIs, the order is Fusion, Web, Data, Mobile, Cloud. Ranked by Completion Ratio stability the order is Data, Cloud, Web, Mobile, Fusion. The Spearman rank correlation between the two is -0.50. Note what this does and does not prove: on its own a negative correlation only says the two metrics disagree, not which one is right. The reason to trust the Velocity vs Capacity ordering is that it is built from two quantities neither of which is a plan, whereas Completion Ratio has a planning decision in its denominator. The Fusion case below makes the disagreement concrete enough to judge.

Metric stability by team · coefficient of variation across PI 8 to PI 10
Lower is more stable. Velocity vs Capacity separates the teams across a sixfold range. Completion Ratio clusters them, ranks them in a different order, and puts the most consistent team last.
Velocity vs Capacity Completion Ratio

Fusion is the clearest case. Its Velocity vs Capacity across three PIs was 104.4%, 104.8% and 109.6%, a coefficient of variation of 2.7%, which is exceptional consistency. Its Completion Ratio was 39%, 52% and 64%, a CV of 23.5%, the worst in the organisation. Both describe the same team in the same period. One says it is the most reliable team we have. The other says it is the least predictable.

A fair objection, and the answer to it. Fusion's high Completion Ratio variation is a genuine improving trend, not noise: 39% to 52% to 64%. Coefficient of variation cannot tell a trend from instability, so it penalises improvement. That is a real limitation of using CV, and it applies to Velocity vs Capacity too. It does not rescue Completion Ratio, because the underlying problem remains that the metric moves with commitment-setting rather than with delivery. But any stability measure adopted should be read alongside direction, not instead of it.

3. Scrum's own position is that a properly calibrated team misses about half the time

The Scrum Patterns group, whose Yesterday's Weather and velocity patterns are the origin of this style of measurement, states the expectation directly. From Notes on Velocity: "the team should expect to finish all items on that backlog only 50 percent of the time. Note that the Scrum Team commits to the Sprint Goal and not their forecast delivery." The same page also states: "The team can use velocity as a forecast of the work it will complete, but it is not a target or a guarantee for stakeholders." The companion Yesterday's Weather pattern makes the distributional point: "the team should expect about half of the Sprints to fall short of achieving Yesterday's Weather, and about half to exceed it."

The Scrum Guide removed the word commitment for exactly this reason in July 2011. The official revision note reads: "Development Teams do not commit to completing the work planned during a Sprint Planning Meeting. The Development Team creates a forecast of work it believes will be done, but that forecast will change as more becomes known throughout the Sprint." The 2020 Guide reintroduced the word, but attached it to the Sprint Goal, the Product Goal and the Definition of Done, never to scope.

The most useful line for our situation is also from Notes on Velocity, and it is about what to target instead: "It is more important to reduce the variance in velocity than to increase its magnitude."

Sources · Scrum PLoP, Yesterday's Weather and Notes on Velocity · Scrum Guide revision history, changes between the 2010 and 2011 Guides, item 1 · Ron Jeffries, Story Points Revisited (2019): "I like to say that I may have invented story points, and if I did, I'm sorry now" and "I think tracking how actuals compare with estimates is at best wasteful."

Proposal: four measures, no single score

Replace the shared 80% Completion Ratio target with a small set that cannot be improved by changing the commitment. Each carries a confidence label: Verified means computed from our own data and checked, High means consistent with primary sources and our data supports it, Assumption means it needs a PI of trial before being trusted.
Primary throughput measure

1. Velocity vs Capacity, targeted as a band Metric verifiedBand is a judgement

Total resolved work, committed plus unplanned, divided by measured capacity. This is already in the warehouse and already on the team briefs. Crucially it removes the commitment decision from the measure: velocity is what got done, capacity is derived from who was available, and neither is a plan. It is not un-gameable in general. There are two known attack surfaces, and both should be watched: story point inflation raises the numerator, and under-declaring capacity raises the ratio while also shrinking the ad hoc reserve that depends on it.

Target a band of 85% to 115% rather than a floor. Below 85% suggests capacity is being lost to something unmeasured. Above 115% suggests capacity is being under-declared, which matters because capacity is the denominator for the ad hoc reserve. A band makes both failure modes visible; a floor only catches one and rewards under-declaring capacity. Be clear what this would have scored historically: 8 of the 15 team-PIs in PI 8 to PI 10 sit inside 85-115, and 7 sit outside (values range from 73% to 133%). The band width is a judgement, not a derivation, and the honest way to set it is to run one PI reporting against it before attaching any consequence.

PI 10 actuals
Cloud 96%, Data 84%, Fusion 110%, Mobile 106%, Web 76%
Discriminates
Yes. Team CVs range 2.7% to 18.4%, a sixfold spread, where Completion Ratio CVs cluster between 6.6% and 23.5% with no clear separation.
Watch for
Both attack surfaces. Capacity is self-declared, so it needs a sanity check against headcount and leave. Estimates are self-set, so a rising velocity with flat item counts and flat release output is the signature of inflation rather than improvement.
Planning discipline measure

2. Ad hoc reserve, measured by fill rate High confidence

Reserve an explicit share of capacity for unplanned work at PI planning, sized from each team's own history. Then measure fill rate: unplanned work actually done divided by the reserve. This is the named practice Illegitimus Non Interruptus from the Scrum Patterns group, and it is what the PI 11 planning change already sets up.

Target fill rate 80% to 100%. Under 80% means the reserve is too large and is hiding slack, so shrink it next PI. Over 100% means the reserve was breached, which is the signal to escalate rather than silently absorb. The pattern's own rule is stronger than what is proposed here: it says the sprint must automatically abort and be replanned. Adapted to a PI cadence, the workable version is that a breach triggers replanning and a conversation with management rather than silent absorption.

The practice is intended to be self-liquidating, which is its most attractive property: "if the team uses Yesterday's Weather to size the buffer and the buffer almost never fills up, the buffer size continuously gets smaller, making the interrupt problem go away." Fusion shows that interrupt load can fall on its own, from 33% of capacity in PI 8 to 16% in PI 10. That is illustrative rather than confirmatory, since Fusion has never run a reserve.

Precedent
Google SRE sets an advertised goal of keeping operational work below 50% of engineer time and surveys actuals quarterly, reporting about 33%. The rationale is verbatim ours: "toil tends to expand if left unchecked and can quickly fill 100% of everyone's time."
Sizing
See the next section. Our own three-PI history gives a mean of 27% and a range of 11% to 50%.
Watch for
A reserve only works if breaching it has a consequence. Without the escalation rule it becomes a second backlog.
Visibility measure

3. Flow distribution across four work types High confidence

Publish the split of each team's output across Features, Defects, Risks and Debt, per Mik Kersten's Flow Framework. Risks in that taxonomy explicitly means compliance and security work, which gives our Part 11 and SOC 2 artefacts a named home instead of an unlabelled deduction from feature capacity.

This is not a target. It is a disclosure. The point is that when 28% of output is regulated documentation and rising, leadership should be choosing that allocation deliberately rather than discovering it as a shortfall in feature delivery.

Current split
By story points: Features 64%, Risks 28%, Debt and operational 8%, Defects under 1%. By item count: 46%, 24%, 18%, 12%. Report both, because defects and support tickets are largely unpointed and the points view understates them by an order of magnitude.
Honest caveat
The Flow Framework publishes no recommended distribution and no thresholds. Anyone quoting a percentage target for it is inventing one. It prescribes that you set and inspect a distribution, not what it should be.
Also needed
Platform Engineering's support queue should appear here. 424 tickets is not a rounding error.
Outcome measure

4. Sprint goal hit rate, with the parsing fixed Verified

Keep this one. It is the measure closest to what Scrum actually says teams should commit to, and it is the only current metric where the shared 80% target is both reachable and being reached: Cloud and Fusion both hit 100% in PI 10.

But fix the mechanics first. Data wrote parseable goals in 2 of 5 sprints and Web in 1 of 5, because both used [] with no space between the brackets, which the warehouse regex does not match. On the goals that did parse, four of five teams scored 100% and Mobile scored 78%, so the metric currently rewards teams whose goals the parser can read and says very little about the rest. Until that is fixed it measures Jira formatting as much as delivery.

Fix
Agree on - [ ] and - [x], or widen the regex to accept []. The second is a one-line change and needs no behaviour change from teams.
Target
Keep 80%. Four of five teams are already at 100% on the goals that parse, so it is demonstrably achievable once the measurement is sound.

Sizing the ad hoc reserve per team

Every credible source prescribes a method rather than a number: size the reserve from your own measured history, then let the observed fill rate retune it. Below is what our history says.
Unplanned work completed, as a percentage of capacity
This is what each team actually spent on work that was not in the plan. It is the empirical basis for the reserve.
Cloud Data Fusion Mobile Web
TeamPI 8PI 9PI 10MeanMaxTrendProposed reserve
Cloud23%23%37%28%37%Rising35%
Data22%23%50%32%50%Rising sharply40%
Fusion33%18%16%22%33%Falling20%
Mobile19%45%28%31%45%Volatile30%
Web15%11%45%24%45%Rising sharply35%
Organisation21%25%36%27%n/aRising30%

Teams with a falling trend take their three-PI mean. Teams with a rising trend take the mean blended with the most recent PI, on the grounds that the recent number is the better predictor of the next one. Fusion gets the smallest reserve because it has earned it, which is the incentive the practice is supposed to create.

These numbers are close to the external precedents, which is reassuring. Google SRE measures actual operational load at 33% against a 50% cap. The Scrum Patterns worked example uses one third, 20 points of a 60 point velocity. Our organisation-wide mean is 27% and our 85th percentile across fifteen team-PIs is 44%. Nobody credible publishes a universal percentage, and we should not either, but landing in the same range as two independent sources suggests the method is sound.

A claim to reject if it comes up. There is a widely repeated assertion that McKinsey research found high-performing agile teams plan to 75-80% of capacity and achieve 35-40% higher sprint completion. No such publication exists as far as we can establish: no title, author, date or link, and the trail ends at a vendor content page citing nothing. If this number is presented in the discussion, ask for the source.

Options tested and not recommended

Two obvious candidates were tested against our own data and do not work here yet. Recording them so they are not proposed again without the prerequisites.

Item throughput instead of story points Tested, rejected

The Kanban Guide's position is that throughput should be "the exact count of work items" rather than a sum of estimates, which would remove estimate inflation from the measure entirely. It is a genuinely attractive idea and it fails on our data. Comparing sprint-level coefficient of variation, item counts are less stable than story points for four of five teams: Cloud 114% against 41%, Data 119% against 52%, Fusion 85% against 41%, Web 57% against 40%. Only Mobile improves, at 30% against 41%.

The cause is item size heterogeneity. A Platform Engineering support ticket and a Software Design Specification both count as one item. Until work items are sliced to comparable size, counting them is noisier than estimating them. Revisit if the organisation adopts a slicing discipline.

Cycle time and a Service Level Expectation Blocked on data

This is the strongest candidate in the literature and the one to aim for. A Service Level Expectation states delivery as a distribution with a confidence level, for example "85% of work items finish within eight days", computed from history. It cannot be improved by committing less, which is precisely the property Completion Ratio lacks.

We cannot compute it yet. The issues table records created_date and resolution_date but no timestamp for when work actually started, so the only available interval is creation to resolution, which is lead time including backlog wait. The distortion is severe enough to invert the ranking: Fusion shows the worst 85th-percentile lead time at 239 days despite being the most consistent delivery team, because its tickets sit in the backlog a long time before being picked up.

Prerequisite: capture the transition into In Progress. That is a Jira change plus a warehouse column. Once it exists, an SLE per team becomes the natural primary predictability measure and Velocity vs Capacity can step back to being a capacity-utilisation check.

Also unavailable: per-team creep attribution from tickets. The issues table has no sprint field, and the board-based team mapping in dim_jira_project is unreliable: it assigns FSN (Fusion), PW (Prolaio Web) and DATA all to Cloud. Reconciling ticket story points against warehouse velocity gives 184% to 397% for Cloud and 18% to 60% for Web. Team-level creep composition therefore cannot be reported from tickets today. Adding sprint identity to the issues feed would fix this and would make the creep analysis in this document reproducible per team rather than org-wide.

Suggested transition

Sequenced so that each step is useful on its own and nothing depends on a change that has not landed yet. This is a proposal for discussion, not a decision.
  1. Before PI 11 planning: fix the sprint goal regex. Widen it to accept [] alongside [ ]. One line, no behaviour change asked of any team, and it immediately makes the one target that works actually measurable for Data and Web.
  2. At PI 11 planning: set an explicit ad hoc reserve per team using the sizes in the table above, and agree the escalation rule that a breach triggers replanning rather than absorption. The reserve is the intervention; the fill rate is just how we tell whether it was sized right.
  3. Retire the 80% Completion Ratio target from team lead OKRs. Keep reporting the ratio as context, because it still says something about planning calibration, but stop scoring anyone on it. No team has come within 15 points of the target in three PIs, and ten of the fifteen team-PIs in that window were arithmetically incapable of reaching it.
  4. Adopt Velocity vs Capacity as the throughput measure with an 85-115% band. It needs no new instrumentation. Pair it with a capacity sanity check so the denominator stays honest.
  5. Publish flow distribution in the PI briefs so that compliance and operational load are visible as deliberate allocations. Bring Platform Engineering into the reporting set at the same time.
  6. Extend the AddedWork labelling to all added work. Today it captures a quarter to a third. If it captured all of it, the question of what creep consists of becomes a standing report rather than a one-off investigation.
  7. Then, once the In Progress timestamp is captured, pilot a Service Level Expectation on one team. Fusion is the natural candidate given its consistency. If it works, it becomes the predictability measure and this whole problem stops being about commitment at all.

What this does not solve

None of the above reduces the 212% demand overload. Better measurement makes the overload visible and stops teams being scored for absorbing it, which is worth doing on its own. But the decision about what not to build sits with Product and Engineering leadership together, and no metric change substitutes for it. The honest framing for the discussion is that this proposal fixes the instrument, not the load.

Confidence and method

Verified. All agile figures come from sprint_metrics and the capacity table, active sprints only, PI 8 to PI 10. Demand, reserve and stability calculations were computed directly from those figures and rechecked. The Completion Ratio arithmetic and the rank correlation are computed from the same source. Flow distribution and the added-work categories come from the issues table.

High confidence. The commitment-creep relationship documented in the trend analysis, where between PI 9 and PI 10 the rank ordering by commitment cut matches the rank ordering by creep rise across all five teams. The same test over the full PI 8 to PI 10 span gives a positive but imperfect association, so the finding is specific to the most recent transition. The external precedents for reserve sizing are primary sources quoted verbatim.

Assumption, needs a PI to validate. That an explicit reserve changes behaviour rather than just relabelling. That the proposed reserve sizes are right; they are derived from three PIs, which is the minimum the Scrum Patterns group considers meaningful for a running average. That the 85-115% band on Velocity vs Capacity is the right width; it is a judgement, not a finding.

Categorisation method. Added-work categories were assigned by rules over issue type and summary text, not by human review of each ticket. The residual 15% is charted as Other, and four categories below 1% are omitted from the chart. The categories are indicative of composition, not an audited taxonomy. The Flow Framework mapping is also a local judgement: Software Design Specifications, Quality Verification Tasks, Software Item Specs, Requirements and Test Cases were all mapped to Risks, which is defensible for a regulated context but is not dictated by the taxonomy and materially affects the 28% figure.

Where no authority exists. We found no published guidance on measuring engineering throughput where regulated documentation is a large fixed overhead. The Flow Framework's Risks flow item and DORA's approach of measuring the compliance path's own wait time are the closest available substitutes, and both are adaptations rather than established practice for this situation.