Engineering Performance · Trend Analysis · PI 8 to PI 10
Three PIs of recalibration · what actually changed
A cross-team look at PI 8 (Oct 2025 to Feb 2026), PI 9 (Feb to Apr 2026) and PI 10 (Apr to Aug 2026). All figures cover active sprints only; IP sprints are excluded throughout. The question this sets out to answer is whether the commitment recalibration that began after PI 8 changed how much work the organisation completes, or only how that work is labelled.
The short version
Commitment swung from 175% of capacity to 125%. Output stayed level with capacity throughout. Velocity vs Capacity was 100%, 99% and 94% across PI 8, 9 and 10, a six-point band, while commitment moved by 49 points. What changed is the label on the work: unplanned work rose from 21% of everything completed in PI 8 to 38% in PI 10. Completion Ratio, the metric the recalibration was meant to improve, sat at 45%, 47% and 46%. It did not move either.
Completion Ratio
46%
45 · 47 · 46 across PI 8-10
Velocity vs Capacity
94%
100 · 99 · 94, stable
Creep Ratio
53%
31 · 31 · 53, jumped in PI 10
Unplanned share of output
38%
21 · 25 · 38, rising
Six patterns
Each is stated with the evidence that supports it. Where a pattern has a plausible alternative explanation, that is noted rather than buried.
1
Output tracks capacity, not commitment
Velocity vs Capacity across the organisation was 100%, 99% and 94% in PI 8, 9 and 10, a six-point band. Over the same window commitment moved from 175% of capacity to 157% to 125%, a swing of 49 points, and in absolute terms rose 123 SP then fell 233 SP. The plan moved roughly eight times as much as the ratio of output to capacity did. Where velocity did change it followed capacity: both rose from PI 8 to PI 9 and both eased in PI 10. Capacity is the better predictor of throughput.
2
Commitment and creep moved in opposite directions, in consistent rank order
Ranked by the size of their commitment cut between PI 9 and PI 10, the teams read Data, Web, Cloud, Fusion, Mobile. Ranked by creep increase they read Data, Web, Cloud, Fusion, Mobile. The two orderings match exactly. Data cut commitment by 135 percentage points relative to capacity and creep rose 102pp. Web cut 44pp and creep rose 93pp. Cloud cut 29pp and creep rose 11pp. Fusion barely moved on either (down 20pp and down 1pp). Mobile went the other way, raising commitment 42pp, and creep fell 15pp. The magnitudes are not proportional, so this is not a fixed exchange rate, but the rank consistency across five teams is hard to attribute to chance.
3
Completion Ratio has not responded to two PIs of recalibration
45.1%, 47.1%, 46.4%. The metric the 80% target is written against has been flat within two percentage points while its denominator rose 13% and then fell 22%. If commitment discipline were the binding constraint, this number should have risen when commitment came down. It did not, which points at the constraint being elsewhere.
4
Bug load is flat organisation-wide but redistributing sharply
Totals of 109, 111 and 98 look stable. Underneath, Mobile halved from 80 to 39 while Cloud went 10 to 21 and Data went 2 to 22. Mobile's improvement and the Cloud and Data increases roughly cancel out. Both rising teams shipped a major platform version in PI 10 (Cloud 3.11.0, Data data-ecosystem-core 1.0.0), which is a plausible and fairly ordinary explanation.
5
Release output is rising while story point output is flat
Setting Cloud aside, whose tag count is inflated by CI firing on every commit, final releases across the other four teams went from 5 in PI 8 to 24 in PI 9 to 30 in PI 10. Fusion went from 0 to 9, Web from 5 to 8, Data from 0 to 4. Teams are shipping considerably more frequently without resolving more points, which points at smaller batch sizes rather than more work.
6
PI 9 was an activity spike across the whole organisation
Merged MRs went 638, 2,108 and 907 org-wide. Commits went 8,389, 17,479 and 9,895. PI 9 roughly tripled the PI 8 merge rate and PI 10 settled back at about 1.4 times the PI 8 baseline. Every team finished PI 10 above its own PI 8 level: Cloud +26%, Data +206%, Fusion +241%, Mobile +20%, Web +17%. The spike was broad rather than concentrated: distinct contributing accounts were 30, 35 and 34, the top account never exceeded 22% of the total, and the median account went from 14 merges to 34 to 19. Review depth held constant at 1.93, 1.93 and 1.98 approvals per merge.
Output holds while the plan shrinks
Organisation-wide totals across active sprints. Capacity is the sum of measured team capacity; commitment is what was planned at sprint start; velocity is everything resolved, committed and unplanned combined.
Capacity, commitment and velocity · organisation total
Story points per PI, active sprints only. Commitment peaks in PI 9 then falls sharply. Velocity stays close to capacity throughout.
Capacity
Commitment
Velocity
Creep
Commitment fell from 952 story points in PI 8 to 842 in PI 10, having peaked at 1,075 in PI 9. Against that, velocity moved from 546 to 630, peaking at 676. Capacity rose 26% from PI 8 to PI 9 and then eased 2%, and velocity followed that shape closely (up 24%, then down 7%). Commitment did not: it rose 13% then fell 22%.
The interpretation that fits: demand on these teams is set by something outside the sprint plan. When the plan is smaller than demand, the difference arrives as creep. When the plan is larger, the excess simply does not get done and shows up as a low Completion Ratio. On either path the amount of work that lands stays close to how many people are available, which is what capacity measures. This is the reading that best fits three PIs of data; it has not been tested against an independent measure of demand.
Where output came from
Committed work completed versus unplanned work completed, as a share of total velocity
Committed done Creep done
The three headline ratios
Completion Ratio and Velocity vs Capacity stay flat while Creep Ratio jumps
Completion Velocity vs Cap Creep
In PI 8 roughly one story point in five that the organisation completed was never planned. In PI 10 it was closer to two in five. That is the single clearest change across the three PIs, and it happened while every headline throughput number stayed flat.
Commitment and creep, team by team
If cutting commitment relabels work rather than reducing it, teams that cut should show creep rising. The PI 9 to PI 10 data is consistent with that on four of five teams, with Fusion close to zero on both measures, including Mobile, which moved commitment the other way and saw creep fall. The rank ordering is also consistent: sorting the teams by the size of their commitment cut produces the same sequence as sorting them by creep increase. The magnitudes are not proportional, so this is a strong association with a consistent rank order rather than a fixed exchange rate.
Commitment change versus creep change · PI 9 to PI 10
Percentage-point change in commitment as a share of capacity (left bar) against percentage-point change in creep ratio (right bar). Teams are ordered by the size of the commitment change. The bars oppose each other on Data, Web, Cloud and Mobile. Fusion is close to zero on both.
Change in commitment vs capacity Change in creep ratio
| Team | Commit:Cap PI 9 | Commit:Cap PI 10 | Change | Creep PI 9 | Creep PI 10 | Change |
| Data | 227% | 92% | -135pp | 24% | 126% | +102pp |
| Web | 113% | 69% | -44pp | 28% | 122% | +93pp |
| Cloud | 146% | 117% | -29pp | 31% | 42% | +11pp |
| Fusion | 168% | 148% | -20pp | 27% | 26% | -1pp |
| Mobile | 158% | 200% | +42pp | 40% | 25% | -15pp |
Web is the cleanest illustration
Web planned S5 with zero commitment. During that sprint the team resolved 41.5 story points, all of it classified as creep, and shipped three prolaio-ui finals (0.13.0 on 26 June, 0.14.0 on 6 July, 0.15.0 on 7 July). The work was real and it was valuable. It simply had no plan attached to it, so it inflates the creep ratio, contributes nothing to Completion Ratio, and technically breaches the "value delivered every sprint" goal because commit_done was zero.
Three separate metrics recorded a failure for a sprint in which the team shipped three releases. That is worth pausing on before drawing conclusions about Web from any of those three numbers.
Mobile is the counter-example that supports the pattern
Mobile is the one team that increased commitment in PI 10, from 158% to 200% of capacity, and it is the one team whose creep ratio fell materially, from 40% to 25%. Mobile also cut its bug count roughly in half. The team's Completion Ratio (39%) is the second lowest in the organisation, which is a direct consequence of committing twice its capacity rather than an indication of low output: Mobile resolved 157.5 points against 149 of capacity.
An alternative reading is worth stating. Mobile's creep fall might reflect better upstream planning by the product owner rather than the commitment level itself, and PI 9 was an unusually high-output PI for the team, so some of the PI 10 movement is regression to the mean. The pattern now holds in direction on all five teams and the rank ordering is consistent on both measures, which is a good deal stronger than a directional read alone. It is still five teams and three PIs. A common cause acting on both variables at once, such as a shift in what product asked for mid-PI, would produce the same picture, so this remains an association rather than an established mechanism.
Quality and release cadence
Bug counts and release output over the same window. These are the metrics least affected by the planning dynamics above.
Bugs completed per team
Bug and Anomaly issues resolved, active sprints. Mobile dominates the total and drives the org trend.
Final releases per team
Excludes RC, alpha, beta and CI tags. Cloud shown as platform deployments, not tag count.
Mobile's bug count fell from 80 in PI 8 to 39 in PI 10, the largest single improvement anywhere in this dataset. Cloud rose from 10 to 21 and Data from 2 to 22 across the same window. Both rising teams shipped a significant platform version in PI 10, so a rise is not surprising, though Data's elevenfold increase from a very low base is worth a closer look at whether the PI 8 figure of 2 was under-recorded rather than the PI 10 figure being high.
On releases, the pattern is unambiguous once Cloud's CI noise is removed. Fusion went from zero final releases in PI 8 to nine in PI 10. Web went from five to eight. Data from zero to four. Teams are shipping more often while resolving roughly the same number of points, which is consistent with smaller batches reaching production faster. That is generally a healthy direction.
Cloud is the exception and it is the one release goal that clearly missed. Platform production deployments went 2, 5, 2 across PI 8 to 10, against a target of roughly one per active sprint. PI 10's two releases were 54 days apart. A third version, 3.10.16.0, was tagged on 3 June and never deployed before being superseded by 3.11.0 on 30 June.
DORA deployment frequency remains far below target across the board. Production deployment frequency is 0.028 per day for Cloud and 0.053 for Web in PI 10, against a DORA target of 0.5. That is 18 times below target for Cloud and 9 times for Web. Only Web's staging tier (0.75 per day) clears the threshold. Deployment tracking still covers only two of five teams, so the organisation cannot currently report DORA metrics for Data, Fusion or Mobile at all.
Activity and review
Merged merge requests, commits and approvals across the window, active sprints only.
Merged MRs per team
Active-sprint totals. PI 9 is a peak for every team. PI 10 sits above the PI 8 baseline for every team.
| Team | Merged PI 8 | Merged PI 9 | Merged PI 10 | PI 10 vs PI 8 | Merged per active day PI 10 | Approvals per merge PI 10 |
| Cloud | 248 | 666 | 312 | +26% | 4.39 | 1.40 |
| Data | 35 | 292 | 107 | +206% | 1.51 | 1.73 |
| Fusion | 29 | 128 | 99 | +241% | 1.39 | 1.43 |
| Mobile | 243 | 836 | 292 | +20% | 4.11 | 3.14 |
| Web | 83 | 186 | 97 | +17% | 1.28 | 1.22 |
| Organisation | 638 | 2,108 | 907 | +42% | n/a | 1.98 |
The shape is the same on every team and on both measures. PI 9 was an unusually high-activity PI, PI 10 came down from it, and PI 10 still sits above where the organisation was in PI 8. Nothing in the story point data shows a comparable PI 9 spike, so more merge requests in PI 9 did not translate into more resolved points. That points at batch size: PI 9 saw more, smaller units of change moving through review.
Review depth is the most stable measure in the entire dataset. Approvals per merged MR were 1.93, 1.93 and 1.98 across the three PIs. Mobile runs conspicuously high at 3.14 in PI 10 and Web low at 1.22, a spread worth understanding, but the org-level figure has not moved at all while every other metric in this report has.
Team trajectories
Every team's headline metrics across the three PIs. The direction column describes the PI 9 to PI 10 move. Completion Ratio and Committed Done vs Capacity are both strongly influenced by commitment size, so read Velocity vs Capacity alongside them. Bug direction is judged on the raw count, so small absolute changes on low-bug teams (Fusion, Web) read as movement even where the practical difference is slight.
| Team | Metric | PI 8 | PI 9 | PI 10 | Direction |
| Cloud | Completion Ratio | 51% | 42% | 50% | Improving |
| Velocity vs Capacity | 121% | 84% | 96% | Improving |
| Creep Ratio | 25% | 31% | 42% | Worsening |
| Bugs | 10 | 17 | 21 | Worsening |
| Merged MRs | 248 | 666 | 312 | Above PI 8 baseline |
| Sprint goals hit | not tracked | 87% | 100% | Improving |
| Data | Completion Ratio | 32% | 35% | 37% | Improving |
| Velocity vs Capacity | 81% | 101% | 84% | Worsening |
| Creep Ratio | 37% | 24% | 126% | Worsening |
| Bugs | 2 | 9 | 22 | Worsening |
| Merged MRs | 35 | 292 | 107 | Above PI 8 baseline |
| Sprint goals hit | not tracked | 86% | 100% (2 of 5 sprints parseable) | Unclear |
| Fusion | Completion Ratio | 39% | 52% | 64% | Improving |
| Velocity vs Capacity | 104% | 105% | 110% | Improving |
| Creep Ratio | 39% | 27% | 26% | Improving |
| Bugs | 12 | 5 | 10 | Worsening |
| Merged MRs | 29 | 128 | 99 | Above PI 8 baseline |
| Sprint goals hit | not tracked | 100% | 100% | Holding |
| Mobile | Completion Ratio | 39% | 56% | 39% | Worsening |
| Velocity vs Capacity | 99% | 133% | 106% | Worsening |
| Creep Ratio | 32% | 40% | 25% | Improving |
| Bugs | 80 | 76 | 39 | Improving |
| Merged MRs | 243 | 836 | 292 | Above PI 8 baseline |
| Sprint goals hit | not tracked | 46% | 78% | Improving |
| Web | Completion Ratio | 64% | 55% | 45% | Worsening |
| Velocity vs Capacity | 88% | 73% | 76% | Holding |
| Creep Ratio | 26% | 28% | 122% | Worsening |
| Bugs | 5 | 4 | 6 | Slightly worse |
| Merged MRs | 83 | 186 | 97 | Above PI 8 baseline |
| Sprint goals hit | not tracked | 86% | 100% (1 of 5 sprints parseable) | Unclear |
Reading the table
Fusion is the only team improving on every planning metric simultaneously. Completion, Velocity vs Capacity and Creep Ratio all moved the right way in both PI 9 and PI 10. It is also the team that changed its commitment least. Two consecutive PIs at a perfect sprint-goal rate (16 of 16, then 17 of 17) with concrete, deliverable-shaped goals. Mobile's creep fell further in PI 10 (down 15pp against Fusion's 1pp), but from a much higher base and alongside a fall in completion.
Cloud has recovered from a weak PI 9 on completion and velocity, and hit every sprint goal in PI 10, but creep has risen in both PIs in this window (25% to 31% to 42%) and release cadence missed its target.
Web and Data both look worse than they probably are. Both cut commitment aggressively and both saw creep rise sharply over the same period, and both have sprint-goal data too sparse to read. Data's Completion Ratio actually improved slightly, from 34.6% to 36.6%. Web's Velocity vs Capacity dropped from 88% to 73% between PI 8 and PI 9 but has held at 73% and 76% since, and Web shipped more releases in PI 10 than in either prior PI. Web's merge activity also finished PI 10 above its PI 8 baseline, 97 against 83.
Mobile's picture is the most internally contradictory and the most improved on the things that matter to users: half the bugs, better goal hit rate, lower creep, four release candidates against a target of three. Its Completion Ratio fell because it committed 200% of capacity.
What the metrics can and cannot tell us
Three PIs in, the measurement system has some structural problems that limit what any of the above can support.
Completion Ratio is not measuring delivery
The org-level figure has been 45.1%, 47.1%, 46.4% while its denominator swung 13% up then 22% down. Per team, the ranking it produces mostly reflects how aggressively each team commits rather than how much it delivers. Mobile resolved more points than its capacity and scores 39%; Web resolved 76% of capacity and scores 45%. For a team committing at Mobile's PI 10 level of 200% of capacity, hitting 80% completion would require resolving committed work worth 160% of capacity. The closest any team has come in three PIs is Web's 64% in PI 8, still 16 points short.
Velocity vs Capacity is the more honest throughput read and it is stable, comparable across teams, and already in the warehouse. It is worth considering whether the shared target should move to it.
Sprint goal hit rate is becoming unreadable
Parseable goals per team went from 15, 7, 16, 22 and 7 in PI 9 to 16, 2, 17, 9 and 2 in PI 10 (Cloud, Data, Fusion, Mobile, Web in order). Cloud and Fusion held steady or rose. Data, Mobile and Web fell sharply. The cause is partly syntax: Data and Web wrote goals using [] with no space between the brackets, which the warehouse regex \[.\] does not match. Cloud and Fusion, the two teams with clean 100% rates, are also the two teams writing goals in the parseable format consistently.
This creates a circularity worth naming: the teams that look best on goal achievement are the teams whose goals the system can read. That is not the same as the teams that hit their goals most often.
DORA coverage is still partial
Deployment tracking covers only Cloud and Web, so DORA deployment frequency and lead time cannot be computed for Data, Fusion or Mobile. Two of five teams is not enough to say anything organisation-wide about deployment practice, and the two that are covered sit an order of magnitude below target, so the metric is not currently discriminating between them either.
Confidence levels for this report. Verified: all story point, capacity, bug, release and deployment figures come directly from the warehouse and have been checked against source tables. Activity figures (merged MRs, commits, approvals) come from the GitLab events feed and cover 34 to 38 distinct contributors per month across the window. High confidence: Velocity vs Capacity is materially more stable than commitment across the window; Completion Ratio has not responded to changes in commitment; and the inverse commitment-creep association, where the rank ordering by commitment cut matches the rank ordering by creep rise exactly. Still a caveat on that last point: five teams and three PIs is a small sample, the magnitudes are not proportional, and causation is not established. A common cause acting on both variables would produce the same picture. Assumption: that demand on these teams is externally set and roughly constant. This is the best fit for the data but has not been tested against any external demand source such as the product roadmap, support load, or the composition of what creep actually consists of.
Questions this raises
Framed for the Product and Engineering leadership conversation. Each is a genuine open question rather than a recommendation in disguise.
-
If commitment discipline was the goal of the last two PIs, did it work?
Organisation-wide commitment fell from 952 story points to 842, and from 175% of capacity to 125%. Commitment came down on every team except Mobile, substantially on Data, Web and Cloud. Completion Ratio did not move: 45%, 47%, 46%. Creep rose on all three teams that cut substantially and fell slightly on Fusion, which barely cut, while Mobile, the one team that raised commitment, saw creep fall. The reading that best fits the evidence is that the recalibration changed the classification of work rather than the flow of it, though three PIs is a short window and other explanations have not been ruled out. The question for the room is whether that was the intended outcome, whether the exercise should continue into PI 11, and what a success criterion would look like that this data could actually detect.
Why this matters · Two PIs of effort went into this. It is worth knowing whether it moved anything before committing a third.
-
Should the shared 80% target move from Completion Ratio to Velocity vs Capacity?
Completion Ratio is bounded by commitment size, so a team can improve it by committing less rather than delivering more. No team has come close to 80% in three PIs. Velocity vs Capacity sits at 76% to 110% across the teams, is comparable between them, and measures total resolved work against available people.
Why this matters · A target nobody can hit stops functioning as a target. Team leads have this in their OKRs.
-
Where is the unplanned work actually coming from?
38% of everything completed in PI 10 was never in a sprint plan, up from 21% in PI 8. The warehouse can tell us how much creep there is but not what it is. Without knowing whether it is production support, unplanned discovery, cross-team dependencies or late requirements, there is no way to decide whether it should be reduced or simply planned for.
Why this matters · This is the largest single change across the three PIs and the one we understand least.
-
What is Fusion doing differently, and is it transferable?
Fusion improved on completion, velocity and creep simultaneously, holds a perfect two-PI goal record, and went from zero to nine final releases. It is also the team that changed its commitment least, which fits the pattern that stability beats correction. Whether this is process, product owner engagement, team size or the nature of the work is not something the metrics can answer.
Why this matters · One team is consistently doing well. Understanding why is cheaper than any process change.
-
Web's S5 recorded three metric failures for a sprint in which it shipped three releases. Is the goal wrong or the sprint?
Zero commitment in S5 meant zero committed-done for that sprint, pushed Web's PI creep ratio to 122%, and technically breached the "value delivered every sprint" goal, all while 41.5 points were resolved in the sprint and three prolaio-ui releases went out. Either the sprint was planned wrongly or the goal is written in a way that punishes a legitimate way of working.
Why this matters · Jeff's team goal is currently unfalsifiable in the useful direction. It can be missed by shipping.
-
Who owns metric pipeline health, and what would tell us the next gap is there?
The warehouse has no freshness or plausibility check on any of its feeds. A feed can silently under-report for months and the only signal is a metric that looks wrong to somebody reading it. The GitLab events feed drives commit counts, MR throughput and review depth, so a gap there moves several metrics at once and in the same direction, which is exactly the pattern most likely to be mistaken for a real trend.
Why this matters · Every agile and DORA metric in this report depends on a pipeline nobody currently monitors. A cheap row-count and contributor-count check per feed, alerting on a step change, would cover most of the risk.
-
What changed about batch size between PI 8 and PI 9, and was the PI 9 pattern better or just busier?
PI 9 carried 3.3 times the merged-MR volume of PI 8 (2,108 against 638) for roughly the same story point output per unit of capacity: Velocity vs Capacity was 100% then 99%. The median contributor went from 14 merges to 34 and back to 19. Either work was being broken into smaller pieces in PI 9, or the same work was moving through more merge requests. The first is generally a good thing and worth keeping; the second is overhead.
Why this matters · Smaller batches are one of the few levers that reliably improves flow, and we appear to have pulled it once without deciding to.
-
Should we extend deployment tracking to Data, Fusion and Mobile?
DORA metrics currently cover two of five teams. Both covered teams sit an order of magnitude below the 0.5 per day production target (Cloud 18x, Web 9x), so the metric is not currently discriminating between good and bad, it is just uniformly red. Extending coverage would either confirm this is an organisation-wide pattern or reveal variation worth acting on.
Why this matters · Deployment frequency is the DORA metric most tied to delivery outcomes, and we can see it for less than half the organisation.
Method
Scope. PI 8 (30 Oct 2025 to 3 Feb 2026), PI 9 (4 Feb to 29 Apr 2026), PI 10 (28 Apr to 5 Aug 2026). The warehouse's dim_pi ranges overlap by a day at each boundary; sprint-level assignment is used throughout, so no sprint is counted twice. Five teams: Cloud, Data, Fusion, Mobile, Web. Platform Engineering is not included as it does not run the same sprint structure.
IP sprints excluded. Every figure covers active sprints only, which is sprints 1 through 5 in each PI. The IP sprint (sprint 6) is excluded from all totals, ratios and per-day rates, following the standard convention. This means figures here will not match any report built on full-PI totals.
Definitions. Velocity is commit_done + creep_done. Completion Ratio is point-weighted at PI level, SUM(commit_done) / SUM(commitment), not an average of sprint ratios. Committed Done vs Capacity is commit_done / capacity. Velocity vs Capacity is velocity / capacity and can exceed 100%. Creep Ratio is creep / commitment. Bugs include both Bug and Anomaly issue types.
Source conventions. Cloud release cadence uses production deployment dates from the deployments table rather than GitLab tag dates, because tags can lag deployment by weeks. Deployment environments are deduplicated by team and environment before counting, to avoid double-counting rows that appear under both the jira and gitlab sources.
Activity metrics. Merged MRs, opened MRs, commits and approvals are counted from the GitLab events feed over each team's active-sprint window, so they are directly comparable across PIs. Per-day rates use active days in that window.