How the Stabilization sub-phase works
Checkpoints a GC digital service has to pass shows where Stabilization comes in the whole journey, checkpoint by checkpoint.
The first weeks are firefighting. Stabilization is the work of reaching the day when there are no fires left to put out.
Two things change on launch day
- The volume is real. Everyone the service is for can reach it, so what arrives is whatever the real world sends, not what a research session arranged.
- What people do now counts. In Create, somebody filled in a form and nothing happened next, because nothing was meant to. Now an application has to be received, assessed by a person, decided, recorded where it can be found again, and answered. If a grant was promised, the money has to reach somebody's account.
So this is the first time anyone finds out whether the whole service holds, and not just the software: whether the team handling the applications is large enough for the volume arriving, and whether anyone has been left out of the arrangement altogether.
Before you start Stabilization
THE MAKE-OR-BREAK QUESTION
Agree how Stabilization ends before it begins
Heightened support feels safe, and that is its danger: a window with no exit test never closes, and the constant patching hides the weaknesses it should be fixing. Before launch day, agree what steady will look like in numbers, who decides the window is over, and what happens to anything still open when it closes. The decision belongs to the business owner of the application, made on the dashboard's evidence.
Running your service in Stabilization
The team you need
Beta's team shrinks into the running shape. The minimum roles (one person can hold more than one):
- Operations keeps the service up and patched, and releases the fixes.
- Support lead helps people through, and reports what the calls are saying.
- Supplier or in-house developers fix defects while the warranty or the assignment lasts, and hand the knowledge over.
- Business owner of the application owns the decision that Stabilization is over.
Keep the knowledge as the people change: runbooks, known errors, and decisions written down as they are learned. Stabilization is short: a few weeks to a couple of months is typical.
CAUTION
When Stabilization goes wrong
Launch was treated as the finish line, so nobody owns the running service.
The old way is switched off while the new service is still surprising people, so there is no way back.
Support is overwhelmed, and what it hears never reaches the team.
The people who built it were gone on launch day: no warranty, no handover.
The heightened support never ends, and constant patching hides the weaknesses it should fix.
A REAL EXAMPLE
On launch day, there was no way back
When Phoenix went live, the old pay system was switched off, and hundreds of the compensation advisors who understood pay had already been let go. So when the first pay runs came out wrong, there was no old system to fall back on, and almost nobody left who could fix a pay file by hand.
Every fault reached real paycheques at full volume, and the queue of broken pay files grew faster than anyone could clear it. Stabilization's on-ramp exists because of launches like that one: the old way still running, on a dated retirement plan, and the people who understand the service still reachable.
How you know Stabilization is finished
Stabilization is finished when the exit test agreed before launch is met, and the service has become boring: incidents are rare and routine, support volume has settled while use keeps growing, performance holds at full load, and the running team resolves and escalates without the people who built it.
What is still broken is owned and accepted
Boring does not mean perfect. The exit test tolerates open faults, provided each one is diagnosed, has a named owner, and stays open because someone decided it could.
The rule for the open list was agreed before launch, as part of the exit test. Apply it at the close: whatever is accepted moves onto the running team's known-errors list, and whoever pays for its eventual fix is named.
The build team is closed out, and the knowledge is kept
For a supplier build, accepting the open list closes out the warranty, the after-launch period when the supplier fixes defects at no extra charge. Each remaining defect is either fixed under it or accepted with a named owner. Closing it settles who pays from then on: the free fixing ends and the support terms take over.
A service built in-house has no warranty to close. The developers' assignment winds down once the running team handles incidents without them.
The knowledge stays with the running team. The runbook and the known-errors list are theirs by the close, and recent incidents are the proof: handled without a call to the people who built it.
when there is real new capability waiting to be built.
when the service already has the scope it needs. Not every service grows, and going straight to the long steady state is a normal path.
Back toward a rebuild,
when weeks of fixing cannot settle the service and the fault is deeper than patches. This is rare, and it is a Create-sized decision.
The window closes on evidence. Before you move on, have ready:
The official instruments in Stabilization
Everything official that has something happening to it during Stabilization, and what that something is. The tag says what stage the instrument reaches here, not that it is finished.
Placing an instrument in a sub-phase is this guide's own editorial choice, anchored where possible on a real deadline in the instrument itself. The full detail, including who does the work and what the business owner personally does, is in the full instruments table.
- Submit
- Sent, filed, registered or published where the rule says.
- Keep current
- The service changed, or the clock came round. Re-run it, re-test it, or refresh the record.
What the tags mean
Every service
- Business impact analysis (BIA)AssessmentKeep current
The exercise that decides how critical the service is, and produces four numbers with it: maximum allowable downtime, minimum service level, recovery time objective and recovery point objective.
Measured against real incidents for the first time. Editorial placement, argued from what the sub-phase already does.
The government-wide register of what services exist, who they serve, how digital they are, and how much volume they handle.
Registered once the service is live. Easy to forget, because nobody chases it.
- Application Portfolio Management (APM)RegisterSubmit
The register of the applications behind the services, rated for business value, technical condition, support cost and criticality, and sorted into tolerate, innovate, mitigate or eliminate.
Rated once live, including its criticality.
The service team's own recovery arrangements: how this system gets back up, in what order its parts are restored, and proof from testing that the restore works.
The first real incidents test whether the restore works under pressure and whether the recovery target is achievable.
The duty to have a way of spotting, containing and reporting a cyber incident before one happens, and to report it up the government-wide chain when it does.
This is when it gets used. Incidents are reported through the departmental route, not sat on.
Only if it applies
- Business continuity plan (BCP)PlanKeep current
The written arrangements for keeping a critical service delivering at a minimum acceptable level during a disruption, and recovering it afterwards.
A real incident is a live test of arrangements that were written on paper.
Applies when Only if the business impact analysis marks the service critical, meaning disruption would cause a high or very high degree of injury. One department reads that as needing to recover to minimum service levels within 72 hours.
- Material privacy breach reportFilingSubmit
The report a department must make when personal information is lost, accessed or disclosed in a way that could reasonably be expected to cause serious injury.
If it happens, it is reported. Editorial placement: the duty is triggered by the event, not by a phase.
Applies when Only when a breach involving personal information is judged material, on sensitivity of the information, number of people affected, and whether it is a systemic problem. A cyber incident touching personal information can trigger both this and the cyber reporting route at once.
The written statement of what good this project is supposed to do, and the later report confirming what was actually delivered and whether the promised benefits arrived.
The funded project ends with a close-out: what was delivered, what is left of the budget, and the delivery record.
Applies above Universal for anything that counts as a project under the projects and programmes directive, with no dollar trigger. Baseline reporting to the Office of the Comptroller General starts at $25 million.
Assumptions this page makes
You are already working to the Government of Canada Digital Standards, design with users, iterate and improve frequently, work in the open, use open standards, address security and privacy, build in accessibility, empower staff, be good data stewards, design ethical services, and collaborate widely, and to the law on privacy, security, official languages, and accessibility. The standards say how the government works in the digital world. The six Government of Canada digital competencies say what every public servant has to be able to do to work that way, and the team page covers them. This guide builds on those.