What Changes When a Firmware Push Is Treated as Reversible

Apple Ko
Apple Ko
October 8, 2026
📖 6 min read min read
What Changes When a Firmware Push Is Treated as Reversible

One assumption changes, and nine things downstream of it change with it. The assumption is that the push might have to be undone. It does not make the push safe — a fleet-wide update to devices nobody can reach carries risk that no amount of preparation removes. What it changes is what is available afterwards.

Under "the update will work"Under "the update might have to be undone"
Go / no-go decided by how it looksGo / no-go decided against numbers fixed before the build was cut
Previous image lives on a serverPrevious image lives on the device
Update strategy chosen for install speed and flash costUpdate strategy chosen for whether it keeps the old image
The forward path is testedForward and return both tested, from a representative field state
First cohort = whoever checks in firstFirst cohort = units that can be physically reached
Success measured by check-ins receivedReceived and missing check-ins both tracked against an expected list
Fleet firmware state = "latest"Fleet firmware state = a distribution, recorded per serial
Rollback is an incidentRollback is a release, with its own version record
Build identifier lives in the device registryBuild identifier lives on every record

What follows is the reasoning per row, in the order the rows appear.

Go / no-go. "If things look bad" is not a criterion. Under pressure an undefined threshold drifts, and "looks manageable" is available to say at the exact moment it stops being manageable. The replacement is two numbers written a week earlier by someone who was calm: what fraction of the cohort must check in within what window, and what reset count ends the push. Written early, because the person writing them at 11 p.m. on the night of the push is not the same person.

Where the previous image lives. A rollback delivered over the air is still a rollback — but it depends on the delivery path still working, and the delivery path is one of the things a bad build can take out. If recovery has to survive losing that link, a bootable previous image has to already be on the device and still be there after the install.

Update strategy. This is worth getting exactly right, because the direction is easy to invert — and worth naming which scheme is being described, because "dual-slot" covers several that behave differently.

The example here is the Legacy A/B mechanism documented by the Android Open Source Project, and nothing else. There, the new image is written to the slot that is not currently running. That is what preserves the way back: the running image is never touched during the install. The documentation puts it as a rule — no partition used by the current slot is updated as part of the update — and states that a device which fails to boot the new slot reverts to the old one and remains usable. Other seamless or dual-slot arrangements may or may not hold to that; each needs reading on its own terms.

The hazard is the small-device bootloader family, and the example here is the MCUboot configurations as documented, again not a general claim. There the strategy is a build option and only some of the options keep a way back. In overwrite-only mode the new image is copied into the primary slot and the old one is not retained: the upgrade is permanent, and no revert exists. Swap mode retains it — but swap comes in two request types, and they behave differently. A test swap boots the new image with its confirm flag unset and reverts on the next boot unless the application writes that flag after whatever self-checks it runs. A permanent swap has the flag written during the swap itself, before the application ever runs, and waits for nothing.

So the question to put to a firmware supplier is not "is there a rollback slot." It is three questions: which upgrade strategy is the bootloader built for, which swap type does the update request, and who writes the confirm flag.

Those answers establish how trial boot and confirmation work. They do not establish that a return is usable, and three further conditions decide that.

There has to be something that triggers the return. An unconfirmed test image reverts on the next boot — but only if a next boot happens. A unit that boots successfully and then cannot communicate will sit there, confirmed or not, until a watchdog, a reset or a person intervenes. Booting and being reachable are different states, and only one of them is visible from the dashboard.

The security policy has to permit it. Signature verification, anti-rollback counters and monotonic security versions exist precisely to stop an older image running, and they do not make an exception because the older image is yours. The recovery path has to be tested inside whatever the policy allows; disabling downgrade protection to make rollback work is not a fix.

And the previous image still has to be bootable and still has to cope with the state the newer build left behind — a migrated configuration, a rewritten record format, a counter that changed meaning.

Both paths tested, from a used unit. This is an easy item to skip, and the reason is understandable: it feels like testing something that already works. It is not. Rolling back exercises a path the forward test never touches — the bootloader selecting the other slot, the configuration surviving or not surviving, the device coming up on a build that has already been superseded in the fleet record. And a unit straight off the bench rolls back from a state no field device is in. The accumulated state is the part that breaks.

First cohort. If the remote recovery fails, the remaining recovery is a site visit. The update reaches the whole cohort at network speed; a van does not. That is what makes the first cohort worth choosing rather than accepting: pick units scattered across three countries because they happened to check in first, and the rollback plan has quietly become a travel plan.

Counting both directions. Check-in count going up is evidence that the units still reporting are still reporting. It is circular, and it is the number on every dashboard. Silence is not the opposite and not proof either — coverage, schedules, power state and a collection outage all produce it. What is worth having is an expected list, by serial and build, and both columns against it: which reported, which did not. Missing reports are a reason to stop and look, not a finding about the firmware.

Fleet firmware state as a distribution. After a staged push and a partial rollback, one programme is served by two builds at once. That is not a transient state to be tidied up later; it is the state, for as long as it takes to reach the rest of the fleet. Per-serial firmware history lets an investigation compare failures by build and by production lot. It does not separate the two on its own — if one lot happened to take the update first, build and lot coincide, and separating them needs units where they do not.

Rollback as a release. Same version record, same change log, same list of affected serials. An unrecorded rollback adds a second unknown on top of the original fault: a fleet whose firmware distribution nobody can state. The configuration needs its own line here, because reverting a build does not necessarily revert what the build wrote.

Build identifier on every record. The registry says what a device is running now. The record has to say what it was running then. The difference stays invisible right up until the week somebody tries to reconcile two months of a mixed fleet — a field that changed name, a timestamp that moved from transmit time to acquisition time, a counter whose blanking parameter shifted so the same physical event produces a different number depending on which build recorded it. None of that arrives as an error. It arrives as data that looks fine and is not comparable.

Nine rows, and none of them caps what a bad push costs. They limit how many devices are exposed first and how many options are left afterwards, which is a smaller claim and the only one on offer.

Six questions that separate a plan from an intention:

Sources for the update-strategy section: Android Open Source Project, A/B (seamless) system updates; MCUboot, design documentation — sections on overwrite-only and swap image-upgrade modes, and on swap types.

Tags
#Firmware #fleet-operations #OTA

Share This Article

If you found this article helpful, please share it with your network

Apple Ko

About Apple Ko