Fleet acceptance goes wrong in four ways, and three of them happen before anyone reads a temperature.
Illustrative, and deliberately ordinary: two hundred trackers land on a Tuesday, the pilot is scheduled for the following Monday, and somewhere in between someone is supposed to establish that these particular units are fit to carry the programme. When nobody owns that week, the checks get improvised — and the improvisation has a pattern.
One. Nobody wrote down what a good unit looks like.
Units get unboxed, powered up, and watched until they appear on a map. Green dots. Tick.
But "it reported" is not an acceptance criterion. It is an observation made under one unstated configuration, at one temperature, against no reference. A unit that reports every fifteen minutes on a bench at 22 °C has demonstrated that it reports every fifteen minutes on a bench at 22 °C.
Here is the working template used in the rest of this article. It is not an industry standard and not a complete acceptance procedure — it is four fields that make a result mean something to someone who was not in the room: the configuration under test, the environment, the reference the reading is compared against, and the limit. A real procedure has to say how the comparison is made as well.
This sounds pedantic until the second batch arrives and someone asks whether it is as good as the first. Without a recorded configuration, environment, reference and limit, that question cannot be settled from the record. It gets settled by whoever is more confident.
Two. The sample is whatever was easiest to reach.
Twenty units out of two hundred sounds like a reasonable check. Then you find out which twenty: the top layer of the first pallet, because the rest were shrink-wrapped.
That is evidence about the top layer of the first pallet. If the shipment contains more than one production lot — and at two hundred units it very well might — a convenience sample can sit entirely inside one of them. What it says about that lot is limited too: twenty units taken because they were reachable are not twenty units taken at random from the lot, so they constrain the accessible part of it and not the lot as a whole. The other lots are simply unchecked, which is a different situation from having checked them and found nothing.
A bigger sample helps, for the ordinary statistical reason. But it does not fix this, because a hundred units off the same accessible pallet face are still a hundred units off one pallet face. What fixes it is knowing what the sample is a sample of. Ask which lots are in the delivery, and draw across them.
Getting lot information is its own small negotiation. It is not always on the carton, it is not always in the packing list, and asking for it after delivery tends to produce a spreadsheet assembled from memory. The reliable version is to ask for it before the units ship, as a line in the order rather than a favour afterwards. Even then the answer can be partial: an assembler may know its own build lots without knowing the component lots underneath them, and those are sometimes the ones that explain a clustered failure. Partial is still worth having. A delivery split across three build lots, checked across all three, is a different piece of evidence from twenty units drawn from wherever the shrink wrap was already open. And if the answer comes back as a single lot for the whole delivery, that is useful too — not because it predicts failure, but because it says nothing in the fleet is insulated by having come from somewhere else. A lot-level problem, if there is one, would be a fleet-level exposure.
Three. Acceptance happens in the office.
Warm room, full signal, mains power, the unit sitting on a desk with nothing around it.
The programme runs somewhere else: inside a steel container, under a metallised liner, at −20 °C, on a cell that has been sitting in a warehouse for four months. Every one of those changes something the office test did not exercise.
The bench test is not wasted — it is a low-cost way to catch units that are dead or running the wrong build, provided the check actually reads the firmware identity back rather than watching for a green dot. It should stay. What it cannot do is stand alone. At least one part of the acceptance should happen in a condition the programme will actually meet: a freezer for an hour, a unit buried in a loaded pallet rather than resting on top of it. Trading the twenty-first clean pass for one programme-relevant condition is not automatically stronger evidence — it is evidence about something the bench sample cannot reach at all.
Which condition to pick is worth five minutes of thought rather than defaulting to cold. Two things decide it, not one: how often the programme meets that condition, and what it costs when a unit fails there. A condition that occurs on every lane and produces a recoverable nuisance may rank below one that occurs on a tenth of them and loses the shipment. Frequency alone will send you to the wrong freezer.
Within that, prefer the condition the programme meets first and most often. For a lane that loads at an ambient dock and runs chilled, the interesting moment is the transition, not the steady state. For a device that will spend its life inside a metallised liner, the interesting condition is the liner, and a freezer tells you very little about it. The rule of thumb is to reproduce whatever the office bench most obviously is not.
Here is the part that is genuinely hard to judge, and worth admitting: a one-hour freezer soak is not a qualification. It does not establish cold performance over a season. What it does is catch the units that are obviously wrong before two hundred of them go into a programme. Those are different jobs, and conflating them is its own failure mode.
Four. The record does not survive the week.
Somebody checks the units carefully, writes the results on a sheet, and the sheet becomes a tick in a project tracker. Accepted, 200 of 200.
That line is doing two jobs and admitting to neither. Twenty units had the detailed check. The other hundred and eighty had, at most, a basic power-up. A decision to accept the delivery is a legitimate thing to record — it just belongs in a different field from a claim that two hundred units individually passed a test, and the tracker collapses them into one tick.
Write them apart. Detailed check: 20 units, results attached. Basic check: 200 units, criterion named. Delivery disposition: accepted. Three facts, three fields, and none of them pretending to be another.
Eighteen months later a subset of that fleet starts behaving oddly. The only surviving fact is that the delivery was accepted. Which units, built when, running what firmware, checked against what, by whom, in what conditions — gone. And the person who ran the week may well have moved on, which is not a criticism of anyone, just how eighteen months works.
What matters is not the pass. It is the values behind it, tied to identifiers that still mean something later.
| Record per unit | What it does later |
|---|---|
| Serial number | Links a field complaint back to a specific acceptance record |
| Production lot, where it can be obtained | Supports comparison across lots and investigation of possible lot-level causes |
| Firmware build and configuration at test | Behaviour can change with either, independently of the other |
| Measured value and unit, criterion revision, and pass/fail where one applies | Makes a distribution across units visible; a column of "pass" alone hides it, and a value without its criterion cannot be re-judged |
| Which check this unit received | Separates the detailed sample from the units that had only a basic check |
| Reference and environment | Preserves part of the context the value has to be compared in |
| Procedure and criterion, by revision | The method is in the document, not in the memory of whoever ran it |
| Date and who ran it | Says when the check happened and who to ask, while there is still someone to ask |
The record format is the cheap part — a spreadsheet with eight columns instead of one. Getting the lot mapping, reaching units inside the packaging and arranging one realistic condition still cost time, access and an owner.
The three failures that arrive before anyone reads a temperature are the criterion, the sample and the conditions. The fourth arrives later and quietly, and it is the one that takes the other three down with it: a retest eighteen months on can establish what the units do now, but nothing establishes what they did during a week that was never written down.
Acceptance testing does not prove the fleet is good. It establishes what was checked, on which units, under what conditions — so that a problem eighteen months out has something to be compared against.