Which of 590 process sensors actually separate failing runs from passing ones?
A semiconductor plant monitors 590 process signals. When a lot fails, an engineer has to decide where to look. A conventional screen says 80 of those signals separate failing runs from passing ones. After correcting for the number of comparisons, grouping near-duplicates, and checking that each effect survives the process changing mid-record, nine do.
Of 590 signals, 464 are testable. The rest never vary, or are too sparsely observed in one outcome group to compare. Testing each against outcome at the usual 5% threshold flags 80. But 464 tests at that threshold are expected to flag around 23 by chance alone, with nothing real present. Each bar below removes a different way the count overstates itself: multiplicity, then redundancy between near-identical channels, then effects that do not reproduce across the record. The gap between 80 and 9 is the finding, not the 9.
Figure 1. Signal counts at five successive filters, from all sensors tested down to those that are significant, independent of near-duplicates, and stable across both halves of the record.
This dataset carries no cost information, so no saving is claimed here. The two counts above are ours; the rates below are yours. Set them to your fab and the arithmetic follows.
Read this as a scoping tool, not a business case. It assumes review time scales with the number of signals checked, that every signal on the list is actually reviewed, and that the nine-signal list finds root cause as often as the eighty-signal one. The third assumption is the load-bearing one and this dataset cannot test it. A fab holding both lists against its own excursion records for one quarter could.
The 33 signals surviving correction, ordered by effect size. Colour marks whether the effect reproduces when the record is split in half. Ten do. Twenty-one appear only when the halves are pooled, and two cannot be assessed. Two patterns are worth noting: the six largest effects are all stable ones, and almost every asterisked signal, meaning one with a near-duplicate in the set, falls in the unstable group. The three filters are not independent objections. They agree about which signals to trust.
Figure 2. Cohen's d for each sensor passing Benjamini–Hochberg at FDR 0.10, coloured by whether the effect holds across both time halves. An asterisk marks a sensor sharing a redundancy group with at least one other.
Every signal lands in exactly one bucket. Two of these are recommendations to monitor less, and both rest only on structural facts: a signal taking one value across all 1,567 runs, or recorded on under 10% of them. Neither is built from a signal failing a significance test. That restraint is deliberate. A false entry on the shortlist costs an engineer an hour. A signal wrongly cut from monitoring stops producing the data that would reveal the mistake. The errors are not symmetric, so the irreversible one is never made on inferential grounds.
Figure 3. Every sensor assigned to one bucket. The two-sensor shortlist_low_coverage segment is drawn to scale, and is nearly invisible by design.
The failure rate is not stationary. It runs between 13% and 23% through mid-August, then between 1% and 5% from September. Because failures concentrate early, a signal that merely drifted over these 90 days will separate failing from passing runs without having anything to do with failure. That is why every effect is rechecked within each half of the record. The decline is also consistent with someone having found and fixed a process problem, in which case some early-only signals were real and are now historical.
Figure 4. Weekly failure rate over the 14-week window, against the overall rate. The shaded band is the early half of the median split used in the stability check.
Effect size is not separation. These are the strongest signals in the set, standardised to share an axis, and the passing and failing distributions still overlap heavily. A Cohen's d of 0.6 is a shift of half a standard deviation. Enough to rank a signal as worth checking first, nowhere near enough to classify a run from it. This analysis produces an ordering for human attention, not a detector.
Figure 5. Per-sensor pass and fail distributions for up to eight shortlist sensors with the largest effect sizes. Values are standardised within each sensor (z-score) so sensors on different scales can share an axis; this is a display transform only.
| Action | Owner | Needs first | Reversible |
|---|---|---|---|
| Check the 33 shortlisted signals first on an excursion, leading with the ten that hold across both halves of the record | Process engineering | Nothing. Usable as a work instruction today | Yes |
| Investigate why sensors 112 and 113 are recorded on a minority of runs, before any change to their measurement families | Process engineering | ID-to-measurement mapping for two families | Yes |
| Review the 122 signals that have never varied, for sampling frequency or retirement | Metrology, equipment engineering | ID-to-measurement mapping; confirmation the constancy is not an acquisition fault | No |
| Review low-coverage signals at family level. The 50 fall into 8 groups recorded on identical runs, so this is 8 decisions and not 50 | Metrology, process engineering | Confirmation that partial coverage is by design rather than by fault | No |
| Change nothing for the remaining 385. They are measured, they vary, and nothing here connects them to outcome | — | — | n/a |
The last column is the reason the two reduction rows carry heavier prerequisites. A wrong entry on the shortlist costs an engineer an hour and corrects itself. A signal wrongly cut from monitoring stops producing the data that would reveal the error. Both reduction recommendations therefore rest only on structural facts about the record, never on a signal failing a significance test.
Source data: UCI SECOM, 1567 runs over 2008-07-19 to 2008-10-17, 924530 readings. Per-sensor Welch's t (fail vs pass), Benjamini–Hochberg at FDR 0.1 (with 0.05 also reported); near-duplicates grouped at |r| ≥ 0.95; stability checked by splitting at the median run timestamp. Reduced-monitoring buckets are assigned on structural grounds only, never on a significance result. Full pipeline and SQL: repository.