Category: Reports People Trust

Operational reporting, measurement, and why numbers lose the room.

  • In-Spec Isn’t Good Enough

    Note: the numbers have been changed for privacy.

    Every quality check was passing. Customers were calling in about a chemical taste in the water anyway. Both of those things were true at the same time.

    I was a continuous improvement leader at a manufacturing plant that produced spring water and drinking water alongside dairy products. We had an increase in customer comment frequency for the spring water. They complained about a chemical taste.

    In the prior year, we’d get one to two complaints in a month. Between August and September, we received half a dozen. This was occurring across shifts and production dates. We checked the quality processing system. All quality tests were green.

    Was it real?

    Before chasing causes, you have to make sure it’s not noise. The main lesson from Six Sigma is that variation happens. If a packaging line is meant to fill to 500 grams, it’s not going to be 500.00 grams in each package. Some will have 490, some 510. Your task is to understand what amount of variation is acceptable compared to how much variation reduction is possible. Process specs vs capability.

    That’s to say, a few extra complaints in a month doesn’t necessarily mean anything meaningfully changed. Clustering illusion has its own page on wikipedia for a reason.

    We checked the control chart. Control charts attempt to quantify the “yeah, I know this product can vary between 480 and 520 grams, but how do I know if there’s an underlying shift.” If you know that 99% of the time that a process will behave one way, then finding something rare indicates you’re really lucky, or something has changed. The underlying principle is that you’re not special, so go with the more likely explanation, and conclude that something has changed. The system we followed was Nelson Rules, a collection of 8 “this data is sufficiently weird” rules to decide if the process had changed.

    There are different types of control charts. Counts, proportions, averages. For tracking rare events like customer complaints, we used a G-chart instead of a standard control chart. A G-chart tracks the time between events, instead of counting the number of events in an interval. When you’re getting one or two complaints a month, a count-based chart needs many months of data before it can detect a shift. A G-chart picks it up faster because the spacing between events is more sensitive to changes than the count.

    Each event gives you more information. It uses some basic probabilities to give you an estimate of what “99% of the time” looks like. In the below chart, based off the average time being 34 days, 99% of the time, you’d see next comment within 196 days. The below one is simulated, but it showed the same trend. Nelson rule 2 violation. Something had changed.

    The spec gap

    We decided to focus on chlorine levels in the finished product, since that was the main cause of the chemical taste associated with tap water. The current process involves testing for chlorine at the final water filter, with an upper spec limit of 0.50 milligrams per liter. The product release records for the past four months indicated that there were no products released with a level above the upper spec limits.

    So we conducted a taste threshold experiment, using the guidelines provided in (American Public Health Association, 1976). Eight people, five concentration levels, blanks mixed in to keep them honest.

    We’d done organoleptics work before, on an orange juice off-flavor project, and one thing we’d learned was that taste detection thresholds follow a lognormal distribution. I’d spent time at a university library digging into why, and the short version is that sensory receptor triggers are exponential, which makes the distribution right-skewed. You can’t just average everyone’s detection point and call it a threshold. You need the geometric mean, which pulls lower than a simple average would. That matters when you’re setting a spec limit against it.

    The experiment revealed an average detection level of 0.152 milligrams of free chlorine per liter, well below the upper spec limit. Comparing this to the final water filter results revealed multiple samples released at a level higher than the average threshold detection level.

    The spec allowed more than three times the amount a person could detect. Product was passing every quality check while containing chlorine that customers could plainly taste. The spec was the problem.

    Number line showing chlorine concentration from 0 to 0.55 mg/L. The taste detection threshold at 0.152 mg/L and new spec at 0.148 mg/L are nearly identical on the left. The old spec limit at 0.50 mg/L is far to the right, with the entire region between labeled as in spec but tasteable.

    We adjusted the chlorine standard down to 0.148 mg/L, from 0.50 mg/L.

    What changed

    Ultimately we discovered that a failing diversion valve was letting small amounts of city water into the line while we were bottling spring water.

    The fact there was a failing valve was almost to the side of the other lesson from the investigation: the spec wasn’t good enough. Two main things went into place: the tightened chlorine spec (0.148 mg/L) and a control plan that monitored both. Customer complaints dropped from six in the investigation period to one over the following four months, a trend that has held in the years since.