Category: Reports People Trust

Operational reporting, measurement, and why numbers lose the room.

  • We Were Developing Great Data Engineers

    In an analytics organization

    During the COVID pandemic the news showed carnage at the grocery store. Empty shelves down the paper aisle and shopping carts full of toilet paper. People were panic buying paper towels and toilet paper while the stores were trying to reassure the public in vain that the product was on its way.

    I worked in supply chain in manufacturing. My boss told me that we knew we had a thousand pallets of toilet paper in our supply chain, but no idea where. The best we could do was reassure people that it was coming.

    Something had to change, and it resulted in a supply chain analytics organization. I came over that year from a creamery, where I’d been a continuous improvement leader, and I stayed with SCA until 2026.

    The island

    We started as an island, a scrappy team air-dropped in to fill a massive gap. We built Tableau and Power BI dashboards, the data engineering and the data warehousing servers behind them. The infrastructure was there but immature. Our mandate was to get the business represented outside of manual spreadsheets, and to get it done yesterday. 2020 was never going to happen again.

    For the business segment I joined, transportation was transitioning from a cost center to a profit center with its own internal financial statements, and a profit center needs reporting that a cost center never had.

    Nobody could quite place where our team was. Operations thought we were the part of technology that spoke English instead of computer. Technology saw us as the business. The data science organization saw us as the business too, the ones who could sometimes convince technology to change their mind. A continuous bridge department is always going to be like that, but as I’ve argued before, strategy falls apart without a clear story.

    Great data engineers

    A good bridge department is movement. Someone could use it to get from one department to another and carry what they’d learned with them. We had that as a nebulous future state concept. People retooling for other roles, or gaining perspectives to bring back to their home department. The recurring cast was a different breed, generalists.

    And people did cross the bridge from us to other departments, the data science organization even, but they only went as data engineers. They were good at it. They became good at it on our team.

    It was a clear sign something was wrong. A bad system will defeat a good person every time. Whatever our org chart said, our team was developing great data engineers, because we were still fundamentally a data engineering team.

    Rebasing in credibility

    As an organization founded out of a crisis of missing information, our strategy was credibility. With leadership that they had numbers to run the business on, with the field that the reports accurately represented what they were doing, with technology that we wouldn’t expose massive risk. I didn’t set that strategy. I helped and influenced it, and I led the integration work that came out of it.

    Transportation was the first team to move its data engineering over to the technology organization. I led it, and guided the data scientists on our team to move data the SQL sandbox we’d always used to the dedicated team. We learned technology’s governance, their production releases, their release schedules for data engineering.

    After strategy sessions with our subteam, we decided that, day to day, the work was meant to be based on 70% operations, 20% technology and 10% data science. My manager wanted data science to grow to 20 or 30% eventually. Lean into cross pollination without losing who we are.

    When the team’s composition would change, I would recruit members with that percentage in mind. One was a CS grad, the other came from a car shop. My job was integration. Our mandate was building data understanding into the organization. Each recruitment, each project, judged against that standard.

    As the data organization matured, the transportation subteam took the lead in moving specializations off to technology or to the data science organization. We kept some data science projects on purpose, for keeping current on a common language (science reviews and governance), or for proving out ideas that didn’t have enough concrete evidence to be worth tens of millions of dollars a year.

    The integration was the part I liked. Strategic alignment, org change, the relationship on the other side of every handoff. A continuous improvement background follows you, and working on a team’s credibility was the hook for me. You don’t get far in technology without strong relationships.

    Leaning into operations

    Most of our work was for operations, and that’s where we had to keep our credibility. SCA was supposed to stay close to the ground and never turn into an ivory tower.

    I pushed for relationships with the field and pushed for a power users space in Power BI, and we got it. People we’d identified in operations could build and deploy their own models and semantic models there. We learned what they were using data for and folded those uses into our own work. They built credibility with their peers. And we got to poach them.

    Two operations people came to SCA. One of them was my hire. He’d been building reporting on his own, and he filled a gap on our team. I had a similar strategy in manufacturing. Operations employees on rotating office assignments for continuous improvement, learning to ask questions from a different perspective, relationships on both sides, moving back to the field armed with tools to effect real change.

    And strengthening the science

    The data science organization was a subsidiary with its own formal process, check-ins and science reviews. The future vision for SCA was tighter integration with the science organization. I was the point for integration on our side of the fence. The first project we ran their way was a survey of the drivers delivering into our distribution centers. The existing dashboard in Qualtrics presented numbers but did nothing to inform the sites what to do.

    SCA volunteered to take it over, building the data pipeline from the Qualtrics API into our Databricks data lake, modeling it into Power BI. After analysis we recommended restructuring the survey from 15 questions to 3, with a free form text box and a language model trained to read the comments, and people decided what to act on.

    NPS went from 15 to 31 while the toolkit was in use. The organizational design we’d pictured from the start was coming into shape.

    Reconciliation

    I got us our own Unity Catalog domain for analytics products and data staging. That meant the Data Architecture Committee, and getting technology, engineering, governance, data science and the business to agree on a setup that followed the data science organization’s standards instead of technology’s defaults. It had stalled before. It went through.

    The data mesh was built to support analytics, so analytics could be purely analytics. We were the people talking to the business, building the reporting and the agents. (The data teams did some of that too. It’s a mesh.) I worked with the data teams to move the pipelines I’d designed, the ones the org depended on, over to them. The next step was moving everything into the data science organization’s data workspaces. I was the owner on our side, working with a director on theirs.

    When I left, two people had come to SCA from operations. A few had gone to the data science organization as data engineers, but we had a clear idea of who we were and a clear game plan on how to get there.

  • In-Spec Isn’t Good Enough

    Note: the numbers have been changed for privacy.

    Every quality check was passing. Customers were calling in about a chemical taste in the water anyway. Both of those things were true at the same time.

    I was a continuous improvement leader at a manufacturing plant that produced spring water and drinking water alongside dairy products. We had an increase in customer comment frequency for the spring water. They complained about a chemical taste.

    In the prior year, we’d get one to two complaints in a month. Between August and September, we received half a dozen. This was occurring across shifts and production dates. We checked the quality processing system. All quality tests were green.

    Was it real?

    Before chasing causes, you have to make sure it’s not noise. The main lesson from Six Sigma is that variation happens. If a packaging line is meant to fill to 500 grams, it’s not going to be 500.00 grams in each package. Some will have 490, some 510. Your task is to understand what amount of variation is acceptable compared to how much variation reduction is possible. Process specs vs capability.

    That’s to say, a few extra complaints in a month doesn’t necessarily mean anything meaningfully changed. Clustering illusion has its own page on wikipedia for a reason.

    We checked the control chart. Control charts attempt to quantify the “yeah, I know this product can vary between 480 and 520 grams, but how do I know if there’s an underlying shift.” If you know that 99% of the time that a process will behave one way, then finding something rare indicates you’re really lucky, or something has changed. The underlying principle is that you’re not special, so go with the more likely explanation, and conclude that something has changed. The system we followed was Nelson Rules, a collection of 8 “this data is sufficiently weird” rules to decide if the process had changed.

    There are different types of control charts. Counts, proportions, averages. For tracking rare events like customer complaints, we used a G-chart instead of a standard control chart. A G-chart tracks the time between events, instead of counting the number of events in an interval. When you’re getting one or two complaints a month, a count-based chart needs many months of data before it can detect a shift. A G-chart picks it up faster because the spacing between events is more sensitive to changes than the count.

    Each event gives you more information. It uses some basic probabilities to give you an estimate of what “99% of the time” looks like. In the below chart, based off the average time being 34 days, 99% of the time, you’d see next comment within 196 days. The below one is simulated, but it showed the same trend. Nelson rule 2 violation. Something had changed.

    The spec gap

    We decided to focus on chlorine levels in the finished product, since that was the main cause of the chemical taste associated with tap water. The current process involves testing for chlorine at the final water filter, with an upper spec limit of 0.50 milligrams per liter. The product release records for the past four months indicated that there were no products released with a level above the upper spec limits.

    So we conducted a taste threshold experiment, using the guidelines provided in (American Public Health Association, 1976). Eight people, five concentration levels, blanks mixed in to keep them honest.

    We’d done organoleptics work before, on an orange juice off-flavor project, and one thing we’d learned was that taste detection thresholds follow a lognormal distribution. I’d spent time at a university library digging into why, and the short version is that sensory receptor triggers are exponential, which makes the distribution right-skewed. You can’t just average everyone’s detection point and call it a threshold. You need the geometric mean, which pulls lower than a simple average would. That matters when you’re setting a spec limit against it.

    The experiment revealed an average detection level of 0.152 milligrams of free chlorine per liter, well below the upper spec limit. Comparing this to the final water filter results revealed multiple samples released at a level higher than the average threshold detection level.

    The spec allowed more than three times the amount a person could detect. Product was passing every quality check while containing chlorine that customers could plainly taste. The spec was the problem.

    Number line showing chlorine concentration from 0 to 0.55 mg/L. The taste detection threshold at 0.152 mg/L and new spec at 0.148 mg/L are nearly identical on the left. The old spec limit at 0.50 mg/L is far to the right, with the entire region between labeled as in spec but tasteable.

    We adjusted the chlorine standard down to 0.148 mg/L, from 0.50 mg/L.

    What changed

    Ultimately we discovered that a failing diversion valve was letting small amounts of city water into the line while we were bottling spring water.

    The fact there was a failing valve was almost to the side of the other lesson from the investigation: the spec wasn’t good enough. Two main things went into place: the tightened chlorine spec (0.148 mg/L) and a control plan that monitored both. Customer complaints dropped from six in the investigation period to one over the following four months, a trend that has held in the years since.