We Were Developing Great Data Engineers

In an analytics organization

During the COVID pandemic the news showed carnage at the grocery store. Empty shelves down the paper aisle and shopping carts full of toilet paper. People were panic buying paper towels and toilet paper while the stores were trying to reassure the public in vain that the product was on its way.

I worked in supply chain in manufacturing. My boss told me that we knew we had a thousand pallets of toilet paper in our supply chain, but no idea where. The best we could do was reassure people that it was coming.

Something had to change, and it resulted in a supply chain analytics organization. I came over that year from a creamery, where I’d been a continuous improvement leader, and I stayed with SCA until 2026.

The island

We started as an island, a scrappy team air-dropped in to fill a massive gap. We built Tableau and Power BI dashboards, the data engineering and the data warehousing servers behind them. The infrastructure was there but immature. Our mandate was to get the business represented outside of manual spreadsheets, and to get it done yesterday. 2020 was never going to happen again.

For the business segment I joined, transportation was transitioning from a cost center to a profit center with its own internal financial statements, and a profit center needs reporting that a cost center never had.

Nobody could quite place where our team was. Operations thought we were the part of technology that spoke English instead of computer. Technology saw us as the business. The data science organization saw us as the business too, the ones who could sometimes convince technology to change their mind. A continuous bridge department is always going to be like that, but as I’ve argued before, strategy falls apart without a clear story.

Great data engineers

A good bridge department is movement. Someone could use it to get from one department to another and carry what they’d learned with them. We had that as a nebulous future state concept. People retooling for other roles, or gaining perspectives to bring back to their home department. The recurring cast was a different breed, generalists.

And people did cross the bridge from us to other departments, the data science organization even, but they only went as data engineers. They were good at it. They became good at it on our team.

It was a clear sign something was wrong. A bad system will defeat a good person every time. Whatever our org chart said, our team was developing great data engineers, because we were still fundamentally a data engineering team.

Rebasing in credibility

As an organization founded out of a crisis of missing information, our strategy was credibility. With leadership that they had numbers to run the business on, with the field that the reports accurately represented what they were doing, with technology that we wouldn’t expose massive risk. I didn’t set that strategy. I helped and influenced it, and I led the integration work that came out of it.

Transportation was the first team to move its data engineering over to the technology organization. I led it, and guided the data scientists on our team to move data the SQL sandbox we’d always used and into the dedicated team. We learned technology’s governance, their production releases, their release schedules for data engineering.

After strategy sessions with our subteam, we decided that, day to day, the work was meant to be based on 70% operations, 20% technology and 10% data science. My manager wanted data science to grow to 20 or 30% eventually. Lean into cross pollination without losing who we are.

When the team’s composition would change, I would recruit members with that percentage in mind. One was a CS grad, the other came from a car shop. My job was integration. Our mandate was building data understanding into the organization. Each recruitment, each project, judged against that standard.

As the data organization matured, the transportation subteam took the lead in moving specializations off to technology or to the data science organization. We kept some data science projects on purpose, for keeping current on a common language (science reviews and governance), or for proving out ideas that didn’t have enough concrete evidence to be worth tens of millions of dollars a year.

The integration was the part I liked. Strategic alignment, org change, the relationship on the other side of every handoff. A continuous improvement background follows you, and working on a team’s credibility was the hook for me. You don’t get far in technology without strong relationships.

Leaning into operations

Most of our work was for operations, and that’s where we had to keep our credibility. SCA was supposed to stay close to the ground and never turn into an ivory tower.

I pushed for relationships with the field and pushed for a power users space in Power BI, and we got it. People we’d identified in operations could build and deploy their own models and semantic models there. We learned what they were using data for and folded those uses into our own work. They built credibility with their peers. And we got to poach them.

Two operations people came to SCA. One of them was my hire. He’d been building reporting on his own, and he filled a gap on our team. I had a similar strategy in manufacturing. Operations employees on rotating office assignments for continuous improvement, learning to ask questions from a different perspective, relationships on both sides, moving back to the field armed with tools to effect real change.

And strengthening the science

The data science organization was a subsidiary with its own formal process, check-ins and science reviews. The future vision for SCA was tighter integration with the science organization. I was the point for integration on our side of the fence. The first project we ran their way was a survey of the drivers delivering into our distribution centers. The existing dashboard in Qualtrics presented numbers but did nothing to inform the sites what to do.

SCA volunteered to take it over, building the data pipeline from the Qualtrics API into our Databricks data lake, modeling it into Power BI. After analysis we recommended restructuring the survey from 15 questions to 3, with a free form text box and a language model trained to read the comments, and people decided what to act on.

NPS went from 15 to 31 while the toolkit was in use. The organizational design we’d pictured from the start was coming into shape.

Reconciliation

I got us our own Unity Catalog domain for analytics products and data staging. That meant the Data Architecture Committee, and getting technology, engineering, governance, data science and the business to agree on a setup that followed the data science organization’s standards instead of technology’s defaults. It had stalled before. It went through.

The data mesh was built to support analytics, so analytics could be purely analytics. We were the people talking to the business, building the reporting and the agents. (The data teams did some of that too. It’s a mesh.) I worked with the data teams to move the pipelines I’d designed, the ones the org depended on, over to them. The next step was moving everything into the data science organization’s data workspaces. I was the owner on our side, working with a director on theirs.

When I left, two people had come to SCA from operations. A few had gone to the data science organization as data engineers, but we had a clear idea of who we were and a clear game plan on how to get there.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *