Skip to content
Upriver
Close menu

Why Data Quality Is a People Problem — A Conversation with Tiankai Feng

An edited Q&A from the Data Splash podcast, hosted by Ido Bronstein and brought to you by Upriver. Our guest is Tiankai Feng, a data and AI leader with deep expertise in data governance and the author of two books, Humanizing Data Strategy and Humanizing AI Strategy. Tiankai has worked on data throughout his career — as a product owner, analyst, data scientist, and engineer — and now focuses on strategy and governance, with a particular interest in the human side of making data and AI work.

ido.png Ido Bronstein
August 13th, 2026

This conversation has been edited for length and clarity.


The 30-Second Splash

Ido: Governance teams — sherpas or police?

Tiankai: Sherpas.

Ido: What breaks first, the model or the metadata?

Tiankai: Metadata.

Ido: Dream vacation destination?

Tiankai: Maldives.

Ido: Which instrument explains data governance best?

Tiankai: The conductor's staff.

Ido: Is AI going to replace data stewards? Yes or no?

Tiankai: No. It won't replace it.


Why leave analytics for governance?

Ido: In 2020 you left analytics for governance — a transition most people are afraid of, because governance is about people, not technology. What attracted you to it?

Tiankai: The reason I switched was driven by my passion for data quality. In my analytics role I was dealing a lot with internal product data, and I realized the quality really was suboptimal — a lot of issues in timeliness, completeness, and accuracy of the data.

I was basically a frequent complainer to the data governance team. Look, there's something wrong with the data quality here. Can we please fix this?

And at some point they said: do you want to help solve the problem, instead of just complaining about it? Why don't you join us and we fix the problem together.

That sounded like a really interesting proposition. Let me see what's happening there and why it's so hard to fix. So I switched to the governance side — and I realized why it's so hard to fix.

Ido: If you cannot win them, join them.

Tiankai: That's true. But it really comes down to mindset. Especially in an enterprise environment, it's so easy to blame others for things not being right and just sit back and wait for someone else to fix it. What's much more valuable and much more effective is that if you know how it should be correct, you use your own expertise to help fix the problem together. Instead of saying it's their fault and it's not my problem, and hiding from it.

Why don't we come together more often and bring our expertise together to fix problems, instead of throwing it over the fence and waiting for it to be fixed?


So why is data quality so hard to fix?

Ido: You said that when you moved to the governance side, you understood why it's hard. Our audience is mostly data engineers, and they're used to complaining too. Help them understand what makes this so difficult.

Tiankai: There are two things. The first one is the overall mismatch between the intended usage of the data and the actual usage of the data.

Most of the data we create in organizations for the first time is created for operational reasons. We set up product data in product management systems, customer data in CRM systems, and so on. All of these things exist because we want to use the data directly to do something with it — logistics, marketing, whatever it is.

But over recent years the use cases for that data became so much more. Suddenly we use analytics to create reports, we create data science models to predict things, we use generative AI to create new things with it. Those are all use cases that were not initially thought of when the data was created.

So the requirements toward the data and the actual users of the data don't really come together nicely. The first step is to understand all the requirements toward the data that actually exist — not just the ones we assumed in the beginning.

Ido: And once you have the requirements?

Tiankai: Then we know exactly how we define good quality. And then we realize the data isn't matching the quality we want. So we do root cause analysis, and it usually boils down into three aspects: human error, process error, and technical error.

Technical error is the most straightforward. A pipeline broke down, a system is down, something is off. We fix it, we prevent it from happening next time, and it goes well.

A process problem is trickier, because it's about time and sequence. For example, the data is only ready two weeks before the deadline, but they actually needed it six weeks before the deadline. So we have a structural problem there.

The last and most difficult one is human error — when the data is originally created as a manual entry by one human being. When that happens and it doesn't match the requirements, the only person who can fix it is the human who put the data in in the first place, because there's no other reference to what is actually correct. We can only say that it's wrong, but only one person knows what the correct record looks like.

That means we have to train people to do it better, and create better interfaces for them to input the data. The more human-centric the source of the data quality problem is, the more difficult it gets. And as data engineers who are really just downstream — how do you train people working in completely different departments how to input data? It becomes a much bigger change management effort.

Ido: That really resonates with how I felt about these problems in my previous work as a data engineer. The things under your control, the technical part, are always easier. It's first of all your fault — okay, I built a pipeline that isn't schedulable, I had an edge case I didn't think about. But when it comes to human things that are out of your control, that's when you need organizational alignment, first of all on the fact that this is important and we need to invest in it. A lot of the time people don't understand that.


What does good governance actually look like?

Ido: You've taken this journey from crawling to running in both enterprise and consulting. What are the transitions a company goes through when governance is working? What does success look like?

Tiankai: The main maturity we can measure is how we move from being reactive to being more proactive.

Reactive means a lot of bad things in data are happening, and data governance comes in and tries to fix it with reactive solutions. Data quality is bad — let's come up with a data cleaning algorithm. There's a compliance issue — okay, let's pay the fine first, and then we're going to figure out a policy to prevent this from happening next time. You're fixing pinpoint issues, trying to fix whatever is burning.

Then slowly it becomes more routine, and at some point further down the maturity line you see that we can prevent things too. Not only reactively fixing stuff — let's build in guardrails, let's build in policy as code, let's build in automatic processes that help us predict what might go wrong so we can prevent it from happening in the first place.

Then we actually move away from reactive firefighting into planning ahead. We can build the right modules and the right building blocks to prevent these issues in the future. And when governance is more embedded — seen as the default, so that every data product and every data set and asset in the future has governance in mind and isn't just reactively repurposed — I think the maturity has really increased.


Why do use cases matter so much?

Ido: Something you said earlier really resonated — you need to start with the use case. Having a clean table with zero nulls is nice, but it's not a goal. The goal is how you drive business value. The organizations I've seen in a very good state from a governance perspective didn't chase fixing data quality here and finding PII there. They asked what the business wants to achieve with the data right now, and made sure all the stakeholders involved in that business use case were aligned.

Tiankai: Absolutely. On a psychological note, use cases have purpose. If you just create something so it exists, it doesn't really have a purpose. And without purpose, it's really hard to convince people.

The good thing about use cases in an organizational context is that they usually come with clear requirements and a clear quantification of the impact as well. So not only can we say this is exactly what we need from this — we can also say we need this so we can make so much money, or save so much money.

Having that kind of surrounding is an undeniable justification. It's not that it's unclear what we want to do and we're all misunderstanding each other. We can say we have to do this, because otherwise we're not able to achieve that impact. So we better all get motivated to really do this.

In a way it becomes an objective truth. We need to do this, and we all have the facts to back this up — let's get this done. That's how governance can really bloom, and not just become a fluffy thing on the side.


How did AI change the governance mandate?

Ido: Where do you see data governance in the role of making an organization AI-ready?

Tiankai: AI and all the hype around it has forced data governance not only to become more the center of attention, but also to evolve.

With AI, we all try to make things work, and the moment we put it into production we realize the data isn't really supporting the use case enough. So what can we do? Suddenly it's all a problem — we should have governed the data better is the conclusion. And then we go to governance.

But then we realize that even with the traditional governance concepts, it's not covering all of the new things we need to do. Most importantly, for example: reporting use cases only needed maybe two to three years of data at maximum. Now suddenly with AI we want a lot more historical, granular data to train models. So we have to go back in time a lot more, but all of that was never in scope for governance. Suddenly we have to go back all that time and realize what data is usable and what isn't. That's a really tricky environment to be in.

The other one is that with generative AI, the focus shifted from structured data a lot more to unstructured data. And data governance in recent years has focused much more on just structured data. So governing unstructured data in documents and pictures and videos is really a new thing. Suddenly we have to understand what the difference is between governing normal tables and actual PDFs on SharePoint. How do we even do this? How do we create a governed way around it?

Those are the main issues — why data governance is now more in focus, but is also forced to evolve to meet the requirements of AI work.


How do you govern unstructured data?

Ido: This feels like a whole new approach we need to understand. How do you think about it?

Tiankai: My first reaction is that a lot of the quality comes from curating unstructured data properly.

We all know how SharePoints and shared drives work. We have fifty different versions of the same document in one folder — PDF version one, PDF version fifty, PDF final. All of these versions are in the same folder.

And if we just let an AI go on it, an agent or a model, it would be completely confused by all of the different versions and the conflicting information in there.

So in that analogy, we as the data experts and the business experts need to curate better. What is actually the source of truth? What are the correct and up-to-date documents? Label them or curate them first. Once that happens, we can create better ways to save them in knowledge products and combine them with semantics and all these hype words around context. That's helpful. But it starts first with even knowing and understanding which are the right documents, and which unstructured data we want governed, to have a clear scope.

Ido: It sounds like exactly the problem we solved ten years ago in structured data. Curating the data, making sure we have one source of truth. We just need to do it with different tools today — maybe semantic tools, maybe different storage behind the scenes. But data governance still has the same ideas about how it enables organizations.

Tiankai: Exactly.


How is AI helping governance teams?

Ido: And the other direction — how does AI help you as a data governance steward?

Tiankai: This is where I'm really optimistic, and I'm seeing really good things in the market already. AI capabilities have started to genuinely support the day-to-day tasks in data governance and data stewardship, simply because they take away a lot of the tedious, repetitive work.

Take basic policy writing. Now that you have generative AI, you can let it draft something up for you and finalize it with some human touch. You don't have to write it all from scratch anymore.

Same with data profiling. There are really smart ways for it to run automatically and surface anomalies and patterns, without you having to run all the queries on your own and then check the results yourself. And the same with monitoring and alerting — it can identify incidents, it can already identify drift for you. You don't have to build all the solutions on your own. Most of these things come out of the box now.

And when we think about agents on top of that: as a steward you don't have to click manually into a dashboard every day anymore. An agent can capture all of that for you, inform you about what you need to know every day, and propose actions for you that you just have to approve.

Governance can become so much leaner, and we as humans can focus much more on making the right decisions and deciding the right things. We're not so occupied anymore with the tedious manual tasks. I think that's a really good thing.


What's next for the data governance songs?

Ido: You have to reveal what's going to be in the next hit on the data governance songs you're making.

Tiankai: To be honest, I don't know yet. Even my first one was an unexpected hit — Governors of Data. I didn't expect it to be still so relevant. Now it's four years later and people are still sharing it with each other and using it to communicate data governance better. I'm very happy that I can use my music to make it more approachable. But I don't really have a plan for my music. It just happens when I have an idea, and then I just do it.


What should a new data steward do today?

Ido: If you had to recommend to a new steward what to do today in order to be the best steward five years from now — what should they do?

Tiankai: Two things.

One is, never forget to focus on what's important. And that will evolve over time, because what is important today might not be important anymore tomorrow. So always keep an eye on business value and on skills — what are the right problems to solve, and do you have the right skills to solve them?

The second one is don't just stagnate in your day-to-day. If you see any way of optimizing your day-to-day, with AI or without AI, go for it. It never hurts to reflect and optimize your workload. It takes some conscious decisions to do that, because it's so easy to get stuck in your day-to-day routine. But we shouldn't afford that, because we can do things so much more efficiently nowadays.


Closing takeaways

Ido: A few takeaways I'm leaving with:

  1. Start with the use case, not the cleanup. A use case comes with clear requirements and a quantified impact, and that combination is what turns governance from a fluffy side project into something the organization can't argue with.

  2. Data quality failures are technical, process, or human — and they get harder in that order. The technical ones are yours to fix. The human ones require change management in departments that have never heard of you.

  3. AI didn't just make governance more urgent, it changed the scope. Models want far more historical data than governance ever tracked, and generative AI moved the problem into unstructured documents most organizations have never curated.


If you enjoyed this conversation, follow the Data Splash podcast. We have a lot more coming. Until next time — keep your data flowing.

For the full episode:

Watch on Youtube: https://www.youtube.com/watch?v=1LbCGw0gxnI
Listen on Spotify: https://open.spotify.com/episode/5XgM9sCvw1QaeTpIAy4Nz1?si=gj18Bf1ZRi63611XFiMKOQ


To follow the Data Splash:

Youtube: https://www.youtube.com/playlist?list=PLJboMtNQp-MtIK1v7fn6eA35Ylh2T85IN
Spotify: https://open.spotify.com/show/4MD8LSNbpzW46qeq3qptpr


Related Articles

AI
Data Quality
How Real-Time Web Data Is Becoming a First-Class Citizen in the AI-Era Stack - A Conversation with Amaury Desrosiers

An edited Q&A from the Data Splash podcast, hosted by Omri and brought to you by Upriver. Our guest is Amaury Desrosiers, Director of Solution Consulting at Nimble, a real-time web data platform that deploys AI-powered web search agents to browse the live internet, verify what they find, and turn it into clean, structured data that AI agents and analytics can actually use. Amaury leads the solution consulting team, translating what customers need into how the product can solve it.

omri.png Omri Lifshitz
June 17th, 2026
Data Quality
Technical
AI
How OpenLineage Is Becoming Infrastructure for the AI Era — A Conversation with Harel Shein

An edited Q&A from the Data Splash podcast, hosted by Ido Bronstein and brought to you by Upriver. Our guest is Harel Shein, Senior Engineering Manager at Datadog, where he works on data observability products. Harel previously led data engineering and R&D teams at WeWork and Astronomer, and is a Technical Steering Committee member and contributor to OpenLineage, the open standard for data pipeline lineage.

ido.png Ido Bronstein
May 19th, 2026

Your coffee can wait.
Your data can’t.

Bring your messiest ticket. Our agent solves it before your coffee gets cold.

By clicking Accept, you agree to the storing of cookies on your device to enhance site navigation and analyze site usage. View our Cookies Notice and Privacy Policy for more information.