Why I Started Building a Governed AI Data Steward
What years inside enterprise data taught me about the real bottleneck in AI.
Much of today’s conversation about AI begins with the model. Which one is most capable, which copilot fits which workflow, how many steps an agent can take on its own.
These are reasonable questions, and the progress behind them is real.
But after years working across enterprise product and data environments, I have become more interested in a question that comes earlier in the chain:
What exactly is the AI reasoning over, and can the enterprise trust it?
That question sounds simple. In practice, it is where a great deal of difficulty lives.
A model can reason fluently, summarize elegantly, and produce a confident recommendation. But if the context it was given is incomplete, inconsistent, or of uncertain authority, the outcome inherits those weaknesses.
Sophisticated reasoning over unreliable enterprise context still produces unreliable outcomes. It simply produces them faster — and often more persuasively.
This essay is about how I arrived at that view, and why it eventually led me to start building something myself.
The problem that kept reappearing
Across years of enterprise product work in healthcare, financial services, insurance, retail, and payments, I repeatedly encountered variations of the same underlying problem.
The technology was rarely the missing piece.
Most large organizations I worked around had capable platforms, skilled teams, and serious investment in data. What they struggled with was something harder to buy:
trusted context.
The patterns were familiar.
The same customer, product, or supplier might be represented differently across several systems, each version reasonable on its own terms and inconsistent with the others.
Records could be incomplete in ways that only became important once someone tried to combine them.
It was not always obvious which source was authoritative for a particular fact — or whether any source truly was.
Metadata might exist, sometimes in impressive volume, while turning that metadata into something that actually shaped day-to-day decisions was another matter.
Lineage could be documented and still be difficult to follow when a real question arrived.
Governance often sat slightly outside the workflows it was intended to guide: policies, reviews, and controls surrounding the work rather than always being embedded in how the work happened.
And much of the most valuable knowledge lived with experienced data stewards and analysts — people who knew which records to distrust, which sources carried more weight, and which exceptions mattered — but whose judgment was rarely encoded anywhere a system could meaningfully use.
Access and ownership rules could be distributed across systems, teams, and processes. Each individual piece made sense. The whole was much harder to reason about.
None of this is a criticism of the organizations or teams involved.
These are natural consequences of large enterprises evolving over many years: acquisitions, regulatory change, new channels, legacy platforms that still work, and replacement platforms that never completely replace them.
Fragmentation is not necessarily a failure of effort.
It is often simply what scale and time do to information.
Then AI arrived
Traditional enterprise applications tend to operate through relatively explicit workflows.
A rule is defined. A screen is designed. A process contains steps and approvals.
When the underlying data is messy, the damage is real, but it is often bounded by the shape of that workflow. Exceptions surface. Someone investigates. A human notices that something does not make sense.
Generative AI and agents change the stakes.
They introduce probabilistic reasoning into places that were previously deterministic. They can use tools, interpret context, make recommendations, and increasingly participate in workflows with some degree of autonomy.
That is precisely what makes them useful.
It is also what makes poor enterprise context more consequential.
An agent reasoning over conflicting customer records, or over a value whose source nobody can confidently establish, may continue reasoning over that uncertainty unless the system is explicitly designed to recognize and preserve it.
So the question I found myself returning to was not whether agents would eventually become capable enough.
It was this:
How should an AI agent reason responsibly when the enterprise itself is uncertain about identity, ownership, evidence, or authority?
I do not think that question has a single answer, and I am wary of anyone who claims it does.
But I became increasingly convinced that it was the right question to work on.
Why I decided to build instead of only study
For most of my career I have worked as a product leader: framing problems, shaping platforms, and working closely with engineering, data, architecture, and business teams to decide what should be built and why.
That vantage point taught me a great deal about enterprise problems.
It also has limits.
You can understand a system’s purpose without fully understanding its failure modes. You can discuss an architecture without having personally encountered every trade-off buried inside it.
I did not want to become someone who simply added “Agentic AI” to a profile.
If I believed the difficult part of enterprise AI involved trust, evidence, and authority, I wanted to understand those questions at the level where the trade-offs become concrete: design decisions, edge cases, tests, failures, and the uncomfortable moments when a system has to do something reasonable with incomplete information.
So I began building, independently and incrementally, what I call an Ontology-Grounded Agentic AI Data Steward.
It is an independent engineering project and, in many ways, a technical apprenticeship — an exploration of how deterministic enterprise evidence, bounded AI reasoning, and human stewardship might coexist within one governed system.
Data stewardship felt like the right setting because so many of the problems converge there:
identity, quality, ownership, evidence, and human judgment.
The project is still evolving. I am building it one capability at a time, and much of what I have learned so far has come from discovering where my initial assumptions were wrong.
One principle changed the direction of the project
One of those discoveries reshaped how I think about the entire problem.
It can be stated in one sentence:
Absence of evidence is not evidence of agreement.
Consider two customer records being compared to determine whether they might represent the same person.
Both records are missing a date of birth.
A naive comparison might treat the two empty fields as matching because, technically, they contain the same thing.
But they are identical only in what they do not tell us.
Two unknowns do not combine into a known.
Treating them as agreement quietly converts missing information into confidence that was never earned.
Put that way, it seems obvious.
Yet the same pattern appears in many forms, and it becomes more dangerous once AI enters the picture.
Language models are extremely good at producing coherent answers.
And coherence can sometimes smooth over uncertainty.
A system that summarizes, recommends, or participates in decisions needs to preserve uncertainty rather than hide it.
It should be capable of saying, in effect:
I do not know this — and that absence matters.
For me, this principle became less a rule about matching records and more a lens for the entire project.
Wherever a system reasons, I now ask whether it is distinguishing what it actually knows from what simply has not been contradicted.
I expect to return to this idea in much greater depth.
The model should not own the facts
A related conviction followed.
Some things in an enterprise are, or should be, deterministic.
Did a value come from an authoritative source?
Was a particular rule satisfied?
Does an authenticated identity hold a certain permission?
These are not necessarily matters of interpretation, and I do not think a probabilistic model should become their source of truth.
That does not make AI less useful.
It clarifies where its usefulness lies.
Deterministic systems can remain authoritative for facts and rules where that is appropriate.
AI can reason over those facts — explaining ambiguity, surfacing what deserves attention, helping prioritize work, and making complicated evidence easier for people to understand.
Humans remain essential where consequential judgment is required, particularly when evidence is genuinely ambiguous or the cost of being wrong is significant.
I have found it unhelpful to frame this as deterministic systems versus AI, as if one eventually needs to defeat the other.
The more interesting problem is deciding which responsibilities belong to which layer.
And making those responsibilities visible.
That is why provenance and auditability increasingly feel fundamental to trust rather than features to bolt on later.
Autonomy is not the goal
Much of the excitement around agents focuses on autonomy — how much an agent can accomplish without a person involved.
I understand the appeal.
But in enterprise environments, especially regulated or consequential ones, I have increasingly found another question more useful.
Instead of asking:
How autonomous can we make the agent?
I prefer asking:
What is the minimum authority this agent needs to create meaningful value?
That reframing changes the design conversation.
It encourages separating concepts that can otherwise become blurred together:
knowing who someone is and knowing what they are permitted to do;
forming a recommendation and executing an action;
a human making a judgment and a system receiving authority to act upon it.
An AI recommendation should not automatically become an action simply because it exists.
Likewise, a human judgment does not necessarily imply permission for every downstream system consequence.
There should be deliberate boundaries between those things.
Bounded agents may sound less impressive than fully autonomous ones.
I suspect they may prove more deployable.
Enterprises are more likely to trust systems whose authority they can understand than systems whose capabilities they cannot fully account for.
Sometimes the most important capability of an intelligent system may be knowing what it cannot do.
What I am exploring now
The project remains a work in progress, and I want to be careful not to describe as finished what is still being built.
At a high level, the areas I am currently exploring include:
- trusted enterprise context as a foundation for AI reasoning;
- bounded reasoning that preserves uncertainty instead of hiding it;
- human stewardship as a designed part of the system rather than an afterthought;
- identity and permissions as distinct, explicit concerns;
- provenance and auditability that make outcomes understandable;
- what policy-aware enterprise agents might eventually look like in practice.
Each of these is substantial enough to deserve its own discussion.
And I expect my thinking on all of them to continue changing as I build.
I started this work because I wanted to understand firsthand a problem I had observed from many different angles over many years.
What I have learned so far has mostly reinforced the instinct that started it.
The hardest problem in enterprise AI may not be getting models to reason more intelligently. It may be creating the evidence, context, and authority boundaries that let enterprises trust what happens when they do.
That is the problem I have chosen to learn about by building.
I’m continuing to document what I learn while building at the intersection of enterprise AI, trusted data, and governed agents.
Follow on LinkedIn · Explore the Agentic Data Steward project
- Enterprise AI
- Agentic AI
- Data Governance
- Product
- Data Platforms
- Human-in-the-Loop