By Selvakumar, Senior Director of Technology, Eucloid Data Solutions
Every few months, a new model tops the benchmarks. Yet in the enterprise systems my teams build, model capability is rarely what limits an AI use case in production. What limits it is how little the model can access the business it serves, and how reliable that access is.
Consider a simple question for an internal assistant: how many active customers did we have last quarter? To answer correctly, the system needs to identify the authoritative data source, understand how “active” is defined, account for duplicate records and respect the user’s access permissions. A more capable model cannot compensate for missing definitions, conflicting records, or unclear access policies.
This is driving a fundamental shift in data engineering. The discipline has traditionally been viewed as plumbing: move data from A to B, clean it up and make it available. Enterprise AI expands that responsibility to making business knowledge understandable and trustworthy enough for machines to use.
I see this as a five-layer trust stack.
The data layer handles ingestion and lineage, establishing where information comes from and how it has changed.
The semantic layer gives each business metric an agreed, governed definition.
The ontology layer connects those metrics to real business entities and relationships, such as how a customer, account, and transaction relate to one another.
The context engineering layer makes relevant business knowledge, instructions and prior interactions available to AI agents, giving them a form of institutional memory.
The self-service layer lets people use that foundation to ask questions in plain language.
Governance, covering access, privacy and auditability, runs through all five layers. And the stack operates as a loop, not a pipeline: the questions people ask reveals missing definitions, relationships and context, helping teams improve the foundation underneath.
Data engineering is no longer just about moving data. It is about building a foundation AI that can be trusted to reason on. Three practical consequences follow.
Business context starts with the data foundation
Prompt engineering and document retrieval are useful, but they cannot reliably resolve underlying data that disagrees with itself.
With a large US home healthcare provider, we consolidated more than 65 source systems and 5,000 data objects, running through over 800 pipelines, into a single governed platform. Before introducing AI use cases, we established consistent definitions for core entities such as patient, caregiver and visit. The conversational assistants that followed relied on that work to ground their answers in governed, documented sources.
This becomes particularly important when people move from dashboards to conversational interfaces. A dashboard can conceal conflicting definitions behind a fixed calculation. A conversational interface exposes them as users ask variations of the same question. Does “active customer” mean someone who placed an order, holds a current contract or used the service within a specific period? The organisation needs to make those distinctions explicit before an assistant can apply them reliably.
Business rules must become machine-readable
Much of an organisation’s operating knowledge lives in places that AI cannot reliably access: an analyst’s SQL, a spreadsheet formula or an unwritten convention within a team.
A lender, for example, might use different delinquency thresholds for different products. A retailer’s finance team might exclude orders reversed within 48 hours from a revenue metric, while the sales team tracks them separately. These distinctions can be legitimate, but the system must know which definition applies to which question.
Metric definitions and business rules need to be encoded, tested, and versioned, with clear owners and review processes. Data contracts should detect breaking upstream changes before they affect downstream answers. Where teams use different definitions, those differences should be documented and exposed explicitly.
The semantic and ontology layers make these distinctions usable, helping an assistant identify both what a metric means and how it relates to the business question being asked.
Context drifts, so it must be monitored
A semantic layer built around last year’s organization can become outdated after a reorganization, acquisition or pricing change, even when the underlying pipelines continue to run successfully.
If two sales regions merge but the region’s mapping remains unchanged, an assistant can produce misleading answers about regional performance while sounding entirely confident. The data arrived on time. The business meaning is wrong.
Context therefore needs the same operational attention as pipelines. Teams need checks for definitions that diverge across systems, broken lineage and access policies that no longer match the data they protect. Questions that users repeatedly correct or cannot get answered should feed back into the stack.
For technology leaders, this calls for broader AI readiness metrics. What share of your top 50 business metrics has an agreed, certified definition? How many are available through a semantic layer an assistant can query, rather than buried in report logic? Is there an evaluation set, built with domain experts, that tests whether answers match what the business would accept? Who is responsible for updating that context when the business changes?
As capable models become more widely available, the more durable advantage will lie in the engineering underneath: how completely, accurately and securely an organization makes its own knowledge available to AI.
That is the expanding responsibility of data engineering, and an essential condition for enterprise AI to earn trust.