The semantic mesh: activating enterprise context beyond MCP connectivity

Thank you to everyone who joined our live session The semantic mesh: activating enterprise context beyond MCP connectivity, part of our AI-Native webinar series with Zhamak Dehghani and Bill Graham!

Over the past two years, MCP has given agents a standard way to find tools, find documentation, authenticate, and reach the right endpoints. The innovation it has unlocked is hard to overstate.

But connection is not comprehension.

Reaching an endpoint tells an agent where information lives. It does not tell the agent what that information means, how to interact with it, or how to validate what comes back. This session was about that gap, and about the semantic capabilities now built into every Nextdata autonomous data product to close it.

A quick recap: two bets that technology caught up with

When we started building Nextdata OS, agents were not yet in common use. We made two bets that felt unusual to people coming from a database or data warehousing background.

The first bet was that the world needs to be organized semantic-first. Storage technology shifts as vendors jockey, but what a customer means to your business, what a supply chain means, how you run operations: that business semantic is the ground truth of how you operate. Meaning matters.

The second bet was that every data product should be a running kernel: an engine independently addressable on the web, owning its own lifecycle from producing to sharing to governing its data.

With agents coming into common use, both bets have become essential features for the market. Agents need semantics to use data correctly. And because each data product is already a running kernel, it can serve MCP endpoints out of the box. We have had that capability for a while. This session showed the next level.

The real problem: connecting agents to what makes you different

Every organization knows its differentiation will not live in the LLM everyone else is using, or in weights trained on publicly available data. It lives in the data, knowledge, documentation, and conversations that are unique to the enterprise.

So there is enormous optimism about agentic workflows and AI assistants moving the top and bottom line. And there is an equally large challenge: getting enterprise context to those agents in a way that is reliable, safe, and sovereign.

Broadly, there are two ways to connect the two.

Path one: the monolith

Data, compute, access control, semantics, models, and the agent harness all come from one stack. If you can standardize everything on that stack, it is a reasonable choice. When it works, it feels like magic.

But it comes with costs:

  • Movement of data or movement of control. Data in SaaS applications has to be extracted. Data under another warehouse or lakehouse can be viewed remotely, but control follows storage. That migration takes quarters or years, not a weekend, and you run two roadmaps while it happens.
  • The landscape moves faster than the migration. By the time your data is under one roof, AI has shifted twice.
  • Deep lock-in. Not just compute and control, but your semantics, your meaning, and your agentic interfaces.
  • Trusting the magic. When the answer is wrong, it is hard to see which semantics, which tables, and which steps produced it.

If your whole organization fits in that one bottle, this path may work for you. Most enterprises do not fit.

Path two: the decentralized hourglass

This session was for organizations with real complexity: many business units, many platforms, many clouds, mergers and acquisitions.

In the decentralized path, your structured data, documents, and unstructured content stay where they are. Any agent harness and any model at the top connect through a narrow waist of standard protocols to any information at the bottom, in whatever form and wherever it lives. No migration, freedom to choose and to change, and no magic to trust.

What belongs in the squishy middle?

When the internet happened and we all moved behind screens to trade, communicate, and get work done, we trusted a narrow waist: TCP/IP and HTTP. The hard question for this wave of technology is what that waist looks like for data.

How do you give any agent, yours or rented, safe and deterministic enterprise context, while keeping your choice and keeping your control?

Five layers, five owners, five systems

Ask the industry how to connect agents to organizational context and every vendor answers through its own prism. A useful map is Sanjeev Mohan's FAQ on metadata, semantics, taxonomy, ontology, knowledge graphs, and context, which breaks the problem into five layers:

  • Data and metadata. The raw data plus data about it: schemas, owners, freshness, lineage.
  • Ontology. A formal model of the business: what kinds of things exist, their attributes, and how they relate. For a retailer: a product is sold through a channel.
  • Knowledge graph. The ontology filled with real entities. SKU 4471 in the catalog is the same shoe as a Shopify product ID and a distributor code.
  • Semantic layer. Metrics, KPIs, and calculation rules, born in the BI era. Sell-through is the share of what you shipped that end customers actually bought: ship 1,000 pairs of trail runners, sell 700, and sell-through is 70%.
  • Context layer. The newest layer: runtime assembly, the moment an agent pulls from all the others to answer a question like are trail runners selling through, or are distributors just stocking up?

None of these layers is wrong. Agents need all of that information.

The problem is that there are five layers, five owners, and five systems, out of sync. Meaning is defined top-down by a central team or system. The pipeline that actually produces the sell-through data has no say in how sell-through is defined. The people who produce the data are not the people who define it.

So agents, or people, are left to stitch the layers together and keep them in sync.

The alternative: slice the layers by domain, and let the producer own the meaning

Instead of five horizontal layers owned by five teams, slice every layer by domain or domain concept. Each slice is owned by an autonomous data product. One data product per domain concept, or however you choose to break it up. One owner per product: the team that actually knows the data.

Inside an autonomous data product:

  • Data and metadata are managed by the product, while the data itself stays on Snowflake, Databricks, OneLake, or wherever you want it.
  • Orchestration of the code that produces the data runs inside the product.
  • Ontology lives in the product's semantic model. Sell-through is defined where sell-through is produced.
  • Relationships between data products are expressed as links. Every element of a semantic model has a unique URL.
  • The semantic layer is simply the set of terms a data product defines and offers to others.

And the context layer? It is not a layer at all. Context emerges from the interconnection of well-defined, semantic-first data products.

We don't move your data. We move the ownership of its meaning to the data products that produce it, the ones best positioned to get it right. There is nothing to sync, because nothing is separate. Agents don't visit five systems to stitch an answer together. They follow the links.

A semantic mesh emerges

Once data products link to each other, what emerges is a semantic mesh: a web of addressable data products in which every field is linked. Semantic models live inside the data products in open file formats and are kept in sync with the data by each product's running kernel.

That mesh is what fills the squishy middle of the hourglass. At the top, choose any agent harness and any model, or all of them at once. In the middle, autonomous data products produce, serve, govern, and semantically answer agent questions through a storage-agnostic, format-agnostic API. At the bottom, your information stays wherever it is, whether in Salesforce, Snowflake, or Workday. Once there is an API in front, nothing has to move.

Inside the semantic model: meaning is the payload, MCP is the carrier

Every autonomous data product owns its semantic model. That model captures:

  • Intent and context. What the data product is about, which domain it belongs to, and the context it operates in.
  • Structure. A storage-agnostic shape of the data, whether structured or unstructured.
  • Relationships. Typed links to other data products. Is my product ID the same as, or derived from, another product's product code?
  • Terms and definitions. What sell-through means. What quarter means.
  • Constraints. The rules used to validate the data itself.
  • Answerable questions. Questions the model supports out of the box, plus narrower ones you define explicitly.

Semantic models live in plain text files, in YAML or a Pythonic DSL, under your source control. AI can generate them from your data, or you can handcraft them.

Meaning is the payload. MCP is the protocol that carries it and handles authentication and authorization. We serve semantic models through REST APIs, with MCP adaptations out of the box. Today the carrier is MCP. Tomorrow it may be something else. What outlives the protocol is your semantic model.

Four functions every agent needs

Every data product, regardless of the storage or format behind it, exposes the same four functions agents need to plan, query, and check their work:

  1. Discover. What semantics does this product serve, and what questions can it answer?
  2. Understand. Describe it to me. Look up the glossary definition.
  3. Use. Formulate a query through a structured JSON interface, not SQL, so complex calculations and non-tabular data can sit behind it.
  4. Validate. Return information that lets the harness and its skills check whether the answer is right.

Access is checked on every call, on every data product, through the same interface, with no additional integration.

What we demonstrated in the webinar

Bill Graham, Nextdata's head of architecture, walked through the demos.

1. An agent that refuses to guess

Using Claude with a lightweight skill pointed at the Nextdata MCP gateway, Bill asked the mesh a business question: compare units sold in the last seven days with the previous seven days, by manufacturer, across all channels, broken out by channel.

The agent said it could not answer. The semantic layer only tied manufacturers to in-store sales. Online and wholesale data were not in one place, and no seven-day trend existed.

That is exactly the behavior you want. Our motto has always been that bad data is worse than no data. One confidently wrong answer and a business owner stops trusting the platform. "I can't answer that" is acceptable.

2. Deploy a data product, and the agent can answer

Next, Bill defined a new data product, Sales Influence Insights, in the Python DSL. It aggregates store, online, and wholesale sales and computes week-over-week trends. Its spec declares transform logic, triggers, inputs, and outputs, with promises on each output model.

The output is multimodal: the same model is written to Databricks, Snowflake, and vector embeddings in Pinecone. Each output model carries human- and machine-readable descriptions, typed fields, dimensions, glossary terms, and semantic links. For example, the product ID in channel sales velocity is the same as the product code in the product catalog data product.

One command with the NXT CLI launched it. The new product appeared in the Nextdata Discovery Tool with its lineage, its model relationships, a trust summary of its promises and expectations, and its MCP functions.

Bill asked the same question again, without naming a data product. The agent found the new channel sales velocity model, recognized it had prior-seven-day aggregations across all channels, and returned the answer by manufacturer and by channel. It also showed the semantic query it used, measures and dimensions included.

The context changed because the data product changed. No prompt rewrite, no separate catalog update.

3. Opening the box

The third demo showed exactly what the agent sees. Bill pointed Claude at the MCP gateway and asked it to draw the semantic graph it received: data products, models, glossary terms, and the links between them.
‍

Then he asked Claude to explain the MCP calls behind its earlier answer: initialize, list tools, list models, describe models, and finally run a semantic query. The agent mapped words in the question to concepts in the graph, concepts to models, and models to a semantic query call.

No magic: what the demos proved

Two points from the demos are worth pausing on.

The answers are deterministic. The LLM translates a human question into concepts, and concepts into the primitives the data product exposes. The service generates and executes the query. The agent never writes SQL. LLMs are very good at writing queries, but they are not reliably deterministic at it. The model chooses what to ask for. It never decides how to fetch it.

Meaning is built bottom-up. People from traditional data management, steeped in OWL, RDF, and knowledge graphs, tend to imagine these layers defined by an observer who studies the information and enforces a model top-down. There are very few examples of that working at scale. In the demo, every data product that produces data defined its own relationships to the adjacent nodes it cares about. The agent saw the whole mesh without a human or a central system defining it. It is a web-based, bottom-up approach: a subtle difference, but a big one.

And because that information is bound to each data product, it stays in sync. It is not an external semantic repository that drifts the moment you create it. It is tied to the product's lifecycle, so if anything drifts, the product's promises and expectations raise alerts.

That is why we call them autonomous data products: they run the code, the pipeline, the agent code, and the prompts that produce the data, and they monitor, create, and share it with its meaning intact.

Security and guardrails when you don't control the harness

The last demo covered identity and guardrails.

Identity: an agent acting on behalf of a person

Before querying, Bill logged in to Nextdata OS and received a personal access token. The skill passes that token with every request. Nextdata maps the identity to whatever credentials the user holds on each platform, such as a temporary lease on Snowflake or an Entra-authenticated grant on Microsoft services. OAuth is supported too, and new identity protocols will be as they arrive.

Agents are not trusted because they sit inside your perimeter, your network, or behind your gateway. Access is enforced at the data product, on every call.

Guardrails as a lean skill

On the agent side, Bill showed the Claude skill that sets the guardrails. It runs in phases: bootstrap, locate the mesh and check authentication, discover models and tools, check catalog coverage, map language to measures, dimensions, and filters, then pass an intent gate before anything runs. The gate checks coverage, relevance, a plain-language restatement in the endpoint's own concepts, and whether the answer can be given authoritatively. Only then does execution happen, deterministically, followed by a review of the result and any caveats.

Claude does the reasoning. Bundled Python tools do the plumbing. The service executes the query.

You buy the harness, rent it, or use one off the shelf. You rarely control it. So we split the work: skills carry procedure, products carry truth. Meaning, access, and contracts are enforced at the endpoint, so a badly behaved agent still can't see what it isn't entitled to.

Not all harnesses, LLMs, or model versions behave the same, so skills must be tuned per harness and per model. That is the cost of optionality, and Nextdata provides it. We keep skills lean, just enough to nudge the agent to do the right thing. The hard work happens in Nextdata OS, a Rust-based kernel with a common interface for whatever front end you use.

Why catalogs, warehouses, and MCP gateways are not enough

You can do some of what we showed today with other technologies. Not all of it in one place.

  • Catalogs describe the data, as they always have, but point somewhere else to get it. Their meaning is after the fact, disconnected from producers, and they govern metadata, not the data itself.
  • Warehouses and lakehouses can serve semantics and MCP where the data lives, which commits you to moving everything under one storage and compute system.
  • MCP gateways and brokers, a new category, are a front door. They provide access and connectivity, but the meaning stays implicit and disconnected from the producer.

Autonomous data products push semantics as close as possible to the producer. The moment data is produced, its meaning travels with it and does not drift. Meaning and structure are validated against the underlying bits. And access is enforced on every call, at every data product. A gateway guards the door. A data product guards the data.

Because each product is a running process, its endpoint is backed by the data, not a proxy for it. It is queryable at the source, open to any agent, and protected. Nobody has to hand-write MCP servers over tables and SaaS apps, the MCP tsunami many organizations are living through now.

From the Q&A

Are the semantic models built on RDF or OWL? No, by design. RDF and OWL assume an outside observer defining triples top-down. In Nextdata, every node is a separate engine describing its own small universe and its own relationships, and the mesh emerges bottom-up. If you need it, the emergent mesh can be translated to RDF.

Can AI derive the domain knowledge? Yes. AI can generate the semantic understanding, perhaps the first 80%, with human review for the rest. The difference is where it lands: bundled into the data product and evaluated continuously as part of its lifecycle. If the semantics drift, that can trigger another review.

Why Rust? The server-side kernel is Rust, exposed over HTTP, so you can query it from any language. We chose it, before it was fashionable for data systems, for speed, compile-time reliability, and a small memory footprint. The client-side skill bundles pre-built Python tools with no code generation, so agent time goes only to LLM reasoning.

What's next

MCP is how we carry semantic context to agents today. If the protocol evolves, we will evolve with it. The semantic model, owned by the data product that produces the data, is what is here to stay.

The AI-Native webinar series continues through the year, with more product announcements to come.

More resources

If you missed The semantic mesh: activating enterprise context beyond MCP connectivity, you can watch the webinar recording here.

Read the previous session recap: Data 3.0: The structural shift from schema-first to semantic-first

And if you're ready to see autonomous data products serve semantic context to your own agents, get started or get in touch at sales@nextdata.com.

‍

Building Data Agents Outside the Lakehouse Monolith - White paper

Download Whitepaper

Nextdata & data mesh resources

Articles, events, videos, podcasts and more that share our thinking and provide insights on data products and implementing data mesh.

Join the movement.

When data empowers everyone, it changes everything.

Let’s change the way data is created, shared, and used, forever.

Nextdata is hiring. We’re looking for pragmatic, empathetic problem-solvers who understand the needs of tomorrow and dare to challenge the ways of the past.

An error occurred while processing your request. Please check the inputted data and try again.
This is a success message.