Contáctenos
Todos los artículos

IA y aprendizaje automático

AI-Ready Data for Agents: Five Contracts

2 de octubre de 20268 min de lecturaPor Equipo de ingeniería de Bayseian

Enterprise agents need more than connected data. Define permissions, business meaning, freshness, action rules and evaluation before giving an agent operational access.

Este artículo está disponible actualmente solo en inglés.

An AI-ready data foundation needs a warehouse or lake, but the storage layer is only the start. Teams also need contracts that tell an agent what it may see, what the data means, how current it is, which actions are allowed and how the result will be judged.

That distinction matters as agents move from answering questions to changing records, launching workflows and creating applications. A retrieval system can find a plausible document. An operational agent also needs permission-aware context, stable business definitions and a safe path from a recommendation to an action.

A small autonomous agent approaches a central semantic map connected to governed data layers

Microsoft's Fabric and SQL announcements on September 28, 2026 make this architectural shift unusually visible. The update connects Fabric IQ, OneLake, semantic models, ontologies, data agents and operational controls. Some capabilities are generally available, while others remain in preview or are still planned. The useful lesson is broader than one platform: enterprise teams need to manage business context as deliberately as they manage tables and files.

What changed in the latest Fabric announcement

Microsoft said Fabric IQ context is generally available in Copilot Chat and Cowork. It also introduced or expanded several preview capabilities, including IQ sharing, a data engineering agent, ontology creation and Fabric observability. Google BigQuery mirroring was described as generally available, while bidirectional Salesforce Data Cloud 360 integration was announced in preview.

These labels are important. A preview can be useful for a controlled pilot, but it should not quietly become a production dependency. Procurement, support, regional availability, data residency and service-level expectations still need separate review.

The announcement also shows why "connect the data" is too vague. The proposed foundation contains several different layers:

| Layer | What it supplies to an agent | Failure when it is missing | |---|---|---| | Source and access | Records, documents and identity-aware permissions | The agent sees too much, too little or a copied dataset with different controls | | Semantic context | Metrics, terms, relationships and business rules | The agent answers with the wrong definition of revenue, customer or risk | | Operational interface | Approved queries, tools and write paths | A correct answer becomes an unsafe or non-idempotent action | | Observability | Traces, lineage, cost and outcome records | Teams cannot explain what happened or improve the system | | Evaluation | Test cases and acceptance thresholds | A polished response passes even when the business result is wrong |

This is the difference between a data connection and an agent-ready contract.

Contract 1: identity and permission propagation

Start with the user and workload identity, not the index. An agent should retrieve only what the requesting person and the agent's service identity are allowed to use for that task. If a platform mirrors or shortcuts data from another system, test whether source permissions, row-level rules and revocations still behave as expected.

A practical acceptance test should cover at least four cases:

  1. An authorized user can retrieve the expected record.
  2. An unauthorized user cannot infer that the record exists.
  3. Revoking access changes the agent's result within the agreed time.
  4. The audit trail identifies the user, agent, source and policy decision.

This is closely related to deciding which systems to connect first. Sources with unclear ownership or entitlements are poor first candidates, even when their content looks valuable.

Contract 2: shared business meaning

Most enterprise questions contain terms that look obvious until two teams answer them differently. "Active customer," "net revenue" and "late delivery" may each have several valid definitions. Embeddings do not resolve that disagreement.

The context layer should record:

  • the canonical definition and its owner;
  • the source fields and transformation logic behind it;
  • valid dimensions, units and time windows;
  • exceptions and regional variations;
  • the date the definition changed.

Microsoft's announcement points to semantic models and ontologies as ways to carry this meaning into agents. The technology choice can vary. The requirement should not: definitions need owners, versioning and tests. An ontology generated from existing assets is a starting point, not proof that the meaning is correct.

Contract 3: freshness, lineage and confidence

Agents often combine a live operational record with slower documents and derived metrics. Without explicit freshness, a confident answer can mix values from different business moments.

For every important source, define the expected update interval, the last successful refresh, the lineage back to the source and what the agent should do when the source is late. Useful responses may need to say, "Inventory is current as of 09:42 UTC, but the finance forecast was last refreshed yesterday."

Do not hide uncertainty behind a single confidence score. Record whether uncertainty comes from stale data, conflicting sources, an incomplete permission path or model interpretation. Those conditions lead to different remediation.

Contract 4: read and action boundaries

Question answering and operational action should not share an undifferentiated connection. A tool that can read an order does not automatically need permission to refund it.

For each action, define:

  • the allowed caller and purpose;
  • required parameters and validation rules;
  • idempotency behavior;
  • monetary, volume or time limits;
  • approval and escalation conditions;
  • a reversible or compensating step;
  • the event written to the audit log.

This is where API design for AI agents becomes part of data architecture. Clear schemas, actionable errors and idempotency keys reduce the chance that an agent turns an ambiguous request into a repeated side effect.

Contract 5: evaluation against business outcomes

An agent-ready foundation needs a fixed evaluation set before a broad rollout. The set should include permission boundaries, ambiguous terminology, stale sources, conflicting records and failed tool calls. Score the outcome that matters, not only the wording of the answer.

For an analytical agent, that may mean the correct metric, filter and citation. For an operational agent, it may mean the correct state change, no duplicate write and a complete audit record. Include cases where the right result is to ask for clarification or route the task to a person.

Keep a small regression set for every data contract. Run it when a source schema, semantic definition, policy or tool changes. Production traces can suggest new cases, but sensitive data should be reviewed and minimized before it enters a test set.

A 30-day implementation sequence

The safest starting point is one bounded workflow, not an enterprise-wide context layer.

Week 1: choose the decision

Write down the user, the decision or action, the systems involved and the cost of a wrong result. Exclude sources without a clear owner.

Week 2: define the contracts

Document permissions, five to ten key business terms, freshness thresholds and the read/write boundary. Turn each one into an acceptance test.

Week 3: connect and observe

Connect the smallest useful source set. Capture source references, policy decisions, tool calls, latency and cost. Keep write actions behind approval.

Week 4: evaluate and release narrowly

Run the fixed test set, review failures by category and release to a small authorized group. Expand only when the team can explain failures and measure the operational result.

How to assess the Microsoft Fabric options

The September 28 release is useful evidence of where enterprise data platforms are heading, but the feature list should not replace architecture review. Confirm the exact status and regional availability of every dependency. Test permission propagation across mirrored or shared data. Decide who owns semantic definitions. Check whether telemetry is available to your security and operations teams. Keep preview services outside critical paths unless the risk is explicitly accepted.

For teams already using Fabric and Power BI, the new context integrations may reduce duplicate plumbing. Mixed-cloud teams should test whether shortcuts, mirroring and shared semantic assets preserve the controls they already rely on. Teams on another platform can apply the same five contracts without reproducing Microsoft's product stack.

The practical goal is to give each agent a governed, testable route to the small amount of context and action it genuinely needs. That rarely requires centralizing every byte. If you are planning that route, Bayseian's AI systems team can help define the contracts, evaluation set and production controls before the implementation grows.

Sources

IA empresarialAgentes de IAData ArchitectureGobernanzaMicrosoft Fabric

¿Trabaja en algo similar?

Sin discursos de venta: solo una conversación práctica con el equipo que construye y opera estos sistemas en producción.

Inicie una conversación