DataFab / Platform / Knowledge Fabric
Layer 1 · the foundation
Inside the Fabric, discovery to ontology.
The foundational data-integration and intelligence layer of DataFab. It discovers and resolves the systems you already run into one usable knowledge graph — the ontology is derived from your estate, the data is read in place, and the knowledge state is kept current by the fabric itself.
01 · foundation
An access enabler, not a repository.
A metadata-driven data fabric that gives unified access to distributed data assets — maintaining mappings and relationships rather than duplicating data. One source of truth, unified analytics, nothing moved.
Knowledge graph
Graph-native storage with entity resolution at its centre.
Entity resolution
Cross-source matching and linking into golden records.
Connectivity
Direct database, API and MCP connectors, generated as well as consumed.
Data flow
Read from and write back, through the same governed connectors.
Discovery
Automated identification and cataloguing of the estate.
Active metadata
Continuous analysis, classification and enrichment.
External sources
Governed open-source and third-party integration, with provenance.
Data lineage
End-to-end tracking of every field from source to graph.
The category, and the difference
Beyond the traditional data fabric.
A traditional data fabric helps you integrate and access distributed enterprise data. It solves reach: one way in to many systems, without another migration. That is real, and DataFab does it — a governed connector layer across databases, APIs, documents and streams, read in place.
DataFab goes further. It discovers and catalogues what those distributed systems describe, resolves the records they hold into entities, derives a semantic model from what it finds, and maintains that as a governed knowledge state — without requiring the underlying data to be moved into another repository. An enterprise data fabric gives you access to the rows. A knowledge fabric gives you the customer those rows are about, the relationships between them, and the provenance for every one.
What a data fabric layer gives you
- Unified access to distributed data across the estate
- A metadata catalogue of assets and where they live
- Virtualised queries without physical consolidation
- Governance policies applied at the access layer
- Answers in the vocabulary of tables and columns
What the Knowledge Fabric adds
- Entity resolution — the twelve records that are one customer, merged and reversible
- A semantic layer derived from your sources, not hand-modelled up front
- A knowledge graph of typed, dated relationships with confidence and source on every edge
- Derived relationships that appear in no source record, carrying the rule that produced them
- Governance enforced at traversal, so reasoning never sees what it may not see
- Answers in the vocabulary of the business — and the evidence for each one
Data fabric, data mesh and data virtualisation are architectural patterns for reaching distributed data. This is the layer above them: what the estate means, kept current and governed, on data that never moves.
The anatomy
Components, output, ecosystem.
The Fabric is not one service. It is a defensible core wrapped in three shells: the components that do the work against your sources, the outputs those components produce, and the ecosystem that puts those outputs to work. Each shell inherits the governance of the one beneath it, so nothing at the edge can act outside what the centre allows.
What runs against your estate
Source connection and open-source ingestion, discovery, lineage, entity resolution, quality metrics, active metadata, schema management and ontology management. Nine services, all read-only against the systems you already operate.
What the components produce
The knowledge graph, reasoning and inference over it, a data catalogue, observability, governance, federated search, data unification and the agentic data layer that exposes all of it safely.
What consumes the output
The Knowledge & Agentic Studio, MCP Studio and vision-language-action capability — the surfaces where people and agencies turn a governed knowledge state into governed work.
The approach
From connection to knowledge.
Point the fabric at a source. The discovery, resolution, ontology and upkeep happen inside it — the knowledge state forms, and maintains, itself. You do not model the enterprise before the fabric can understand it; the model emerges as the fabric understands the enterprise.
Connect
Point the fabric at a source you already run — database, API, event stream, document store or MCP endpoint. Read-only, with your credentials, over the connection pattern your network allows.
Discover
It catalogues the assets and relationships across the estate. Types, keys, distributions, anomalies and undocumented join paths — read rather than described in a document nobody has updated since the last migration.
Resolve
Records are matched and merged into golden records. Deterministic and probabilistic scoring, active learning and structural similarity decide that six records are one real entity — thresholded, reversible and logged with its basis.
Derive
A schema-bounded ontology emerges from the sources — entity types, relationship types and attributes, mapped source to graph. The estate proposes its own model, and review starts from evidence rather than a blank page.
Maintain
Change detection keeps the knowledge state current, continuously. A hand-built ontology goes stale the moment a system changes; a derived one forms in place and stays current.
The modelling, the mapping and the upkeep happen inside the fabric, drawn from the sources you already run. That is the whole difference between value in weeks and value at the end of a transformation programme.
02 · architecture
Six governed layers.
Each layer carries its own controls. Requests flow down to the source; data is read in place and resolved on the way back up. The semantic layer is the Fabric’s primary output — the derived model everything above consumes.
| Layer | Name | Components | Controls |
|---|---|---|---|
| 01 | Presentation | Search UI · API gateway · platform integration APIs | Authentication, rate limiting, input validation |
| 02 | Service | Search · lineage · quality · discovery · governance | Service-to-service authentication and authorisation, mTLS |
| 03 | Semantic layer | Derived ontology · entity, relationship and attribute types · source-to-graph mappings | Schema-bounded, human-approved, versioned — the model every agency reasons over |
| 04 | Knowledge graph | Entity store · relationship store · query engine | Encryption at rest, access-control lists |
| 05 | Connectivity | Database · API · MCP · event streams | Credential vault, secure connections, sampling only |
| 06 | Source systems | Customer-managed data — read in place | Customer-managed, customer credentials |
03 · entity resolution
A name match is never an identity.
Records from every source are blocked, scored, clustered, reviewed and merged into one authoritative golden record — with a confidence score and full provenance. Uncertain matches go to a person, not a guess.
Blocking
Reduce the comparison space to candidate pairs.
Matching
Deterministic and probabilistic scoring across exact, fuzzy-name, phonetic and address-standardised methods.
Clustering
Group records at configurable thresholds you control.
Human review
Uncertain matches are routed to a person and the decision is logged.
Golden record
One merged, authoritative entity — provenanced, scored and reversible.
04 · cross-lingual resolution
One identity, in every language.
Real estates do not speak one language. The same identity arrives rendered across scripts and spellings, and different people arrive sharing one name. The Fabric handles both at the ontology level: it reasons over entities, not strings.
This is not translation. A name written in Arabic, Cyrillic, Han or Latin script — in any of a dozen renderings — resolves to one entity. Transliteration and cross-script matching are native, seamless and governed, so the picture is whole no matter what language the source speaks.
Variants and aliases merged
Spelling variants, aliases and transliterations collapse into one golden record — reversible and provenanced.
Same name, different person
Different entities that share a name are held apart by context: biography, relationships, geography, time and source.
100+ languages and every major script
Matched phonetically and semantically rather than by spelling.
At the ontology level
Resolution is a property of the graph, not a feature of one tool — so every layer and every utility inherits it.
A name-match that merges two different people is worse than no match at all.
Every merge and every split is provenanced, scored and reversible05 · data flow
Resolve in place.
Direct access without extract-transform-load or replication. The fabric reads from source systems and resolves on the way back — the records never leave, and there is no integration layer to engineer and then maintain forever.
Remains in source — authoritative records
- Actual records — names, addresses, identifiers
- Document files — case files, contracts, evidence
- Financial data — in the source systems that own it
- Communication logs — in their native systems
- Operational data — all business data, where it lives
Stored in the fabric — mappings and derived insight
- Entity-resolution mappings — A = A across systems
- Cross-system relationships — the edges discovered between them
- Investigation annotations — tags and notes
- Derived insight — conclusions, with the rule that produced them
- Temporal snapshots — point-in-time state so an answer survives the source moving
06 · discovery & active metadata
It catalogues itself, continuously.
Crawlers and profilers catalogue the estate you already run — metadata and statistics only, no manual inventory — while active metadata classifies, monitors quality, tracks lineage and detects change, so the knowledge state maintains itself.
Discovery
| Crawler engine | Traverses source structures |
| Schema extractor | Read-only table and column metadata |
| Profiler | Statistical profiles, sampling only |
| Classifier | Machine-learning detection of personal and sensitive data |
Classification
| Personal data | Restricted access, masking |
| Legally privileged | Highly restricted |
| Financial | Confidential, encrypted |
| Health | Restricted, health-data controls |
| Custom | Configurable handling you define |
Lineage
| Technical | Impact analysis |
| Business | Compliance mapping |
| Column | Sensitive-data tracking |
| Operational | Audit trail |
Downstream assets inherit upstream restrictions, and the lineage itself serves as compliance and audit evidence. Inherited access-control lists are the floor — stricter, never looser — reconciled into one permission model the enterprise authors and signs off, not the vendor.
07 · the semantic layer
The ontology builds itself.
You do not hand-build the model. It emerges from what the fabric has already extracted — source structures, data profiles and resolved entities become a governed domain model. A hand-built ontology is a specialist programme measured in months; a derived one forms in place and stays current.
Extract
Read source structures — tables, columns, keys and foreign-key relationships. Read-only; metadata and sampling only.
Profile & classify
Data types, cardinality and distributions inferred; personal and sensitive fields flagged in place.
Resolve
Records matched and merged into golden records across every source.
Derive
The domain model is inferred — entity, relationship and attribute types — as a schema-bounded ontology, mapped source to graph.
Govern & refine
Seed schemas bound it, a person reviews and approves it, and active metadata keeps it current as sources change.
Derivation is constrained by seed schemas — entity, relationship and attribute types — so the fabric can enrich and extend the model but never invent entities outside it. Hallucination prevention by construction, with uniform structure and validation across every source.
The same resolved entities can be viewed for monitoring, for investigation or for reporting without re-modelling. One graph carries many logical interpretations, and each is governed.
The knowledge graph
Entities and typed, dated relationships.
The graph is where the semantic layer stops being a model and becomes an answer. Person, organisation, account, matter, document, address, identifier and event nodes, connected by typed relationships that each carry a validity window, a confidence score and the source they came from — so a traversal returns not just a path but the reason the path exists.
Nothing selected
Click any node to resolve it: its attributes, every typed relationship it carries, and the confidence and validity window on each. Derived edges are shown in gold and carry the rule that produced them.
Illustrative entities and values. Derived edges appear in no source record — they carry the rule that produced them and the premises it consumed.
Not a join guessed at query time
Entity store, relationship store and query engine, purpose-built for traversal. Relationships are first-class objects with their own type, properties and lineage rather than a foreign key somebody inferred.
Valid-from, valid-to
Every edge carries a validity window and point-in-time snapshots are retained, so “who controlled this company in March” is a question the graph answers rather than one you reconstruct.
Inference that shows its working
Relationships that appear in no source record carry the rule that produced them and the premises the rule consumed. A derived edge is never presented as a source fact.
The path respects the wall
Entitlement is enforced as the graph is walked, so reasoning runs over a view that already excludes what it may not see. Nothing is retrieved and then filtered.
| Traversal | The question it answers | What comes back with the answer |
|---|---|---|
| Ownership chain | Who ultimately controls this counterparty? | Every hop, every confidence, and the level at which it halted |
| Shortest path | How are these two entities connected at all? | The path, the edge types crossed, and what was excluded by policy |
| Neighbourhood | What sits one and two hops from this party? | Ranked by edge strength and recency, not alphabetically |
| Temporal slice | What did this picture look like on a given date? | The graph as it stood, from validity windows and snapshots |
| Community | Which clusters does this entity belong to? | Hierarchical clustering with the basis for membership |
| Pattern | Does this structure match a known typology? | The match, the rule version, and the evidence for each leg |
08 · persistent knowledge graph
Corporate memory.
A durable knowledge base the fabric derives and accumulates from your sources over time, organised as a four-level knowledge tree. Extraction is constrained by seed schemas, and every extracted element carries the record of where it came from.
Attributes
Entity property information — names, dates, values.
Relations
Entity-to-entity relationship triples.
Keywords
Semantic keyword indexing for search.
Communities
Hierarchical clustering for global context.
| Provenance element | What is recorded |
|---|---|
| Connector ID | Source-system reference |
| Document reference | Document identifier, title, location |
| Precise location | Page, paragraph, character offset, cell range |
| Extracted text | The exact text the entity came from |
| Confidence score | 0–1, carried on the edge |
| Timestamp | When extraction occurred |
09 · external sources
Enrichment, with provenance.
A registry of approved sources feeds an enrichment engine — each field schema-bound and traced from source to graph. Only approved sources, with per-source credential isolation, rate limiting, data minimisation, time-limited caching and full lineage.
Document intelligence & OCR
The documents the graph could not read.
Most of what an enterprise knows sits in documents no query can reach — contracts, policies, filings, correspondence, scanned files, handwriting. The extraction service makes them legible on request and returns the text together with the evidence of how that text was produced. It is an extraction service, not an ingestion pipeline: the source of record is never migrated or duplicated, and no secondary corpus accumulates.
Text layer, OCR, handwriting
Where a document carries a text layer it is read directly. Where it does not, optical character recognition runs with layout analysis so tables, columns and marginalia survive the transition. Handwriting is recognised where it is legible and flagged where it is not.
Resolved before dispatch
A document cannot be inspected in an external service to decide whether it was too sensitive to send there. Classification resolves from metadata the organisation already holds — before any byte moves. Unclassifiable is treated as sensitive, and the most restrictive input wins.
Per field, not per document
A single accuracy figure for a whole document tells a caseworker nothing. Confidence is carried per extracted field and per region, rolled up on the case, so the thin parts are visible before anyone relies on them.
Every extraction is defensible
Direct extraction or OCR and by which backend; the classification basis and the routing decision it drove; model version, confidence, source hash, timestamp and duration — captured at processing time rather than reconstructed later.
Extracted entities, bound to schema
What comes out is aligned to your seed schemas and enters the graph as entities and relationships with a precise location on the source page — document identifier, page, paragraph, offset, cell range — and the exact text the entity came from.
It flags rather than invents
Where confidence falls below the acceptance threshold the service marks the region and routes it for review rather than producing plausible text. Reprocessing is supported, and the second run is recorded as a second run.
Language detection, translation, document classification, redaction and semantic extraction beyond the bound schema are scoped per deployment. Throughput and accuracy targets come from the service specification and are targets, not measured results.
MCP Studio & the agentic data layer
The estate, exposed as governed tools.
Connectivity is MCP First in both directions. Inbound, the Fabric consumes Model Context Protocol connectors to reach any system you run. Outbound, it generates them — every governed dataset and every Fabric service published as a native tool an agency can call, with access still enforced at the source rather than at the tool.
Authoring, when nothing off the shelf fits
Every enterprise has a system nobody else has. MCP Studio is where your engineers define a connector against the protocol contract, upload the command, arguments, tool definitions and credential schema, have it validated for protocol conformance, schema correctness, tool behaviour and security, activate it for your tenant, and publish it into your own catalogue. No waiting on a vendor roadmap for the system that matters most.
The agentic data layer
What the Fabric generates is not a raw database endpoint dressed up as a tool. It is a governed surface: the resolved entity rather than the table, the permitted slice rather than the whole, the derived edge with its rule attached rather than a bare join. An agency calling it inherits entitlement, classification and audit automatically, because those are properties of the graph beneath the tool.
Quality, observability & federated search
The Fabric reports on itself.
A knowledge state you cannot inspect is a knowledge state you cannot trust. The Fabric tracks its own governance health against targets you set, surfaces where the estate disagrees with itself rather than smoothing it over, and makes everything it has catalogued discoverable through federated, metadata-driven search.
Find it without moving it
Metadata-driven search across every connected source, returning the asset, its owner, its classification and its lineage — then resolving to the record only if the caller is entitled to see it.
Built by crawling, not by asking
The catalogue is produced by discovery rather than by a documentation project. Assets, owners, classifications and relationships are populated automatically and kept current by change detection.
Freshness, drift and disagreement
How far behind source a resolved value is, which distributions are moving, which joins are breaking, and where two systems disagree about the same entity — surfaced as findings, not averaged away.
The same connectors, in reverse
Where an agency is authorised to act, the write goes back through the connector it read from — versioned, replayable, reversible, and behind a deterministic gate.
One view, many interpretations
The same resolved entities can be viewed for monitoring, for investigation or for reporting without re-modelling. One graph carries many logical interpretations, each governed.
Minimised by construction
Sampling rather than extraction during profiling, masking and tokenisation by classification tier, time-limited caching on enrichment, and differential-privacy techniques where an aggregate could otherwise disclose an individual.
Ontology governance · ethical walls
The wall has to hold at the ontology, not the folder.
Once an estate is resolved into a graph, the traditional information barrier stops working. Blocking a matter file does not block the conclusion, because the conclusion was never in the file — it was distributed across a billing entry, a calendar, a conflicts search and a time narrative, each of which looks harmless on its own.
Entity and relationship classes
A barrier is declared over the ontology, not over a folder: which entity types a role may see at all, and which relationship types it may traverse — in which direction. The wall is part of the model rather than a list of exceptions maintained beside it.
Derived edges are suppressed too
The fabric produces relationships that appear in no source record. Where the premise of a derived edge sits behind a wall, the edge is withheld along with it — otherwise the inference quietly leaks the thing the barrier existed to protect.
Counts can be a disclosure
“How many matters do we hold for parties in this sector?” can reveal a walled engagement without returning a single restricted record. Aggregate responses are bounded by the same policy that bounds the records beneath them.
Excluded before planning
Enforcement happens as the graph is traversed, so reasoning runs over a view that already excludes what it may not see. Nothing is retrieved and then filtered, because a filtered answer still had to be computed from the restricted thing.
By the firm, not the vendor
Inherited source entitlements are the floor — stricter, never looser — reconciled into one consolidated permission model your risk function authors, signs off and inspects rule by rule.
The refusal is evidence
Every exclusion is written to the tamper-evident trail, including the ones a user never sees. When the question later becomes whether the barrier held, the record already answers it.
Conflicts and information barriers in professional services; compartmentation and need-to-know in national security; segregation between first and second line in a regulated bank; and any estate where two teams work adjacent matters on one resolved graph. One fabric never means one undivided view.
10 · security
Governed at every layer.
Security is woven through the stack — authentication and role-based access at the top, encrypted credentials at the connectors, and a tamper-evident trail throughout.
Authentication
OAuth 2.0 with OIDC, mTLS, SAML 2.0 and API keys.
Access control
Viewer, contributor, editor, admin and auditor — with attribute-based scoping at the graph.
Credentials
AES-256, hardware-security-module backed, just-in-time and scoped. Never stored in an agent definition.
Data protection
Public through highly restricted, with masking and tokenisation applied by classification.
Audit trail
Hash-chained and encrypted, auditor role only, with no delete capability.
Governance is measured, not asserted: the fabric tracks its own governance health against targets you set. Target values are agreed per deployment.
What it builds
The resolved entity is the unit of the asset.
A customer that eleven systems could not agree on, now one governed record with its provenance, its confidence and its history attached. No competitor can obtain it, no vendor can license it to you, and no amount of model capability substitutes for it. Every source you connect adds to a holding that only you own.
Next step
Point it at one source and watch the model appear.
A data plane stands up in hours. Connecting the first sources is typically same-day. Governance is authored as code and evolves from there.